Large model reasoning method, device and equipment, storage medium and computer program product

By generating and deleting high-quality thinking chains, the resource cost of large language models during reasoning is reduced, and the problem of increased inference cost caused by excessively long thinking chains is solved.

CN120106209AInactive Publication Date: 2025-06-06BEIJING HONGTENG INTELLIGENT TECH CO LTD
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202411845343.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-13
Publication Date
2025-06-06
Estimated Expiration
Not applicable · inactive patent

Smart Images

  • Figure CN120106209A_ABST
    Figure CN120106209A_ABST
Patent Text Reader

Abstract

The invention discloses a large model reasoning method, device and equipment, a storage medium and a computer program product, and relates to the technical field of data processing, and the method comprises the steps: generating a plurality of high-quality thinking chains corresponding to a to-be-reasoned task based on a high-quality sample data set, the high-quality sample data set is a data set with a thinking chain meeting a preset high-quality condition; determining optimal thinking chain data in the high-quality thinking chain; performing thinking chain deletion on the optimal thinking chain data to obtain a target thinking chain; and outputting a reasoning result corresponding to the to-be-reasoned task based on the target thinking chain through the large language model. According to the method, the optimal thinking chain data in the high-quality thinking chain corresponding to the to-be-reasoned task can be deleted, and then model reasoning is carried out through the large language model based on the target thinking chain obtained after deletion, so that the problem that in the prior art, when the large language model carries out reasoning based on the thinking chain with the too long length, the reasoning efficiency is high is solved; and the reasoning cost is increased.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of data processing technology, and in particular to large-model reasoning methods, devices, equipment, storage media and computer program products. Background Art

[0002] Large Language Model (LLM) is a deep learning model trained on massive text data. After inference, the trained large model can not only generate natural language text, but also deeply understand the meaning of the text and handle various natural language tasks, such as text summarization, classification, question answering, translation, and document summarization. Therefore, it is increasingly widely used. In practical applications, various complex tasks such as mathematical logic operations, common sense, and symbolic reasoning require strong reasoning capabilities from large language models. The current main solution is to improve the performance of large language models on multiple complex reasoning tasks through the Chain-of-Though technology. Among them, the Chain-of-Though can decompose complex problems step by step into multiple sub-problems, so that the large model can gradually solve each sub-problem, thereby improving the accuracy of the final generated results.

[0003] At present, when large language models perform reasoning based on thought chains, they usually generate multiple candidate answers at one time through a beam search model, and then randomly select the answer with the lowest confusion and thought chain content from the correct candidate answers as training data to guide the model to perform reasoning. However, when the length of the selected thought chain is too long, the content that the model needs to generate will increase accordingly, resulting in an increase in the resource cost of large model reasoning. Summary of the invention

[0004] The main purpose of this application is to provide a large model reasoning method, device, equipment, storage medium and computer program product, aiming to solve the technical problem in the prior art that when a large language model is reasoned based on a thought chain that is too long, the reasoning cost will increase.

[0005] To achieve the above objectives, the present application proposes a large model reasoning method, the method comprising:

[0006] Generate a number of high-quality thought chains corresponding to the task to be reasoned based on a high-quality sample data set, wherein the high-quality sample data set is a data set containing thought chains that meet preset high-quality conditions;

[0007] Determining optimal thinking chain data among the high-quality thinking chains;

[0008] Performing thought chain deletion on the optimal thought chain data to obtain a target thought chain;

[0009] The reasoning result corresponding to the task to be reasoned is output based on the target thinking chain through the large language model.

[0010] In one embodiment, the step of generating a plurality of high-quality thought chains corresponding to the task to be reasoned based on the high-quality sample data set includes:

[0011] Acquire an initial data set corresponding to the task to be inferred, wherein the initial data set is a data set of task answers corresponding to the reasoning subtasks in the task to be inferred without a thought chain or without a thought chain;

[0012] Determining text prompt information corresponding to the reasoning subtask in the initial data set;

[0013] Based on the text prompt information and the high-quality sample data set, a plurality of high-quality thought chains corresponding to the task to be reasoned are generated.

[0014] In one embodiment, the step of generating a plurality of high-quality thought chains corresponding to the task to be reasoned based on the text prompt information and the high-quality sample data set includes:

[0015] Acquire contextual example information from a high-quality sample dataset based on the reasoning subtask;

[0016] Constructing task prompt information corresponding to the inference sample in the initial data set based on the text prompt information and the context example information;

[0017] A large language model is used to generate a number of high-quality thought chains based on the task prompt information.

[0018] In one embodiment, the step of determining the optimal thought chain data in the high-quality thought chain includes:

[0019] Filtering the high-quality thought chains according to the maximum length value of the large language model to obtain a first thought chain;

[0020] Filter the first chain of thinking based on the number of steps in the first chain of thinking to obtain a second chain of thinking;

[0021] The optimal chain of thinking data is acquired based on the example sequence corresponding to the context example information and the second chain of thinking.

[0022] In one embodiment, the step of filtering the high-quality thought chains according to the maximum length value of the large language model to obtain the first thought chain includes:

[0023] Determining the thinking chain length value of the high-quality thinking chain;

[0024] Determine the data to be filtered in the high-quality thought chain based on the maximum length value of the large language model and the thought chain length value;

[0025] The data to be filtered is filtered to obtain a first thinking chain.

[0026] In one embodiment, the step of filtering the first chain of thought based on the number of steps in the first chain of thought to obtain the second chain of thought includes:

[0027] Determine the target number of steps for the first thought chain;

[0028] comparing the number of steps in the first thought chain with the target range of steps;

[0029] The first chain of thinking is filtered according to the comparison result to obtain a second chain of thinking.

[0030] In one embodiment, the step of acquiring optimal thought chain data based on the example sequence corresponding to the context example information and the second thought chain includes:

[0031] Adjusting the example order corresponding to the context example information in real time;

[0032] Generate a plurality of rounds of thought chain data based on the adjusted context example and the second thought chain by the large language model;

[0033] Remove the farthest thinking chain data corresponding to each round of thinking chain data to obtain the optimal thinking chain data.

[0034] In one embodiment, before the step of removing the furthest thought chain data corresponding to each round of thought chain data to obtain the optimal thought chain data, the step further includes:

[0035] Building a candidate thought chain pool based on the reasoning samples;

[0036] The step of removing the furthest thought chain data corresponding to each round of thought chain data to obtain the optimal thought chain data includes:

[0037] Determine the furthest thought chain data which is farthest from each round of thought chain data from the candidate thought chain pool;

[0038] The furthest thought chain data is removed to obtain the optimal thought chain data.

[0039] In one embodiment, the step of performing thought chain deletion on the optimal thought chain data to obtain a target thought chain includes:

[0040] Determine the fragments to be removed in the optimal thought chain data by using an adaptive thought chain deletion algorithm;

[0041] The optimal thinking chain data is subjected to thinking chain deletion based on the fragment to be removed to obtain a target thinking chain.

[0042] In one embodiment, the step of determining the to-be-removed segments in the optimal thought chain data by using an adaptive thought chain deletion algorithm comprises:

[0043] Constructing a target thinking chain generation function based on the optimal thinking chain data and the removable fragments in the optimal thinking chain data;

[0044] Determining the function target of the target thinking chain generation function;

[0045] The target thinking chain generation function is trained based on the adaptive thinking chain deletion algorithm and the function target to determine the fragments to be removed in the optimal thinking chain data.

[0046] In addition, to achieve the above purpose, the present application also proposes a large model reasoning device, the device comprising:

[0047] A thought chain generation module, used to generate a number of high-quality thought chains corresponding to the task to be inferred based on a high-quality sample data set, wherein the high-quality sample data set is a data set containing thought chains that meet preset high-quality conditions;

[0048] A thinking chain determination module, used to determine the optimal thinking chain data in the high-quality thinking chain;

[0049] A thinking chain deletion module is used to delete the optimal thinking chain data to obtain a target thinking chain;

[0050] The large model reasoning module is used to output the reasoning result corresponding to the task to be reasoned based on the target thinking chain through the large language model.

[0051] In addition, to achieve the above-mentioned objectives, the present application also proposes a large model reasoning device, which includes: a memory, a processor, and a computer program stored on the memory and executable on the processor, and the computer program is configured to implement the steps of the large model reasoning method described above.

[0052] In addition, to achieve the above-mentioned purpose, the present application also proposes a storage medium, which is a computer-readable storage medium, and a computer program is stored on the storage medium. When the computer program is executed by the processor, the steps of the large model reasoning method described above are implemented.

[0053] In addition, to achieve the above-mentioned purpose, the present application also provides a computer program product, which includes a computer program, and when the computer program is executed by a processor, it implements the steps of the large model reasoning method described above.

[0054] The present application provides a large model reasoning method, which discloses generating several high-quality thinking chains corresponding to the task to be reasoned based on a high-quality sample data set, wherein the high-quality sample data set is a data set containing thinking chains that meet preset high-quality conditions; determining the optimal thinking chain data in the high-quality thinking chain; performing thinking chain deletion on the optimal thinking chain data to obtain a target thinking chain; outputting the reasoning result corresponding to the task to be reasoned based on the target thinking chain through a large language model; compared with the prior art, when the length of the thinking chain randomly selected from the correct candidate answers generated by the bundle search model is too long, the resource cost of the large model reasoning will increase. Since the present invention can delete the optimal thinking chain data in the high-quality thinking chain corresponding to the task to be reasoned, obtain the target thinking chain, and then perform model reasoning based on the target thinking chain through the large language model, the technical problem in the prior art that the reasoning cost will increase when the large language model performs reasoning based on a thinking chain that is too long. BRIEF DESCRIPTION OF THE DRAWINGS

[0055] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present application and, together with the description, serve to explain the principles of the present application.

[0056] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, for ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative labor.

[0057] Figure 1 A flowchart of the first embodiment of the large model reasoning method of the present application is provided;

[0058] Figure 2 A flowchart diagram of the second embodiment of the large model reasoning method of this application;

[0059] Figure 3 This is a flowchart of the thinking chain screening enhancement algorithm in the large model reasoning method of this application;

[0060] Figure 4 A flowchart diagram of the third embodiment of the large model reasoning method of this application;

[0061] Figure 5 This is the overall flow chart of the large model reasoning method in this application;

[0062] Figure 6 This is a schematic diagram of the module structure of the large model reasoning device of the embodiment of the present application;

[0063] Figure 7Schematic diagram of the device structure of the hardware operating environment involved in the large model reasoning method in the embodiment of the present application.

[0064] The purpose, features and advantages of this application will be further described in conjunction with the embodiments and with reference to the accompanying drawings. DETAILED DESCRIPTION

[0065] It should be understood that the specific embodiments described herein are only used to explain the technical solutions of the present application and are not used to limit the present application.

[0066] In order to better understand the technical solution of the present application, a detailed description will be given below in conjunction with the accompanying drawings and specific implementation methods.

[0067] The main solution of the embodiment of the present application is: based on a high-quality sample data set, generate several high-quality thinking chains corresponding to the task to be reasoned, and the high-quality sample data set is a data set that contains thinking chains that meet preset high-quality conditions; determine the optimal thinking chain data in the high-quality thinking chains; perform thinking chain deletion on the optimal thinking chain data to obtain a target thinking chain; and output the reasoning result corresponding to the task to be reasoned based on the target thinking chain through a large language model.

[0068] In the prior art, when a large language model performs reasoning based on thought chains, it usually generates multiple candidate answers at one time through a bundle search model, and then randomly selects the answer with the lowest confusion and thought chain content from the correct candidate answers as training data to guide the model to perform reasoning. However, when the length of the selected thought chain is too long, the content that the model needs to generate will increase accordingly, thereby increasing the resource cost of large model reasoning.

[0069] The present application provides a solution that can delete the optimal thinking chain data in the high-quality thinking chain corresponding to the reasoning task, obtain the target thinking chain, and then perform model reasoning based on the target thinking chain through a large language model, thereby solving the technical problem in the prior art that when a large language model performs reasoning based on a thinking chain that is too long, it will lead to increased reasoning costs.

[0070] It should be noted that the execution subject of this embodiment can be a computing service device with data processing, network communication and program running functions, such as a tablet computer, a personal computer, a mobile phone, etc., or an electronic device capable of realizing the above functions, a large model reasoning device, etc. The following takes the large model reasoning device as an example (hereinafter referred to as the device) to illustrate this embodiment and the following embodiments.

[0071] Based on this, the embodiment of the present application provides a large model reasoning method, referring to Figure 1 , Figure 1 This is a flow chart of the first embodiment of the large model reasoning method of the present application.

[0072] In this embodiment, the large model reasoning method includes steps S10 to S40:

[0073] Step S10: Generate a number of high-quality thought chains corresponding to the task to be reasoned based on the high-quality sample data set, wherein the high-quality sample data set is a data set containing thought chains that meet preset high-quality conditions.

[0074] It should be noted that the above-mentioned high-quality sample data set may be a data set composed of thought chains that meet the preset high-quality conditions, wherein the thought chains that meet the preset high-quality conditions may be thought chains with high accuracy, strong logic and complete content. c Contains several data samples D i , each data sample D i = {q i , (r 1 ,…,r n ), a i}, where q i For the problems in the sample, (r 1 ,…,r n ) is a thinking chain that meets the preset high-quality conditions, a i For the answer to the question.

[0075] It should be understood that the above-mentioned tasks to be inferred may be any tasks that require a large language model for reasoning. The tasks to be inferred in this embodiment may include but are not limited to text comprehension tasks, logical reasoning tasks, mathematical calculation tasks, etc.

[0076] It should be noted that the above-mentioned high-quality thought chain can be a thought chain generated by the large language model based on the prompt information in the high-quality sample data set. In this embodiment, the device can obtain prompt information related to the task to be inferred from the high-quality sample data set, so that the large language model can generate a high-quality thought chain corresponding to the task to be inferred based on the prompt information.

[0077] Furthermore, the step S10 includes:

[0078] Step S101: obtaining an initial data set corresponding to the task to be inferred, wherein the initial data set is a data set in which there is no thought chain or there is no thought chain and task answers corresponding to the reasoning subtasks in the task to be inferred.

[0079] It is understandable that the above-mentioned reasoning subtasks may be subtasks obtained by decomposing the task to be reasoned into steps. Accordingly, the initial data set may be a data set lacking a thought chain or lacking a thought chain and an answer to a question. In this embodiment, the initial data set corresponding to the task to be reasoned is a data set lacking a thought chain or lacking a thought chain and an answer to the subtask to be reasoned.

[0080] Step S102: Determine the text prompt information corresponding to the reasoning subtask in the initial data set.

[0081] It should be understood that the above text prompt information may be prompt information for guiding the large language model to complete the derivation of the reasoning subtasks, that is, specific prompt words for each subtask in the task to be reasoned.

[0082] Step S103: generating a plurality of high-quality thought chains corresponding to the task to be reasoned based on the text prompt information and the high-quality sample data set.

[0083] Specifically, the step S103 includes: obtaining context example information from a high-quality sample data set based on the reasoning subtask; constructing task prompt information corresponding to the reasoning sample in the initial data set based on the text prompt information and the context example information; and generating a number of high-quality thought chains based on the task prompt information through a large language model.

[0084] It should be noted that the above-mentioned context example information can be the context example information with the highest relevance to the reasoning subtask in the high-quality sample data set. In practical applications, the context example information can be the example information used by the large language model when performing task derivation, including: task description information, example input and output pairs, etc., wherein the task description information can be description information of the specific type of task that the large language model currently needs to perform; the example input and output pairs can be prompt information to help the large language model understand the specific requirements of the task to be reasoned. In this embodiment, the device can determine the degree of correlation between all context examples in the high-quality sample data set and the reasoning subtask, and recall the context example information with the highest degree of correlation with the reasoning subtask in the high-quality sample data set, thereby obtaining the context example information.

[0085] It should be noted that the task prompt information may be a prompt word used to guide the large language model to complete the derivation of the reasoning task. In this embodiment, the device may be for each reasoning sample D in the initial data set D that lacks a thought chain or lacks a thought chain and an answer to the question. j Build task prompt j In this embodiment, the task prompt word Prompt j = {prompt, D c ,qj}, where prompt is the text prompt information corresponding to the reasoning subtask, D c is the context example information, q j is the reasoning subtask.

[0086] In this embodiment, after determining the task prompt information corresponding to the reasoning sample, the task prompt information can be input into the large language model, so that the large language model can generate several high-quality thought chains according to the task prompt information. In addition, in order to further improve the quality of the thought chain answers generated by the large language model, answer labels can be added to the task prompt information, and the beam search algorithm can be used to allow the large language model to perform multiple samplings to generate multiple thought chain answers. Among them, the beam search algorithm can be a width-limited optimal search method. The beam search algorithm refers to selecting several candidate texts with the highest probability (usually a fixed number N, i.e., beam width) in the candidate sequence of each step in the decoding process of the large model for expansion, thereby generating high-quality output, and the sequences that are not selected will be discarded.

[0087] In practical applications, if the reasoning samples in the initial data set contain answer labels or answers, the wrong answers can be directly compared and screened out; if there are no answer labels or answers, the most likely answer can be selected through a majority voting mechanism, and then the above steps of generating the thought chain answer are repeated. Among them, the majority voting mechanism can be a method of eliminating label noise by combining the prediction results of multiple models. If a certain label receives more than half of the votes, it is predicted as the label, otherwise the prediction is rejected. In this embodiment, the bundle search algorithm always selects the N tokens (i.e., text fragments, including words, punctuation marks, numbers, etc.) with the highest current probability as candidates at each step. Compared with the greedy algorithm, it alleviates the problem of local optimal solutions to a certain extent, but it still faces the possibility that the candidate is not the optimal answer. Therefore, in this embodiment, the bundle search algorithm can retain tokens based on the target difference decoding strategy, wherein the target difference decoding strategy can be:

[0088]

[0089] In the formula, is the candidate sequence, and is the first N tokens with the highest probability at step t, and a=b+1, t is the length of the candidate sequence currently generated, q j is the inference subtask in the task to be inferred. In this embodiment, the goal of the bundle search is to find a candidate sequence that maximizes the conditional probability p, thereby avoiding the local optimal dilemma of the greedy algorithm.

[0090] Step S20: Determine the optimal thinking chain data in the high-quality thinking chain.

[0091] It should be noted that the above-mentioned optimal thinking chain data can be more robust thinking chain data obtained by filtering the data in the high-quality thinking chain.

[0092] In this embodiment, the device can first perform a preliminary filtering of the data in the high-quality thought chain according to the length of the large language model, and then perform a secondary filtering of the steps that are too short or too long in the thought chain after the preliminary filtering, and finally filter the thought chain through the model's sensitivity to contextual learning examples and the differences between the model-generated thought chain and the human-written thought chain, and finally obtain more robust optimal thought chain data.

[0093] Step S30: Perform thought chain deletion on the optimal thought chain data to obtain a target thought chain.

[0094] It should be noted that the above-mentioned target thinking chain can be a thinking chain obtained by removing the content with lower influence in the optimal thinking chain data.

[0095] It should be noted that when the large language model performs task reasoning based on thought chains, the thought chain (i.e., the model’s thinking process) content is usually added to the generation part of the model, while the model’s training task, i.e., predicting the next token, remains unchanged. At this time, the training goal of the large model can be:

[0096]

[0097] In the formula, π θ is the parameter that needs to be optimized in the large language model, y is the final prediction result, and r 1 ,…,r n are the calculation steps and results in the middle of the thinking chain, and x is the problem input.

[0098] It can be seen that if the length of the thought chain is longer, the content that the model needs to generate will also increase accordingly, which will inevitably lead to more resource consumption. Among them, reducing the length of the thought chain steps in the generated results is the most direct way to reduce resource costs. In practical applications, the large language model itself has a certain thinking ability. For some questions, it can directly give the correct answer without outputting the thinking steps. Therefore, retaining only a part of the thought chain content or no thought chain content at all in the training data of the large model has no effect on the correctness of the output answer of the large model. Therefore, this embodiment can remove text fragments with lower influence in the optimal thought chain data, and finally obtain the target thought chain, so as to shorten the length of the generated thought chain and reduce the cost of generating the thought chain by the large model.

[0099] It should be noted that although the optimal thinking chain data has been screened for length and quality, it still requires a high training cost to train a large language model using large-scale data containing thinking chains. At the same time, if the large language model needs to output a complete thinking chain in each reasoning process, it will greatly reduce the reasoning efficiency of the model. In practical applications, since the large model itself has a certain reasoning ability, and the thinking chain is mainly used to split a complex task into multiple simple and continuous subtasks and gradually complete them to obtain the final answer, the large model itself can only use a certain number of reasoning steps as prompt information and output the correct answer. Therefore, in order to reduce the reasoning cost of the large language model, this embodiment can use an adaptive thinking chain deletion algorithm to delete the optimal thinking chain data, wherein the algorithm can remove a certain number of tokens from the back of the thinking chain in order to ensure the coherence and fluency of the thinking chain content, and finally obtain the target thinking chain. However, removing the content in the thinking chain will inevitably lead to a decrease in the accuracy of the large model reasoning, so this embodiment can maximize the removal of the content in the thinking chain while stabilizing the reasoning ability of the large model as much as possible, thereby achieving the purpose of reducing the cost of model training and reasoning.

[0100] Step S40: Outputting the reasoning result corresponding to the task to be reasoned based on the target thinking chain through the large language model.

[0101] It should be understood that the large language model is a natural language processing (NLP) model built using deep learning technology, which aims to simulate the processing and generation capabilities of human language as much as possible. The large language model in this embodiment may include but is not limited to GPT-3, GPT-4, etc.

[0102] In practical applications, after maximizing the removal of the content in the optimal thinking chain data while ensuring the accuracy of the large model's reasoning results, the target thinking chain can be obtained. Then, the large language model can decompose the task to be reasoned into multiple sub-tasks of consecutive steps based on the target thinking chain, and solve these sub-tasks separately, thereby realizing the reasoning of the reasoning task and obtaining the final reasoning result.

[0103] The present embodiment provides a large model reasoning method, which discloses generating several high-quality thinking chains corresponding to the task to be reasoned based on a high-quality sample data set, where the high-quality sample data set is a data set containing thinking chains that meet preset high-quality conditions; determining the optimal thinking chain data in the high-quality thinking chain; performing thinking chain deletion on the optimal thinking chain data to obtain a target thinking chain; outputting the reasoning result corresponding to the task to be reasoned based on the target thinking chain through a large language model; compared with the prior art, when the length of the thinking chain randomly selected from the correct candidate answers generated by the bundle search model is too long, it will lead to an increase in the resource cost of large model reasoning. Since the present embodiment can delete the optimal thinking chain data in the high-quality thinking chain corresponding to the task to be reasoned to obtain the target thinking chain, and then perform model reasoning based on the target thinking chain through the large language model, it solves the technical problem in the prior art that when the large language model performs reasoning based on a thinking chain that is too long, it will lead to an increase in the reasoning cost.

[0104] Based on the first embodiment of the present application, in the second embodiment of the present application, the same or similar contents as those in the above-mentioned embodiment 1 can be referred to the above introduction, and will not be repeated in the following. Figure 2 , Figure 2 A flow chart of the second embodiment of the large model reasoning method of this application.

[0105] In this embodiment, step S20 includes steps S201 to S203:

[0106] Step S201: filtering the high-quality thought chains according to the maximum length value of the large language model to obtain a first thought chain.

[0107] It should be understood that the above maximum length value may be the maximum text length value that the large language model can accept when processing input information. In this embodiment, the maximum length value may refer to the maximum number of characters or tokens that the model can receive at one time. For example, the maximum input limit of GPT-3 is 4096 tokens, which means that the model can process up to 4096 tokens in one processing task.

[0108] It can be understood that the above-mentioned first thought chain can be a thought chain obtained by filtering the data exceeding the maximum length value of the large model in the high-quality thought chain.

[0109] Specifically, the step S201 includes: determining the thinking chain length value of the high-quality thinking chain; determining the data to be filtered in the high-quality thinking chain based on the maximum length value of the large language model and the thinking chain length value; filtering the data to be filtered to obtain the first thinking chain.

[0110] It should be understood that the above-mentioned thought chain length value is the length value of the high-quality thought chain. The above-mentioned data to be filtered can be data in the high-quality thought chain that exceeds the maximum length value of the large language model. In this embodiment, after obtaining several high-quality thought chains, the device can determine the length value corresponding to each high-quality thought chain, and compare the length value corresponding to each high-quality thought chain with the maximum text length value that the large language model can accept, and then filter out the part of the data in the high-quality thought chain that exceeds the maximum allowable length of the large language model according to the comparison result to obtain the first thought chain.

[0111] Step S202: Filter the first chain of thought based on the number of steps in the first chain of thought to obtain a second chain of thought.

[0112] It is understandable that the above number of steps can be the number of steps in the first chain of thought. In practical applications, the number of steps in a chain of thought is related to the complexity of the task. Therefore, in order to reduce the complexity of the large model in performing task reasoning, this embodiment can filter out steps with too much content in the first chain of thought. In addition, in order to ensure the accuracy of the large model in performing task reasoning, steps with too little content in the first chain of thought can also be filtered out, and finally the second chain of thought is obtained.

[0113] Specifically, the step S202 includes: determining a target step number range of the first thinking chain; comparing the number of steps of the first thinking chain with the target step number range; filtering the first thinking chain according to the comparison result to obtain a second thinking chain.

[0114] It should be understood that the above-mentioned target step number range can be the range in which the number of steps of the first thinking chain is expected to be. In practical applications, when reasoning with a large language model, if the thinking chain has too many steps, the complexity of the model reasoning will be increased; if the thinking chain has too few steps, the accuracy of the large model reasoning will be low. Therefore, in order to reduce the complexity of the large model reasoning while ensuring the accuracy of the model reasoning, the device can pre-set an expected range of the number of steps of a thinking chain, that is, the above-mentioned target step number range, and then compare the number of steps of the first thinking chain with the target step number range. If the number of steps of the first thinking chain exceeds the target step number range, the steps with too much or too little content in the first thinking chain can be filtered until the number of steps of the first thinking chain is within the target step number range, and finally the second thinking chain is obtained.

[0115] Step S203: Acquire optimal thinking chain data based on the example sequence corresponding to the context example information and the second thinking chain.

[0116] It can be understood that the above example order is the order of examples in context learning.

[0117] In actual applications, in context learning, changes in the order of examples have a greater impact on the prediction results of the large model, while in the thinking chain, the order of steps in the thinking chain has a smaller impact on the manually written thinking chain and a greater impact on the thinking chain generated by the model. Therefore, in order to make the second thinking chain data as inclined as possible to the thinking chain content generated by humans, so as to achieve the robustness of the model thinking process as much as possible, this embodiment can continuously adjust the order of examples in context learning, let the large model repeatedly generate multiple rounds of thinking chain results, and remove the thinking chain farthest from the current round in the thinking chain pool in each round to obtain the optimal thinking chain data, wherein the thinking chain pool can be a storage space for storing multiple thinking chains. In this embodiment, the thinking chain pool can be composed of each reasoning sample D in the initial data set D. j Build to get.

[0118] Furthermore, the step S203 includes:

[0119] Step S203a: adjusting the example order corresponding to the context example information in real time.

[0120] Step S203b: Generate several rounds of thought chain data based on the adjusted context example and the second thought chain through the large language model.

[0121] It should be understood that the above-mentioned adjusted context examples may be context examples obtained after adjusting the order of the examples. In this embodiment, the device may continuously adjust the order of the context examples so that the large language model can repeatedly generate multiple rounds of thinking chain data based on the context examples after adjusting the order of the examples and the reasoning steps in the second thinking chain.

[0122] Step S203c: Remove the farthest thinking chain data corresponding to each round of thinking chain data to obtain the optimal thinking chain data.

[0123] It should be noted that the above-mentioned optimal thinking chain data can be the thinking chain data in the thinking chain pool that is farthest from the thinking chain data of each round. In this embodiment, the distance between them can be determined by calculating the cosine similarity between each thinking chain data in the thinking chain pool and the thinking chain data generated by the large language model in each round, and the thinking chain data in the thinking chain pool that is farthest from the thinking chain data generated in this round is determined as the farthest thinking chain data of this round, and it is removed. After removing the farthest thinking chain data of each round, the optimal thinking chain data is finally obtained. In addition to determining the distance between thinking chains by the cosine similarity between thinking chains, this embodiment can also replace the calculation of cosine similarity by Euclidean distance, Manhattan distance, Minkowski distance, Pearson correlation coefficient, etc.

[0124] Specifically, before step S203c, it also includes: constructing a candidate thought chain pool based on the reasoning sample.

[0125] It can be understood that the candidate thought chain pool can be a storage space for storing thought chains. j Construct a candidate thought chain pool, wherein the candidate thought chain pool may contain multiple different thought chains, which may be used multiple times and may be updated and modified as needed.

[0126] Accordingly, the step S203c includes: determining the furthest thought chain data which is farthest from each round of thought chain data from the candidate thought chain pool; removing the furthest thought chain data to obtain the optimal thought chain data.

[0127] In a specific implementation, this embodiment can filter thought chains through a thought chain screening enhancement algorithm to obtain a more robust and high-quality thought chain. Figure 3 , Figure 3 This is a flowchart of the thought chain screening enhancement algorithm in the large model reasoning method of this application. Figure 3 As shown, the generation result of the self-robust thinking chain screening and enhancement algorithm provided in this embodiment is the thinking chain R j = {r 1 ,…,r n}, the data used by the algorithm include: context candidate example set {Dc j}, current prediction sample D j (i.e. reasoning samples), thinking chain pool j , large model π θ , current question j, and answer x j ,y j When using the thinking chain screening enhancement algorithm to filter the thinking chain, you can first initialize the data (i.e., initialization), and then judge | Pool j |Is it greater than 1? If so, adjust the context candidate example set Dc j Example d k and d l The adjusted context example D′c is obtained in the order j , so that the large model can be based on the adjusted context example D′c j 、The answer to the current question x j and j And the second thought chain generates several rounds of thought chains P′, and then the thought chain pool Pool can be calculated j Zhongsi Thinking Chain P zThe distance ω between the thought chain P generated by the big model in each round, and these distances ω are sorted to select the thought chain pool Pool according to the sorting results j Determine the furthest thought chain data P t , and the furthest thought chain data P t From the thinking chain pool Pool j After removing the farthest thinking chain data from the thinking chain data generated by the large model in each round, the optimal thinking chain data is finally obtained.

[0128] In this embodiment, it is disclosed that high-quality thought chains are filtered according to the maximum length value of the large language model to obtain a first thought chain; the first thought chain is filtered based on the number of steps in the first thought chain to obtain a second thought chain; and the optimal thought chain data is obtained based on the example order corresponding to the context example information and the second thought chain. Since this embodiment can perform preliminary filtering of the length and quality of the thought chain according to the maximum length of the large language model and the number of steps of the first thought chain, and perform robustness filtering of the thought chain through the sensitivity of the large language model to the context examples, it is possible to obtain optimal thought chain data with higher robustness, which is beneficial to improving the stability and reliability of the large language model in task reasoning.

[0129] Based on the first embodiment and / or the second embodiment of the present application, in the third embodiment of the present application, the same or similar contents as those in the above embodiments can be referred to the above introduction, and will not be described in detail later. Figure 4 , Figure 4 A flow chart of the third embodiment of the large model reasoning method of this application.

[0130] In this embodiment, step S30 includes steps S301 to S302:

[0131] Step S301: Determine the segments to be removed in the optimal thought chain data through an adaptive thought chain deletion algorithm.

[0132] It should be noted that the above adaptive thought chain deletion algorithm can be an algorithm for dynamically adjusting the length of the thought chain according to different usage scenarios of the large model. In this embodiment, the adaptive thought chain deletion algorithm can adjust the length of the thought chain by removing less influential content in the thought chain.

[0133] It can be understood that the above-mentioned fragments to be removed can be content that can be removed from the optimal thinking chain data, wherein the fragments to be removed in this embodiment can be content in the optimal thinking chain data that has no effect or little effect on the accuracy of the answer output by the large language model.

[0134] Further, the step S301 includes: constructing a target thinking chain generation function based on the optimal thinking chain data and removable fragments in the optimal thinking chain data; determining the function target of the target thinking chain generation function; training the target thinking chain generation function based on the adaptive thinking chain deletion algorithm and the function target to determine the fragments to be removed in the optimal thinking chain data.

[0135] It should be noted that the target thought chain generation function can be a function for maximizing the removable content in the thought chain while ensuring the accuracy of the large model reasoning result. Correspondingly, the function objective can be to maximize the length of the removable generated content in the thought chain.

[0136] In this embodiment, the device can train the target thinking chain generation function based on the adaptive thinking chain deletion algorithm to maximize the length of the removable generated content in the thinking chain, and finally determine the fragments to be removed in the optimal thinking chain data that have no effect or little effect on the accuracy of the answer output by the large model, and remove the fragments to be removed from the optimal thinking chain data, and finally obtain the target thinking chain to shorten the length of the generated thinking chain, thereby reducing the cost of subsequent large models for task reasoning. Among them, the target thinking chain generation function in this embodiment can be:

[0137]

[0138] In the formula, π θ is the parameter that needs to be optimized in the large language model, y is the final prediction result, {r 1 ,…,r n} is the thinking chain, x is the question input, R u is the removable part of the thought chain, len(R u ) is the length of the removable part, and the generation goal of the function is to maximize the length of the removable generated content.

[0139] It should be noted that the adaptive thought chain pruning algorithm in this embodiment can dynamically remove a certain number of tokens in the thought chain from the back to the front during reasoning. At this time, the state of the current large language model can be evaluated by the loss of the verification sample. Specifically, for a large language model in the current epoch (epoch can be the samples in all training data sets that the large language model has traversed and processed once) Its loss l on the validation sample v′ can be estimated as:

[0140]

[0141] In the formula, v′ is the verification sample, θ is all the parameters of the large language model, R is the thought chain in the training set, and R uIt is a removable part of a thought chain, and its longest length is the length of the thought chain itself.

[0142] In practical applications, when a large model performs task reasoning, the length of the thought chain removed from each reasoning sample may be different. Therefore, this embodiment can dynamically adjust the part of the thought chain that should be saved by saving the removed part of the thought chain in each sample, thereby optimizing the training and reasoning costs. In this embodiment, the part to be removed in the thought chain can be introduced as a trainable parameter τ into the optimization process of SGD (Stochastic Gradient Descent), and its loss value is the current training validation set loss l. At this time, the thought chain of the current training sample i The impact of the change on this epoche inf can be:

[0143]

[0144] Where N is the size of the training data set, j represents the data in the training data set, ∈ represents the allowable error range, or the interval unit used to adjust the model evaluation (such as the frequency of changes in the length of the thought chain can be adjusted based on the training step size, epoch, etc.), and η represents the learning rate.

[0145] In practical applications, since the large model itself has a certain reasoning ability, if the large model is trained using an explicit thinking chain training method, the model will also output the model thinking process for those questions that can be directly answered each time it is reasoning, which will greatly reduce the efficiency of model reasoning. The adaptive thinking chain deletion algorithm proposed in this embodiment can continuously adjust the length of the thinking chain to be output by dynamically learning changes. Among them, when no token is removed from the thinking chain, the training process of the model is explicit thinking chain training; when all thinking chains are removed, the training process of the model is implicit thinking chain training, that is, the model can directly give the answer, thereby avoiding the large model from overthinking during task reasoning, resulting in increased reasoning costs, and at the same time improving the model reasoning efficiency. It can be seen that this solution can adjust the length of the thinking chain in actual training according to different usage scenarios to meet different business needs.

[0146] It should be noted that the data label prediction tasks during the adaptive thinking chain training in this solution may include: classification label prediction tasks and text generation class label prediction tasks, wherein the evaluation index of the classification label prediction task can be calculated by the accuracy of the prediction result, and the calculation method can be:

[0147]

[0148] In the formula, Accuracy represents the accuracy of the prediction result, #Correct represents the number of samples with accurate prediction, and #Total represents the total number of samples.

[0149] The evaluation index of the text generation class label prediction task can be calculated by Rouge-N, which is mainly calculated based on N-gram (i.e., a continuous word sequence of length N). The calculation formula is as follows:

[0150]

[0151] In the formula, Rouge-N is a feature used to measure the similarity between the generated text and the real text, Count match (gram n ) represents the recall rate, which is used to indicate the ratio of correctly matched N-grams in the system output to all N-grams in the reference text. n ) is the number in the reference text, S is the system text, and ground-truth is the reference text.

[0152] Step S302: Perform thought chain deletion on the optimal thought chain data based on the fragment to be removed to obtain a target thought chain.

[0153] It can be understood that after determining the to-be-removed segments in the optimal thinking chain data, the to-be-removed segments in the optimal thinking chain data can be removed to finally obtain the target thinking chain.

[0154] In the specific implementation, refer to Figure 5 , Figure 5 This is the overall flow chart of the large model reasoning method in this application. Figure 5As shown, first, a high-quality sample data set with a complete thinking chain can be obtained, and the high-quality sample data set contains questions, thinking chain content and question answers. For example, Example 1: Question: There are nine computers in the server room; Answer: 29; Reasonableness: There are 4 days from Monday to Thursday, and 5 computers are added every day, which means that a total of 4*5=20 computers are added; Example N: Question: A has 23 yuan, she bought 5 cakes, each for 3 yuan, how much money does she have left? Answer: 8; Reasonableness: She bought 5 cakes, each for 3 yuan, which means that she spent 5*3 yuan on cakes. She had 23 yuan at the beginning, so she now has 23-15=8 yuan. Then the large language model can generate several high-quality thinking chains based on the text prompt information corresponding to each subtask in the task to be inferred and the high-quality sample data set. In this embodiment, the large language model can perform multiple samplings through the bundle search algorithm to generate multiple thinking chain answers, wherein the bundle search algorithm in this embodiment can retain tokens based on the difference decoding strategy, so that the problem of local optimal solution can be avoided. Through the bundle search process of the large model, each candidate sample can build a candidate thought chain pool. Then, by continuously adjusting the order of examples in context learning, the large language model can repeatedly generate multiple rounds of thought chain results, and remove the thought chain farthest from the thought chain generated in this round in the candidate thought chain pool in each round to achieve robust filtering of the thought chain and obtain a filtered thought chain pool. At this time, the thought chains in the filtered thought chain pool have higher robustness. Finally, the thought chains in the filtered thought chain pool can be used as training samples to continue adaptive thought chain optimization training to remove the content in these thought chains that has no effect or little effect on the accuracy of the answer output by the large model, and obtain the target thought chain. Finally, the large language model can reason on the reasoning task based on the target thought chain and obtain the corresponding reasoning results.

[0155] In this embodiment, it is disclosed that the fragments to be removed in the optimal thinking chain data are determined by an adaptive thinking chain deletion algorithm; the optimal thinking chain data is subjected to thinking chain deletion based on the fragments to be removed to obtain a target thinking chain. Since in this embodiment, the content in the optimal thinking chain data that has no effect or little effect on the accuracy of the answer output by the large model can be determined as the fragments to be removed, and it is removed from the optimal thinking chain to obtain the target thinking chain, the large language model can perform reasoning on the reasoning task based on the target thinking chain, thereby reducing the reasoning cost of the large language model and improving the reasoning efficiency.

[0156] It should be noted that the above examples are only used to understand the present application and do not constitute a limitation on the large model reasoning method of the present application. More simple transformations based on this technical concept are all within the scope of protection of the present application.

[0157] This application also provides a large model reasoning device, please refer to Figure 6 , the large model reasoning device comprises:

[0158] A thought chain generation module 10 is used to generate a plurality of high-quality thought chains corresponding to the task to be inferred based on a high-quality sample data set, wherein the high-quality sample data set is a data set containing thought chains that meet preset high-quality conditions;

[0159] A thinking chain determination module 20, used to determine the optimal thinking chain data in the high-quality thinking chain;

[0160] A thinking chain deletion module 30 is used to delete the optimal thinking chain data to obtain a target thinking chain;

[0161] The large model reasoning module 40 is used to output the reasoning result corresponding to the task to be reasoned based on the target thinking chain through the large language model.

[0162] The large model reasoning device provided by the present application adopts the large model reasoning method in the above embodiment, which can solve the technical problem in the prior art that when the large language model is reasoned based on a thought chain that is too long, the reasoning cost will increase. Compared with the prior art, the beneficial effects of the large model reasoning device provided by the present application are the same as the beneficial effects of the large model reasoning method provided by the above embodiment, and the other technical features in the large model reasoning device are the same as the features disclosed in the above embodiment method, which will not be repeated here.

[0163] The present application provides a large model reasoning device, which includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the large model reasoning method in the above-mentioned embodiment one.

[0164] Reference below Figure 7 , which shows a schematic diagram of the structure of a large model reasoning device suitable for implementing the embodiment of the present application. The large model reasoning device in the embodiment of the present application may include but is not limited to mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (Portable Application Descriptions), PMPs (Portable Media Players), vehicle-mounted terminals (such as vehicle-mounted navigation terminals), etc., and fixed terminals such as digital TVs, desktop computers, etc. Figure 7 The large model reasoning device shown is merely an example and should not bring any limitation to the functions and scope of use of the embodiments of the present application.

[0165] like Figure 7 As shown, the large model reasoning device may include a processing device 1001 (such as a central processing unit, a graphics processor, etc.), which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM: Read Only Memory) 1002 or a program loaded from a storage device 1003 to a random access memory (RAM: Random Access Memory) 1004. In RAM1004, various programs and data required for the operation of the large model reasoning device are also stored. The processing device 1001, ROM1002 and RAM1004 are connected to each other through a bus 1005. An input / output (I / O) interface 1006 is also connected to the bus. Generally, the following systems can be connected to the I / O interface 1006: an input device 1007 including, for example, a touch screen, a touchpad, a keyboard, a mouse, an image sensor, a microphone, an accelerometer, a gyroscope, etc.; an output device 1008 including, for example, a liquid crystal display (LCD: Liquid Crystal Display), a speaker, a vibrator, etc.; a storage device 1003 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 1009. The communication device 1009 can allow the large model reasoning device to communicate with other devices wirelessly or by wire to exchange data. Although the figure shows a large model reasoning device with various systems, it should be understood that it is not required to implement or have all the systems shown. More or fewer systems can be implemented or have alternatively.

[0166] In particular, according to the embodiments disclosed in the present application, the process described above with reference to the flowchart can be implemented as a computer software program. For example, the embodiments disclosed in the present application include a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program includes a program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from a network through a communication device, or installed from a storage device 1003, or installed from a ROM 1002. When the computer program is executed by the processing device 1001, the above-mentioned functions defined in the method of the embodiment disclosed in the present application are executed.

[0167] The large model reasoning device provided by this application adopts the large model reasoning method in the above embodiment to solve the technical problems of large model reasoning. Compared with the prior art, the beneficial effects of the large model reasoning device provided by this application are the same as the beneficial effects of the large model reasoning method provided by the above embodiment, and the other technical features in the large model reasoning device are the same as the features disclosed in the method of the previous embodiment, which will not be repeated here.

[0168] It should be understood that the various parts disclosed in this application can be implemented by hardware, software, firmware or a combination thereof. In the description of the above embodiments, specific features, structures, materials or characteristics can be combined in any one or more embodiments or examples in a suitable manner.

[0169] The above is only a specific implementation of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art who is familiar with the present technical field can easily think of changes or substitutions within the technical scope disclosed in the present application, which should be included in the protection scope of the present application. Therefore, the protection scope of the present application should be based on the protection scope of the claims.

[0170] The present application provides a computer-readable storage medium having computer-readable program instructions (ie, computer programs) stored thereon, wherein the computer-readable program instructions are used to execute the large model reasoning method in the above-mentioned embodiment.

[0171] The computer-readable storage medium provided in the present application may be, for example, a USB flash drive, but is not limited to electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, systems or devices, or any combination of the above. More specific examples of computer-readable storage media may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In this embodiment, the computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in combination with an instruction execution system, system or device. The program code contained on the computer-readable storage medium may be transmitted using any appropriate medium, including but not limited to: wires, optical cables, RF (Radio Frequency), etc., or any suitable combination of the above.

[0172] The above-mentioned computer-readable storage medium may be included in the large model reasoning device; or it may exist independently without being assembled into the large model reasoning device.

[0173] The above-mentioned computer-readable storage medium carries one or more programs. When the above-mentioned one or more programs are executed by the large model reasoning device, the large model reasoning device: generates several high-quality thinking chains corresponding to the task to be reasoned based on the high-quality sample data set, and the high-quality sample data set is a data set that contains thinking chains that meet the preset high-quality conditions; determines the optimal thinking chain data in the high-quality thinking chain; performs thinking chain deletion on the optimal thinking chain data to obtain the target thinking chain; and outputs the reasoning result corresponding to the task to be reasoned based on the target thinking chain through the large language model.

[0174] Computer program code for performing the operations of the present application may be written in one or more programming languages ​​or a combination thereof, including object-oriented programming languages ​​such as Java, Smalltalk, C++, and conventional procedural programming languages ​​such as "C" or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, as a separate software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0175] The flow chart and block diagram in the accompanying drawings illustrate the possible architecture, function and operation of the system, method and computer program product according to various embodiments of the present application. In this regard, each square box in the flow chart or block diagram can represent a module, a program segment or a part of a code, and the module, the program segment or a part of the code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the square box can also occur in a sequence different from that marked in the accompanying drawings. For example, two square boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each square box in the block diagram and / or flow chart, and the combination of the square boxes in the block diagram and / or flow chart can be implemented with a dedicated hardware-based system that performs a specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.

[0176] The modules involved in the embodiments described in this application may be implemented by software or hardware, wherein the name of the module does not constitute a limitation on the unit itself in some cases.

[0177] The readable storage medium provided in this application is a computer-readable storage medium, which stores computer-readable program instructions (i.e., computer programs) for executing the above-mentioned large model reasoning method, and can solve the technical problem in the prior art that when a large language model is reasoned based on a thought chain that is too long, the reasoning cost will increase. Compared with the prior art, the beneficial effects of the computer-readable storage medium provided in this application are the same as the beneficial effects of the large model reasoning method provided in the above-mentioned embodiment, and will not be repeated here.

[0178] The present application also provides a computer program product, including a computer program, which implements the steps of the large model reasoning method as described above when executed by a processor.

[0179] The computer program product provided by the present application can solve the technical problem in the prior art that when a large language model performs reasoning based on a thought chain that is too long, the reasoning cost increases. Compared with the prior art, the beneficial effects of the computer program product provided by the present application are the same as the beneficial effects of the large model reasoning method provided by the above embodiment, and will not be repeated here.

[0180] The above descriptions are only some embodiments of the present application, and are not intended to limit the patent scope of the present application. All equivalent structural changes made using the contents of the present application specification and drawings under the technical concept of the present application, or direct / indirect applications in other related technical fields are included in the patent protection scope of the present application.

[0181] The present invention discloses A1, a large model reasoning method, the method comprising:

[0182] Generate a number of high-quality thought chains corresponding to the task to be reasoned based on a high-quality sample data set, wherein the high-quality sample data set is a data set containing thought chains that meet preset high-quality conditions;

[0183] Determining optimal thinking chain data among the high-quality thinking chains;

[0184] Performing thought chain deletion on the optimal thought chain data to obtain a target thought chain;

[0185] The reasoning result corresponding to the task to be reasoned is output based on the target thinking chain through the large language model.

[0186] A2. As described in A1, the step of generating a plurality of high-quality thought chains corresponding to the task to be reasoned based on the high-quality sample data set comprises:

[0187] Acquire an initial data set corresponding to the task to be inferred, wherein the initial data set is a data set of task answers corresponding to the reasoning subtasks in the task to be inferred without a thought chain or without a thought chain;

[0188] Determining text prompt information corresponding to the reasoning subtask in the initial data set;

[0189] Based on the text prompt information and the high-quality sample data set, a plurality of high-quality thought chains corresponding to the task to be reasoned are generated.

[0190] A3. As described in A2, the step of generating a plurality of high-quality thought chains corresponding to the task to be reasoned based on the text prompt information and the high-quality sample data set comprises:

[0191] Acquire contextual example information from a high-quality sample dataset based on the reasoning subtask;

[0192] Constructing task prompt information corresponding to the inference sample in the initial data set based on the text prompt information and the context example information;

[0193] A large language model is used to generate a number of high-quality thought chains based on the task prompt information.

[0194] A4. As described in A3, the step of determining the optimal thinking chain data in the high-quality thinking chain comprises:

[0195] Filter the high-quality thought chains according to the maximum length value of the large language model to obtain a first thought chain;

[0196] Filter the first chain of thinking based on the number of steps in the first chain of thinking to obtain a second chain of thinking;

[0197] The optimal chain of thinking data is acquired based on the example sequence corresponding to the context example information and the second chain of thinking.

[0198] A5. As described in A4, the step of filtering the high-quality thought chains according to the maximum length value of the large language model to obtain the first thought chain comprises:

[0199] Determining the thinking chain length value of the high-quality thinking chain;

[0200] Determine the data to be filtered in the high-quality thought chain based on the maximum length value of the large language model and the thought chain length value;

[0201] The data to be filtered is filtered to obtain a first thinking chain.

[0202] A6. As described in A4, the step of filtering the first chain of thought based on the number of steps in the first chain of thought to obtain the second chain of thought comprises:

[0203] Determine the target number of steps for the first thought chain;

[0204] comparing the number of steps in the first thought chain with the target range of steps;

[0205] The first chain of thinking is filtered according to the comparison result to obtain a second chain of thinking.

[0206] A7. As described in the method of A4, the step of obtaining the optimal thinking chain data based on the example sequence corresponding to the context example information and the second thinking chain comprises:

[0207] Adjusting the example order corresponding to the context example information in real time;

[0208] Generate a plurality of rounds of thought chain data based on the adjusted context example and the second thought chain by the large language model;

[0209] Remove the farthest thinking chain data corresponding to each round of thinking chain data to obtain the optimal thinking chain data.

[0210] A8. The method as described in A7, before the step of removing the furthest thought chain data corresponding to each round of thought chain data to obtain the optimal thought chain data, further includes:

[0211] Building a candidate thought chain pool based on the reasoning samples;

[0212] The step of removing the furthest thought chain data corresponding to each round of thought chain data to obtain the optimal thought chain data includes:

[0213] Determine the furthest thought chain data which is farthest from each round of thought chain data from the candidate thought chain pool;

[0214] The furthest thought chain data is removed to obtain the optimal thought chain data.

[0215] A9. In the method as described in any one of A1 to A8, the step of performing thought chain deletion on the optimal thought chain data to obtain a target thought chain comprises:

[0216] Determine the fragments to be removed in the optimal thought chain data by using an adaptive thought chain deletion algorithm;

[0217] The optimal thinking chain data is subjected to thinking chain deletion based on the fragment to be removed to obtain a target thinking chain.

[0218] A10. As described in A9, the step of determining the fragments to be removed in the optimal thought chain data by using an adaptive thought chain deletion algorithm comprises:

[0219] Constructing a target thinking chain generation function based on the optimal thinking chain data and the removable fragments in the optimal thinking chain data;

[0220] Determining the function target of the target thinking chain generation function;

[0221] The target thinking chain generation function is trained based on the adaptive thinking chain deletion algorithm and the function target to determine the fragments to be removed in the optimal thinking chain data.

[0222] The present invention also discloses B11, a large model reasoning device, the device comprising:

[0223] A thought chain generation module, used to generate a number of high-quality thought chains corresponding to the task to be inferred based on a high-quality sample data set, wherein the high-quality sample data set is a data set containing thought chains that meet preset high-quality conditions;

[0224] A thinking chain determination module, used to determine the optimal thinking chain data in the high-quality thinking chain;

[0225] A thinking chain deletion module is used to delete the optimal thinking chain data to obtain a target thinking chain;

[0226] The large model reasoning module is used to output the reasoning result corresponding to the task to be reasoned based on the target thinking chain through the large language model.

[0227] B12. In the device as described in B11, the thought chain generation module is also used to obtain an initial data set corresponding to the task to be reasoned, wherein the initial data set is a data set of task answers corresponding to the reasoning subtasks in the task to be reasoned where there is no thought chain or no thought chain exists; determine the text prompt information corresponding to the reasoning subtask in the initial data set; and generate several high-quality thought chains corresponding to the task to be reasoned based on the text prompt information and the high-quality sample data set.

[0228] B13. In the device as described in B12, the thought chain generation module is also used to obtain context example information from a high-quality sample data set based on the reasoning subtask; construct task prompt information corresponding to the reasoning sample in the initial data set based on the text prompt information and the context example information; and generate a number of high-quality thought chains based on the task prompt information through a large language model.

[0229] B14. In the device as described in B13, the thought chain determination module is also used to filter the high-quality thought chain according to the maximum length value of the large language model to obtain a first thought chain; filter the first thought chain based on the number of steps in the first thought chain to obtain a second thought chain; and obtain the optimal thought chain data based on the example order corresponding to the context example information and the second thought chain.

[0230] B15. In the device as described in B14, the thought chain determination module is also used to determine the thought chain length value of the high-quality thought chain; determine the data to be filtered in the high-quality thought chain based on the maximum length value of the large language model and the thought chain length value; filter the data to be filtered to obtain the first thought chain.

[0231] B16. In the device as described in any one of B11 to 15, the thought chain deletion module is also used to determine the fragments to be removed in the optimal thought chain data through an adaptive thought chain deletion algorithm; and perform thought chain deletion on the optimal thought chain data based on the fragments to be removed to obtain a target thought chain.

[0232] B17. In the device as described in B16, the thinking chain deletion module is also used to construct a target thinking chain generation function based on the optimal thinking chain data and the removable fragments in the optimal thinking chain data; determine the function target of the target thinking chain generation function; and train the target thinking chain generation function based on the adaptive thinking chain deletion algorithm and the function target to determine the fragments to be removed in the optimal thinking chain data.

[0233] The present invention also discloses C18, a large model reasoning device, which includes: a memory, a processor, and a large model reasoning program stored in the memory and executable on the processor, wherein the large model reasoning program is configured to implement the steps of the large model reasoning method described above.

[0234] The present invention also discloses D19, a storage medium, on which a large model reasoning program is stored. When the large model reasoning program is executed by a processor, the steps of the large model reasoning method described above are implemented.

[0235] The present invention also discloses E20, a computer program product, which includes a computer program, and when the computer program is executed by a processor, the steps of the large model reasoning method described above are implemented.

Claims

1. A large model reasoning method, characterized in that: The method includes: Generate a number of high-quality thought chains corresponding to the task to be reasoned based on a high-quality sample data set, wherein the high-quality sample data set is a data set containing thought chains that meet preset high-quality conditions; Determining the optimal thinking chain data among the high-quality thinking chains; Performing thought chain deletion on the optimal thought chain data to obtain a target thought chain; The reasoning result corresponding to the task to be reasoned is output based on the target thinking chain through the large language model.

2. The method according to claim 1, characterized in that The step of generating a plurality of high-quality thought chains corresponding to the task to be reasoned based on the high-quality sample data set includes: Acquire an initial data set corresponding to the task to be inferred, wherein the initial data set is a data set of task answers corresponding to the reasoning subtasks in the task to be inferred without a thought chain or without a thought chain; Determining text prompt information corresponding to the reasoning subtask in the initial data set; Based on the text prompt information and the high-quality sample data set, a plurality of high-quality thought chains corresponding to the task to be reasoned are generated.

3. The method according to claim 2, characterized in that The step of generating a plurality of high-quality thought chains corresponding to the task to be reasoned based on the text prompt information and the high-quality sample data set includes: Acquire contextual example information from a high-quality sample dataset based on the reasoning subtask; Constructing task prompt information corresponding to the inference sample in the initial data set based on the text prompt information and the context example information; A large language model is used to generate a number of high-quality thought chains based on the task prompt information.

4. The method according to claim 3, characterized in that The step of determining the optimal thinking chain data in the high-quality thinking chain includes: Filtering the high-quality thought chains according to the maximum length value of the large language model to obtain a first thought chain; Filter the first chain of thinking based on the number of steps in the first chain of thinking to obtain a second chain of thinking; The optimal chain of thinking data is acquired based on the example sequence corresponding to the context example information and the second chain of thinking.

5. The method according to claim 4, characterized in that The step of acquiring the optimal thinking chain data based on the example sequence corresponding to the context example information and the second thinking chain comprises: Adjusting the example order corresponding to the context example information in real time; Generate a plurality of rounds of thought chain data based on the adjusted context example and the second thought chain by the large language model; Remove the farthest thinking chain data corresponding to each round of thinking chain data to obtain the optimal thinking chain data.

6. The method according to any one of claims 1 to 5, characterized in that The step of performing thought chain deletion on the optimal thought chain data to obtain a target thought chain includes: Determine the fragments to be removed in the optimal thought chain data by using an adaptive thought chain deletion algorithm; The optimal thinking chain data is subjected to thinking chain deletion based on the fragment to be removed to obtain a target thinking chain.

7. A large model reasoning device, characterized in that: The device comprises: A thought chain generation module, used to generate a number of high-quality thought chains corresponding to the task to be inferred based on a high-quality sample data set, wherein the high-quality sample data set is a data set containing thought chains that meet preset high-quality conditions; A thinking chain determination module, used to determine the optimal thinking chain data in the high-quality thinking chain; A thinking chain deletion module is used to delete the optimal thinking chain data to obtain a target thinking chain; The large model reasoning module is used to output the reasoning result corresponding to the task to be reasoned based on the target thinking chain through the large language model.

8. A large model inference device, characterized in that: The device comprises: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the computer program is configured to implement the steps of the large model reasoning method according to any one of claims 1 to 6.

9. A storage medium, characterized in that: The storage medium is a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the large model reasoning method according to any one of claims 1 to 6 are implemented.

10. A computer program product, characterized in that The computer program product comprises a computer program, and when the computer program is executed by a processor, the steps of the large model reasoning method according to any one of claims 1 to 6 are implemented.

Citation Information

Cited By

  • Large language model reasoning method and device based on diffusion model, terminal and medium

    CN120806155A

  • Model training method and device, storage medium and program product

    CN121072783A