Adaptive inference deep training method and device for large language model and medium
Patent Information
- Application Number
- CN202611281265.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-08-21
- Publication Date
- 2026-09-29
AI Technical Summary
[0003]然而,现有技术普遍采用固定长度的推理范式,即无论任务难易程度,均要求大语言模型生成较长的推理过程
[0009]借由上述技术方案,本发明提供的一种面向大语言模型的自适应推理深度训练方法、装置及介质,通过本发明的技术方案,首先,通过小规模高质量的第一问题样本集,以及应用大语言模型具备根据问题输出对应回答的能力,将思考度量化为对应的思考度向量,使得思考度作为可学习的量,避免人为设定使思考度标准模糊,导致后续的最优回答训练数据不是与问题复杂度匹配的推理深度,进而使训练出的目标大语言模型无法学会按问题复杂度准确匹配推理深度,其次,在第一模型的引导下,训练轻量化,成本低,速度快的初始第二模型,训练完成后,第二模型可以代替第一模型执行根据问题输出对应回答的任务,提高了最优回答训练数据生成的效率,降低了成本,最后,在应用第二模型的基础上,对大规模的第二问题样本集,生成所有思考度向量下,第二问题样本集对应的最优回答,将第二问题样本集与第二问题样本集对应的最优回答构成最优回答训练数据,该最优回答本身就包括推理过程和答案,该推理过程是与问题复杂度匹配的推理深度,从而训练第一模型时,无需人为指定思考度向量,使得训练完成的目标大语言模型,具备按问题复杂度自适应分配推理深度的能力,在保持高推理准确率的前提下,最少的思考,从而降低推理成本,提高推理速度,另外,大规模的第二问题样本集对应大规模的最优回答训练数据,使得目标大语言模型具备较高的泛化能力和推理稳定性。
Smart Images

Figure CN122840260A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of artificial intelligence technology and is applied in scenarios such as finance or healthcare. In particular, it relates to an adaptive inference deep training method, device, and medium for large language models. Background Technology
[0002] Large language models demonstrate powerful capabilities in complex reasoning tasks, effectively processing multi-source information and performing logical deduction and structured analysis. Taking the reasoning of large language models in complex tasks such as financial analysis and healthcare as examples, in financial analysis scenarios, they can be used for knowledge retrieval and question answering, insurance contract risk analysis, business consulting, financial report analysis, investment research, and customer credit assessment. In healthcare scenarios, they can be used for medical examination report analysis, knowledge retrieval and question answering, medical record structuring, and remote consultation assistance.
[0003] However, existing technologies generally employ a fixed-length inference paradigm, requiring large language models to generate lengthy inference processes regardless of task difficulty. This approach leads to problems such as inference redundancy, error accumulation, and decreased inference stability in a large number of low-to-medium complexity tasks. Summary of the Invention
[0004] In view of this, the present invention provides an adaptive inference depth training method, device and medium for large language models, which enables large language models to adaptively allocate inference depth according to problem complexity, significantly reducing overall computational cost and improving inference speed while ensuring inference accuracy.
[0005] According to one aspect of the present invention, an adaptive inference deep training method for large language models is provided, the method comprising: Obtain a first question sample set, use the large language model as the first model, and generate a thinking degree vector corresponding to each thinking degree based on the first question sample set, the first model and the thinking degree discriminator, wherein the thinking degree discriminator is a trained classification model; Based on the first model, and according to the first problem sample set and the different thought degree vectors, an initial second model is trained to obtain the trained second model. Obtain a second question sample set with the same question distribution as the first question sample set. Based on the second question sample set, the different thought degree vectors, and the second model, determine the optimal answer training data. Train the first model based on the optimal answer training data to obtain the trained target large language model.
[0006] According to another aspect of the present invention, an adaptive inference deep training device for large language models is provided, the device comprising: The generation module is used to obtain a first question sample set, use a large language model as the first model, and generate a thinking degree vector corresponding to each thinking degree based on the first question sample set, the first model and the thinking degree discriminator, wherein the thinking degree discriminator is a trained classification model. The first training module is used to train an initial second model based on the first model, according to the first question sample set and different thought degree vectors, to obtain a trained second model. The second training module is used to obtain a second question sample set with the same question distribution as the first question sample set, determine the optimal answer training data based on the second question sample set, the different thought degree vectors and the second model, and train the first model based on the optimal answer training data to obtain the trained target large language model.
[0007] According to another aspect of the present invention, a storage medium is provided on which a computer program is stored, which, when executed by a processor, implements the above-described adaptive inference deep training method for large language models.
[0008] According to another aspect of the present invention, a computer device is provided, including a storage medium, a processor, and a computer program stored on the storage medium and executable on the processor, wherein the processor, when executing the program, implements the above-described adaptive inference deep training method for large language models.
[0009] By employing the above technical solutions, this invention provides an adaptive inference depth training method, apparatus, and medium for large language models. Through this invention's technical solution, firstly, by using a small-scale, high-quality first question sample set and leveraging the ability of a large language model to output corresponding answers based on questions, the thinking metric is quantified into a corresponding thinking degree vector. This makes the thinking degree a learnable quantity, avoiding the artificial setting that blurs the thinking degree standard, which could lead to subsequent optimal answer training data not matching the inference depth with the question complexity. Consequently, the trained target large language model would be unable to learn to accurately match the inference depth according to question complexity. Secondly, guided by the first model, a lightweight, low-cost, and fast initial second model is trained. After training, the second model can replace the first model in performing the task of outputting corresponding answers based on questions, improving the generation of optimal answer training data. This improves efficiency and reduces costs. Finally, based on the application of the second model, for a large-scale second question sample set, the optimal answer corresponding to the second question sample set under all thinking degree vectors is generated. The second question sample set and the optimal answer corresponding to the second question sample set constitute the optimal answer training data. The optimal answer itself includes the reasoning process and the answer. The reasoning process is a reasoning depth that matches the complexity of the question. Therefore, when training the first model, there is no need to manually specify the thinking degree vector. This enables the trained target large language model to have the ability to adaptively allocate reasoning depth according to the complexity of the question. While maintaining high reasoning accuracy, it requires minimal thinking, thereby reducing reasoning costs and improving reasoning speed. In addition, the large-scale second question sample set corresponds to a large-scale optimal answer training data, which enables the target large language model to have high generalization ability and reasoning stability.
[0010] The above description is merely an overview of the technical solution of the present invention. In order to better understand the technical means of the present invention and to implement it in accordance with the contents of the specification, and in order to make the above and other objects, features and advantages of the present invention more apparent and understandable, specific embodiments of the present invention are described below. Attached Figure Description
[0011] The accompanying drawings, which are included to provide a further understanding of the invention and form part of this invention, illustrate exemplary embodiments of the invention and are used to explain the invention, but do not constitute an undue limitation of this application. In the drawings: Figure 1 The diagram illustrates a flowchart of an adaptive inference deep training method for large language models provided by an embodiment of the present invention. Figure 2 This diagram illustrates another adaptive inference deep training method for large language models provided by an embodiment of the present invention. Figure 3This diagram illustrates the structure of an adaptive inference deep training device for large language models provided in an embodiment of the present invention. Figure 4 This diagram illustrates the structure of another adaptive inference deep training device for large language models provided in an embodiment of the present invention. Detailed Implementation
[0012] The present invention will be described in detail below with reference to the accompanying drawings and embodiments. It should be noted that, unless otherwise specified, the embodiments and features described in the present invention can be combined with each other.
[0013] This embodiment provides an adaptive inference deep training method for large language models, such as... Figure 1 As shown, the method includes: 101. Obtain the first question sample set, use the large language model as the first model, and generate a thinking degree vector corresponding to each thinking degree based on the first question sample set, the first model and the thinking degree discriminator, wherein the thinking degree discriminator is a trained classification model.
[0014] In this embodiment, a small-scale and high-quality first problem sample set is obtained, which is used to generate the thinking degree vector corresponding to each thinking degree and the initial second model for training in step 102 of the embodiment, so as to obtain the trained second model.
[0015] The first question sample set includes multiple first question samples. The modality of any first question sample is not limited, including but not limited to: any one or more combinations of text, image, audio, video, table, and chart.
[0016] Among them, the reasoning degree reflects the reasoning depth of the internal reasoning process of the large language model when dealing with the current complex task. This reasoning depth is determined by the number of reasoning steps of the large language model, the granularity of the reasoning chain, and the degree of computational resource investment.
[0017] It should be noted that the large language model has the ability to output corresponding answers based on the question. The answer includes the reasoning process and the answer. In this embodiment, this ability is used to quantify the degree of thinking and generate a thinking degree vector corresponding to each degree of thinking. The degree of thinking is used as a learnable quantity that can be scaled up. This avoids the problem of artificially setting the standard of thinking degree to be vague, which would cause the subsequent optimal answer training data to be not a reasoning depth that matches the complexity of the question. As a result, the trained target large language model will not be able to learn to accurately match the reasoning depth according to the complexity of the question.
[0018] However, large language models do not yet have the ability to adaptively allocate inference depth according to problem complexity. That is, regardless of problem complexity, they generate a long inference process, i.e., a long inference depth. In a large number of low-complexity problems, this will lead to inference redundancy, error accumulation and decreased inference stability. Therefore, through steps 101-103 of the implementation example, the target large language model that has been trained is obtained, which has the ability to adaptively allocate inference depth according to problem complexity.
[0019] 102. Based on the first model, and according to the first problem sample set and the different thought degree vectors, train an initial second model to obtain a trained second model.
[0020] In this embodiment, the first model is usually large, costly, and slow. Therefore, under the guidance of the first model, a lightweight, low-cost, and fast initial second model is trained. After training, the second model can replace the first model to determine the optimal answer training data.
[0021] 103. Obtain a second question sample set with the same question distribution as the first question sample set. Based on the second question sample set, the different thought degree vectors, and the second model, determine the optimal answer training data. Train the first model based on the optimal answer training data to obtain the trained target large language model.
[0022] In this embodiment, the second question sample set includes multiple second question samples, and the modality of any second question sample is not limited, including but not limited to: any one or more combinations of text, image, audio, video, table, and chart.
[0023] The thought level vector needs to be manually input into the model, meaning the model outputs a response corresponding to the thought level, preventing the model from autonomously determining the reasoning depth. The ultimate goal of this application is to enable the target large language model to adaptively allocate reasoning depth based on question complexity. Therefore, based on the second model, the optimal response corresponding to the second question sample set is generated. The second question sample set and its corresponding optimal response constitute the optimal response training data. This optimal response itself includes the reasoning process and the answer, and the reasoning process represents a reasoning depth matched to the question complexity. Thus, when training the first model, there is no need to manually specify the thought level vector, enabling the trained target large language model to adaptively allocate reasoning depth based on question complexity. This minimizes thinking while maintaining high reasoning accuracy, thereby reducing reasoning costs and increasing reasoning speed.
[0024] The second question sample set is large-scale. Therefore, the optimal answer training data obtained through the above method is also large-scale, providing large-scale training data for training the large language model, so that the trained target large language model has high generalization ability and reasoning stability.
[0025] This invention provides an adaptive inference depth training method, apparatus, and medium for large language models. Through the technical solution of this invention, firstly, by using a small-scale, high-quality first question sample set and leveraging the ability of a large language model to output corresponding answers based on questions, the thinking metric is quantified into a corresponding thinking degree vector. This makes the thinking degree a learnable quantity, avoiding the ambiguity caused by artificially setting the thinking degree standard, which would lead to subsequent optimal answer training data not matching the inference depth with the question complexity. Consequently, the trained target large language model would be unable to learn to accurately match the inference depth according to question complexity. Secondly, guided by the first model, a lightweight, low-cost, and fast initial second model is trained. After training, the second model can replace the first model in performing the task of outputting corresponding answers based on questions, improving the efficiency of generating optimal answer training data. This reduces costs. Finally, based on the application of the second model, for a large-scale second question sample set, the optimal answer corresponding to the second question sample set under all thinking degree vectors is generated. The second question sample set and the optimal answer corresponding to the second question sample set constitute the optimal answer training data. The optimal answer itself includes the reasoning process and the answer. The reasoning process is a reasoning depth that matches the complexity of the question. Therefore, when training the first model, there is no need to manually specify the thinking degree vector. This enables the trained target large language model to have the ability to adaptively allocate reasoning depth according to the complexity of the question. While maintaining high reasoning accuracy, it requires minimal thinking, thereby reducing reasoning costs and improving reasoning speed. In addition, the large-scale second question sample set corresponds to a large-scale optimal answer training data, which enables the target large language model to have high generalization ability and reasoning stability.
[0026] Furthermore, as a refinement and extension of the specific implementation methods described above, and to fully illustrate the specific implementation process in this embodiment, another adaptive inference deep training method for large language models is provided, such as... Figure 2 As shown, the method includes: 201. Obtain the first question sample set, use the large language model as the first model, and generate a thinking degree vector corresponding to each thinking degree based on the first question sample set, the first model and the thinking degree discriminator, wherein the thinking degree discriminator is a trained classification model.
[0027] In this embodiment, a small-scale (e.g., thousands of) first problem sample set is extracted from problem types such as mathematics, logical reasoning, financial calculation, and symbolic judgment. ( (This is the total number of first problem samples). It should be noted that the first problem sample set covers first problem samples of different difficulties. For example, the first problem sample set covers first problem samples of three difficulties: easy, medium, and hard. The first problem samples are high-quality seed problems. For any first problem sample, two conditions are satisfied: (1) There is at least one verifiable correct answer (a question without a verifiable correct answer is not a first question sample). (2) It has a clear input, output and intermediate reasoning structure (that is, the first problem sample itself is decomposable, the input is the first problem sample including all known conditions required to solve the goal, the output is the goal to be solved, and the intermediate reasoning structure is the first problem sample itself from input to output, there is a clear and decomposable reasoning process).
[0028] Large language models have the ability to output corresponding answers based on questions, specifically: In this embodiment, generating a thinking degree vector corresponding to each thinking degree based on the first problem sample set, the first model, and the thinking degree discriminator includes: For any level of thought, generate a vector corresponding to that level of thought, until the first answer obtained based on the vector corresponding to that level of thought, the first question sample set, and the first model is correct; Based on the first question sample set, the first answer, and the thinking degree discriminator, the probability that the first answer belongs to the thinking degree is obtained, and the first loss value is calculated based on the probability. Determine whether the first loss value is less than or equal to a first preset threshold. If not, update the vector corresponding to the thinking degree based on the first loss value until the first loss value is less than or equal to the first preset threshold, and then determine the vector corresponding to the thinking degree as the thinking degree vector.
[0029] As one implementation method, let's take three levels of consideration as an example. For low thinking, i.e., extremely short path, For intermediate thinking, that is, including necessary intermediate steps, For high-level thinking, it includes complete, fine-grained step-by-step calculations.
[0030] To accurately quantify the level of thought, a thought level vector is generated for each level. Specifically, for any given level of thought, a vector corresponding to that level is randomly generated. This vector, along with the first question sample set, is input into the first model. The first model concatenates the vector corresponding to the level of thought with the sequence of first question sample vectors corresponding to the first question sample set to obtain a concatenated vector. Based on this concatenated vector, the first answer is output. The first answer corresponds to the vector corresponding to that level of thought and is a set. Each first sub-answer corresponds to a first question sample. Each first sub-answer includes a reasoning process and an answer. The answer in each first sub-answer is extracted and compared with the correct answer of the corresponding first question sample to determine whether the answer in the corresponding first sub-answer is correct. If all answers in the first sub-answers are correct, the obtained first answer is correct. If at least one answer in the first sub-answer is incorrect, the obtained first answer is incorrect. The vector corresponding to that level of thought is adjusted until the obtained first answer is correct.
[0031] The thought level discriminator is a trained classification model whose input is a question and an answer, and whose output is a probability distribution of different thought levels, for example { 0.85, 0.1, Then, the first question sample set and the first answer are input into the thinking degree discriminator to obtain the first sub-probability of each first sub-answer belonging to that thinking degree. The set of all first sub-probabilities is taken as the probability of the first answer belonging to that thinking degree. The first sub-loss value is calculated based on the first sub-probability. The average of all first sub-loss values is calculated to obtain the first loss value. For example, , The first sub-loss value, Given the first sub-probability, determine whether the first loss value is less than or equal to the first preset threshold. If not, update the vector based on the first loss value. For example, calculate the gradient based on the first loss value, and calculate the updated vector corresponding to the current thinking degree based on the gradient, the current vector corresponding to the thinking degree, and the preset learning rate. Continue until the first loss value is less than or equal to the first preset threshold to obtain the thinking degree vector corresponding to the thinking degree. Therefore, the thinking degree vector corresponding to each thinking degree is learned.
[0032] Finally, we obtain the thought level vector corresponding to each thought level, for example... , , .
[0033] 202. Based on the first model, and according to the first problem sample set and the different thought degree vectors, train an initial second model to obtain a trained second model.
[0034] In this embodiment, the step of training an initial second model based on the first model and according to the first question sample set and different thought degree vectors to obtain a trained second model includes: for any thought degree, inputting the first question sample set and the thought degree vector corresponding to that thought degree into the first model, and outputting the standard answer corresponding to that thought degree; inputting the first question sample set and the thought degree vector corresponding to that thought degree into the initial second model, and outputting the second answer corresponding to that thought degree; calculating a second loss value based on the standard answers and the second answers under all thought degrees, determining whether the second loss value is less than or equal to a second preset threshold, and if not, updating the parameters of the initial second model until the second loss value is less than or equal to the second preset threshold, thereby obtaining a trained second model.
[0035] It should be noted that the first model is usually large, costly, and slow. Therefore, under the guidance of the first model, a lightweight, low-cost, and fast initial second model is trained. After training, the second model can replace the first model to perform the process of step 203 of the embodiment, which is given a second question sample set and a thinking degree vector, and outputs the third answer with the corresponding thinking degree.
[0036] Any thought vector The first problem sample set Input the first model and output the thought degree vector. Corresponding standard answer In any thought degree vector Below, each first problem sample Corresponding to a standard sub-answer For example, if there are three thought degree vectors, then there are three standard sub-answers under all thought degrees. This allows the reasoning depth to be explicitly modeled as a learnable control variable. The same principle applies to the second answer, and will not be elaborated upon here.
[0037] The second loss value is calculated based on the standard answer and the second answer under all the stated consideration levels. Specifically: for the same consideration level, the initial second model predicts multiple words for each position, and each word corresponds to a second sub-probability. This is compared with the standard word at the same position in the standard answer. The probability corresponding to the word that matches the standard word is taken as the target second sub-probability, thus obtaining the target second sub-probability for each position. For each target second sub-probability, the corresponding second sub-loss value is calculated, for example... , This is the second sub-loss value. Given the second sub-probability, calculate the average of all second sub-loss values to obtain the average value corresponding to that level of thought. Similarly, obtain the average value corresponding to all levels of thought, and continue to calculate the average of the average values corresponding to all levels of thought to obtain the second loss value.
[0038] 203. Obtain a second question sample set with the same question distribution as the first question sample set. Based on the second question sample set, the different thought degree vectors, and the second model, determine the optimal answer training data. Train the first model based on the optimal answer training data to obtain the trained target large language model.
[0039] It should be noted that the sample set for the second question ( (The total number of samples in the second problem) and the first problem sample set in step 201 of the embodiment. The difference is that the first question sample set is small, such as thousands, while the second question sample set is large, such as tens of thousands or hundreds of thousands. The purpose of the first question sample set is to obtain the thought degree vector and the second model, while the purpose of the second question sample set is to determine the massive amount of optimal answer training data to train the first model.
[0040] The same problem distribution means that the proportions of the first and second problem sample sets are the same across multiple dimensions. For example, if the first problem sample set covers first problem samples of three difficulty levels (easy, medium, and hard) with the same proportion of each difficulty level, then the second problem sample set will also cover second problem samples of three difficulty levels (easy, medium, and hard) with the same proportion of each difficulty level. If the first problem sample set includes first problem samples of four problem types (mathematics, logical reasoning, financial calculation, and symbolic judgment) with the same proportion of each problem type, then the second problem sample set will also cover second problem samples of four problem types (mathematics, logical reasoning, financial calculation, and symbolic judgment) with the same proportion of each problem type.
[0041] The reason why the second problem sample set has the same problem distribution as the first problem sample set is that, since the thinking degree vector and the second model are both determined based on the first problem sample set, in order to ensure that the thinking degree vector and the second model are equally effective on the second problem sample set, the second problem sample set should have the same problem distribution as the first problem sample set.
[0042] In this embodiment, determining the optimal answer training data based on the second question sample set, different thought degree vectors, and the second model includes: for any thought degree, inputting the second question sample set and the thought degree vector corresponding to that thought degree into the second model to generate a third answer corresponding to that thought degree; filtering out the optimal answer corresponding to the second question sample set from the third answers under all thought degrees; and determining the second question sample set and the optimal answer corresponding to the second question sample set as the optimal answer training data.
[0043] The step of selecting the optimal answer corresponding to the second question sample set from the third answers under all the aforementioned levels of consideration includes: Based on the correctness verification, the correct answers corresponding to the second question sample set are selected from the third answers under all the aforementioned thinking levels; The answer is removed from the correct answer to obtain the reasoning process. Based on the preset segmentation rule, the number of steps in the reasoning process is determined. The number of steps is used as the length of the reasoning chain. The correct answer with the shortest reasoning chain length is selected as the optimal answer corresponding to the second question sample set.
[0044] Wherein, in any thought degree vector Below is the third answer corresponding to the second question sample set. In any thought degree vector Below, each second problem sample Corresponding to a third sub-answer For example, if there are three thought degree vectors, then each second question sample... The three third sub-answers corresponding to all levels of consideration are: .
[0045] For the same sample of the second question, among all the third sub-answers under all levels of thought, the shorter the reasoning chain length, the fewer steps the second model takes to obtain the correct answer. This reduces computational costs, increases reasoning speed, and reduces the accumulation of errors in redundant reasoning processes. Therefore, the correct answer with the shortest reasoning chain length is the optimal answer.
[0046] The correctness verification process uses the correct answer as a benchmark to judge and filter the answers. Specifically, the correctness verification is implemented as follows: each third sub-answer includes a reasoning process and an answer; the answers corresponding to the three third sub-answers are extracted and compared with the second question sample. The system matches the corresponding correct answer, removes incorrect third-sub-answers, filters out correct third-sub-answers, and selects the correct third-sub-answer as the correct answer. For example... The answer is incorrect. and If the answer is correct, then and As a sample of the second problem The two corresponding correct answers.
[0047] For each correct answer, including the reasoning process and the answer, the answer is removed from each correct answer. Based on preset segmentation rules, such as sequence number segmentation and symbol segmentation, the number of steps corresponding to the two reasoning processes is determined, which is the length of the corresponding reasoning chain. The correct answer with the shortest reasoning chain length is used as the sample for the second question. The corresponding optimal answer ,For example, This is the second problem sample. The corresponding optimal answer.
[0048] The optimal answer to the second question sample is: .
[0049] It should be noted that the thought level vector needs to be manually input into the model, meaning that the model outputs answers corresponding to the thought level, and cannot autonomously determine the reasoning depth. However, the ultimate goal of this application is to enable the target large language model to adaptively allocate reasoning depth according to question complexity. Therefore, based on the application of the second model, the optimal answer corresponding to the second question sample set is generated. The second question sample set and the optimal answer corresponding to the second question sample set constitute the optimal answer training data. The optimal answer itself includes the reasoning process and the answer. The reasoning process is a reasoning depth that matches the question complexity. Thus, when training the first model, there is no need to manually specify the thought level vector, so that the trained target large language model has the ability to adaptively allocate reasoning depth according to question complexity. While maintaining high reasoning accuracy, it requires minimal thinking, thereby reducing reasoning costs and improving reasoning speed.
[0050] In this embodiment, training the first model based on the optimal answer training data to obtain the trained target large language model includes: inputting the second question sample set into the first model, obtaining a fourth answer from the first model, calculating a third loss value based on the fourth answer and the optimal answer corresponding to the second question sample set, determining whether the third loss value is less than or equal to a third preset threshold, and if not, updating the parameters of the first model until the third loss value is less than or equal to the third preset threshold to obtain the target large language model.
[0051] The process of calculating the third loss value is the same as that of calculating the second loss value, and will not be repeated here.
[0052] 204. Use the trained target large language model as the first model for updating.
[0053] 205. Obtain a third problem sample set with a different problem distribution than the first problem sample set, and generate an updated thinking degree vector corresponding to each thinking degree based on the third problem sample set, the updated first model, and the thinking degree discriminator.
[0054] 206. Obtain a fourth question sample set with the same question distribution as the third question sample set. Based on the fourth question sample set, the different updated thinking degree vectors, and the second model, determine the updated optimal answer training data. Train the updated first model based on the updated optimal answer training data to obtain the trained updated target large language model.
[0055] For steps 204-206 of the embodiment, the third question sample set includes multiple third question samples, and the modality of any third question sample is not limited, including but not limited to: any one or more combinations of text, image, audio, video, table, and chart. Similarly, the fourth question sample set includes multiple fourth question samples, and the modality of any fourth question sample is not limited, including but not limited to: any one or more combinations of text, image, audio, video, table, and chart.
[0056] The self-evolution of the target large language model is achieved, that is, the ability to adaptively allocate inference depth and accuracy are continuously jointly optimized through multiple rounds of iteration. Specifically, in order to ensure that the target large language model can still maintain the ability to adaptively allocate inference depth when facing new problem distributions and avoid overfitting to the distribution of the first problem sample set, step 205 of the embodiment uses a third problem sample set with a different problem distribution than the first problem sample set to regenerate an updated thought degree vector adapted to the new problem distribution, thereby driving the continuous evolution of the target large language model. Specifically, the method for generating the thought degree vector corresponding to each thought degree in step 205 of the embodiment is the same as that in step 201 of the embodiment, and will not be repeated here. The method for training the first model and obtaining the trained target large language model in step 206 of the embodiment is the same as that in step 203 of the embodiment, and will not be repeated here.
[0057] This invention provides an adaptive inference depth training method, apparatus, and medium for large language models. Through the technical solution of this invention, firstly, by using a small-scale, high-quality first question sample set and leveraging the ability of a large language model to output corresponding answers based on questions, the thinking metric is quantified into a corresponding thinking degree vector. This makes the thinking degree a learnable quantity, avoiding the ambiguity caused by artificially setting the thinking degree standard, which would lead to subsequent optimal answer training data not matching the inference depth with the question complexity. Consequently, the trained target large language model would be unable to learn to accurately match the inference depth according to question complexity. Secondly, guided by the first model, a lightweight, low-cost, and fast initial second model is trained. After training, the second model can replace the first model in performing the task of outputting corresponding answers based on questions, improving the efficiency of generating optimal answer training data. This reduces costs. Finally, based on the application of the second model, for a large-scale second question sample set, the optimal answer corresponding to the second question sample set under all thinking degree vectors is generated. The second question sample set and the optimal answer corresponding to the second question sample set constitute the optimal answer training data. The optimal answer itself includes the reasoning process and the answer. The reasoning process is a reasoning depth that matches the complexity of the question. Therefore, when training the first model, there is no need to manually specify the thinking degree vector. This enables the trained target large language model to have the ability to adaptively allocate reasoning depth according to the complexity of the question. While maintaining high reasoning accuracy, it requires minimal thinking, thereby reducing reasoning costs and improving reasoning speed. In addition, the large-scale second question sample set corresponds to a large-scale optimal answer training data, which enables the target large language model to have high generalization ability and reasoning stability.
[0058] Furthermore, as Figure 1 and Figure 2 The specific implementation of the method shown in this invention provides an adaptive inference deep training device for large language models, such as... Figure 3 As shown, the device includes: a generation module 31, a first training module 32, and a second training module 33; The generation module 31 is used to obtain a first question sample set, use a large language model as the first model, and generate a thinking degree vector corresponding to each thinking degree based on the first question sample set, the first model and the thinking degree discriminator, wherein the thinking degree discriminator is a trained classification model. The first training module 32 is used to train an initial second model based on the first model, according to the first problem sample set and different thought degree vectors, to obtain a trained second model. The second training module 33 is used to obtain a second question sample set with the same question distribution as the first question sample set, determine the optimal answer training data based on the second question sample set, the different thought degree vectors and the second model, and train the first model based on the optimal answer training data to obtain the trained target large language model.
[0059] Accordingly, in order to generate a thinking degree vector corresponding to each thinking degree based on the first question sample set, the first model, and the thinking degree discriminator, the generation module 31 is specifically used to generate a vector corresponding to any thinking degree until a first answer is obtained based on the vector corresponding to the thinking degree, the first question sample set, and the first model; to obtain the probability that the first answer belongs to the thinking degree based on the first question sample set, the first answer, and the thinking degree discriminator; to calculate a first loss value based on the probability; to determine whether the first loss value is less than or equal to a first preset threshold; if not, to update the vector corresponding to the thinking degree based on the first loss value until the first loss value is less than or equal to the first preset threshold, and to determine the vector corresponding to the thinking degree as the thinking degree vector corresponding to the thinking degree.
[0060] Accordingly, in order to train an initial second model based on the first model and according to the first question sample set and different thought degree vectors, and obtain a trained second model, the first training module 32 is specifically used to: input the first question sample set and the thought degree vector corresponding to the thought degree into the first model for any thought degree, and output the standard answer corresponding to the thought degree; input the first question sample set and the thought degree vector corresponding to the thought degree into the initial second model, and output the second answer corresponding to the thought degree; calculate a second loss value based on the standard answers and the second answers under all thought degrees, determine whether the second loss value is less than or equal to a second preset threshold, and if not, update the parameters of the initial second model until the second loss value is less than or equal to the second preset threshold, and obtain a trained second model.
[0061] Accordingly, in order to determine the optimal answer training data based on the second question sample set, the different thought degree vectors, and the second model, the second training module 33 is specifically used to, for any thought degree, input the second question sample set and the thought degree vector corresponding to that thought degree into the second model to generate a third answer corresponding to that thought degree; filter out the optimal answer corresponding to the second question sample set from the third answers under all thought degrees; and determine the second question sample set and the optimal answer corresponding to the second question sample set as the optimal answer training data.
[0062] Accordingly, in order to select the optimal answer corresponding to the second question sample set from the third answers under all the aforementioned thinking levels, the second training module 33 is specifically used to select the correct answer corresponding to the second question sample set from the third answers under all the aforementioned thinking levels based on automatic correctness verification; remove answers from the correct answers to obtain the reasoning process; determine the number of steps in the reasoning process based on preset segmentation rules; use the number of steps as the reasoning chain length; and select the correct answer with the shortest reasoning chain length as the optimal answer corresponding to the second question sample set.
[0063] Accordingly, in order to train the first model based on the optimal answer training data and obtain the trained target large language model, the second training module 33 is specifically used to input the second question sample set into the first model, obtain a fourth answer from the first model, calculate a third loss value based on the fourth answer and the optimal answer corresponding to the second question sample set, determine whether the third loss value is less than or equal to a third preset threshold, and if not, update the parameters of the first model until the third loss value is less than or equal to the third preset threshold to obtain the target large language model.
[0064] In specific application scenarios, an adaptive inference deep training device for large language models, such as... Figure 4 As shown, the device further includes: an update module 34, specifically used to use the trained target large language model as the update first model; obtain a third question sample set with a different question distribution than the first question sample set; generate an updated thinking degree vector corresponding to each thinking degree based on the third question sample set, the update first model, and the thinking degree discriminator; obtain a fourth question sample set with the same question distribution as the third question sample set; determine the update optimal answer training data based on the fourth question sample set, the different update thinking degree vectors, and the second model; train the update first model based on the update optimal answer training data to obtain the trained updated target large language model.
[0065] It should be noted that other corresponding descriptions of the functional units involved in the adaptive inference deep training device for large language models provided in this embodiment can be found in [reference needed]. Figures 1 to 2 The corresponding description will not be repeated here.
[0066] Based on the above, Figures 1 to 2 Accordingly, this embodiment also provides a storage medium, which may be volatile or non-volatile, storing a computer program that, when executed by a processor, implements the above-described method. Figures 1 to 2 The method shown is an adaptive inference deep training method for large language models.
[0067] Based on this understanding, the technical solution of the present invention can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (such as a CD-ROM, USB flash drive, portable hard drive, etc.) and includes several instructions to cause a computer device (such as a personal computer, server, or network device, etc.) to execute the methods of various implementation scenarios of the present invention.
[0068] Based on the above, Figures 1 to 2 The method shown and Figure 3 , Figure 4 To achieve the above objectives, the present application also provides a computer device, specifically a personal computer, server, network device, etc., as shown in the illustrated embodiment. This computer device includes a storage medium and a processor; the storage medium stores a computer program; the processor executes the computer program to achieve the above-described objectives. Figure 1 and Figure 2 The method shown is an adaptive inference deep training method for large language models.
[0069] Optionally, the computer device may also include a user interface, a network interface, a camera, radio frequency (RF) circuitry, sensors, audio circuitry, a Wi-Fi module, etc. The user interface may include a display screen, input units such as a keyboard, etc., and optional user interfaces may also include USB interfaces, card reader interfaces, etc. The network interface may optionally include standard wired interfaces, wireless interfaces (such as Wi-Fi interfaces), etc.
[0070] Those skilled in the art will understand that the computer device structure provided in this embodiment does not constitute a limitation on the physical device, and may include more or fewer components, or combine certain components, or have different component arrangements.
[0071] The storage medium may also include an operating system and a network communication module. The operating system is a program that manages the hardware and software resources of the aforementioned computer device, supporting the operation of information processing programs and other software and / or programs. The network communication module is used to enable communication between the various components within the non-volatile storage medium, as well as communication with other hardware and software in the information processing entity device.
[0072] Through the above description of the embodiments, those skilled in the art can clearly understand that the present invention can be implemented by means of software plus necessary general-purpose hardware platform, or it can be implemented by hardware.
[0073] This invention provides an adaptive inference depth training method, apparatus, and medium for large language models. Through the technical solution of this invention, firstly, by using a small-scale, high-quality first question sample set and leveraging the ability of a large language model to output corresponding answers based on questions, the thinking metric is quantified into a corresponding thinking degree vector. This makes the thinking degree a learnable quantity, avoiding the ambiguity caused by artificially setting the thinking degree standard, which would lead to subsequent optimal answer training data not matching the inference depth with the question complexity. Consequently, the trained target large language model would be unable to learn to accurately match the inference depth according to question complexity. Secondly, guided by the first model, a lightweight, low-cost, and fast initial second model is trained. After training, the second model can replace the first model in performing the task of outputting corresponding answers based on questions, improving the efficiency of generating optimal answer training data. This reduces costs. Finally, based on the application of the second model, for a large-scale second question sample set, the optimal answer corresponding to the second question sample set under all thinking degree vectors is generated. The second question sample set and the optimal answer corresponding to the second question sample set constitute the optimal answer training data. The optimal answer itself includes the reasoning process and the answer. The reasoning process is a reasoning depth that matches the complexity of the question. Therefore, when training the first model, there is no need to manually specify the thinking degree vector. This enables the trained target large language model to have the ability to adaptively allocate reasoning depth according to the complexity of the question. While maintaining high reasoning accuracy, it requires minimal thinking, thereby reducing reasoning costs and improving reasoning speed. In addition, the large-scale second question sample set corresponds to a large-scale optimal answer training data, which enables the target large language model to have high generalization ability and reasoning stability.
[0074] Those skilled in the art will understand that the accompanying drawings are merely schematic diagrams of a preferred embodiment, and the modules or processes shown in the drawings are not necessarily essential for implementing the present invention. Those skilled in the art will understand that the modules in the apparatus of the embodiment can be distributed within the apparatus of the embodiment as described, or they can be located in one or more apparatuses different from this embodiment, with corresponding changes. The modules of the above-described embodiment can be combined into one module, or further divided into multiple sub-modules.
[0075] The serial numbers used above are for descriptive purposes only and do not represent the superiority or inferiority of the implementation scenarios. The above disclosures are merely a few specific implementation scenarios of the present invention; however, the present invention is not limited thereto, and any variations conceived by those skilled in the art should fall within the protection scope of the present invention.
Claims
1. An adaptive inference deep training method for large language models, characterized in that, The method includes: Obtain a first question sample set, use the large language model as the first model, and generate a thinking degree vector corresponding to each thinking degree based on the first question sample set, the first model and the thinking degree discriminator, wherein the thinking degree discriminator is a trained classification model; Based on the first model, and according to the first problem sample set and the different thought degree vectors, an initial second model is trained to obtain the trained second model. Obtain a second question sample set with the same question distribution as the first question sample set. Based on the second question sample set, the different thought degree vectors, and the second model, determine the optimal answer training data. Train the first model based on the optimal answer training data to obtain the trained target large language model.
2. The method according to claim 1, characterized in that, The step of generating a thinking degree vector corresponding to each thinking degree based on the first problem sample set, the first model, and the thinking degree discriminator includes: For any level of thought, generate a vector corresponding to that level of thought, until the first answer obtained based on the vector corresponding to that level of thought, the first question sample set, and the first model is correct; Based on the first question sample set, the first answer, and the thinking degree discriminator, the probability that the first answer belongs to the thinking degree is obtained, and the first loss value is calculated based on the probability. Determine whether the first loss value is less than or equal to a first preset threshold. If not, update the vector corresponding to the thinking degree based on the first loss value until the first loss value is less than or equal to the first preset threshold, and then determine the vector corresponding to the thinking degree as the thinking degree vector.
3. The method according to claim 1, characterized in that, The step of training an initial second model based on the first model, according to the first problem sample set and different thought degree vectors, to obtain a trained second model includes: For any level of consideration, input the first question sample set and the consideration vector corresponding to that level of consideration into the first model, and output the standard answer corresponding to that level of consideration; Input the first question sample set and the thought degree vector corresponding to the thought degree into the initial second model, and output the second answer corresponding to the thought degree; Calculate a second loss value based on the standard answer and the second answer under all the aforementioned thinking degrees, determine whether the second loss value is less than or equal to a second preset threshold, if not, update the parameters of the initial second model until the second loss value is less than or equal to the second preset threshold, and obtain the trained second model.
4. The method according to claim 1, characterized in that, The step of determining the optimal answer training data based on the second question sample set, different thought degree vectors, and the second model includes: For any level of thought, input the second question sample set and the thought level vector corresponding to that level of thought into the second model to generate a third answer corresponding to that level of thought; Select the optimal answer corresponding to the sample set of the second question from all the third answers under all the aforementioned levels of consideration; The second question sample set and the optimal answer corresponding to the second question sample set are determined as the optimal answer training data.
5. The method according to claim 4, characterized in that, The step of selecting the optimal answer corresponding to the second question sample set from the third answers under all the aforementioned levels of consideration includes: Based on the correctness verification, the correct answers corresponding to the second question sample set are selected from the third answers under all the aforementioned thinking levels; The answer is removed from the correct answer to obtain the reasoning process. Based on the preset segmentation rule, the number of steps in the reasoning process is determined. The number of steps is used as the length of the reasoning chain. The correct answer with the shortest reasoning chain length is selected as the optimal answer corresponding to the second question sample set.
6. The method according to claim 4, characterized in that, The step of training the first model based on the optimal answer training data to obtain the trained target large language model includes: The second question sample set is input into the first model, and the first model obtains the fourth answer; Based on the optimal answer corresponding to the fourth answer and the second question sample set, calculate the third loss value and determine whether the third loss value is less than or equal to the third preset threshold. If not, update the parameters of the first model until the third loss value is less than or equal to the third preset threshold to obtain the target large language model.
7. The method according to claim 1, characterized in that, The method further includes: Use the trained target large language model as the first model for updating; Obtain a third question sample set with a different question distribution from the first question sample set; and generate an updated thinking degree vector corresponding to each thinking degree based on the third question sample set, the updated first model, and the thinking degree discriminator. Obtain a fourth question sample set with the same question distribution as the third question sample set. Based on the fourth question sample set, the different updated thinking degree vectors, and the second model, determine the updated optimal answer training data. Train the updated first model based on the updated optimal answer training data to obtain the trained updated target large language model.
8. An adaptive inference deep training device for large language models, characterized in that, The device includes: The generation module is used to obtain a first question sample set, use a large language model as the first model, and generate a thinking degree vector corresponding to each thinking degree based on the first question sample set, the first model and the thinking degree discriminator, wherein the thinking degree discriminator is a trained classification model. The first training module is used to train an initial second model based on the first model, according to the first question sample set and different thought degree vectors, to obtain a trained second model. The second training module is used to obtain a second question sample set with the same question distribution as the first question sample set, determine the optimal answer training data based on the second question sample set, the different thought degree vectors and the second model, and train the first model based on the optimal answer training data to obtain the trained target large language model.
9. A storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, it implements the adaptive inference deep training method for large language models as described in any one of claims 1 to 7.
10. A computer device comprising a memory, a processor, and a computer program stored on a storage medium and executable on the processor, characterized in that, When the processor executes the program, it implements the adaptive inference deep training method for large language models as described in any one of claims 1 to 7.