Mathematical application question automatic answering method based on large model knowledge distillation

Through a method based on large-scale knowledge distillation, high-quality training data is generated using diversity problem generators and data filters, and targeted data enhancement is performed by evaluating the learning status of the model, which solves the problems of deployment difficulties of large-scale language models and time-consuming model training in the existing technology, and realizes the efficient solution ability of student models.

CN120067261AActive Publication Date: 2025-05-30HUAZHONG NORMAL UNIV

Patent Information

Application Number
CN202510148440.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-11
Publication Date
2025-05-30
Estimated Expiration
2045-02-11

AI Technical Summary

Technical Problem

At this stage, the automatic answering method of mathematical application questions based on large language models has problems such as huge parameters, difficulty in deployment, time-consuming model training and high computing resources, difficulty in understanding the reasoning steps of the teacher model of the student model, and lack of feedback mechanism.

Method used

An automatic solution method for mathematical application problems based on large model knowledge distillation is proposed. By designing diversity sub-problem generators, diversity analog problem generators and data filters, a high-quality training data set is generated, and targeted data enhancement is performed by evaluating the learning state of the model, and the solution ability of students' models is dynamically improved.

Benefits of technology

The student model has achieved the ability to solve mathematical application problems close to the large language model with very few parameters, which reduces the cost of practical application and improves the learning efficiency and solution accuracy of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120067261A_ABST
    Figure CN120067261A_ABST
Patent Text Reader

Abstract

The invention relates to a mathematical application question automatic answering method based on large model knowledge distillation, comprising: inputting a mathematical application question to be answered into a large model-based knowledge distillation model, generating a mathematical expression as an equation solution sequence, the large model-based knowledge distillation model comprising a teacher model and a student model, the teacher model is constructed based on a large language model LLMs, and the student model is obtained by training data generated by the teacher model. A large language model is used as a teacher model, firstly, multi-aspect mathematical application questions are designed and generated to perform actual enhancement of student models, secondly, learning states of the student models are dynamically evaluated so as to generate various mathematical application question variants in a targeted manner, the understanding ability of the student models for semantics and scenes of the mathematical application questions is improved, and the teaching efficiency is improved. The problem that a model with few parameters is poor in solving capacity is solved, and a student model obtains the mathematical application problem solving capacity close to a large language model with few parameters.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of natural language processing, and particularly to an automatic solution method for math word problems based on large model knowledge distillation. Background Art

[0002] Math Word Problem (MWP) solving is the intersection of multiple natural language processing technologies, involving popular intelligent technologies such as machine reading comprehension, knowledge Q&A, and human-machine intelligent dialogue. It is not only a core task of mathematical reasoning but also an important test benchmark for machine intelligence, and has long been widely concerned by researchers. With the emergence of PLMs and LLMs, in order to better apply to downstream tasks, the knowledge distillation paradigm has become popular in natural language processing. The basic idea behind it is to extract a model with a small number of parameters for specific task services from a language model with a huge number of parameters. Due to the relatively small scale of the MWP dataset, when fine-tuning a large general network for downstream processing of word problem solving, knowledge distillation enables the model to concentrate the knowledge of a general model into a smaller and more focused model. When performing knowledge distillation based on LLMs, most previous studies have emphasized using the CoT method. CoT uses the inference steps generated by a large model as the teacher model as the distilled knowledge to fine-tune a smaller student model. These methods have all successfully improved the performance of the student model, but also shown certain limitations.

[0003] At the current stage, there are still certain defects in the research on MWP solving based on LLMs: First, LLMs have tens of billions or hundreds of billions of parameters and cannot be deployed on a large scale. Some require API availability to be reproduced. When performing targeted fine-tuning on downstream tasks, a large amount of time and computing power resources are also required. These further limit the popularization and use of LLMs in professional fields. Second, in knowledge distillation based on LLMs, it is difficult for small student models to understand the inference steps of CoT, which will affect their learning efficiency. There is research evidence showing that only particularly large language models have the ability to perform CoT during the reasoning process. Therefore, many student models trained with CoT inference steps have not achieved satisfactory accuracy in the MWP solving task. Finally, the existing knowledge distillation methods based on LLMs lack feedback from the student model to the LLMs teacher model, ignoring the evaluation of the state of the student's learned knowledge, which further affects the performance of the student model. To solve the above problems, the present invention proposes an automatic solution method for math word problems based on large model knowledge distillation. Summary of the Invention

[0004] The object of the present invention is to provide an automatic solution method for math word problems based on large model knowledge distillation, enabling the student model to obtain the math word problem solving ability close to that of the large language model with extremely few parameters.

[0005] To achieve the above object, the present invention provides the following solutions:

[0006] A method for automatically solving mathematical word problems based on large model knowledge distillation, comprising:

[0007] Obtain the mathematical word problem to be solved;

[0008] Input the mathematical word problem to be solved into a large model-based knowledge distillation model to generate a mathematical expression as a solution sequence, wherein the large model-based knowledge distillation model includes a teacher model and a student model, the teacher model is constructed based on the large language model LLMs, and the student model is obtained by training with the data generated by the teacher model.

[0009] Optionally, the teacher model includes a diversity sub-problem generator, a diversity analogy problem generator, and a data filter;

[0010] The diversity sub-problem generator is used to input the initial source data set to generate multiple problems with the same semantic description and sub-equations and corresponding answers;

[0011] The diversity analogy problem generator is used to input the initial source data set to generate mathematical word problems with the same problem and different semantic descriptions;

[0012] The data filter is used to screen and filter the problems generated by the diversity sub-problem generator and the diversity analogy problem generator to obtain a training data set.

[0013] Optionally, screening and filtering the problems generated by the diversity sub-problem generator and the diversity analogy problem generator includes: filtering the problems with incorrect generation formats, complementing the incomplete problems generated based on the zero-shot prompting technology of LLMs, and eliminating the problems that cannot be answered and have incorrect formats to obtain the training data set.

[0014] Optionally, training the student model with the data generated by the teacher model includes:

[0015] S1. After training the student model with the training data set for a preset number of times, use the evaluation data set to evaluate the learning state of the student model after training for the preset number of times to obtain the problems that cannot be solved, wherein the evaluation data set is selected from the data generated by the teacher model;

[0016] S2. Call the diversity sub-problem generator and the diversity analogy problem generator in the teacher model to generate targeted data for the problems that cannot be solved to obtain an enhanced data set;

[0017] S3. Incorporate the enhanced dataset into the training dataset, and return to S1 until a preset stopping condition is reached.

[0018] Optionally, the student model includes a model based on an encoder-decoder architecture and a model based on a decoder architecture.

[0019] Optionally, the model based on the encoder-decoder architecture includes: using GTS as the decoder, the backbone network LSTM, RoBERTa-base and RoBERTa-large as the encoders, where the encoder is used to model the input sequence and return the hidden state of the input math problem, and the decoder is used to generate the distribution of solving equations according to the output features of the encoder.

[0020] Optionally, the model based on the decoder architecture is a large language model with a decoder architecture, where the large language model with a decoder architecture is fine-tuned using a reparameterization method during training.

[0021] The beneficial effects of the present invention are as follows: (1) Targeted Knowledge Distillation (TKD) uses LLMs as the role of the teacher model, designs a generation module for generating diverse MWP variants, and mainly generates two types of variants from the source MWP: sub-problems with the same text description but diverse problems, and similar problems with the same problem description but diverse text descriptions and scenarios. At the same time, a question filter is designed to clean and screen the generated questions to obtain high-quality MWP. TKD does not require the small model to have the CoT ability, enabling the small model to learn more effectively. (2) TKD designs an evaluation module for evaluating the learning state of the model, uses the generated comprehensive and diverse questions to evaluate the student model, finds the weak points of the solving ability according to the learning state of the student model, and feeds them back to the LLMs for targeted dynamic generation. TKD dynamically makes up for the deficiencies in the areas where the student model is not good at learning, further improving the learning efficiency and solving ability of the student model. (3) Compared with existing work, TKD is based on the API application of GPT-4. The student model is designed in two categories. One category uses GTS as the decoder, and the encoder part uses three different backbone networks LSTM, RoBERTa-base and RoERTa-large. The other category uses the less parameter large models Llama2-7b and Qwen2.5-7b. Experimental results show that under the condition of reducing the actual application cost, TKD not only has significantly better accuracy than the fine-tuning baseline, but also can achieve a solving ability comparable to that of LLMs with a student model with very few parameters. Description of the Drawings

[0022] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the accompanying drawings required in the embodiments. Obviously, the accompanying drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can be obtained based on these drawings.

[0023] Figure 1 It is the framework diagram of the knowledge distillation model based on the large model for the embodiments of the present invention;

[0024] Figure 2 It is the comparison diagram of the ablation experiment on the MAWPS dataset for the embodiments of the present invention;

[0025] Figure 3 It is the comparison diagram of the ablation experiment on the SVAMP dataset for the embodiments of the present invention;

[0026] Figure 4 It is the comparison diagram of the ablation experiment on the MathQA dataset for the embodiments of the present invention. Detailed implementation manners

[0027] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, rather than all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.

[0028] To make the above objects, features, and advantages of the present invention more obvious and understandable, the present invention will be further described in detail below in conjunction with the accompanying drawings and specific implementation manners.

[0029] This embodiment provides a method for automatically solving math word problems based on large model knowledge distillation, including:

[0030] Input the math word problem to be solved into the knowledge distillation model based on the large model to generate a math expression as a solution sequence. Among them, the knowledge distillation model based on the large model includes a teacher model and a student model. The teacher model is constructed based on the large language model LLMs, and the student model is obtained by training with the data generated by the teacher model.

[0031] Further, the teacher model includes a diversity sub-problem generator, a diversity analogy problem generator, and a data filter;

[0032] The diversity sub-problem generator is used to input the initial source dataset to generate multiple problems with the same semantic description and sub-equations and corresponding answers;

[0033] A diversity analogy question generator for generating math word problems with the same questions and different semantic descriptions from an initial source dataset input;

[0034] A data filter for screening and filtering the questions generated by the diversity sub-question generator and the diversity analogy question generator to obtain a training dataset.

[0035] Specifically, in this embodiment, in combination with the characteristics of the LLM, two question generators are designed for data generation and one data filter is designed for data filtering, using MWP source to represent the initial source dataset, which is the initial input. Before starting training, diversity data generation is performed on MWP source .

[0036] In the diversity sub-question generator G 1 , MWP with the same semantic description and different questions will be generated. To verify the model's scenario understanding ability, when the model understands the scenario of a math problem and its underlying logic, they should be able to solve the perceived problems in that scenario and generate corresponding sub-questions. This research is based on BERT to decompose semantics and equations to generate diverse sub-equations and sub-questions. Now, due to the powerful functionality of LLMs, in this embodiment, through a small number of example and prompt designs, multiple questions and corresponding answers with the same semantic description and sub-equations can be generated.

[0037] In the diversity analogy question generator G 2 , MWP with the same questions and different semantic descriptions will be generated. For MathCheck to verify the model's question understanding, the proposed rewrite variant is: transform the original question into a new question, using different wordings or different sentence structures, but without changing its original mathematical logic. It focuses on semantic robustness and aims to test whether the model can correctly reason when dealing with different descriptions of the same mathematical logic. Naturally, such rewrite variants are used for model training and can also enhance the model's further learning and understanding of semantics. Similarly, in this embodiment, using LLMs, multiple analogy MWPs with the same questions and different semantic descriptions as the source question can be generated through simple prompt designs with zero-shot.

[0038] This embodiment designs a data filter to screen the questions generated by the above two generators, and filtering out harmful augmented data can improve data quality and the performance of downstream solvers.

[0039] Furthermore, screening and filtering the questions generated by the diversity sub-question generator and the diversity analogy question generator includes: filtering out questions with incorrect generation formats, and based on the zero-shot prompting technology of LLMs, completing the questions with incomplete generation, eliminating questions that cannot be answered and have incorrect formats, to obtain a training dataset.

[0040] Furthermore, training the student model with the data generated by the teacher model includes:

[0041] S1. After training the student model with the training dataset for a preset number of times, use the evaluation dataset to evaluate the learning state of the student model after the preset number of training times, and obtain the problems that cannot be solved. Among them, the evaluation dataset is selected from the data generated by the teacher model;

[0042] S2. Call the diversity sub-problem generator and the diversity analogy problem generator in the teacher model to generate targeted data for the problems that cannot be solved, and obtain the enhanced dataset;

[0043] S3. Merge the enhanced dataset into the training dataset, and return to S1 until the preset stop condition is reached.

[0044] Specifically, in this embodiment, MWP train represents the part of the dataset used to train the student model, and MWP val represents the part of the dataset used to evaluate the learning state of the student model. A certain proportion of data is randomly selected from the data generated by the data generation module based on LLMs as MWP val for subsequent evaluation, and the remaining part is used as MWP train for training. At the beginning of training, in this embodiment, the student model solver in the S representation framework is used, and the student model S is made to learn on MWP train After learning a certain amount of knowledge, that is, after training a certain number of times, evaluate the learning state of S. According to the performance of S on the specifically generated evaluation-specific dataset MWP val for the (X, Y) problem pairs that it cannot solve, which reflects the deficiencies in the knowledge learned by S during training, call the data generation module G based on LLMs to generate targeted data, and the generated (X, Y) form the enhanced dataset. In this embodiment, MWP aug is used to represent. After each round of targeted data enhancement, MWP aug is merged into MWP train to help S in the new round of training. Since during the training process, whenever S learns a sufficient amount of knowledge, S will be evaluated and targeted data enhancement will be performed, and added to the subsequent training.

[0045] Furthermore, the student model includes a model based on the encoder-decoder architecture and a model based on the decoder architecture.

[0046] Furthermore, the model based on the encoder-decoder architecture includes: using GTS (Goal-driven TreeStructured decoder) as the decoder, the backbone network LSTM, RoBERTa-base, and RoBERTa-large as the encoders. The encoder is used to model the input sequence and return the hidden state of the input math problem, and the decoder is used to generate the distribution of solving equations based on the output features of the encoder.

[0047] Furthermore, the model based on the decoder architecture is a large language model with a decoder architecture. During training, the large language model with a decoder architecture is fine-tuned using the reparameterization method.

[0048] Specifically, in this embodiment, the models based on the decoder architecture are Llama2-7b and Qwen2.5-7b that only have the decoder architecture.

[0049] The method of this embodiment will be further described below:

[0050] In this embodiment, the numbers in the problem X and the expression Y are represented using {N 0 , N 1 ,..., N j}, where N i refers to the i-th number, and j is the maximum value of the number of digits in X and Y. The numbers in math problems are theoretically impossible to be exhausted. Therefore, the general quantifier V num is constructed by the mapping function preprocessed by the arithmetic expression. This function stores all the numbers that appear in X and maps them to the digital token list X q ={N 0 , N 1 ,..., N j} in order. During the process of the system generating the expression Y, if the specified number N i is generated, the corresponding number will be automatically retrieved from the mapping function. The advantages of this definition are: (1) It is applicable to the data preprocessing techniques widely used in MWP solving, can unify the representation of quantities, and reduce the vocabulary. (2) During the data generation process, since the goal of the generator is to generate MWP variants, digital mapping can prevent the generation of variants that only change the value of the quantity.

[0051] The evaluation criteria for the MWP solving task adopt the prediction results of the accuracy evaluation model, which are divided into two criteria: answer accuracy and expression accuracy. Answer accuracy means that when the predicted calculated value is equal to the answer, the generated expression is regarded as correct. Expression accuracy is used to measure whether the generated expression tree matches the target expression tree. There is a fundamental difference between the two: for expression accuracy, only when the generated expression is correct, the answer will be considered correct, and any expression different from the target expression is regarded as incorrect. For example, the answers of generating " 0 [N] 1 +N 2 -N" and " 0 (N 1 +N 2 )-N" are the same, but the expressions are different. In this case, the accuracy of the answer is true, while the accuracy of the expression is false. 0 +N 1 -N 2 0 +N 1 )-N 2

[0052] The targeted knowledge distillation framework TKD based on large models proposed in this embodiment. The whole model can be divided into three modules: the diversity data generation module based on LLMs, the student model learning state evaluation and targeted generation module, and the teacher model and student model selection module based on LLMs. Among them, the diversity data generation module based on LLMs generates diverse variants of math word problems according to the input original math word problems, constitutes a dataset for data augmentation and evaluation, and further improves the solving ability of the student model; the student model learning state evaluation and targeted generation module mainly uses the evaluation dataset generated by the diversity data generation module based on LLMs to comprehensively evaluate the student model that has learned certain knowledge, feedback the learning state of the student model, and achieve further targeted data generation; the teacher model and student model module based on LLMs applies LLMs as the teacher role to train the appropriate student model. Figure 1 For the TKD model architecture, each module will be described in detail below.

[0053] 1. Data generation module based on LLMs:

[0054] In recent research work, a new paradigm for mathematical reasoning evaluation, MathCheck, was proposed, which requires the model to understand semantics, resist interference, and understand scenarios. It uses four question forms (including the original question and its three rewritten variants) to check the reasoning robustness of the model. The question variants it proposed are comprehensive and diverse, which is not only a new evaluation paradigm for MWP, but also can be used for model training. Inspired by the above work, this embodiment combines the characteristics of LLM, designs two question generators for data generation and one data filter for data filtering.

[0055] This embodiment uses MWP source ​​denotes the initial source dataset, serving as the very first input. Before starting the training, for MWP source diverse data generation is performed.

[0056] In the diverse sub-question generator G 1 of this embodiment, MWP with the same semantic description but different questions will be generated. To verify the model's scenario understanding ability, when the model understands the scenario of a math problem and its underlying logic, it should be able to solve other problems in that scenario and can generate rewritten variants that change its questions. In previous work, for example, the diverse equation generator represents the source MWP as (text description, question, equation), decomposes the original equation through the diverse equation generator to generate multiple sub-equations, and each generated sub-equation and the original text description are input into the equation-aware question generator to generate corresponding sub-questions. This research is based on decomposing semantics and equations by BERT to generate diverse sub-equations and sub-questions. Now, due to the powerful functionality of LLMs, in this embodiment, based on the few-shot prompting technique of LLMs, examples of variant questions and corresponding answers derived from three source questions are designed as samples, combined with "generate variants for math word problems, the answers of the variants should be correct and different from the source question, but have exactly the same format as the source question, including not only operands but also operators, excluding the modulo operator, and using N0, N1,... to represent numbers without other numerical representation forms." as a prompt, then a variety of variant questions with the same semantic description but different sub-equations and corresponding answers can be generated.

[0057] In the diverse analogy question generator G 2 of this embodiment, MWP with the same question but different semantic descriptions will be generated. MathCheck proposed the rewritten variant to verify the model's question understanding: transform the original question into a new question using different wordings or different sentence structures without changing its original mathematical logic. It focuses on semantic robustness and aims to test whether the model can reason correctly when dealing with different descriptions of the same mathematical logic. Naturally, such rewritten variants are used for model training and can also enhance the model's further learning and understanding of semantics. Similarly, in this embodiment, based on the zero-shot prompting technique of LLMs, "generate a math word problem and its answer with exactly the same format as the previous one, without explanation, using N0, N1,... to represent numbers without other numerical representation forms." is designed as a prompt, and multiple analogy math word problems with the same question and equation as the source question but different semantic descriptions and corresponding answers can be generated without giving samples.

[0058] In this embodiment, a data filter is designed to screen the problems generated by the above two generators. Filtering out harmful augmented data can improve data quality and the performance of downstream solvers. Due to the existence of large model hallucinations, the data generated by LLMs contains various problems. For example, the generated problems are incomplete, the second half is missing, or the representation format of the generated problems is inaccurate, etc., which are all difficulties to be solved. In this embodiment, format checking is first performed, and regular expressions are used to filter out problems with incorrect generation formats. Secondly, answer proofreading is performed based on the zero-shot prompting technique of LLMs. Design "generate variant problems and correct answers, without explanation, use N0, N1,... to represent numbers, without other numerical representations. Follow the same format as the following format" as the prompt. For the problems generated by the above two generators, based on LLMs, incomplete problems are completed, and problems that cannot be answered and have incorrect formats are excluded, and finally a high-quality training dataset is obtained.

[0059] Compared with the previous use of human manual annotation or ordinary models, the diverse sub-problem generator and diverse analogy problem generator designed in this embodiment achieve the automation and enrichment of problem variants and ensure high quality, thus improving the upper limit of the solver's solution.

[0060] 2. State Evaluation and Targeted Generation Module:

[0061] Evaluating requires diverse data to reflect the learning state of the solver. The solver may memorize the MWP in the training set rather than fully understand them, so the training set cannot be directly used to verify the student solver. Therefore, in this embodiment, a certain proportion of data is randomly selected from the high-quality dataset obtained by the data generation module to form an evaluation dataset to evaluate the solution ability of the student model.

[0062] During the training of the student model, in this embodiment, it is designed to evaluate the learning state of the student model using the evaluation dataset after the student model has acquired a certain amount of knowledge. The gap between the student model and the ideal state is reflected by the problems that the student model cannot solve, and the data generation module based on LLMs is called to generate targeted augmented data for the weak points of the student model, and the augmented data is supplemented into the next round of training. During the training of the student model, multiple evaluations and multiple targeted data augmentations can be performed. The pseudo-code of its algorithm is shown in Algorithm 1. In this embodiment, MWP train represents the part of the dataset used to train the student model, and MWP val represents the part of the dataset used to evaluate the learning state of the student model. A certain proportion of data is randomly extracted from the data generated by the data generation module based on LLMs as MWP val for subsequent evaluation, and the remaining part is used as MWP trainTraining is carried out. At the beginning of training, in this embodiment, S is used to represent the student model solver in the framework. Let the student model S learn on MWP train After learning a certain amount of knowledge, that is, after training a certain number of times, evaluate the learning state of S. According to the performance of S on the specifically generated evaluation dedicated dataset MWP val For the (X, Y) problem pairs that it cannot solve, which reflect the deficiencies in the knowledge learned by S during training, call the data generation module G based on LLMs to perform targeted data generation. The generated (X, Y) form an enhanced dataset, which is represented by MWP aug in this embodiment. After each round of targeted data augmentation, MWP aug is merged into MWP train to help S in the new round of training. Since during training, whenever S learns a sufficient amount of knowledge, S will be evaluated and targeted data augmentation will be performed, and added to the subsequent training. This module can be iterated multiple times during the training of S. Therefore, TKD is a dynamic evaluation and generation, with the characteristics of cycling and continuous iteration, until the data generation is accurate and reaches the given number of generations. The evaluation and targeted generation algorithms are shown in Table 1.

[0063] Table 1

[0064]

[0065]

[0066] The state evaluation and targeted generation module proposed in this embodiment gradually increases the scale of the training set, covers more and broader knowledge, and specifically makes up for the short board of the student model, thus generating a more powerful solver.

[0067] 3. Student model based on LLMs:

[0068] The TKD framework proposed in this embodiment, based on the idea of knowledge distillation, needs to select and design a suitable student model for training. In this embodiment, based on the API call of GPT-4, solvers with different parameter sizes are selected as the student model, and the number of parameters ranges from 20M to 7B, which fully reflects the generalization of the TKD framework. The design of the student model is divided into two categories, which will be introduced in detail below.

[0069] 3.1 Small models with millions of parameters:

[0070] Adopt a small model with few parameters, follow the "encoder-decoder" architecture, select GTS as the decoder, and use three different backbone networks, LSTM, RoBERTa-base, and RoBERTa-large, in the encoder part. Use the original structure of the encoder-decoder model for pre-training on the MWP dataset. Given a math problem X, the student model first uses the encoder to model the input sequence and returns the hidden state of the input math problem, which is expressed as:

[0071] h = Encoder(X) (1)

[0072] where h is the output hidden state of the encoder, Encoder ∈ {LSTM, RoBERTa-base, RoERTa-large} is one of the existing text encoder models, and the output features of the encoder are input into the tree-based decoder to generate the distribution of solving equations, which is expressed as:

[0073] P(Y|X) = TreeDecoder(h) (2)

[0074] 3.2 Large language models with billions of parameters:

[0075] Adopt few-parameter LLMs, Llama2-7b and Qwen2.5-7b with only the decoder architecture. When using Llama2-7b to train on the MWP dataset, due to the large number of model parameters, this embodiment uses the parameter-efficient fine-tuning (PEFT) method for fine-tuning. This embodiment selects the LoRA method based on reparameterization.

[0076] The method based on reparameterization aims to use low-rank techniques to transform network weights. This method effectively reduces the number of trainable parameters while retaining the ability to process high-dimensional matrices. LoRA introduces a simple method to update the parameters of the weight matrix by decomposing the weight matrix into the product of two low-rank matrices. It can be expressed as:

[0077] H 0 = H i W 0 + H i ΔW = H i W 0 + H i BA (3)

[0078] where H i ∈ R t×d and H 0 ∈ R t×d are the input and output of a specific layer, such as the attention layer or the fully connected layer, respectively. Where W 0 ∈ R d×dIt can be any pre-trained weight matrix, including the weights in the MLP or Attention layer. B ∈ R r×d and A ∈ R r×d are low-rank matrices used to cover the pre-trained parameter ΔW. r << d is an important hyperparameter of LoRA. During training, the original dense weight matrix of the pre-trained model is frozen, and additional trainable low-rank matrices can be introduced in each layer of the transformer architecture to represent the part of the weight change. This method preserves the general language features captured by the pre-trained model and, due to the use of low-rank matrices, greatly reduces the storage and computational requirements, thus enabling lightweight fine-tuning of the Llama2-7b and Qwen2.5-7b models.

[0079] Experiments:

[0080] 1. Datasets:

[0081] In this embodiment, the performance of TKD is evaluated on three English MWP datasets: MAWPS, SVAMP, and MathQA. The dataset statistics are shown in Table 2.

[0082] Table 2

[0083]

[0084] (1) MAWPS: MAWPS is a commonly used dataset with 2,373 arithmetic word problems. In this embodiment, the sub-dataset MAWPS-S provided by the paper is used. This dataset converts the equation-form answers into arithmetic expressions and removes invalid equation groups, resulting in a total of 1,987 questions. In this embodiment, a 5-fold cross-validation experiment is conducted on this dataset.

[0085] (2) SVAMP: SVAMP is a dataset sampled from existing datasets and manually modified. It carefully fine-tunes some types of questions in the original data to test the robustness of the model and its sensitivity to questions. Since this dataset has no test split, in this embodiment, the experimental settings of the original paper are followed, with a total of 4,138 questions.

[0086] (3) MathQA: MathQA is a comprehensive dataset that can be regarded as both a dataset of math application problems and a multi-step reasoning dataset. In the original MathQA dataset, a certain number of instances with annotated equations cannot obtain correct numerical answers. In this embodiment, the experimental settings of Tan et al. are followed and processed to remove questions with incorrect answers, resulting in a total of 20,207 questions.

[0087] 2. Baseline models:

[0088] To evaluate the effectiveness of the TKD method, three types of advanced models with different numbers of parameters are selected as baseline models for comparison in this embodiment. The considered baseline models are as follows:

[0089] The first type is large language models with hundreds of billions of parameters. The best LLMs and the best open-source LLMs are selected for comparison in this type of model. In this embodiment, the closed-source GPT4, which is the best in multiple evaluations, and the open-source PaLM and its performance in the CoT method are selected for comparison. The number of parameters of PaLM is 540B.

[0090] The second type is large language models with billions of parameters. In this type of models with large-scale parameters, LLMs with relatively fewer parameters are selected. In this embodiment, Llama2-7b and Qwen-7b are selected, and the number of parameters is 7B.

[0091] The third type is small models with millions of parameters. This type of model is based on deep learning and PLMs methods before the emergence of LLMs, and improves the accuracy of tasks through the complex structure of neural networks and the representation enhancement of PLMs. In this embodiment, three methods with GTS as the decoder and LSTM, RoBERTa-base, and RoERTa-large as the encoders are selected, and the number of parameters is 20M, 144M, and 377M respectively.

[0092] 3. Experimental settings:

[0093] In this work, all experiments are carried out using NVIDIA RTX A6000 graphics cards in this embodiment and implemented in Python using the PyTorch framework. For small models with millions of parameters, in this embodiment, the AdamW optimizer is designed to be used, with an epoch of 50, a batch size of 16, an initial learning rate of 8e-6, and halved every 20 epochs. The dropout during training is 0.01. The evaluation and targeted generation methods will be carried out at the 10th epoch and the 30th epoch. For large language models with billions of parameters, in this embodiment, r of the LoRA parameter is designed to be 16, lora_alpha is 32, dropout is 0.1, the initial learning rate is 1e-4, and the epoch is 3. The evaluation and targeted generation methods will be carried out at the 2nd epoch. The LLMs data generator and filter call the GPT-4 API. The evaluation metric used in this embodiment is the answer accuracy.

[0094] Experimental analysis:

[0095] Table 2 shows the experimental results of the TKD method proposed in this embodiment on the MAWPS, SVAMP, and MathQA datasets. Among them, Llama2-7b and Qwen2.5-7b are the results obtained by this embodiment using open-source code under the same fine-tuning parameter settings, and GPT-4 on the MathQA dataset is the result obtained by this embodiment using API testing. The results of other methods are those reported in the original text. The following analysis is carried out based on the results in the table:

[0096] (1) Compared with small models with millions of parameters, the TKD method proposed in this embodiment has achieved better results on all datasets. The "LSTM+GTS+TKD" proposed in this embodiment has improved by 8.1% (90.7% vs. 82.6%), 11.6% (42.4% vs. 30.8%), and 7.6% (78.9% vs. 71.3%) on the three datasets respectively compared with "LSTM+GTS"; the "RoBERTa-base+GTS+TKD" proposed in this embodiment has improved by 4.1% (92.6% vs. 88.5%), 20.5% (61.5% vs. 41.0%), and 6.0% (80.1% vs. 74.1%) on the three datasets respectively compared with "RoBERTa-base+GTS";

[0097] The "RoBERTa-large+GTS+TKD" proposed in this embodiment has improved by 2.9% (93.3% vs. 90.4%), 21.3% (70.8% vs. 49.5%), and 4.8% (81.7% vs. 76.9%) on the three datasets respectively compared with "RoBERTa-large+GTS". This demonstrates the effectiveness of the method proposed in this embodiment.

[0098] (2) Compared with the large language models Llama2-7b and Qwen2.5-7b with billions of parameters, the TKD method proposed in this embodiment has achieved better results on all datasets. The "Llama2-7b+TKD" proposed in this embodiment has improved by 11.3% (70.1% vs. 58.8%), 4.7% (51.1% vs. 46.4%), and 5.3% (27.9% vs. 22.6%) on the three datasets respectively compared with "Llama2-7b"; the "Qwen2.5-7b+TKD" proposed in this embodiment has improved by 2.2% (95.4% vs. 93.2%), 2.6% (76.6% vs. 74.0%), and 3.9% (61.6% vs. 57.7%) on the three datasets respectively compared with "Qwen2.5-7b". This further proves the effectiveness of the method proposed in this embodiment.

[0099] (3) Compared with large language models GPT-4 and PaLM with hundreds of billions of parameters, the TKD method proposed in this embodiment achieves results that are either better than or close to those of the comparison models with fewer parameters on all datasets. On the MAWPS dataset, "Qwen2.5-7b+TKD" proposed in this embodiment, with 7B parameters, is 2.1% higher (95.4% vs. 93.2%) than PaLM+CoT with 540B parameters, approaching the results of GPT-4; on the SVAMP dataset, "RoBERTa-large+GTS+TKD" and "Qwen2.5-7b+TKD" proposed in this embodiment, with 377M and 7B parameters respectively, are 1.4% (70.8% vs. 69.4%) and 7.2% (76.6% vs. 69.4%) higher than PaLM with 540B parameters respectively; on the MathQA dataset, "RoBERTa-large+GTS+TKD" proposed in this embodiment, with 377M parameters, is 0.8% higher (81.7% vs. 80.9%) than GPT-4. This further demonstrates the effectiveness of the method proposed in this embodiment.

[0100] The experimental results on the datasets MAWPS, SVAMP, and MathQA are shown in Table 3. From the results in Table 3 and the above analysis, it can be shown that the TKD method generates diverse data based on LLMs, and the diverse data improves the solving performance of the student model; at the same time, the evaluation and targeted generation modules of the TKD method not only feedback the learning state of the solver but also further generate data targeted, thus obtaining a more powerful student model.

[0101] Table 3

[0102]

[0103] Ablation experiment:

[0104] The model structure of this embodiment mainly includes two modules: a diverse data generation module based on LLMs and a module for evaluating the learning state of the student model and generating data targeted. To verify the effectiveness of each module, this embodiment designs two types of ablation experiments:

[0105] (1) To study the effectiveness of the data generation module: A represents removing the diverse data generation module based on LLMs.

[0106] (2) To study the effectiveness of the evaluation and targeted generation module: B represents removing the evaluation and targeted generation method.

[0107] The experimental results of the ablation experiment are as Figure 2 、 Figure 3 、 Figure 4 shown. The following inferences can be drawn from the results in the figure:

[0108] (1) TKD is superior to A, from which it can be inferred that the diversity data generation based on LLMs can effectively generate comprehensive and diverse data, improving the problem-solving ability of the student model;

[0109] (2) TKD is superior to B, from which it can be inferred that evaluating the student model during training can feedback the learning state of the student model, thereby further generating data targeted, and thus improving the problem-solving ability of the student model.

[0110] Case study:

[0111] In addition, this embodiment further selects several examples to illustrate the superiority of the TKD method. The problem generation example is shown in Table 4. For the selected source problem, the diversity data generation module based on LLMs in the TKD method generates data, obtaining two types of data.

[0112] (1) Compared with the source problem, the generated diverse sub-problem types show that LLMs have successfully generated questions and answers with the same semantic description but different questions;

[0113] (2) Compared with the source problem, the generated diverse analogy problem types show that LLMs have successfully generated questions and answers with the same question but different descriptions.

[0114] Table 4

[0115]

[0116]

[0117] In order to obtain high-quality data, this embodiment further filters the generated questions and answers.

[0118] The data filtering example is shown in Table 5. The answer format obtained by the generation module is incomplete and incorrect, and the answer format after passing through the data filter is complete and correct.

[0119] Table 5

[0120]

[0121] In this embodiment, aiming at the task of solving mathematical application problems, the previous knowledge distillation framework is limited to the model structure, unable to effectively improve the performance of the student model, and lacking learning feedback for the student model. A method for solving mathematical application problems based on targeted knowledge distillation of large language models is proposed. The TKD method proposed in this embodiment performs data augmentation in the diversity data generation module based on LLMs. The generated diverse data is effectively filtered, which can further improve the solving performance of the student model. At the same time, in order to better feedback the learning state of the student model, diverse evaluation data sets are introduced in the state evaluation and targeted generation module to evaluate the learning state of the student model, and then targeted data is generated to supplement the deficiencies in the student model's learning. The experimental results on the MAWPS, SVAMP, and MathQA data sets demonstrate the effectiveness of the method proposed in this embodiment.

[0122] The embodiments described above are only descriptions of the preferred embodiments of the present invention, and do not limit the scope of the present invention. Without departing from the design spirit of the present invention, various deformations and improvements made by those of ordinary skill in the art to the technical solutions of the present invention shall fall within the protection scope determined by the claims of the present invention.

Claims

1. A method for automatically solving mathematical word problems based on large model knowledge distillation, characterized in that: include: Get math word problems to be solved; The mathematical word problems to be solved are input into a knowledge distillation model based on a large model to generate mathematical expressions as a sequence of equation solutions, wherein the knowledge distillation model based on a large model includes a teacher model and a student model, the teacher model is constructed based on a large language model LLMs, and the student model is obtained by training with data generated by the teacher model.

2. The method for automatically solving mathematical word problems based on large model knowledge distillation according to claim 1 is characterized in that: The teacher model includes a diversity sub-question generator, a diversity analogy question generator and a data filter; The diversity sub-question generator is used to input the initial source data set to generate multiple questions with the same semantic description and sub-equations and corresponding answers; The diversity analogy problem generator is used to generate mathematical word problems with the same problem and different semantic descriptions by inputting the initial source data set; The data filter is used to screen and filter the questions generated by the diversity sub-question generator and the diversity analogy question generator to obtain a training data set.

3. The method for automatically solving mathematical word problems based on large model knowledge distillation according to claim 2 is characterized in that: Screening and filtering the questions generated by the diversity sub-question generator and the diversity analogy question generator include: filtering the questions with incorrect format, and using the LLMs-based zero-sample prompt technology to complete the incomplete questions, eliminating the questions that cannot be answered and have incorrect format, and obtaining the training data set.

4. The method for automatically solving mathematical word problems based on large model knowledge distillation according to claim 2 is characterized in that: Training the student model using the data generated by the teacher model includes: S1. After the student model is trained for a preset number of times using the training data set, the learning state of the student model after the preset number of training is evaluated using the evaluation data set to obtain unsolvable problems, wherein the evaluation data set is selected from the data generated by the teacher model; S2. Calling the diversity sub-problem generator and the diversity analogy problem generator in the teacher model to generate targeted data for the unsolvable problem to obtain an enhanced data set; S3. Merge the enhanced data set into the training data set, and return to S1 until a preset stop condition is reached.

5. The method for automatically solving mathematical word problems based on large model knowledge distillation according to claim 1 is characterized in that: The student model includes a model based on an encoder-decoder architecture and a model based on a decoder architecture.

6. The method for automatically solving mathematical word problems based on large model knowledge distillation according to claim 5 is characterized in that: The model based on the encoder-decoder architecture includes: using GTS as a decoder, the backbone network LSTM, RoBERTa-base and RoBERTa-large as encoders, wherein the encoder is used to model the input sequence and return the hidden state of the input mathematical problem, and the decoder is used to generate the distribution of the solution equation according to the output features of the encoder.

7. The method for automatically solving mathematical word problems based on large model knowledge distillation according to claim 5 is characterized in that: The model based on the decoder architecture is a large language model with a decoder architecture, wherein the large language model with the decoder architecture is fine-tuned using a re-parameterization method during training.

Citation Information

Patent Citations

  • Autonomous evolution method and system for online education intelligent teaching assistance

    CN114896975A

  • Arithmetic character question automatic answering method and system based on variational knowledge distillation

    CN117521812A

  • Anomaly and fraud detection with fake event detection using machine learning

    EP3761227A1

  • Automatic compression method and platform for multilevel knowledge distillation-based pre-trained language model

    WO2022126797A1

Cited By

  • Process reward model training method and system

    CN120430424A