Data processing method for improving advanced arithmetic capability of small-sized large language model

By constructing multiple data sets and adopting three-stage fine-tuning strategies, the advanced arithmetic computing capabilities of small and large language models are improved, and the shortcomings of existing technology small and medium-sized LLMs in advanced arithmetic computing are solved, and the comprehensive performance in a multi-task environment is achieved.

CN119940466AActive Publication Date: 2025-05-06JINAN UNIVERSITY
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510015889.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-06
Publication Date
2025-05-06
Estimated Expiration
2045-01-06

AI Technical Summary

Technical Problem

The prior art is difficult to significantly improve its advanced arithmetic computing capabilities while maintaining the original natural language processing capabilities of small large language models (LLMs), especially in complex arithmetic operations and multitasking environments.

Method used

By constructing instruction datasets, including arithmetic expression datasets, natural language processing datasets, and mathematical application datasets, adopting three-stage supervised fine-tuning strategies and the introduction of special markers, small LLMs are fine-tuned to enhance their advanced arithmetic skills.

Benefits of technology

It significantly improves the advanced arithmetic computing power of small LLMs, allowing them to handle a variety of complex arithmetic operations including logarithmic, trigonometric functions and composite functions, while maintaining their original natural language understanding and generation capabilities, avoiding catastrophic forgetting problems.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119940466A_ABST
    Figure CN119940466A_ABST
Patent Text Reader

Abstract

The invention discloses a data processing method for improving the advanced arithmetic capability of a small-sized large language model. The method comprises the following steps: constructing an instruction data set and carrying out data standardization processing; performing fine tuning on the small LLM by using different quantities of equation expression data in the instruction data set, evaluating the fine-tuned small LLM on the arithmetic test set, and obtaining the data quantity of the equation expression used by the model with the highest score; performing fine tuning on the small LLM by adopting the equation expression data of the equation expression data volume used by the model with the highest score and different quantities of natural language processing data to obtain the natural language processing data volume used by the model with the highest average score; and performing fine adjustment on the LLM by adopting the equation expression data of the equation expression data volume used by the model with the highest score, the natural language processing data of the natural language processing data volume used by the model with the highest average score and different quantities of mathematical application data, and obtaining the mathematical application data volume used by the model with the highest average score.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of computer technology, and in particular to a data processing method for improving the high-level arithmetic capabilities of a small large language model. Background Art

[0002] Large language models (LLMs) have demonstrated remarkable capabilities in understanding and generating language and are widely used in finance, education, psychology, healthcare, and other fields. These models are trained on large-scale, diverse text data to generate high-quality natural language text and perform well in a variety of tasks. However, despite the remarkable progress of LLMs in many fields, there are still significant challenges in the performance of advanced arithmetic calculations, especially for small LLMs.

[0003] Advanced arithmetic calculations involve not only the basic four operations, but also logarithms, exponentiations, trigonometric functions, and mixed operations of multiple advanced operations. These complex computing tasks place higher demands on LLMs because they require the model to have stronger mathematical reasoning capabilities and precise calculation capabilities. Currently, a variety of methods have been developed and successfully applied to improve the mathematical reasoning capabilities of LLMs, but these methods still have many limitations in practical applications.

[0004] First, decomposing the problem into natural language steps (Chain-of-Thought, CoT) or generating code / program integrated reasoning steps (Program-of-Thought, PoT) is a common solution. These methods help the model gradually solve complex mathematical problems by generating intermediate reasoning steps. However, these methods usually require large-scale LLMs to generate and process these intermediate steps. Although large-scale LLMs excel in reasoning capabilities, their high computational cost and resource requirements make them difficult to economically deploy and apply in actual production environments.

[0005] Secondly, it is also a common method to improve the computational capabilities of LLMs by calling calculator tools. For example, Toolformer enhances the ability of LLMs in processing basic arithmetic operations (such as addition, subtraction, multiplication and division) by integrating external calculator tools. However, these methods rely on few-shot learning or prompting, that is, the model needs to receive explicit prompts to call the calculator tool every time it calculates. This approach not only increases the complexity of use, but also limits the autonomy and flexibility of the model. In addition, these methods can usually only handle basic arithmetic operations and cannot accurately calculate advanced arithmetic operations.

[0006] In addition, while existing methods improve mathematical reasoning capabilities, they often cause catastrophic forgetting of the model's original capabilities (such as common sense reasoning and reading comprehension). That is, when the model focuses on learning mathematical reasoning tasks, it may lose its performance in other tasks. This catastrophic forgetting problem makes it difficult for the model to take into account the needs of multiple tasks in practical applications, limiting its potential for widespread application.

[0007] In summary, although existing solutions have made some progress in improving the mathematical reasoning ability of LLMs, they still face many challenges in practical applications. Especially for small LLMs, how to significantly improve their advanced arithmetic computing capabilities while maintaining their original natural language processing capabilities is still an urgent problem to be solved. Therefore, developing a data processing technology that can enable small LLMs to have advanced arithmetic skills has important practical significance and application value. Summary of the invention

[0008] To solve the above technical problems, the present invention proposes a data processing method for improving the high-level arithmetic capabilities of small large language models, thereby enhancing the high-level arithmetic skills of LLMs.

[0009] To achieve the above object, the present invention provides a data processing method for improving the high-level arithmetic capability of a small large language model, comprising:

[0010] Constructing an instruction data set, and performing standardization processing on the instruction data set to obtain standardized data; the instruction data set includes an arithmetic expression data set, a natural language processing data set, and a mathematical application data set;

[0011] The normalized arithmetic expression data is input into a small LLM fine-tuned using different amounts of arithmetic expression data, and the fine-tuned small LLM is evaluated on an arithmetic test set to obtain the amount of arithmetic expression data used by the model with the highest score as the first stage of fine-tuning;

[0012] The small LLM is fine-tuned using the arithmetic expression data of the model with the highest score and the natural language processing data after normalization of different quantities. The fine-tuned small LLM is evaluated on the joint test set of arithmetic, MWP and NLP tasks, and the amount of natural language processing data used by the model with the highest average score in arithmetic, natural language processing and mathematical word problem data sets is obtained as the second stage of fine-tuning;

[0013] The small LLM is fine-tuned by using the formula expression data of the formula expression data used by the model with the highest score, the natural language processing data of the natural language processing data used by the model with the highest average score, and different amounts of the mathematical application data after standardization. The fine-tuned small LLM is evaluated on the joint test set of arithmetic, MWP and NLP tasks, and the amount of mathematical application data used by the model with the highest average score is obtained as the third stage of fine-tuning.

[0014] Optionally, the arithmetic expression data set includes 4 basic operators and 8 advanced arithmetic operators.

[0015] Optionally, the natural language processing dataset includes a multilingual dataset.

[0016] Optionally, the mathematics application dataset includes several MWP instances collected from various educational websites.

[0017] Optionally, a special mark for calling a calculator is added to the arithmetic expressions of the arithmetic expression data set and the mathematical application data set.

[0018] Optionally, the construction process of the addition operation in the basic operation expression in the first stage fine-tuning process includes:

[0019] Randomly generate the number of items;

[0020] Randomly generate a number of digits for each item;

[0021] Generate a random number of corresponding digits for each item number;

[0022] Combine each term using addition to complete the construction of the basic operation expression in the addition process.

[0023] Optionally, the construction process of the mixed operation in the basic operation expression in the first stage fine-tuning process includes:

[0024] Randomly generate the number of combinations;

[0025] Randomly generate arithmetic operator types for each combination type;

[0026] The subcombination expression is constructed by the corresponding arithmetic operator type constructor;

[0027] All sub-combinations are combined through randomly generated basic operators to complete the construction of basic operation expressions in the mixed operation process.

[0028] Optionally, the special mark is " <thought>"and" <api>”.

[0029] Technical effect of the invention: The present invention discloses a data processing method for improving the high-level arithmetic ability of a small large language model, which can significantly improve the high-level arithmetic computing ability of a small large language model (LLMs), so that it can handle a variety of complex arithmetic operations including logarithms, trigonometric functions and composite functions, while maintaining its original natural language understanding and generation capabilities, avoiding the problem of catastrophic forgetting. Through the introduction of a three-stage supervised fine-tuning strategy and special tags, the model can not only autonomously call the calculator API and seamlessly perform complex arithmetic calculations in natural language sentences, but also significantly improve its comprehensive performance in arithmetic tasks, mathematical word problem tasks and natural language processing tasks. The present invention has broad application prospects in multiple fields, especially in scenarios requiring precise calculations. By significantly improving the high-level arithmetic computing ability of small large language models (LLMs), the method can help students solve complex mathematical problems in the field of education, thereby improving learning effects and teaching quality. In scientific research, the method can be used to process and analyze experimental data, perform complex mathematical calculations, and significantly improve research efficiency and the reliability of results. Through these applications, the present invention demonstrates its strong adaptability and efficiency in a multi-task environment, proving its feasibility and great potential in practical applications. BRIEF DESCRIPTION OF THE DRAWINGS

[0030] The drawings constituting a part of the present application are used to provide a further understanding of the present application. The illustrative embodiments and descriptions of the present application are used to explain the present application and do not constitute an improper limitation on the present application. In the drawings:

[0031] Figure 1 A schematic diagram of a data processing method for improving the high-level arithmetic capability of a small large language model according to an embodiment of the present invention;

[0032] Figure 2 A schematic diagram of a process of constructing a basic operation arithmetic expression according to an embodiment of the present invention;

[0033] Figure 3 The figure is a schematic diagram of the process of constructing a mixed operation arithmetic expression according to an embodiment of the present invention. DETAILED DESCRIPTION

[0034] It should be noted that, in the absence of conflict, the embodiments and features in the embodiments of the present application can be combined with each other. The present application will be described in detail below with reference to the accompanying drawings and in combination with the embodiments.

[0035] It should be noted that the steps shown in the flowcharts of the accompanying drawings can be executed in a computer system such as a set of computer executable instructions, and that, although a logical order is shown in the flowcharts, in some cases, the steps shown or described can be executed in an order different from that shown here.

[0036] like Figure 1 As shown, this embodiment provides a data processing method for improving the high-level arithmetic capability of a small large language model, including:

[0037] Constructing an instruction data set, and performing standardization processing on the instruction data set to obtain standardized data; the instruction data set includes an arithmetic expression data set, a natural language processing data set, and a mathematical application data set;

[0038] The normalized arithmetic expression data is input into a small LLM fine-tuned using different amounts of arithmetic expression data, and the fine-tuned small LLM is evaluated on an arithmetic test set to obtain the amount of arithmetic expression data used by the model with the highest score as the first stage of fine-tuning;

[0039] The small LLM is fine-tuned using the arithmetic expression data of the model with the highest score and the natural language processing data after normalization of different quantities. The fine-tuned small LLM is evaluated on the joint test set of arithmetic, MWP and NLP tasks, and the amount of natural language processing data used by the model with the highest average score in arithmetic, natural language processing and mathematical word problem data sets is obtained as the second stage of fine-tuning;

[0040] The small LLM is fine-tuned by using the formula expression data of the formula expression data used by the model with the highest score, the natural language processing data of the natural language processing data used by the model with the highest average score, and different amounts of the mathematical application data after standardization. The fine-tuned small LLM is evaluated on the joint test set of arithmetic, MWP and NLP tasks, and the amount of mathematical application data used by the model with the highest average score is obtained as the third stage of fine-tuning.

[0041] Furthermore, the arithmetic expression data set includes 4 basic operators and 8 advanced arithmetic operators.

[0042] Specifically, existing research works mainly focus on basic arithmetic operations of low-digit numbers (such as addition and subtraction), and these simple forms and structures cannot fully reflect the computational capabilities of LLMs. Therefore, the present invention systematically reviews the types of arithmetic operators and arithmetic tasks, and creates a comprehensive arithmetic expression dataset. The dataset contains 4 basic operators (such as addition and multiplication) and 8 advanced arithmetic operators (including logarithms, trigonometric functions, and composite functions). All of the above arithmetic operators are able to process low-digit numbers, high-digit numbers, or decimals.

[0043] In addition, the dataset of the present invention also supports mathematical operation instructions for complex numbers. In actual dialogue scenarios, arithmetic operations may be embedded in natural language, such as "the square root of 9" or "the remainder of 412 divided by 57". In order to improve the generalization and robustness of LLMs for arithmetic operators, the present invention generates arithmetic alias expressions as an important supplement to the AI-1 dataset. These alias expressions can help LLMs effectively solve and process problems containing different descriptions or terms, and improve the ability to interpret and perform mathematical operations in different language environments. The detailed information of the arithmetic expression dataset is shown in Table 1.

[0044] Table 1

[0045]

[0046] Furthermore, the construction process of the addition operation in the basic operation expression in the first stage fine-tuning process includes:

[0047] Randomly generate the number of items;

[0048] Randomly generate a number of digits for each item;

[0049] Generate a random number of corresponding digits for each item number;

[0050] Combine each term using addition to complete the construction of the basic operation expression in the addition process.

[0051] Specifically, for each arithmetic type, the present invention uses a constructor to generate an expression from low to high. For basic operations, such as constructing an addition operation, first randomly generate the number of terms to represent the number of terms in the expression, then randomly generate the number of digits for each term to represent the number of digits in the term, and finally combine each term using addition. The construction process of the basic operation expression is as follows: Figure 2 shown.

[0052] Furthermore, the construction process of the mixed operation in the basic operation expression in the first stage fine-tuning process includes:

[0053] Randomly generate the number of combinations;

[0054] Randomly generate arithmetic operator types for each combination type;

[0055] The subcombined expression is constructed by the corresponding arithmetic type constructor;

[0056] All sub-combinations are combined through randomly generated basic operators to complete the construction of basic operation expressions in the mixed operation process.

[0057] Specifically, for advanced operations, expressions are constructed according to the characteristics of the operations. For example, for log, the base and the number of digits of the real number are first determined, and then the base and the real number of the corresponding digits are randomly generated. For another example, for trigonometric functions, first determine whether the angle is expressed in radians (in π) or angle values ​​(in degrees), then randomly generate the number of terms in the trigonometric function, then randomly generate the number of digits in each term and the basic operators between each term, and finally combine all the generated contents.

[0058] For mixed operations, the number of combinations is randomly generated first, that is, the number of combination types. For each combination type, an arithmetic operator type is randomly generated. Then, the corresponding arithmetic type constructor is used to construct a subcombination expression. Finally, all subcombinations are combined using the randomly generated basic operators. The construction process is as follows: Figure 3 shown.

[0059] For the text embedding type, the present invention first constructs an expression without text embedding, and then replaces the mathematical arithmetic symbols in the expression with the text type arithmetic symbols. For example, first construct 5×3, and then use "multiply" or "times" to replace the "×" arithmetic symbol to obtain: 5 times 3 or 5 times 3.

[0060] Furthermore, the natural language processing dataset includes a multilingual dataset.

[0061] Specifically, in order to alleviate the catastrophic forgetting phenomenon that occurs when LLMs are fine-tuned using AI-1 and AI-2, it is necessary to combine dialogue and common sense knowledge, namely the natural language processing dataset (NLP), during supervised training. In the present invention, the well-known multilingual dataset Alpaca (dataset name) is selected as the NLP dataset because of its diversity and complexity. The Alpaca dataset covers various fields such as literature, science, and social sciences. 52,002 English data and 48,818 Chinese data were randomly selected from the original Alpaca dataset to form the AI-3 dataset. Examples are shown in Table 2.

[0062] Table 2

[0063]

[0064] Furthermore, the mathematics application dataset includes 100,000 MWP instances collected from various educational websites.

[0065] Specifically, in order to further enhance the problem-solving and reasoning abilities of LLMs in the field of complex mathematical tasks, a comprehensive mathematical word problem (MWP) dataset, namely AI-2, was constructed. The dataset consists of 100,000 MWP instances collected from various educational websites, each of which is accompanied by a detailed analysis of the solution process. The AI-2 dataset covers the difficulty levels of elementary school, junior high school, and high school, thus ensuring that LLMs' abilities in dealing with mathematical problems of different complexity are comprehensively improved. Examples of AI-2 are shown in Table 3.

[0066] Table 3

[0067]

[0068]

[0069] Furthermore, a special mark for calling the calculator needs to be added during fine-tuning in the first stage and the third stage, and the special mark is " <thought>"and" <api>Specifically, in order to enable small LLMs to autonomously call calculators without explicit instructions and to seamlessly perform any complex arithmetic calculations in natural language sentences, the present invention uses special tags in the AI-1 and AI-2 data sets involved in the calculations. <thought>"and" <api>" is used to encapsulate arithmetic expressions, thereby achieving accurate recognition of mathematical expressions in natural language.

[0070] Since all data entries in AI-1 are simulated, special tags designed are automatically encapsulated in each arithmetic expression. For each MWP in AI-2, it is hoped that special tags will be applied to the corresponding part of its problem analysis. Through few-shot learning and prompting, GPT4 is used to generate instructions containing special tags for calling the computational API. The prompt words are shown below.

[0071] I will give you a math problem and the corresponding solution process. You need to encapsulate the corresponding arithmetic expression in the solution process with a calculator token. The output format of the encapsulated expression is:

[0072] {expression} <thought>This involves calculations, and I need to call the calculator API here <api> [{\"ActionName\":\"Calculator\",\"Args\":{\"equation\":{expression}}}]< / api> =>{result}

[0073] Here is an example:

[0074] Original answer:

[0075] Cost price: $15\\div(1+50\\%)$$=15\\div 1.5$$=10$(yuan) Profit: $15\\times0.8-10$$=12-10$$=2$(yuan) Answer: Selling at 20% off the price can make a profit of $2$ yuan. So the answer is: $2$.

[0076] Answer package:

[0077] In this question, the cost price of each book is $15$, and the profit after sale is $50\\%$. Therefore, the cost price of each book can be calculated as $15\\div(1+50\\%)=$ <thought>I need to calculate division here, so I need to call the calculator API here. <api> [{\"ActionName\":\"Calculator\",\"Args\":{\"equation\":\"15 / 1.5"}}]< / api> =>10< / thought> $10$ yuan. Then, we can calculate the profit of each book as $15\\times0.8-10=$ <thought>Here I need to calculate an algebraic expression, I can call the Calculator API here <api> [{\"ActionName\":\"Calculator\",\"Args\":{\"equation\":\"15*0.8-10\"}}]< / api> =>2< / thought> $2$. Therefore, if the price of each book is set at $80\\%$ of the original price, a profit of $2$ can be obtained. Therefore, the answer is $2$.

[0078] Please encapsulate the answers to the following questions based on the above requirements and examples:

[0079] {Questions and Answers}

[0080] These encapsulated instructions from AI-1 and AI-2 not only improve LLMs’ understanding of the problem, but also teach them how to autonomously call precise calculations when faced with similar problems. Examples of special mark encapsulation are shown in Table 4. In addition, to ensure the correctness of the constructed data of AI-1 and AI-2, the encapsulated expressions are calculated using a calculator to filter out erroneous data that cannot be calculated. The processing of these data sets is designed to retain the inherent capabilities of small LLMs while enhancing their ability to handle complex arithmetic and large numbers.

[0081] Table 4

[0082]

[0083]

[0084] The constructed AI-1, AI-2 and AI-3 belong to three different types of data, namely, arithmetic expression datasets, mathematical word problems (MWP) datasets and natural language processing (NLP) datasets. In order to ensure that these datasets can be efficiently processed and utilized by small LLMs, all data are formatted and the processed data is used to fine-tune the model. Standardization not only helps to improve the readability and consistency of the data, but also significantly improves the efficiency and accuracy of the model during training and reasoning. Processed AI-1 training data. The training examples of AI-1, AI-2 and AI-3 are shown in Table 5 below.

[0085] Table 5

[0086]

[0087]

[0088] The construction and standardization of the arithmetic instruction dataset ArithInstruct provides a solid foundation for model training. This paper proposes a three-stage supervised fine-tuning strategy, which can not only give small LLMs arithmetic capabilities but also maintain their original natural language understanding capabilities through reasonable data proportion and order.

[0089] Inspired by curriculum learning theory, the supervised fine-tuning strategy of this application attempts to mimic the process of human mathematical learning, that is, first learning to recognize mathematical symbols and numerical expressions, and then learning to understand logical reasoning steps. In addition, in this learning process, in order to avoid the classic catastrophic forgetting problem in LLMs, a natural language understanding fine-tuning task is incorporated into the mathematical learning process. Specifically, this application first uses AI-1 data to fine-tune small LLMs to enable them to recognize various mathematical expressions. In the second stage, it combines AI-3 data with AI-1 data to help the fine-tuned model maintain its natural language ability. Finally, it uses all available data (AI-1, AI-2, and AI-3) to enhance the reasoning ability of small LLMs through detailed annotated step-level MWP data.

[0090] In addition to the order of the datasets used for supervised fine-tuning in the present invention, it is also very important to determine the appropriate amount of training data. The view that more data is better is not always correct. The quality and relevance of the training examples are often more important than the quantity. When performing supervised fine-tuning, the training samples must not only provide the correct solution, but also teach the model how to get the answer. Since different LLMs are trained from different mixtures of data, it is difficult to come up with a general method to determine the appropriate amount of data to improve the high-level arithmetic and reasoning capabilities of small LLMs.

[0091] In the present invention, the above three-stage model training is performed by a greedy search heuristic method. Specifically, small LLMs are first fine-tuned using different amounts of AI-1 data, such as 1,000, 5,000, 10,000, 20,000, 50,000, and 100,000. Only the model that achieves the highest accuracy on the arithmetic validation set is retained, and the corresponding arithmetic training amount is recorded as A. * , for example 50,000.

[0092] Then, in the second round of fine-tuning, the present invention keeps the size of the arithmetic training data as A * , and add different amounts of NLP training data from AI-3. We evaluate different fine-tuned LLMs on a joint validation set of arithmetic, MWP, and NLP tasks. Since natural language training data can enhance LLMs’ understanding of MWP, we select the model that achieves the highest average score on the joint validation dataset and record the corresponding amount of natural language training data as L * .

[0093] In the third stage of fine-tuning, the present invention fixes the size of arithmetic and natural language training data to A * and L * , and change different amounts of MWP training data, i.e. AI-2. Finally, the present invention selects the fine-tuning model that obtains the highest average score on the joint evaluation dataset, and records the corresponding amount of MWP training data as M.

[0095] Finally, the best model is A * AI-1 data of quantity, L * Quantity of AI-3 data and M * The AI-2 data is fine-tuned, and for Baichuan2-13B-chat, the optimal training data volume A for AI-1, AI-2 and AI-3 is finally obtained. * 、M * and L * They are 50,000, 10,000 and 5,000 respectively.

[0096] Through this three-stage fine-tuning strategy based on curriculum learning, the present invention is able to gradually enhance the high-level arithmetic skills of small LLMs while maintaining their original natural language processing capabilities and avoiding the catastrophic forgetting problem, thereby achieving the comprehensive performance of the model in multiple tasks.

[0097] The performance of the Baichuan2-13B-chat fine-tuned by the three-stage fine-tuning strategy proposed in this invention, hereinafter referred to as Arithmetic-Baichuan2-13B, and the original Baichuan2-13B-chat in two arithmetic test sets: Arith3K and math401, four MWPs (math problem solving): MAWPS, ASDiv-A, GSM8K, SVAMP, and four NLP (natural language processing) datasets: PIQA, BoolQ, HellaSwag, and RACE are compared. The comparison results are shown in Table 6 below:

[0098] Table 6

[0099]

[0100] From the above results, we can see that after the three-stage fine-tuning strategy proposed in this invention, the model can significantly improve its arithmetic ability and mathematical reasoning ability, while maintaining its original natural language processing ability without catastrophic forgetting. The following is an output example of Arithmetic-Baichuan2-13B and Baichuan2-13B-chat on the arithmetic test set Arith3K:

[0101] Input Example 1:

[0102] <reserved_106> Calculate the following arithmetic problem:-581682167-119927\nAnswer:<reserved_107>

[0103] Correct answer:

[0104] -581802094

[0105] Output of Baichuan2-13B-chat:

[0106] The answer to the given arithmetic problem is:-581704175

[0107] Output of Arithmetic-Baichuan2-13B:

[0108] The result of -581682167-119927 is <thought>Involving calculations, I need to call the Calculator API here <api> [{\"ActionName\":\"Calculator\",\"Args\":{\"equation\":\"-581682167-119927\"}}]< / api> >-581802094< / thought> -581802094

[0109] In the above example, the original Baichuan2-13B-chat cannot correctly answer the mathematical arithmetic problem, while Arithmetic-Baichuan2-13B can correctly answer the mathematical arithmetic problem and get the correct answer even when faced with more complex compound operations, as shown below.

[0110] Input Example 2:

[0111] <reserved_106> Calculate the following math problem: sin(8.51π)^(1 / 9)+sqrt(sqrt(358))\nAnswer:<reserved_107>

[0112] Correct answer: 5.3498

[0114] Output of Arithmetic-Baichuan2-13B:

[0115] Involving calculations, I need to call the Calculator API here <api> [{\"ActionName\":\"Calculator\",\"Args\":{\"equation\":\"sin(8.51*pi)^(1 / 9)+sqrt(sqrt(358))\"}}]< / api> >5.3498< / thought> 5.3498

[0116] The present invention discloses a method for improving the high-level arithmetic ability of a small large language model, which can significantly improve the high-level arithmetic computing ability of a small large language model (LLMs), so that it can handle a variety of complex arithmetic operations including logarithms, trigonometric functions and composite functions, while maintaining its original natural language understanding and generation capabilities, avoiding the problem of catastrophic forgetting. Through the introduction of a three-stage supervised fine-tuning strategy and special tags, the model can not only autonomously call the calculator API and seamlessly perform complex arithmetic calculations in natural language sentences, but also significantly improve its comprehensive performance in arithmetic tasks, mathematical word problem tasks and natural language processing tasks. The present invention has broad application prospects in multiple fields, especially in scenarios requiring precise calculations. By significantly improving the high-level arithmetic computing ability of small large language models (LLMs), the method can help students solve complex mathematical problems in the field of education, thereby improving learning effects and teaching quality. In scientific research, the method can be used to process and analyze experimental data, perform complex mathematical calculations, and significantly improve research efficiency and the reliability of results. Through these applications, the present invention demonstrates its strong adaptability and efficiency in a multi-task environment, proving its feasibility and great potential in practical applications.

[0117] The above are only preferred specific implementations of the present application, but the protection scope of the present application is not limited thereto. Any changes or substitutions that can be easily thought of by a person skilled in the art within the technical scope disclosed in the present application should be included in the protection scope of the present application. Therefore, the protection scope of the present application should be based on the protection scope of the claims.< / api> < / thought> < / api> < / thought> < / api> < / thought>

Claims

1. A data processing method for improving the high-level arithmetic capabilities of a small large language model, characterized in that: include: Constructing an instruction data set, and performing standardization processing on the instruction data set to obtain standardized data; The instruction data set includes an arithmetic expression data set, a natural language processing data set and a mathematical application data set; The normalized arithmetic expression data is input into a small LLM fine-tuned using different amounts of arithmetic expression data, and the fine-tuned small LLM is evaluated on an arithmetic test set to obtain the amount of arithmetic expression data used by the model with the highest score as the first stage of fine-tuning; The small LLM is fine-tuned using the arithmetic expression data of the model with the highest score and the natural language processing data after normalization of different quantities. The fine-tuned small LLM is evaluated on the joint test set of arithmetic, MWP and NLP tasks, and the amount of natural language processing data used by the model with the highest average score in arithmetic, natural language processing and mathematical word problem data sets is obtained as the second stage of fine-tuning; The small LLM is fine-tuned by using the formula expression data of the formula expression data used by the model with the highest score, the natural language processing data of the natural language processing data used by the model with the highest average score, and different amounts of the mathematical application data after standardization. The fine-tuned small LLM is evaluated on the joint test set of arithmetic, MWP and NLP tasks, and the amount of mathematical application data used by the model with the highest average score is obtained as the third stage of fine-tuning.

2. The data processing method for improving the high-level arithmetic capability of a small large language model as claimed in claim 1, characterized in that: The arithmetic expression data set includes 4 basic operators and 8 advanced arithmetic operators.

3. The data processing method for improving the high-level arithmetic capability of a small-scale large language model as claimed in claim 1, characterized in that: The natural language processing dataset includes a multilingual dataset.

4. The data processing method for improving the high-level arithmetic capability of a small-scale large language model according to claim 1, characterized in that: The Mathematical Applications Dataset includes several MWP examples collected from various educational websites.

5. The data processing method for improving the high-level arithmetic capability of a small-scale large language model as claimed in claim 1, characterized in that: A special mark for calling a calculator is added to the arithmetic expressions of the arithmetic expression data set and the mathematical application data set.

6. The data processing method for improving the high-level arithmetic capability of a small-scale large language model according to claim 1, characterized in that: The construction process of the addition operation in the basic operation expression during the first stage fine-tuning process includes: Randomly generate the number of items; Randomly generate a number of digits for each item; Generate a random number of corresponding digits for each item number; Combine each term using addition to complete the construction of the basic operation expression in the addition process.

7. The data processing method for improving the high-level arithmetic capability of a small-scale large language model according to claim 1, characterized in that: The construction process of mixed operations in the basic operation expressions in the first stage fine-tuning process includes: Randomly generate the number of combinations; Randomly generate arithmetic operator types for each combination type; The subcombination expression is constructed by the corresponding arithmetic operator type constructor; All sub-combinations are combined through randomly generated basic operators to complete the construction of basic operation expressions in the mixed operation process.

8. The data processing method for improving the high-level arithmetic capability of a small-scale large language model as claimed in claim 5, characterized in that: The special mark is " <thought>"and" <api> ”。< / api> < / thought>

Citation Information

Patent Citations

  • Mathematical big language model fine tuning method, system and equipment with cooperation of data enhancement method and prediction enhancement method and medium

    CN118014056A

  • Large language model processing method, system and device, medium and program product

    CN118297107A

  • Generation method of commodity copywriting information

    CN118941349A

  • Large language model screening system and methods thereof

    US20240427986A1