Construction method of scientific language model for solving scientific problem of university level
By constructing instruction data sets suitable for scientific problems at the university level and supervising fine-tuning the basic models, the insufficient performance of existing language models on university scientific problems has been solved, and significant performance improvement has been achieved, and it is suitable for multiple scientific question-solving tasks.
Patent Information
- Application Number
- CN202311795840.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-25
- Publication Date
- 2025-07-04
AI Technical Summary
Existing language models perform poorly in solving university-level scientific problems, mainly due to the scarce training set and limited scientific models, resulting in low accuracy on university mathematics and science tasks.
Build instruction data sets for scientific questions at the university level, and optimize the basic model through supervised fine-tuning, including data set collection, filtering and optimization, use ChatGLM3-6B-Base as the basic model, and use GPT-3.5-Turbo and GPT-4 to generate the correct answer process for self-correction.
The performance of scientific language models on university-level mathematical and scientific tasks has been significantly improved, and experimental results show that they show strong solution capabilities on multiple evaluation datasets.
Smart Images

Figure CN120258126A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of large models, and relates to a method for constructing a scientific language model, in particular to a method for constructing a scientific language model for solving scientific problems at the university level. Background Art
[0002] Nowadays, after being pre-trained on a large number of primary school-level mathematics data sets, most language models have certain mathematical reasoning abilities and achieve good performance on common mathematical evaluation sets (such as GSM8K and MATH). However, for scientific problems at the university level, the current large language models have the following two limitations: 1. Scarce training sets. 2. Limited trained scientific models.
[0003] Moreover, existing research mainly focuses on the ability of large language models to solve mathematical problems or proposes evaluation benchmarks for scientific problems. According to the evaluation results of existing research, ChatGPT or CPT-4 cannot solve scientific problems at the university level well either. For example, the accuracy rate in university mathematics is only about 25%.
[0004] Therefore, in view of the above defects existing in the prior art, it is necessary to develop a new scientific language model for solving scientific problems at the university level. Summary of the Invention
[0005] In order to overcome the defects of the prior art, the present invention proposes a method for constructing a scientific language model for solving scientific problems at the university level, which can improve the ability of the scientific language model to solve scientific problems at the university level and is significantly superior to the basic model in multiple mathematical tasks and scientific tasks.
[0006] To achieve the above object, the present invention provides the following technical solutions:
[0007] A method for constructing a scientific language model for solving scientific problems at the university level, characterized by comprising the following steps:
[0008] 1) Select a basic model;
[0009] 2) Construct an instruction data set suitable for scientific problems at the university level;
[0010] 3) Perform supervised fine-tuning on the basic model with the instruction data set suitable for scientific problems at the university level to obtain the scientific language model.
[0011] Preferably, in step 1), ChatGLM3-6B-Base is selected as the basic model.
[0012] Preferably, step 2) specifically includes:
[0013] 2.1), Instruction dataset collection;
[0014] 2.2), Instruction dataset filtering;
[0015] 2.3), Instruction dataset optimization.
[0016] Preferably, the instruction dataset collection in step 2.1) is specifically as follows: First, collect a large number of college-level scientific questions and answers from multiple sources including college textbooks, college exercise sets, and college lecture notes; then, extract useful text from the collected college-level scientific questions and answers through optical character recognition technology; finally, supplement the thought chain process in LaTeX format for the scientific questions and answers in the extracted useful text.
[0017] Preferably, the instruction dataset filtering in step 2.2) is specifically as follows: Train a data classifier and filter out noisy data through the data classifier, where the positive samples of the data classifier include high-quality standard answer data answered by annotators, as well as correct answers generated by GPT-3.5-Turbo and GPT-4, and the negative samples include wrong answers generated by the base model; based on the data classifier, select the dataset with a higher probability ranking as the filtered dataset.
[0018] Preferably, the instruction dataset optimization in step 2.3) is specifically as follows: Use GPT-3.5-Turbo and GPT-4 to generate a complete step-by-step solution process for scientific questions. If data with wrong answers is encountered, adjust the prompt again and correct the mistakes by itself until a correct solution process is generated.
[0019] Compared with the prior art, the method for constructing a scientific language model for solving college-level scientific questions of the present invention has one or more of the following beneficial technical effects:
[0020] 1. The present invention obtains a high-quality instruction dataset through three stages, thereby being able to improve the performance of the scientific language model in scientific tasks.
[0021] 2. The present invention selects ChatGLM3-6B-Base as the base model, standardizes all instruction datasets into the format of a chatbot, and performs supervised fine-tuning on the base model, so that the fine-tuned scientific language model is significantly better than the base model in multiple mathematical tasks and scientific tasks.
[0022] 3. After supervised fine-tuning, the final experimental results of the present invention show that the constructed scientific language model outperforms the base model on the mathematics evaluation dataset and also exhibits strong performance on multiple scientific evaluation datasets. With these core capabilities, the scientific language model can be feasibly applied to different scientific question-answering tasks. BRIEF DESCRIPTION OF THE DRAWINGS
[0023] Figure 1 is a flowchart of a method for constructing a scientific language model for solving scientific problems at the university level according to the present invention.
[0024] Figure 2 is a flowchart of constructing an instruction dataset suitable for scientific problems at the university level in the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0025] Before detailing any embodiments of the present invention, it should be understood that the present invention is not limited in its application to the details of the construction and arrangement of components set forth in the following description or illustrated in the following drawings. The present invention is capable of other embodiments and of being practiced or carried out in various ways. Additionally, it should be understood that the language and terms used herein are for the purpose of description and should not be regarded as restrictive. As used herein, the terms "including" or "having" and their variants are intended to cover the listed items and their equivalents as well as additional items.
[0026] Also, in the disclosure of the present invention, the term "a" should be understood as "at least one" or "one or more". That is, in one embodiment, the number of an element can be one, while in other embodiments, the number of the element can be multiple. The term "a" should not be construed as a limitation on the number.
[0027] The present invention proposes a method for constructing a scientific language model for solving scientific problems at the university level. By constructing an instruction dataset for scientific problems at the university level and using it for supervised fine-tuning of the base model, the ability of the scientific language model to solve scientific problems at the university level can be improved, and it is significantly superior to the base model in multiple mathematical tasks and scientific tasks.
[0028] Figure 1 shows a flowchart of a method for constructing a scientific language model for solving scientific problems at the university level according to the present invention. As Figure 1 shown, the method for constructing a scientific language model for solving scientific problems at the university level according to the present invention includes the following steps:
[0029] I. Select a base model.
[0030] In the present invention, a variety of large language models can be selected as the base model. Preferably, ChatGLM3-6B-Base is selected as the base model.
[0031] II. Construct an instruction dataset suitable for college-level scientific questions.
[0032] Most of the existing instruction datasets (e.g., MathInstruction, AgentInstruction) have been collected and widely used for pre-training large language models on multiple tasks, such as mathematical reasoning tasks and agent tasks. However, collecting and constructing an instruction dataset suitable for college-level scientific questions is challenging, and these challenges are mainly reflected in three aspects:
[0033] 1. Richer question types. Compared with mathematical tasks, scientific tasks cover multiple disciplines, such as mathematics, physics, and chemistry, etc.
[0034] 2. More complex questions. Compared with the difficulty of questions in primary, junior high, and senior high schools, college-level scientific questions contain more complex and specialized knowledge points.
[0035] 3. Fewer training sets: Due to the challenges in the above two aspects, there are few publicly available training datasets for college-level scientific tasks.
[0036] Based on this, the present invention proposes to obtain a high-quality instruction dataset and improve the performance of the scientific language model on scientific tasks through three stages. Specifically, as Figure 2 shown, the instruction dataset construction suitable for college-level scientific questions of the present invention includes:
[0037] 1. Instruction dataset collection.
[0038] The goal of the present invention is to construct a broad scientific representation and complex scientific tasks. To ensure that the constructed scientific language model covers a diverse dataset, this instruction dataset should contain rich scientific knowledge. Based on this goal, the present invention narrows the domain scope and constructs a small number of high-quality datasets. This dataset includes multiple domains and complex tasks, such as mathematics, physics, and chemistry. To collect these instruction datasets, a large number of scientific questions and answers were first collected from multiple sources, such as college textbooks, problem sets, lecture notes, etc. Then, useful text was extracted through optical character recognition technology. Finally, the thought chain process in LaTeX format was supplemented for scientific questions and answers.
[0039] 2. Instruction dataset filtering.
[0040] Except for a small amount of high-quality data, most of the data in the real world is noisy. To improve data quality and retain more useful data, the present invention trains a data classifier and strictly filters noisy data. Among them, the positive samples include high-quality standard answer data answered by annotators, as well as correct answers generated by GPT-3.5-Turbo and GPT-4. The negative samples include incorrect data generated by the base model. Based on this data classifier, the present invention automatically selects the data set with a higher probability ranking.
[0041] 3. Instruction data set optimization.
[0042] For a large number of data sets with only answers, the present invention uses GPT-3.5-Turbo and GPT-4 to generate a complete step-by-step solution process. If incorrect answer data is encountered, the present invention will adjust the prompt again to correct the error by itself until a correct solution process is generated, that is, Self-reflection.
[0043] Thus, the instruction data set constructed by the present invention is applicable to scientific questions at the university level.
[0044] III. Perform supervised fine-tuning on the base model with the instruction data set applicable to scientific questions at the university level to obtain the scientific language model.
[0045] To improve the ability of the language model to answer scientific questions, the present invention performs supervised fine-tuning on the base model based on the above-mentioned constructed instruction data set applicable to scientific questions at the university level, and trains a scientific language model applicable to scientific questions at the university level.
[0046] After supervised fine-tuning, the present invention conducts experiments with the scientific language model applicable to scientific questions at the university level after supervised fine-tuning. The final experimental results show that the scientific language model applicable to scientific questions at the university level of the present invention exceeds the base model on the mathematics evaluation data set, and also shows strong performance on multiple scientific evaluation data sets. With these core capabilities, the scientific language model applicable to scientific questions at the university level of the present invention can be feasibly applied to different scientific question answering tasks.
[0047] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than limiting the protection scope of the present invention. Those skilled in the art can modify or equivalently replace the technical solutions of the present invention according to the idea of the present invention, without departing from the essence and scope of the technical solutions of the present invention.
Claims
1. A method for constructing a scientific language model for solving scientific problems at the university level, characterized in that, It includes the following steps: 1). Select a base model; 2). Construct an instruction dataset suitable for scientific questions at the university level; 3). Use the instruction dataset suitable for scientific questions at the university level to perform supervised fine-tuning on the base model to obtain the scientific language model.
2. The method for constructing a scientific language model for solving scientific problems at the university level according to claim 1, wherein, In step 1), ChatGLM3-6B-Base is selected as the base model.
3. The method for constructing a scientific language model for solving scientific problems at the university level according to claim 2, wherein, Step 2) specifically includes: 2.1). Instruction dataset collection; 2.2). Instruction dataset filtering; 2.3). Instruction dataset optimization.
4. The method for constructing a scientific language model for solving scientific problems at the university level according to claim 3, characterized in that, The instruction dataset collection in step 2.1) is specifically as follows: First, collect a large number of scientific questions and answers at the university level from multiple sources including university textbooks, university exercise sets, and university lecture notes; then, extract useful text from the collected scientific questions and answers at the university level through optical character recognition technology; finally, supplement the thought chain process in LaTeX format for the scientific questions and answers in the extracted useful text.
5. The method for constructing a scientific language model for solving scientific problems at the university level according to claim 4, wherein, The instruction dataset filtering in step 2.2) is specifically as follows: Train a data classifier and filter out noisy data through the data classifier. Among them, the positive samples of the data classifier include high-quality standard answer data answered by annotators, as well as correct answers generated by GPT-3.5-Turbo and GPT-4, and the negative samples include incorrect answers generated by the base model; based on the data classifier, select the dataset with a higher probability ranking as the filtered dataset.
6. The method for constructing a scientific language model for solving scientific problems at the university level according to claim 5, characterized in that, The instruction dataset optimization in step 2.3) is specifically as follows: Use GPT-3.5-Turbo and GPT-4 to generate a complete step-by-step solution process for scientific questions. If incorrect answer data is encountered, adjust the prompt again and perform self-correction until the correct solution process is generated.