Construction method, system and equipment of primary and middle school education tutoring large model and medium
By performing two-stage fine-tuning of the Qwen2.5-3B-Instruct model, including reinforcement learning and supervision instructions, the problem of high computing resources, lack of teacher style and insufficient ability to handle complex problems in primary and secondary education is solved, and efficient teacher style tutoring under low resource requirements is achieved.
Patent Information
- Application Number
- CN202510872963.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-27
- Publication Date
- 2025-07-25
- Estimated Expiration
- 2045-06-27
AI Technical Summary
The existing large language model has problems such as high demand for computing resources, lack of teacher style and insufficient ability to handle complex problems in the field of primary and secondary education.
Qwen2.5-3B-Instruct is used as the basic model, and the model is improved by constructing the training data set, reward function and reinforcement learning adjustment parameters, combining the teacher style data set and personalized prompt words for two-stage fine-tuning, including reinforcement learning and supervision instructions fine-tuning to improve the model's reasoning ability and teacher style.
While maintaining low computing resource requirements, the model's ability to handle complex problems is significantly improved, and it is given a clear teacher style, so that the generation and problem-solving process is detailed and clear, making it easy for students to understand.
Smart Images

Figure CN120372300A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology, and specifically, to a method, system, device, and medium for constructing a large model for primary and secondary school education tutoring. Background Art
[0002] Although current large language models perform well in general tasks, in specific teaching scenarios, especially in the field of primary and secondary school education, they have the following defects: Lack of private corpus in the vertical field of primary and secondary school education, and the question and answer effect fails to meet expectations Existing high-performance large models have a large number of parameters and extremely high requirements for computing resources, making it difficult to be deployed and implemented General large models lack the unique teaching style of teachers when answering questions, making it difficult to meet the requirements of a "virtual teacher" (speaking style) and not being suitable for serving groups such as primary and secondary school students.
[0003] Therefore, a new model construction method is needed, which can improve the model's processing ability for complex problems and endow it with a clear teacher style while maintaining low computing resource requirements. Summary of the Invention
[0004] Aiming at the problems of existing large language models in educational applications, such as insufficient ability to handle complex problems, lack of teacher teaching style, and excessive computing resource requirements, the present invention proposes a method, system, device, and medium for constructing a large model for primary and secondary school education tutoring; the method first uses Qwen2.5-3B-Instruct as the base model and constructs a training dataset according to the obtained teaching dataset; then, based on the training dataset, a reward function is constructed to obtain a reward value, and the reinforcement learning method is called to adjust the model parameters to obtain a strengthened base model; finally, based on the constructed teacher style dataset and the set personalized prompt words, the strengthened base model is fine-tuned in a supervised instruction fine-tuning manner to obtain a large model for primary and secondary school education tutoring. Through two-stage fine-tuning, the reasoning ability of the model is improved, and while maintaining low computing resource requirements, the ability of the old model to handle complex problems is improved and it is endowed with a clear teacher style.
[0005] The specific implementation content of the present invention is as follows: A method for constructing a large model for primary and secondary school education tutoring specifically includes the following steps: Step S1: Use Qwen2.5-3B-Instruct as the base model and construct a training dataset according to the obtained teaching dataset; Step S2: Based on the training dataset, construct a reward function to obtain a reward value, and call the reinforcement learning method to adjust the model parameters to obtain a strengthened base model; Step S3: Fine-tune the enhanced base model in the way of supervised instruction fine-tuning according to the constructed teacher style dataset and the set personalized prompt words to obtain a large model for primary and secondary school education tutoring.
[0006] To better implement the present invention, further, the step S2 specifically includes the following steps: Step S21: Construct a reward function according to the answer generated by the base model, the standard answer, the similarity function, and the set reward parameters, and calculate the reward value; Step S22: Calculate the probability distribution according to the set regularization coefficient, the set clipping threshold, and the model parameters; Step S23: Calculate the KL divergence according to the reward value and the probability distribution; Step S24: Adjust the regularization coefficient according to the upper limit value and the lower limit value of the KL divergence to obtain the enhanced base model.
[0007] To better implement the present invention, further, the step S22 specifically includes the following steps: Step S221: Calculate the policy ratio according to the questions obtained from the training dataset, the model sampled answers, and the large model strategy; Step S222: Calculate the advantage estimate according to the reward value and the average reward of all the questions in the current buffer; Step S223: Calculate the probability that the base model generates an answer with the model parameter θ according to the set regularization coefficient, the set clipping threshold, the advantage estimate, and the policy ratio to obtain the probability distribution.
[0008] To better implement the present invention, further, the specific operation of the step S24 is: adaptively adjust the regularization coefficient according to the upper limit value and the lower limit value of the KL divergence. If the KL divergence is greater than the KL divergence upper limit value, then adjust the regularization coefficient λ KL to ; if the KL divergence is greater than the KL divergence upper limit value, then adjust the regularization coefficient λ KL to .
[0009] To better implement the present invention, further, the step S3 specifically includes the following steps: Step S31: Construct a teacher style enhanced dataset according to the obtained typical test questions and the Q&A pair dataset of the corresponding teacher style explanations; Step S32: Convert the teacher style enhanced dataset into data pairs in the Alpaca instruction form and call Role-Play to construct personalized prompt words; Step S33: Fine-tune the enhanced base model in the way of supervised instruction fine-tuning to obtain a large model for primary and secondary school education tutoring.
[0010] To better implement the present invention, further, the specific steps of step S33 include the following steps: Step S331: Insert LoRA adapters into several linear transformation matrices of the base model; Step S332: Keep the original linear transformation matrices frozen and train the newly added LoRA adapter matrices; Step S333: According to the Alpaca instruction-form data pairs, call the Adam optimizer to perform gradient updates on the LoRA adapter matrices to obtain the base model after supervised instruction fine-tuning, that is, the large model for primary and secondary school education tutoring.
[0011] To better implement the present invention, further, the specific content of step S1 is: Use Qwen2.5-3B-Instruct as the base model, and construct a composite mathematical reasoning training dataset according to the obtained high-quality mathematical reasoning dataset and the local private primary and secondary school mathematics dataset.
[0012] Based on the above-mentioned construction method of the large model for primary and secondary school education tutoring, to better implement the present invention, further, a construction system for the large model for primary and secondary school education tutoring is proposed, which is used to execute the above-mentioned construction method of the large model for primary and secondary school education tutoring; it includes an initial unit, a reinforcement unit, and a supervised fine-tuning unit; The initial unit is used to use Qwen2.5-3B-Instruct as the base model and construct a training dataset according to the obtained teaching dataset; The reinforcement unit is used to construct a reward function according to the training dataset to obtain a reward value, and call the reinforcement learning method to adjust the model parameters to obtain the strengthened base model; The supervised fine-tuning unit is used to fine-tune the strengthened base model in the way of supervised instruction fine-tuning according to the constructed teacher-style dataset and the set personalized prompt words to obtain the large model for primary and secondary school education tutoring.
[0013] Based on the above-mentioned construction method of the large model for primary and secondary school education tutoring, to better implement the present invention, further, an electronic device is proposed, which includes a memory and a processor; a computer program is stored on the memory; when the computer program is executed on the processor, the above-mentioned construction method of the large model for primary and secondary school education tutoring is realized.
[0014] Based on the above-mentioned construction method of the large model for primary and secondary school education tutoring, to better implement the present invention, further, a computer-readable storage medium is proposed, and a computer instruction is stored on the computer-readable storage medium; when the computer instruction is executed on the above-mentioned electronic device, the above-mentioned construction method of the large model for primary and secondary school education tutoring is realized.
[0015] The present invention has the following beneficial effects: (1) Through two-stage fine-tuning, the present invention improves the model's ability to handle complex problems and endows it with a clear teacher style while maintaining low computational resource requirements.
[0016] (2) The present invention constructs a large model for primary and secondary school education counseling based on a two-stage fine-tuning framework, effectively solving the problems of existing large language models in specific educational applications, such as insufficient ability to handle complex problems, lack of teacher's explanation style, and excessive computational resource requirements. Compared with models at the same level, it has stronger capabilities and a distinct teacher style.
[0017] (3) By using the reinforcement learning method to continuously adjust the model parameters, the present invention gradually improves the model's performance in complex mathematical reasoning tasks and significantly enhances the model's mathematical reasoning ability.
[0018] (4) After supervised instruction fine-tuning, the model of the present invention can not only solve specific subject problems, but also significantly exhibit a professional explanation style similar to that of primary and secondary school teachers. The problem-solving process generated by the model is detailed, clear, and easy for students to understand. Description of the Drawings
[0019] Figure 1 It is a schematic block diagram of the specific process provided by the present invention.
[0020] Figure 2 It is a schematic diagram of the chain of thought provided by the present invention. Detailed Embodiments
[0021] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. It should be understood that the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments, and therefore should not be regarded as a limitation of the protection scope. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the protection scope of the present invention.
[0022] In the description of the present invention, it should be noted that unless otherwise clearly defined and limited, the terms "set", "connected", and "connected" should be understood in a broad sense. For example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can also be directly connected, or indirectly connected through an intermediate medium, and it can be the internal communication of two components. For those of ordinary skill in the art, the specific meanings of the above terms in the present invention can be understood according to specific situations.
[0023] Embodiment 1: This embodiment proposes a method for constructing a large model for primary and secondary school education tutoring, specifically including the following steps: Step S1: Use Qwen2.5-3B-Instruct as the base model and construct a training dataset based on the obtained teaching dataset.
[0024] Specifically, step S1 is as follows: Use Qwen2.5-3B-Instruct as the base model and construct a composite mathematical reasoning training dataset based on the obtained high-quality mathematical reasoning dataset and the local private primary and secondary school mathematics dataset.
[0025] Step S2: Based on the training dataset, construct a reward function to obtain a reward value, and call the reinforcement learning method to adjust the model parameters to obtain a strengthened base model.
[0026] Specifically, step S2 includes the following steps: Step S21: Construct a reward function according to the answer generated by the base model, the standard answer, the similarity function, and the set reward parameters, and calculate the reward value; Step S22: Calculate the probability distribution according to the set regularization coefficient, the set clipping threshold, and the model parameters; Further, step S22 specifically includes the following steps: Step S221: Calculate the policy ratio according to the questions obtained from the training dataset, the model sampled answers, and the large model policy; Step S222: Calculate the advantage estimate according to the reward value and the average reward of all questions in the current buffer; Step S223: Calculate the probability that the base model generates an answer with model parameter θ according to the set regularization coefficient, the set clipping threshold, the advantage estimate, and the policy ratio to obtain the probability distribution.
[0027] Step S23: Calculate the KL divergence according to the reward value and the probability distribution; Step S24: Adjust the regularization coefficient according to the upper and lower limit values of the KL divergence to obtain a strengthened base model.
[0028] Further, the specific operation of step S24 is as follows: Adaptively adjust the regularization coefficient according to the upper and lower limit values of the KL divergence. If the KL divergence is greater than the upper limit value of the KL divergence, then adjust the regularization coefficient λ KL to ; if the KL divergence is greater than the upper limit value of the KL divergence, then adjust the regularization coefficient λ KL to .
[0029] Step S3: Fine-tune the enhanced base model in a supervised instruction fine-tuning manner according to the constructed teacher style dataset and the set personalized prompts to obtain a large model for primary and secondary school education tutoring.
[0030] The specific steps of step S3 are as follows: Step S31: Construct a teacher style enhanced dataset according to the obtained typical test questions and the Q&A pair dataset of corresponding teacher style explanations; Step S32: Convert the teacher style enhanced dataset into data pairs in the form of Alpaca instructions, and call Role-Play to construct personalized prompts; Step S33: Fine-tune the enhanced base model in a supervised instruction fine-tuning manner to obtain a large model for primary and secondary school education tutoring.
[0031] The specific steps of step S33 are as follows: Step S331: Insert LoRA adapters into several linear transformation matrices of the base model; Step S332: Keep the original linear transformation matrix frozen and train the newly added LoRA adapter matrix; Step S333: According to the data pairs in the form of Alpaca instructions, call the Adam optimizer to perform gradient update on the LoRA adapter matrix to obtain the base model after supervised instruction fine-tuning, that is, the large model for primary and secondary school education tutoring.
[0032] Working principle: In this embodiment, through two-stage fine-tuning, while maintaining low computational resource requirements, the model's ability to handle complex problems is improved and it is given a clear teacher style; a large model for primary and secondary school education tutoring is constructed based on the two-stage fine-tuning framework, effectively solving the problems of insufficient ability to handle complex problems, lack of teacher explanation style, and excessive computational resource requirements existing in existing large language models in specific educational applications, and having stronger capabilities and a distinct teacher style compared to models at the same level.
[0033] Embodiment 2: Based on the above Embodiment 1, as Figure 1 、 Figure 2 shown, a specific embodiment is used for detailed description.
[0034] Step S1: Selection of the base model.
[0035] In this embodiment, Qwen2.5-3B-Instruct is selected as the base model, that is, the base model. This model has a moderate number of parameters of 3B scale, has good language understanding and generation capabilities, and is suitable for applications in the primary and secondary school education field.
[0036] Selecting a large model with a small number of parameters such as Qwen2.5-3B-Instruct solves the problem of "excessive computing resource requirements" raised above. Then, on this basis, two-stage fine-tuning is carried out to make this small model have stronger capabilities and a distinct teacher style compared with models at the same level.
[0037] Step S2: The first stage: Reinforcement learning fine-tuning, the stage of enhanced reasoning ability.
[0038] Step S21: Dataset construction.
[0039] In this embodiment, a publicly available high-quality mathematical reasoning dataset such as S1K and a local private primary and secondary school mathematics dataset are mixed to form a composite mathematical reasoning training dataset EdMath. This dataset covers mathematical reasoning problems of different difficulties and has a high coverage and representativeness. The EdMath dataset contains approximately 3,000 items and includes three fields: questions, chains of thought, and answers.
[0040] Step S22: Reinforcement learning implementation method.
[0041] In this embodiment, a reward-based reinforcement learning method is adopted, and the accuracy of the answers given by the model is used as the reward signal. Specifically, after the model generates an answer, it is compared with the standard answer. A correct answer receives a positive reward, while an incorrect answer is given a negative reward or zero reward. This embodiment also takes into account the Chain-of-Thought (CoT) process and continuously adjusts the model parameters through the reinforcement learning method to gradually improve the performance of the model in complex mathematical reasoning tasks.
[0042] This is the reasoning process, also known as the chain of thought. Before the launch of DeepSeek-R1, only ChatGPT-o1 had reasoning capabilities (such models are also called reasoning models). This thinking process can significantly improve the model's capabilities, as Figure 2 shown.
[0043] Step S221: Reward function setting.
[0044] Among them, A is the standard answer, Â is the answer given by the large model, that is, to make the answer given by the large model closer to the expected value, with α = 0.7 and β = 0.3 by default; among them, α and β are weight coefficients, and α > β indicates that the "answer" is more important than the "reasoning process".
[0045] EM: Exact Match, which determines whether the answer is correct. Only when the final result is correct can a high score be given, a hard indicator; Sim: Text similarity, similarity in the reasoning process is also considered, encouraging "correct thinking", even if there are calculation errors; this can quickly increase the accuracy rate while not discarding the value of the thought chain.
[0046] Step S222: Given the question Q~D, the model samples the answer  and obtains the reward ; D contains questions and answers. Here, Q is randomly selected from the answer column of the dataset D. Then the large model will answer this question Q to get an answer, and then go back to the dataset D to compare it with its corresponding standard answer (calculate the reward), that is, to see if the large model answers correctly.
[0047] Update the parameter θ by maximizing the following formula: (1) Given the question Q, the probability that the model generates the answer  with the parameter θ is the probability distribution; is the policy optimization objective function used in reinforcement learning, which represents the expected total reward that the model can obtain when optimizing the policy by adjusting the model parameter θ.
[0048] is the policy ratio; represents the probability value that the model generates the same predicted answer  for the same question Q when the parameter is the old policy in the previous round or batch. Its role is to serve as a benchmark in the policy update process to measure the improvement degree of the model.
[0049] A t is the advantage estimate and can be calculated by GAE; is the clipping threshold to suppress the excessive policy step size; λ KL is the KL regularization coefficient, which is dynamically adjusted to prevent the policy from diverging; Use "clipping" to protect the training stability and use KL to lock the language fluency; Step S23: Training loop.
[0050] Step S231: Sampling.
[0051] Randomly select BS questions from D and use the current policy π θ to generate the reasoning chain + answer; calculate the reward r immediately; "Thought chain" and "reasoning chain" are the same thing. This thought chain is constrained by the prompt words, that is, after the training in step S235 is completed <reasoning>Inference process…< / reasoning> , through this prompt word, the large model can understand to think first and then give an answer.
[0052] When using DeepSeek, Doubao, or Qianwen, they will think for a long time on their own before answering questions and then give answers. This is the "reasoning process", "chain of thought", or "reasoning chain".
[0053] Step S232: Advantage estimation.
[0054] After accumulating K batches of samples, the advantage A of each trajectory is obtained using the GAE formula t ; Sampling continuously multiple times and accumulating them for a one-time gradient update.
[0055] General formula: : Immediate advantage; : Time discount; λ GAE : Smoothing coefficient; : State value function, where t is the time step; L: The remaining number of steps from the current time t to the end of the trajectory; l : Count, taking values 0, 1, 2... L - 1; However, in the scenario of this embodiment, it can be simplified as: Since it is a single-step question and answer, there is no problem with the t-th step, so the subscript t can be omitted. At the same time, L = 1, and the formula is simplified as follows.
[0056] r: The immediate reward obtained by the model for this question; : The average reward for all questions in the current buffer; In the reinforcement learning training process of this embodiment, each batch contains a certain number of question samples. Assume that each batch contains N samples. After each sample is answered by the model, a reward value r can be obtained i (automatically calculated by the reward function) Then the batch average reward The calculation formula is: r i represents the reward value obtained by the i-th sample in the current batch; N represents the total number of all samples in the current batch; The significance of this average reward value is to establish a baseline for all reward values in the current batch to determine whether the reward value of a single sample is higher or lower than the overall average level, so as to obtain a more stable and effective training signal.
[0057] Step S233: Perform gradient ascent on Equation (1) for these K batches of samples; repeat for E epochs to fully utilize the samples.
[0058] Step S234: Adaptive KL.
[0059] Monitor the average KL during training. If KL > KL max , multiply λ KL by 1.5; if KL < KL min , divide λ KL by 1.5. Kullback-Leibler Divergence (KL divergence) is used to measure how dissimilar two probability distributions are, the difference between the old and new models.
[0060] Some preferred values of hyperparameters: Batch size BS: 32; Number of sampled batches per round K: 8; Epoch: 6; Clip threshold ε: 0.2; KL min / max : 0.01 / 0.05; Initial λ KL : 1.0; Parameter explanation: , a total of M question-answer pairs; π θ , the basic large model strategy to be optimized; π0, the original parameter snapshot, used for KL constraint; R(Q, Â), the reward function, scoring the model answer Â; EM, Exact-Match, 1 if the answer values match exactly, otherwise 0; Sim , semantic similarity, range [0,1]; Step S235: Training completed.
[0061] Output the final strategy π θ ; Output format: <reasoning>Inference process…< / reasoning> <answer>Answer< / answer> Training effect: After the reinforcement learning fine-tuning stage, the mathematical reasoning ability of the model has been significantly enhanced, especially in the performance on complex questions, which is significantly better than the base model without reinforcement learning fine-tuning.
[0062] In the model evaluation stage, this embodiment prepared 1,000 primary and secondary school math problems as the evaluation dataset (not involved in training). The correct answer rate of the model without reinforcement learning was about 0.63, and the correct answer rate of the model after reinforcement learning was about 0.77.
[0063] Step S3: The second stage: supervised instruction fine-tuning (teacher style formation stage).
[0064] The purpose of this step is to enable the model to transition from "knowledge" to "personalization".
[0065] Step S31: Teacher style enhanced dataset.
[0066] Construct a Q&A pair dataset containing about 10,000 typical test questions from various primary and secondary school subjects and their corresponding teacher style explanations, called the "teacher style enhanced dataset", which reflects the common expression methods and teaching logics used by teachers in classroom teaching.
[0067] Step S32: Supervised instruction fine-tuning.
[0068] Based on the enhanced reasoning ability in the first stage, fine-tuning is carried out in the way of supervised instruction fine-tuning (SFT). First, convert the teacher style dataset into data pairs in the Alpaca instruction form, with the input being the question and the output being the test question analysis.
[0069] On this basis, this embodiment designs a personalized prompt based on Role-Play, which can better enable the large model to learn the expected role characteristics and behavior patterns, and realize a personalized interaction experience while maintaining the reasoning ability.
[0070] #Role You are a senior primary and secondary school teacher. With years of in-depth work in the front line of education, you are proficient in the knowledge systems of various primary and secondary school subjects. You can not only accurately answer various subject questions, but also be good at analyzing the thinking trajectories of humans when dealing with complex problems, and clearly present this progressive reasoning process in a vivid, natural and human-cognitive-law-compliant way.
[0071] #Requirements 1. After receiving the subject questions given by the students, use progressive thinking, starting from the basic concepts, and gradually derive to the final answer, clearly explaining the logical basis of each step.
[0072] 2. Give play to the ability of multi-angle thinking, explore various different problem-solving strategies, and comprehensively show different solution paths and applicable scenarios of the problem.
[0073] 3. Incorporate intuitive heuristic elements, share the inspiration that may suddenly occur during the problem-solving process, and how these inspirations guide the direction of problem-solving. On this basis, generate a new analysis process that is natural and conforms to the cognitive characteristics of humans.
[0074] During actual fine-tuning, LoRA efficient parameter fine-tuning is adopted to reduce the consumption of computing resources. In this embodiment, LoRA adapters are first inserted into several linear transformation matrices of the base model and its update can be expressed as where the rank , the original W remains frozen, and only the newly added A and B are trained. A and B are two matrices. For example, a 3*3 original matrix split into AB is 3*1 and 1*3. Then the training amount is reduced from 9 to 6.
[0075] The training process is to send the "question - analysis" instruction pair into the model to generate predicted text , measure the gap with the "standard answer" (here the "standard answer" refers to the text sent in) using cross-entropy, and perform gradient updates on A and B using the Adam optimizer.
[0076] Step S33: Formation of teacher style; The model after supervised instruction fine-tuning can not only solve specific subject problems, but also significantly show a professional teaching style similar to that of primary and secondary school teachers. The problem-solving process generated by the model is detailed and clear, and is easy for students to understand.
[0077] Working principle: This embodiment creates a two-stage fine-tuning framework of "small parameter quantity + reinforcement learning + LoRA-SFT" for the three major deficiencies existing in the existing general large models and fine-tuning schemes: 1. High computing / deployment costs, 2. Limited accuracy for complex problems, 3. Lack of personalized style, which fundamentally improves the usability and teaching value of the model in the primary and secondary school education scenarios.
[0078] Other parts of this embodiment are the same as those of the above Embodiment 1, so they will not be elaborated here.
[0079] Embodiment 3: Based on any one of the above Embodiment 1 - Embodiment 2, this embodiment proposes a construction system for a large model for primary and secondary school education tutoring, which is used to execute the above-mentioned construction method for a large model for primary and secondary school education tutoring; it includes an initial unit, a reinforcement unit, and a supervised fine-tuning unit; The initial unit is used to use Qwen2.5-3B-Instruct as the base model and construct a training dataset according to the obtained teaching dataset; The reinforcement unit is used to construct a reward function based on the training data set to obtain a reward value, and call a reinforcement learning method to adjust the model parameters to obtain a strengthened basic model. The supervised fine-tuning unit is used to fine-tune the strengthened basic model in the way of supervised instruction fine-tuning according to the constructed teacher-style data set and the set personalized prompt words to obtain a large model for primary and secondary school education tutoring.
[0080] This embodiment provides an electronic device, including a memory and a processor; a computer program is stored on the memory; when the computer program is executed on the processor, the above-mentioned method for constructing a large model for primary and secondary school education tutoring is implemented.
[0081] This embodiment also provides a computer-readable storage medium, on which a computer instruction is stored; when the computer instruction is executed on the above-mentioned electronic device, the above-mentioned method for constructing a large model for primary and secondary school education tutoring is implemented.
[0082] Other parts of this embodiment are the same as any one of the above-mentioned Embodiment 1 - Embodiment 2, so they will not be elaborated here.
[0083] The above are only the preferred embodiments of the present invention, and do not impose any form of limitation on the present invention. Any simple modification or equivalent change made to the above embodiments based on the technical essence of the present invention falls within the protection scope of the present invention.
Claims
1. A method for constructing a large model for primary and secondary school education tutoring, characterized in that, Specifically, it includes the following steps: Step S1: Use Qwen2.5-3B-Instruct as the base model and construct a training dataset according to the obtained teaching dataset; Step S2: According to the training dataset, construct a reward function to obtain a reward value, and call the reinforcement learning method to adjust the model parameters to obtain a strengthened base model; Step S3: According to the constructed teacher style dataset and the set personalized prompt words, fine-tune the strengthened base model in the way of supervised instruction fine-tuning to obtain a large model for primary and secondary school education counseling.
2. The construction method of a large model for primary and secondary school education tutoring according to claim 1, wherein The specific steps of step S2 include the following steps: Step S21: According to the answers generated by the base model, the standard answers, the similarity function, and the set reward parameters, construct a reward function and calculate the reward value; Step S22: Calculate the probability distribution according to the set regularization coefficient, the set clipping threshold, and the model parameters; Step S23: Calculate the KL divergence according to the reward value and the probability distribution; Step S24: Adjust the regularization coefficient according to the upper and lower limit values of the KL divergence to obtain a strengthened base model.
3. The construction method of a large model for primary and secondary school education counseling according to claim 2, characterized in that, The specific steps of step S22 include the following steps: Step S221: Calculate the policy ratio according to the questions obtained from the training dataset, the model sampled answers, and the large model strategy; Step S222: Calculate the advantage estimate according to the reward value and the average reward of all questions in the current buffer; Step S223: Calculate the probability that the base model generates an answer with the model parameter θ according to the set regularization coefficient, the set clipping threshold, the advantage estimate, and the policy ratio to obtain the probability distribution.
4. The construction method of a large model for primary and secondary school education counseling according to claim 3, wherein, The specific operation of step S24 is as follows: adaptively adjust the regularization coefficient according to the upper and lower limits of the KL divergence. If the KL divergence is greater than the upper limit of the KL divergence, then the regularization coefficient λ KL is adjusted to . If the KL divergence is greater than the upper limit of the KL divergence, then the regularization coefficient λ KL is adjusted to .
5. A method for constructing a large model for primary and secondary school education tutoring according to claim 1, characterized in that, The specific steps of step S3 include the following steps: Step S31: According to the obtained typical test questions and the Q&A pair dataset with corresponding teacher style explanations, construct a teacher style enhanced dataset; Step S32: Convert the teacher style enhanced dataset into data pairs in the Alpaca instruction form and call Role-Play to construct personalized prompt words; Step S33: Fine-tune the strengthened base model in the way of supervised instruction fine-tuning to obtain a large model for primary and secondary school education counseling.
6. The construction method of a large model for primary and secondary school education counseling according to claim 5, wherein The specific steps of step S33 include the following steps: Step S331: Insert LoRA adapters into several linear transformation matrices of the base model; Step S332: Keep the original linear transformation matrix frozen and train the newly added LoRA adapter matrix; Step S333: According to the data pairs in the Alpaca instruction form, call the Adam optimizer to perform gradient update on the LoRA adapter matrix to obtain the base model after supervised instruction fine-tuning, that is, the large model for primary and secondary school education counseling.
7. A method for constructing a large model for primary and secondary school education counseling according to claim 1, characterized in that The specific content of step S1 is: Use Qwen2.5-3B-Instruct as the base model and construct a composite mathematical reasoning training dataset according to the obtained high-quality mathematical reasoning dataset and the local private primary and secondary school mathematics datasets.
8. A construction system for a large model of primary and secondary school education tutoring, which is used to execute a construction method for a large model of primary and secondary school education tutoring as described in claim 1; characterized in that, It includes an initial unit, a reinforcement unit, and a supervised fine-tuning unit; The initial unit is used to use Qwen2.5-3B-Instruct as the base model and construct a training dataset according to the obtained teaching dataset; The reinforcement unit is used to construct a reward function based on the training data set to obtain a reward value, and call the reinforcement learning method to adjust the model parameters to obtain a reinforced basic model; The supervised fine-tuning unit is used to fine-tune the reinforced basic model in the way of supervised instruction fine-tuning according to the constructed teacher-style data set and the set personalized prompt words to obtain a large model for primary and secondary school education tutoring.
9. An electronic device, characterized in that, It includes a memory and a processor; a computer program is stored on the memory; when the computer program is executed on the processor, the method for constructing the large model for primary and secondary school education tutoring according to any one of claims 1-7 is implemented.
10. A computer-readable storage medium, characterized in that, A computer instruction is stored on the computer-readable storage medium; when the computer instruction is executed on the electronic device according to claim 9, the method for constructing the large model for primary and secondary school education tutoring according to any one of claims 1-7 is implemented.
Citation Information
Patent Citations
Method for adjusting robot task neural network behaviors based on human feedback
CN118061178A
Mahjong matching game level generation method based on reinforcement learning, medium and equipment
CN119425055A
Zero-sample cross-modal image retrieval method and device based on low-rank adaptation
CN119719405A
Automatic Feature Subset Selection based on Meta-Learning
US20200327357A1
Cited By
Construction method of breast cancer patient nutrition tutoring large model
CN121306427A
Generative virtual tutoring teacher model training method based on preference classification
CN121936502A