Teaching, research and training scene three-stage iterative large model training method
Through the three-stage iterative training method, the generalization ability and adaptability of large models in teaching and scientific research scenarios is improved, and the shortcomings of existing models in processing multi-source heterogeneous data and multimedia content are solved, and more efficient and accurate application in the field of education is achieved.
Patent Information
- Application Number
- CN202510120130.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-25
- Publication Date
- 2025-05-30
AI Technical Summary
Existing large models are difficult to meet complex and changeable needs in teaching, scientific research and training scenarios, especially when processing multi-source heterogeneous data and multimedia content, and lack adaptability and generalization capabilities.
A three-stage iterative large-scale model training method is adopted, including data collection and preprocessing, model training and model optimization and application stages. The first stage uses initial model training and fine-tuning, the second stage introduces reinforcement learning and transfer learning, and the third stage adopts adversarial training to further improve the generalization ability and robustness of the model.
It significantly improves the generalization ability and adaptability of the large model in teaching, scientific research and training scenarios, and can handle diversified tasks in the education field more accurately and effectively, and improves teaching quality, scientific research level and training results.
Smart Images

Figure FT_1 
Figure SMS_1
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of large model training, and particularly to a three-stage iterative large model training method for teaching, scientific research, training and cultivation scenarios. Background Art
[0002] In the current era of rapid development of artificial intelligence, large language models have demonstrated powerful capabilities such as language understanding and content generation, and have performed excellently in many tasks of natural language processing, such as translation, dialogue, information extraction, etc., and also have extremely broad application prospects in general fields.
[0003] However, in the current field of education and scientific research, the application demand for artificial intelligence technology is increasing day by day. Especially when dealing with a large number of documents and data, traditional models often struggle to meet the requirements of complex and changing scenarios. Therefore, in order to improve processing efficiency and accuracy, the Education and Science Institute considered the following key factors when constructing its large model: First, the education and scientific research field involves a wide variety of document types, including academic papers, textbooks, research reports, etc., and these documents are usually long in length and large in information volume.
[0004] In addition, during the education and scientific research process, not only text information needs to be processed, but also multimedia content such as charts and images is often involved. How to effectively integrate and utilize this data to train the large model is also an urgent problem to be solved.
[0005] In summary, in order to overcome the deficiencies of existing large models in teaching, scientific research, training and cultivation scenarios and meet the requirements of the teaching, scientific research, training and cultivation fields for model professionalism, reliability and adaptability, this solution aims to make full use of multi-source heterogeneous data in the teaching, scientific research, training and cultivation fields through innovative training methods, integrate abstract experience, observational experience and practical experience, so that the large model can better serve teaching, scientific research, training and cultivation work, and improve teaching quality, scientific research level and training effect. Summary of the Invention
[0006] The present invention provides a three-stage iterative large model training method for teaching, scientific research, training and cultivation scenarios to solve the above problems existing in the prior art.
[0007] In a first aspect, the present invention provides a three-stage iterative large model training method for teaching, scientific research, training and cultivation scenarios, including: Data collection and preprocessing stage: Collect multi-source data in teaching, scientific research, training and cultivation scenarios, including but not limited to teaching materials, scientific research achievements, training records, academic papers, etc.; clean the collected data to remove noise data, duplicate data and invalid data; label the cleaned data, including text classification, entity recognition, relationship extraction, etc., and the labeling work can be carried out in a combination of manual labeling and automatic labeling; Model Training Phase: Construct an initial large model with a multi-layer neural network structure; divide the preprocessed data into a training set, a validation set, and a test set according to a certain ratio; use the training set to perform the first-phase training on the initial large model, and adopt the gradient descent algorithm to optimize the model parameters. During the training process, adjust the training parameters according to the feedback of the validation set to avoid overfitting; after completing the first-phase training, evaluate the model. If the evaluation result does not meet the preset standard, enter the second-phase training; in the second-phase training, introduce a reinforcement learning mechanism. By setting a reward function, encourage the model to generate more accurate and useful results in the teaching, research, training, and cultivation scenarios. At the same time, combine transfer learning to transfer the pre-trained model knowledge in other related fields to the current model; after the second-phase training, evaluate the model again. If it still does not meet the preset standard, enter the third-phase training; the third-phase training adopts an adversarial training method to construct an adversarial network and let the generator and the discriminator perform an adversarial game to further improve the generalization ability and robustness of the model. Model Optimization and Application Phase: Optimize the model after three-phase training, including operations such as model compression and pruning to reduce the storage space and computational amount of the model; apply the optimized model to the teaching, research, training, and cultivation scenarios, such as intelligent teaching assistance, scientific research achievement prediction, training effect evaluation, etc., and continuously optimize and update the model according to the feedback data in the actual application.
[0008] Optionally, in the data collection step, data encryption processing is also included. In the data preprocessing step, data augmentation is included. The data augmentation expands the scale and diversity of the data set through data enhancement techniques such as text rewriting and data synthesis.
[0009] Optionally, in the first-phase training, an adaptive learning rate adjustment strategy is adopted to dynamically adjust the learning rate according to the change of the loss function during the training process.
[0010] Optionally, in the second-phase training, the design of the reward function is based on the actual needs and evaluation indicators in the teaching, research, training, and cultivation scenarios, including but not limited to teaching effect improvement, scientific research achievement innovation, training compliance rate, etc.
[0011] Optionally, in the third-phase training, the generator and the discriminator of the adversarial network adopt different neural network architectures and are alternately trained during the training process.
[0012] Optionally, in the model optimization step, quantization technology is adopted to convert the model parameters from a high-precision data type to a low-precision data type to reduce the storage and computational costs of the model.
[0013] Optionally, when applying the model to the teaching, research, training, and cultivation scenario, a user feedback interface is set up to facilitate users to evaluate the output results of the model and put forward improvement suggestions.
[0014] In a second aspect, the present invention provides a system for a three-stage iterative large model training method in a teaching, research, training, and cultivation scenario, including: A data collection module for collecting multi-source data in the teaching, research, training, and cultivation scenario; A data preprocessing module for cleaning, annotating, and encrypting the collected data; A model training module for performing a three-stage model training process; A model evaluation module for evaluating the model during the training process; A model optimization module for optimizing the trained model; An application module for applying the optimized model to the teaching, research, training, and cultivation scenario and receiving user feedback.
[0015] The present invention has the following beneficial effects compared with the prior art: Through the three-stage iterative training method, the generalization ability of the large model in the education field is effectively improved. In the basic model training in the first stage, through the pre-training and fine-tuning of large-scale text data in the education field, the model has solid knowledge in the education field and basic language understanding ability, laying a solid foundation for subsequent task-oriented training and continuous optimization. In the task-oriented training in the second stage, corresponding training tasks and labels are designed for specific education tasks and scenarios, enabling the model to better adapt to different education task requirements and enhancing the model's performance in specific tasks. In the third stage of continuous optimization and iteration, through online learning and incremental training, the model can continuously adapt to new education scenarios and task requirements and maintain a high generalization ability. Experimental results show that in education tasks such as intelligent teaching, automatic grading, and education content generation, the model trained by this solution has higher accuracy and stability than traditional methods and can effectively cope with the diversity and complexity in the education field. Description of the Drawings
[0016] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0017] Figure 1 It is a flowchart of a three-stage iterative large model training method in a teaching, research, training, and cultivation scenario of the present invention. Detailed Embodiments
[0018] The technical solution of the present invention will be clearly and completely described below in conjunction with the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative work fall within the protection scope of the present invention.
[0019] Unless otherwise specified, the raw materials used in the present invention are all conventional products purchased from the market. Embodiment
[0020] As Figure 1 shown, the embodiment of the present invention provides a three-stage iterative large model training method for teaching, scientific research, training and cultivation scenarios, including: Data collection and preprocessing stage: Collect multi-source data in teaching, scientific research, training and cultivation scenarios, including but not limited to teaching materials, scientific research achievements, training records, academic papers, etc.; clean the collected data to remove noise data, duplicate data and invalid data; label the cleaned data, including text classification, entity recognition, relationship extraction, etc., and the labeling work can adopt a combination of manual labeling and automatic labeling. Model training stage: Construct an initial large model, the initial large model has a multi-layer neural network structure; divide the preprocessed data into a training set, a validation set and a test set according to a certain ratio; use the training set to perform the first stage of training on the initial large model, and use the gradient descent algorithm to optimize the model parameters. During the training process, adjust the training parameters according to the feedback of the validation set to avoid overfitting; after completing the first stage of training, evaluate the model. If the evaluation result does not meet the preset standard, enter the second stage of training; in the second stage of training, introduce a reinforcement learning mechanism, set a reward function to encourage the model to generate more accurate and useful results in teaching, scientific research, training and cultivation scenarios, and at the same time combine transfer learning to transfer the knowledge of pre-trained models in other related fields to the current model; after the second stage of training, evaluate the model again. If it still does not meet the preset standard, enter the third stage of training; the third stage of training adopts an adversarial training method, constructs an adversarial network, and allows the generator and the discriminator to perform adversarial games to further improve the generalization ability and robustness of the model.
[0021] In the first stage of training, continued pre-training and supervised instruction fine-tuning are adopted. The output quality of intelligent writing, literature review, etc. is more accurate, which can effectively meet the needs of specific fields. However, due to the current limitation of the number of preference data pairs, preference optimization training is not carried out. The training goal of this stage is to optimize the performance of the model on these tasks according to the needs of specific downstream application scenarios. In the process of preparing training data, a total of 2,772 documents related to the training tasks are obtained through collection and data augmentation methods in this stage. These documents mainly come from the fields of AI teaching and research, covering various scenario requirements. We screened according to the relevance of the documents and finally retained 510 documents. These documents were further cleaned and organized and finally used for continued pre-training. The 510 selected documents contain nearly 1M tokens, and these tokens form the core data basis for continued pre-training. Through screening, the high quality and relevance of the data are ensured, which helps to improve the performance of the model in the target field. In order to carry out supervised instruction fine-tuning, a large amount of instruction data annotation work has also been carried out. Based on the actual application requirements in teaching and research, we set multiple specific task scenarios, such as: writing tasks in AI teaching and research (including intelligent writing, intelligent rewriting and intelligent modification of teaching and research / event plan, intelligent writing of teaching and research / event notice). During the annotation process, 32,842 instruction data related to the requirement points were successfully generated, and these data will be used for supervised instruction fine-tuning in the first stage. These instruction data cover the task requirements under a variety of complex scenarios and are classified and organized in detail according to different task requirements.
[0022] In continued pre-training, starting from the existing pangu-8B model, continuous pre-training is carried out to further improve the generalization ability of the model and its performance in specific fields. The main purpose of pre-training is to make the model better adapt to downstream tasks by continuing to train on a large amount of domain-related data. The AdamW optimizer is used for training. This optimizer is an improved version of Adam, which can better handle the weight decay problem, thus improving the training effect of the model. The initial learning rate is set to 2e-5. This relatively low learning rate can ensure that the model makes fine adjustments on the existing basis and avoids over-updating the weights. The warmup ratio is set to 0.03, that is, in the initial stage of training, by gradually increasing the learning rate, it is prevented that the training of the model falls into a local optimal solution prematurely.
[0023] The number of training epochs is set to 1 epoch, which means the model will iterate through all the training data once. Since the data scale is moderate and there is a basic model available, training for 1 epoch can achieve good results in a relatively short time. The hardware configuration uses 8 80GB graphics cards, with a batch size of 4 set for each card, and a total batch size of 32. This configuration ensures the parallel training speed on large-scale data while avoiding problems such as out-of-memory errors.
[0024] During the second-stage training process, additional training document corpora and corresponding instruction fine-tuning data were added. In terms of the training method, a preference optimization learning stage was added based on the first round.
[0025] For the requirements of comprehensive and knowledge-based tasks, the project team first collected 12,335 documents, all of which were used for the continued pre-training of the large model. 86,293 instruction data were generated through an automatic construction method, among which 16,416 were comprehensive and 69,877 were knowledge-based. In the second round of training, supervised instruction fine-tuning was carried out to optimize the model's performance on specific tasks. The team collected a total of 12,335 documents related to the training tasks, including those in the fields of teaching and research, teaching, scientific research, publicity, and education policies. After cleaning and organizing, they were used for instruction data annotation. The classification of these data is summarized as follows:
[0026] Meanwhile, for supervised instruction fine-tuning, the team also carried out a large amount of instruction data annotation work. Based on the actual application requirements in teaching and research, scientific research, and continuing education, we set multiple specific task scenarios, such as: Knowledge-based tasks in AI teaching and research (teaching strategy suggestions, classroom interaction and Q&A, virtual teaching and research assistants).
[0027] It should be noted that based on the consideration of optimizing the model performance at multiple levels, the second-stage training selected three strategies: "continued pre-training", "instruction supervised fine-tuning", and "preference alignment optimization".
[0028] In the teaching, research, training, and education scenario, the model may need to process a large amount of content related to courses, teaching materials, training, etc. Although the pre-trained model has mastered a wide range of general knowledge, continued pre-training in a specific domain (such as education) can significantly improve the model's performance in that domain. By further training on education-related corpora (such as teaching materials, educational literature, courseware, etc.), the model can better understand the language characteristics, teaching terms, and professional concepts in the education field; continued pre-training enhances the large model's understanding of the proprietary knowledge in the education field, ensuring that the model can effectively process the language and knowledge system in the education field. For example, the model needs to generate learning content that meets educational goals. In this case, continued pre-training can help the model better understand teaching documents and produce outputs that conform to specific educational scenarios. The model usually needs to perform specific tasks according to the instructions of teaching and research personnel or teachers, such as explaining a certain knowledge point, generating administrative documents, interpreting policies, etc. Through instruction supervision and fine-tuning, the model can learn how to perform reasonable task execution according to specific instructions, improving the controllability and practicality of the model. Instruction fine-tuning can also help the model learn to understand diverse instruction expressions and improve its response ability in the teaching, research, training, and education scenario. For example, a teacher can instruct the model to generate a teaching design for a certain subject, or a teaching and research personnel can ask it to provide a detailed explanation for a certain policy. Through instruction supervision and fine-tuning, the model can more efficiently and accurately meet these task requirements. When different users such as teachers and teaching and research personnel often have different preferences for the model's output. Therefore, the preference alignment optimization strategy is adopted to adjust the model's output through human feedback to make it more in line with the needs and expectations of the target users. This not only improves the interactive experience of the model but also avoids inappropriate or unexpected outputs in the preference alignment optimization strategy.
[0029] The third stage of training mainly focuses on improving weak capabilities. In the first two rounds of training, it was found that the trained model was slightly weak in comprehensive tasks. Therefore, in the third round, relevant data was specifically labeled for enhanced training. For the training data in the third stage, the data annotation method in the second stage of training was used to generate more instruction data. Through this continuous data annotation method, the project team can maintain the consistency and accuracy of data annotation. The project team added a total of 5,381 comprehensive instructions. The unsupervised document corpus and human preference data used for training are consistent with those in the second stage of training. A comprehensive training process was carried out, covering three stages: continued pre-training, supervised instruction fine-tuning, and preference alignment optimization, aiming to improve the performance and generalization ability of the large model in diverse tasks.
[0030] Model Optimization and Application Phase: Optimize the model after three-stage training, including operations such as model compression and pruning to reduce the storage space and computational volume of the model; apply the optimized model to teaching, research, training, and other scenarios, such as intelligent teaching assistance, scientific research achievement prediction, training effect evaluation, etc., and continuously optimize and update the model according to the feedback data in actual applications.
[0031] In this embodiment, data augmentation is a process of enriching the text dataset by generating, collecting, and processing additional data to improve the generalization ability and accuracy of the model. The main work of this project is divided into the following two aspects: (1) Data augmentation method based on existing data; The data augmentation method based on existing data mainly uses natural language processing-related technologies for key information extraction, information rewriting, and content augmentation. These technologies can help increase the diversity and richness of data, improve the quality and usability of the dataset. Key information extraction effectively extracts valuable information from the text by using toolkits such as spaCy, Stanford NLP, Hugging Face Transformers, etc., mainly including entity recognition, relationship extraction, and event extraction, etc. When evaluating the quality and consistency of the rewritten data, metrics such as accuracy, recall rate, and F1 score are used to measure, and at the same time, standards such as BLEU and ROUGE are used to evaluate the quality of the generated data. Through manual review and consistency check, it can be ensured that the extracted data meets the expected standards.
[0032] (2) Data augmentation method based on large model self-distillation; The data augmentation method based on large model self-distillation extracts a document from the document pool as a reference example during each generation to guide the generation of new documents, thereby forming a two-stage generation strategy of "selecting existing documents + generating new documents". This process contains the prior knowledge of the large model and can provide theoretical support for the training of the new model.
[0033] For data annotation, data annotation is the automatic and manual processing of data based on governance standards, including quality assessment, cleaning, and format standardization, generating standardized training sets and test sets, and improving the performance of the model and the reliability of data processing. Spot-checking data annotation is to conduct spot-checks and evaluations on the governed data to ensure the effectiveness of data governance. By examining the integrity, consistency, and accuracy of the data, providing feedback to improve data governance strategies, the quality of data annotation can be quantified and evaluated through various metrics and methods to ensure accuracy and consistency. These include recall, precision, and F1-score, which help evaluate the matching degree of the annotation results with the true labels. To ensure consistency, cross-validation and manual review can be used to detect annotation errors and biases. In addition, data review and regular annotator training are also important means to ensure annotation quality. Through these quantitative and qualitative evaluation methods, the quality of data annotation can be effectively controlled, ensuring the reliability and consistency of the final data when used for model training.
[0034] During the model optimization process, model fine-tuning is to use a rich data training set to perform supervised fine-tuning on the model parameters to adapt to the requirements of new tasks. The methods include full-parameter fine-tuning and efficient parameter fine-tuning to improve the accuracy and generalization ability of the model on specific tasks. In this embodiment, a data governance standard is also provided, which is formulated to ensure data quality, security, and compliance, covering aspects such as data collection, storage, access and sharing, data quality, and cleaning, ensuring the standardized and standardized implementation of data governance work. First, it is necessary to conduct a detailed analysis of the distribution of existing data, which includes aspects such as the type and content length of the documents. Based on the evaluation of the existing data, the characteristics and structure of the data can be better understood, providing an important reference for further analysis and application. In response to the challenges of multi-modal data and tabular data that are prevalent in current documents, clear governance standards and specifications should be established. These standards will help standardize the way of data collection, collation, and presentation, improve the readability and interpretability of the data, and thus more effectively respond to the diverse data forms in the document, ensuring the quality and consistency of the data and enabling it to be efficiently used by large language models.
[0035] In this embodiment, during the model training phase, the basic training model consists of two parts: the Pangu MoE NLP large model and the Pangu multi-modal large model, which can not only ensure the computing power of the model but also maintain a reasonable balance in resource consumption. In terms of design, both the NLP large model and the multi-modal large model adopt the Transformer architecture. The multi-modal version of the Pangu large model is also of the 8B scale. It uses a 300M Vision Transformer pre-trained model as the visual encoder and the 7B-parameter Pangu large model as the language model base. While achieving efficient inference, the 8B version can show good response speed and accuracy when dealing with complex tasks. Through the method of the MoE network architecture, only a small part of the network needs to be modified and updated to adapt to new knowledge and tasks, while the original network structure can remain frozen. This method involves incrementally adding new sub-networks and training a small number of additional parameters instead of retraining the entire large language model, making the whole process more efficient and reducing the computing cost and the time required for retraining. This flexibility is particularly advantageous in the education scenario because the curriculum and knowledge requirements are constantly changing. When there are major model changes and a large amount of data updates, educators can easily integrate updates or new information into the model without making large-scale changes to the entire system, thus more flexibly and effectively keeping the educational content updated and ensuring that learners obtain the latest and most relevant information.
[0036] Although the embodiments of the present invention have been shown and described, for those of ordinary skill in the art, it can be understood that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and spirit of the present invention. The scope of the present invention is defined by the appended claims and their equivalents.
Claims
1. A three-stage iterative large model training method for teaching, scientific research and training scenarios, characterized in that: It includes the following three stages: Data collection and preprocessing stage: Collect multi-source data in teaching, scientific research and training scenarios, including but not limited to teaching materials, scientific research results, training records, academic papers, etc.; clean the collected data to remove noise data, duplicate data and invalid data; annotate the cleaned data, including text classification, entity recognition, relationship extraction, etc. The annotation work adopts a combination of manual annotation and automatic annotation; Model training stage: constructing an initial large model, which has a multi-layer neural network structure; dividing the preprocessed data into a training set, a validation set, and a test set according to a certain ratio; using the training set to perform the first stage of training on the initial large model, using the gradient descent algorithm to optimize the model parameters, and adjusting the training parameters according to the feedback of the validation set during the training process; After completing the first phase of training, the model is evaluated. If the evaluation result does not meet the preset standard, the second phase of training will be entered. In the second phase of training, the reinforcement learning mechanism is introduced. By setting the reward function and combining transfer learning, the pre-trained model knowledge in other related fields is transferred to the current model. After the second phase of training, the model is evaluated again. If it still does not meet the preset standards, it enters the third phase of training. The third phase of training adopts adversarial training to build an adversarial network, allowing the generator and discriminator to play adversarial games. Model optimization and application stage: Optimize the model after three-stage training, including model compression, pruning and other operations, and apply the optimized model to teaching, scientific research and training scenarios, such as intelligent teaching assistance, scientific research results prediction, training effect evaluation, etc., and continuously optimize and update the model based on feedback data from actual applications.
2. According to claim 1, a three-stage iterative large model training method for teaching, scientific research and training scenarios is characterized in that: The data collection step also includes encrypting the data, and the data preprocessing step includes data expansion. The data expansion expands the scale and diversity of the data set through data enhancement techniques such as text rewriting and data synthesis.
3. According to claim 1, a three-stage iterative large model training method for teaching, scientific research and training scenarios is characterized in that: In the first stage of training, an adaptive learning rate adjustment strategy is adopted to dynamically adjust the learning rate according to the changes in the loss function during training.
4. According to claim 1, a three-stage iterative large model training method for teaching, scientific research and training scenarios is characterized in that: In the second stage of training, the design of the reward function is based on the actual needs and evaluation indicators in the teaching, scientific research and training scenarios, including but not limited to the improvement of teaching effects, the innovation of scientific research results, and the training achievement rate.
5. According to claim 1, a three-stage iterative large model training method for teaching, scientific research and training scenarios is characterized in that: In the third stage of training, the generator and discriminator of the adversarial network adopt different neural network architectures and are trained alternately during the training process.
6. The three-stage iterative large model training method for teaching, scientific research and training scenarios according to claim 1 is characterized in that: In the model optimization step, quantization technology is used to convert model parameters from high-precision data types to low-precision data types to reduce the storage and calculation costs of the model.
7. The three-stage iterative large model training method for teaching, scientific research and training scenarios according to claim 1 is characterized in that: When the model is applied to teaching, scientific research and training scenarios, a user feedback interface is set up to facilitate users to evaluate the model's output results and make improvement suggestions.
8. A system for implementing the three-stage iterative large model training method for teaching, scientific research and training scenarios as described in any one of claims 1 to 7, characterized in that: include: Data collection module, used to collect multi-source data in teaching, scientific research and training scenarios; Data preprocessing module, used to clean, label and encrypt the collected data; Model training module, used to perform the three-stage model training process; Model evaluation module, used to evaluate the model during training; Model optimization module, used to optimize the trained model; The application module is used to apply the optimized model to teaching, scientific research and training scenarios and receive user feedback.
9. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, it implements the three-stage iterative large model training method for teaching, scientific research and training scenarios as described in any one of claims 1-7.
10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, it implements the three-stage iterative large model training method for teaching, scientific research and training scenarios as described in any one of claims 1 to 7.
Citation Information
Cited By
Construction method of breast cancer patient nutrition tutoring large model
CN121306427A
Construction method of breast cancer patient nutrition guidance large model
CN121306427B