Preference disturbance reinforcement learning data generation method oriented to teaching, research and training scenes
Through the preferred disturbance of reinforced learning data generation method for teaching, scientific research and training scenarios, the problem of insufficient data quality and diversity in the existing technology is solved, the high adaptability and generalization capabilities of the model are achieved, and the security and coverage of data are ensured.
Patent Information
- Application Number
- CN202510188871.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-20
- Publication Date
- 2025-05-30
AI Technical Summary
The existing technology is difficult to generate high-quality and diverse training data, resulting in limited generalization capabilities of machine learning models in practical applications. Especially in teaching, scientific research and training scenarios, existing data sets lack representation and coverage, and data privacy and security issues are prominent.
Adopt the preferred disturbance reinforcement learning data generation method for teaching, scientific research and training scenarios, and through steps such as data collection and evaluation, preference disturbance design, strategy exploration and learning, feedback loops and optimization and iteration, a wider decision space and diversified data samples are generated to ensure the quality and security of the data.
It significantly improves the adaptability and generalization capabilities of the model, and the generated data sets are more diverse and cover-oriented, adapting to the needs of different learners, while ensuring the safe and compliant use of data, reducing the cost of data collection and labeling, and promoting educational equity.
Smart Images

Figure FT_1
Abstract
Description
Technical Field
[0001] The present invention relates to the fields of computer science and artificial intelligence technology, and particularly to a preference perturbation reinforcement learning data generation method for teaching, scientific research, training and cultivation scenarios. Background Art
[0002] Data-driven personalized learning has become the key to improving teaching quality and efficiency. Machine learning models, especially deep learning models, play a crucial role in this process. However, the performance and accuracy of these models largely depend on the quality and diversity of the training data. Traditional data generation strategies, such as rule-based methods or simple data augmentation techniques, often fail to fully capture the complex behaviors and preferences of learners, resulting in limited generalization ability of the models in practical applications.
[0003] In addition, with the growing demand for educational personalization, the need for data generation strategies that can adapt to different learner characteristics and needs is becoming increasingly urgent. Existing datasets often lack sufficient representativeness and cannot cover a wide range of learning scenarios and individual differences, which limits the application of models in diverse teaching environments. At the same time, data privacy and security issues also pose challenges to data collection and use, especially when it comes to educational data involving minors.
[0004] To address these issues, researchers have begun to explore reinforcement learning methods that incorporate user feedback, aiming to generate higher-quality training data by simulating the decision-making processes and preferences of users. This strategy, known as Reinforcement Learning from Human Feedback (RLHF), uses the judgments and preferences of users as signals to guide the model to learn and progress in complex decision-making problems. However, how to effectively collect and utilize user feedback, and how to design mechanisms to guide the model to explore the data space, remain the main challenges in this field.
[0005] Therefore, the present invention aims at a preference perturbation reinforcement learning data generation strategy for teaching, scientific research, training and cultivation scenarios. By introducing perturbations to adjust the decision boundary of the model, it enables the model to explore and learn a wider decision space, while ensuring that the generated data not only meets the quantity requirements, but also reflects the complexity and diversity of the real world in terms of quality. Summary of the Invention
[0006] The present invention provides a preference perturbation reinforcement learning data generation method for teaching, scientific research, training and cultivation scenarios, aiming to solve the limitations in the prior art, improve the adaptability and generalization ability of the model, and provide more accurate and personalized data support for the field of teaching, scientific research, training and cultivation.
[0007] The present invention provides a preference perturbation reinforcement learning data generation method for teaching, scientific research, training scenarios, comprising the following steps: Data collection and evaluation, collecting user feedback data and evaluating the current performance state of the model to define the required data types and quality; Preference perturbation design, designing a perturbation mechanism to adjust the model decision boundary according to user feedback; Policy exploration and learning, the model uses perturbation to explore the new data space and evaluates the exploration results through reinforcement learning; Feedback loop, feeding back the new discoveries and learning results of the model to user evaluation to form an iterative optimization loop; Optimization and iteration, optimizing the model learning strategy and perturbation strategy according to the feedback results to improve the model performance and generalization ability.
[0008] Optionally, the perturbation mechanism generates diverse data samples by simulating different preferences of multiple users, and then uses the diverse data samples to train the model. The perturbation mechanism perturbs the input data by adding noise or performing random transformations on the samples to simulate different preference changes.
[0009] Optionally, the method further includes requirement analysis and preference definition, clarifying the research objectives and data requirements, and defining the preferences of the researcher.
[0010] Optionally, the strategy further includes initial data collection, obtaining basic data from an existing database or on-site collection based on preliminary requirements.
[0011] Optionally, the method further includes evaluation and optimization of synthetic data, evaluating the quality of the generated synthetic data and optimizing the data generation strategy according to the evaluation results.
[0012] Optionally, the data sources in the data collection stage include collecting operation data of online learning platforms, collecting offline classroom performance data, and integrating students' learning history data. The data collected in the data collection stage is detailedly annotated, and the annotation content includes students' emotional states, difficulty perception of learning tasks, and students' homework situations.
[0013] Optionally, it further includes a model construction stage, and the model construction stage is divided into: Model architecture selection: For student data with sequential data characteristics, recurrent neural networks and their variants are selected; for multi-modal data, an architecture combining a convolutional neural network and a fully connected layer is adopted. The convolutional neural network is used to extract local features of different modal data, and then feature fusion and information interaction are carried out through the fully connected layer. Or a model using the Transformer architecture is used to achieve efficient parallel processing and deep fusion of different modal data through the multi-head attention mechanism. Personalized feature embedding: A dedicated embedding layer is designed in the model to process the personalized features of students. Features such as students' learning styles, hobbies, and knowledge levels are converted into low-dimensional vectors and embedded into the input layer or hidden layer of the model, enabling the model to directly utilize this personalized information during the learning process; for categorical data, such as learning style categories, one-hot encoding or an embedding lookup table is used for embedding to ensure accurate representation of category information; for continuous data, such as learning time, score, etc., after normalization, it is directly used as input or embedded through a linear layer.
[0014] Optionally, after the model architecture is selected, the generated training data is used for model training. It is pre-trained in combination with the labeled data in supervised learning, and then reinforcement learning is used to optimize the strategy in a simulated environment, and the training direction is continuously adjusted according to the real-time feedback of the task.
[0015] On the other hand, this solution also provides a model suitable for training by this method, and this model can generate customized learning materials according to the personalized needs of students.
[0016] Compared with the prior art, the beneficial effects of the present invention are: By combining user feedback and reinforcement learning techniques, the generation of personalized learning materials is realized, significantly improving the effects of education and scientific research training. This strategy enhances the generalization ability of the model by introducing a perturbation mechanism, making the generated data set more diverse and comprehensive, thus adapting to the needs of different learners. In addition, this strategy also includes an automated data cleaning and verification process, improving the accuracy and usability of data, while dynamically adapting to changes in education policies and teaching methods, ensuring the timeliness and relevance of data analysis. The application of privacy protection technology ensures the secure and compliant use of data. By reducing the dependence on the collection of a large amount of real data, this strategy reduces the costs of data collection and annotation, while promoting educational fairness. The continuous feedback loop and model optimization mechanism enable this strategy to continuously adapt to the changes of learners and provide a continuously improved learning experience Description of the Drawings
[0017] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the accompanying drawings required for the description of the embodiments or the prior art. Obviously, the accompanying drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0018] Figure 1 This is the flowchart of the method of the present invention.
[0019] The following will clearly and completely describe the technical solutions of the present invention in conjunction with the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the protection scope of the present invention.
[0020] Unless otherwise specified, the raw materials used in the present invention are all conventional products purchased from the market. Embodiment
[0021] A preference perturbation reinforcement learning data generation method for teaching, scientific research, training and cultivation scenarios, comprising the following steps: 1. Data preprocessing 1.1. Data cleaning Conduct a comprehensive review of the original teaching, scientific research, training and cultivation data, and use statistical analysis methods to identify noise data such as learning time records that deviate significantly from the normal range (for example, outliers with thousands of hours of learning duration), illogical score entries (such as full marks exceeding the normal score setting), etc., and delete them.
[0022] For missing values, for numerical data such as students' homework scores, if the missing proportion is small, use mean filling; if there are a small number of missing points in time series data, perform linear interpolation filling based on the trends of the previous and subsequent data. When text data is missing, refer to the common expressions of the same type of data or use a pre-trained language model for reasonable completion.
[0023] Correct obvious data entry errors, such as typos and format irregularities in text data, unit conversion errors in numerical data, etc., to ensure the preliminary accuracy of the data.
[0024] 1.2. Data standardization For numerical features, such as students' exam scores, use the normalization method to map them to the 0 - 1 interval. The formula is, where is the original value, and and are the minimum and maximum values of this feature respectively.
[0025] For features such as learning time and operation frequency, uniformly convert them into standard time units (such as minutes) or standardized counting units to eliminate the impact of different dimensions on model training, enabling the model to learn based on an equal scale when processing different features.
[0026] 1.3. Feature extraction From text data (such as students' homework, papers, feedback comments), use natural language processing techniques (such as bag-of-words model, TF-IDF algorithm, or deep learning word embedding model) to extract keywords, themes, and semantic features to characterize the key points of students' knowledge mastery and problem areas.
[0027] Based on learning behavior data (such as course access order, learning material stay time), construct learning pattern features, and use sequence pattern mining algorithms to discover students' learning habit paths and preferred learning processes, providing support for the subsequent model to understand students' behavior logic.
[0028] 2. Data collection and evaluation 2.1. Collection from diverse data sources Online learning platform: Deeply capture students' operation logs, including the start time of each class viewing, pause times and durations, fast forward and rewind behaviors, and the clicking of interactive buttons (such as like, collect, question, participate in voting, etc.), accurately reflecting students' immediate responses and interest points to knowledge. At the same time, completely collect the details of homework submissions, covering the submission time, correct or wrong situation, and the number of revisions, for analyzing students' learning progress and knowledge mastery difficulties.
[0029] Offline classroom: Arrange observers or use intelligent classroom devices to record the content, frequency, and initiative of students' participation in classroom discussions, the depth, breadth, and pertinence of questions, the role performance in group cooperation activities (such as leader, coordinator, recorder, contributor, etc.), and the degree of concentration reflected by eye focus and body language (such as positive gestures like raising hands, nodding, or negative states like lowering the head, being distracted).
[0030] Learning history: Integrate past performance data in the school's teaching management system, subdivide and statistically analyze it by subject, semester, and knowledge module, draw students' knowledge growth curves, and locate long-term weak links; sort out students' competition awards and scientific research project practice experiences they have participated in, and explore their potential advantages and areas of expertise.
[0031] 2.2. Questionnaire surveys and psychological tests Learning Style Questionnaire: Regularly distribute questionnaires such as the VARK questionnaire to accurately determine whether students are visual learners (preferring to learn through pictures, charts, and videos), auditory learners (good at absorbing knowledge through lectures and audio materials), reading / writing learners (having a high acceptance of reading and writing exercises), or kinesthetic learners (liking to learn through hands-on operations, experiments, and simulation experiences) through multi-dimensional questions, providing guidance for the design of personalized learning materials.
[0032] Interest and Hobby Survey: Design a questionnaire with interest options covering multiple disciplines and fields (such as science, art, sports, humanities and social sciences) to understand the directions of students' extracurricular preferences, so that the model can associate teaching content with interest points and enhance learning attraction.
[0033] Psychological State Test: Periodically use professional psychological assessment scales (such as learning anxiety scale, learning motivation scale, self-rating scale of emotional state) to monitor the emotional fluctuations, stress load, and changes in learning motivation of students during the learning process, ensure that data generation takes into account students' psychological factors, and promote humanized teaching.
[0034] 2.3, Fine-grained Data Annotation Emotional State Annotation: Real-time analyze the facial expressions in the online learning video images of students, use a deep learning-based facial expression recognition model to identify emotions such as confusion (frowning, blank look in the eyes), interest (wide eyes, upturned corners of the mouth), boredom (pouting, wandering eyes), etc., and at the same time combine text data (such as emotional words in learning platform comments and homework annotations) for comprehensive annotation.
[0035] Perception of Learning Task Difficulty: Based on the knowledge levels set in the teaching syllabus, combined with the actual answering accuracy rate and completion time of students, label learning tasks as easy (high accuracy rate, short completion time), moderate (medium accuracy rate, regular time), difficult (low accuracy rate, long struggle or giving up), to assist the model in understanding the ability boundaries of students.
[0036] Deep Annotation of Homework Answers: In addition to the regular right / wrong judgment, for wrong questions, analyze in detail the root causes of the errors, whether it is concept confusion, careless calculation, reasoning logic error, or misunderstanding of the question, providing accurate knowledge diagnosis basis for the model and facilitating the formulation of personalized tutoring strategies.
[0037] 2.4, Define Data Types and Quality According to Requirements Based on the teaching syllabus, research project objectives, and training themes, clarify the modalities of the required data. For example, text materials are needed to explain theoretical knowledge, video materials are required to demonstrate experimental operations, and image resources are indispensable for assisting art appreciation. At the same time, according to the model training stage and application scenarios, stipulate data quality standards. For example, academic literature data requires authoritative sources and accurate citations, and exercise data needs to cover different difficulty levels and have accurate answers to ensure that the data can effectively drive the model to learn.
[0038] 3. Preference perturbation design 3.1 Parameter Perturbation For key parameters in deep learning models, such as the weight matrix and bias term of the neural network, set a fine perturbation range based on the current performance and stability of the model. For models in the initial training stage and with poor stability, randomly perturb the weight parameters in a narrow range (such as ±0.01) with a very small step size (such as 0.0001), observe the changes in model output, and avoid model collapse due to excessive perturbation. As the model converges and stability improves, appropriately expand the perturbation step size (such as 0.001) and range (such as ±0.1) to explore a wider parameter space and tap into potential optimization directions.
[0039] Adjust model parameters that are closely related to preference decisions. For example, in a learning resource recommendation model, if it is found that students have recently frequently browsed a certain type of specific topic material, appropriately increase the disturbance amplitude of the recommendation weight parameter corresponding to the topic, and guide the model to adapt more flexibly to students' interest drift in subsequent decisions.
[0040] 3.2 Input Data Perturbation Text data: Use the natural language processing toolkit to perform synonym replacement on teaching texts, case analyses, academic paper abstracts, etc., and ensure that the replaced text is fluent and natural based on the synonym dictionary and semantic understanding without changing the core semantics; implement word order adjustment, follow grammatical rules and logical coherence principles, reorganize sentence structures, simulate different writing styles or explanation orders, and broaden students' adaptability to knowledge expression.
[0041] Image data: Use the image processing library to rotate (small angles, such as ±10°), scale (between 0.8 and 1.2 times), and crop (keep key subjects and remove redundant background) experimental demonstration images, teaching illustrations, and artwork images to simulate different observation perspectives, display key points, and size adaptation requirements, thereby enhancing the diversity of students' visual perception.
[0042] Audio data: With the help of audio processing software, the volume of lecture recordings, language learning audio, music appreciation materials, etc. is adjusted (within the range of ±3dB) to simulate the listening effects under different environmental volume levels; the pitch is fine-tuned (±0.5 semitones) to change the timbre characteristics, such as making the explanation voice more friendly or professional, enriching students' auditory experience.
[0043] 4. Model construction 4.1 Selection of Appropriate Model Architecture Sequence data processing: When dealing with students' learning process data (such as the course learning trajectory recorded in chronological order and the sequence of assignment completion), recurrent neural networks (RNNs) and their variants are preferably used. If the data sequence is long and there are long-term dependencies, such as tracking the knowledge accumulation process of students throughout the semester, the long short-term memory network (LSTM) can effectively capture and remember distant information with its unique gating mechanism, preventing the problem of gradient vanishing or explosion; if the data is relatively short, has high real-time requirements, and limited computing resources, the gated recurrent unit (GRU) simplifies the structure, improves the operation efficiency while maintaining a certain sequence processing ability, and quickly responds to students' short-term learning dynamics.
[0044] Multi-modal data fusion: For complex scenarios that simultaneously contain multi-modal data such as learning records (text), classroom performance videos, and psychological assessment audio, an architecture combining a convolutional neural network (CNN) and a fully connected layer is adopted. The CNN is responsible for efficiently extracting local features in images and video frames, generating feature maps through the sliding convolution of convolutional kernels; the fully connected layer then splices and fuses the features output by the CNN with text and audio features to achieve cross-modal information interaction. Or use a model based on the Transformer architecture, which uses the multi-head attention mechanism to process different modal data in parallel, dynamically allocates attention weights, deeply fuses multi-modal information, and accurately grasps the overall learning state of students.
[0045] 4.2, Personalized feature embedding Design a dedicated embedding layer: Embed the personalized feature vectors of students in the front end or hidden layer of the model, map features such as learning styles (e.g., visual encoded as [1,0,0,0], auditory as [0,1,0,0], etc.), interests and hobbies (transforming interest fields through word embedding or one-hot encoding), and knowledge levels (quantifying according to grade segments and knowledge mastery modules) into low-dimensional vectors, and integrate them into the model calculation process, so that personality factors are incorporated into the model decision-making from the beginning.
[0046] Embedding of categorical and continuous data: For categorical data, such as learning style categories, if the number of categories is small, one-hot encoding is used to intuitively present category differences; if the number of categories is large, an embedding lookup table is used to convert the high-dimensional sparse one-hot encoding into a low-dimensional dense vector to reduce the computational burden. For continuous data, such as learning time and assignment scores, first normalize them to the 0 - 1 interval, and then directly input them into the model or embed them through a linear transformation layer to adjust the feature scale and the adaptability of the model input.
[0047] 5. Strategy exploration and learning 5.1. Simulation environment construction Build a highly realistic virtual environment for teaching, research, and training, covering teaching scenarios (simulating classroom lectures, group discussions, online Q&A, etc.), research scenarios (virtual laboratory operations, literature retrieval and analysis, project collaboration platforms), and training scenarios (skill practice fields, case review rooms, assessment stations). Set a variety of tasks and challenges in each scenario. For example, in the teaching scenario, arrange tasks such as knowledge point explanation, classroom questioning, and after-class homework correction; in the research scenario, set problems such as experimental design, data collection and analysis, and paper writing; in the training scenario, arrange challenges such as skill operation, troubleshooting, and solution optimization, providing sufficient space for the model to display and explore.
[0048] 5.2. Reinforcement Learning Interaction The model loads perturbed data and enters the simulation environment. It formulates action strategies according to reinforcement learning algorithms (such as the Proximal Policy Optimization algorithm based on policy gradients or the Deep Q-Network algorithm based on value functions). During the interaction with the environment, every time the model executes an action (such as recommending a learning resource, giving a research suggestion, or designing a training process), the environment immediately feedbacks reward or punishment signals. Rewards are quantitatively set according to indicators such as the improvement of students' knowledge mastery (measured by test scores and skill assessment results), the enhancement of learning interest (reflected by learning activity and active participation), the output of research results (the quantity and quality of published papers and patent applications), and the improvement of training satisfaction (student evaluations and pass rates); punishments correspond to negative situations such as an increase in errors (rising error rates and increasing numbers of experimental failures) and a decline in learning enthusiasm (shortening of learning time and decreasing participation frequency). The model continuously adjusts its own policy parameters based on these feedbacks, optimizes action selections, and gradually learns to make optimal decisions in complex teaching, research, and training scenarios.
[0049] 6. Model Training 6.1 Model Selection Make a precise selection according to the characteristics of specific teaching, research, and training tasks. For tasks involving complex knowledge reasoning and policy optimization, such as scientific research project planning and high-end technology training, deep reinforcement learning models are the first choice. Taking the Deep Q-Network (DQN) and its extensions as an example, it can use the powerful representation ability of deep neural networks to approximate the optimal policy and handle high-dimensional and non-linear decision spaces; for scenarios that need to quickly adapt to new students and new scenarios and frequently switch task modes, such as online education platforms receiving a large number of students with different backgrounds and short-term vocational training facing changing demands, meta-learning models are more suitable. It pre-trains on multiple similar small tasks, extracts general learning patterns, enables rapid knowledge transfer when facing new tasks, converges quickly, and saves training time and resources.
[0050] 6.2 Hyperparameter Tuning Using cross-validation technology, the dataset is divided into training set, validation set and test set (the common ratio is 6:2:2). The model is trained on the training set, and the model performance is evaluated on the validation set. The optimal configuration is found by repeatedly testing different hyperparameter combinations. For the key hyperparameter of learning rate, a logarithmic scale search is used, such as trying from 0.0001 to 0.1 with a 10-fold step size; for batch size, combined with the dataset size and memory limitations, it is screened from 16 to 128 with a 2-fold step size; the number of network layers and neurons is based on the complexity of the task, gradually expanding from a simple structure (such as 2-3 layers, 32-64 neurons per layer) to a complex structure (5-10 layers, 128-512 neurons per layer), observing the changes in indicators such as the accuracy and loss function value of the model on the validation set, and determining the best hyperparameter combination.
[0051] 6.3 Training Execution First, use labeled data for supervised learning pre-training to allow the model to quickly grasp basic data patterns and knowledge features. For example, use exercises with labeled answers to train the model to identify knowledge weaknesses, and use correctly classified academic literature to train the model to understand subject classification. Then switch to reinforcement learning mode, place the model in a simulation environment, and continuously interact with the environment according to the rules established in the strategy exploration and learning stages, and continuously receive feedback to adjust the strategy. During the entire training process, monitor model performance indicators (such as accuracy, loss, and reward value) in real time, and adjust the training direction in a timely manner according to the indicator trend. If the model is found to be stuck in a bottleneck in learning a certain type of knowledge, increase the supply of relevant data or adjust the perturbation strategy in a targeted manner to guide the model to break through the difficulties.
[0052] 7. Model Evaluation 7.1 Verification From the historically accumulated teaching, scientific research and training data, according to the data distribution characteristics similar to the training set (such as subject ratio, student level, task type distribution), a certain proportion (usually 20% - 30%) is randomly selected as the validation set. Run the trained model on the validation set to observe the generalization ability of the model for unseen data, focusing on indicators such as accuracy, recall rate, and F1 score. If the performance of the model is found to be significantly reduced on the validation set, such as the accuracy rate is more than 10% lower than the training set, it may indicate that the model has an overfitting problem, and it is necessary to retrospectively adjust the model complexity, increase the diversity of data perturbations, or optimize the training strategy.
[0053] 7.2 Testing Build a dedicated test set that covers a wider range of educational research, training scenarios, and diverse student types to ensure a comprehensive examination of the model's actual application performance. The data sources for the test set can include actual cases from different schools, regions, and subject fields, and are carefully screened and annotated. Place the model on the test set for final performance evaluation. In addition to conventional metrics, in combination with the actual needs of educational research, training, introduce metrics such as the depth of knowledge mastery (evaluated through complex knowledge quizzes and project practice results), and the stimulation of innovative thinking (observing the frequency and quality of students' new ideas and methods proposed in scientific research and creation) to comprehensively judge the effectiveness of the model in real scenarios.
[0054] 7.3. Calculation of Performance Metrics Accuracy: In classification tasks (such as determining students' preference types for learning resources and grading their knowledge mastery levels), calculate the proportion of correctly classified samples in the total number of samples. The formula is , where TP is true positive, TN is true negative, FP is false positive, and FN is false negative. It is commonly used to measure the accuracy of the model's judgment. For example, when predicting whether a student has mastered a certain knowledge point, a high accuracy indicates a high probability that the model's judgment is correct.
[0055] Precision and Recall: Precision refers to the proportion of truly positive samples among the samples predicted as positive. The formula is ; Recall is the proportion of correctly predicted positive samples among all positive samples. The formula is . When evaluating the model's capture of students' preferences, precision measures the accuracy of the model when predicting that students like a certain learning resource, and recall measures the model's ability to correctly identify students who truly like the resource. The combination of the two can more comprehensively reflect the model's performance.
[0056] Mean Squared Error: In regression tasks (such as predicting the improvement in students' grades and their learning time requirements), calculate the average of the squares of the differences between the predicted values and the true values. The formula is , where is the true value, is the predicted value, and is the number of samples. It is used to measure the deviation degree between the model's predicted value and the actual value. The smaller the MSE, the more accurate the model's prediction.
[0057] Preference Matching Rate: Obtain students' true preferences, such as their preferences for different subject contents and teaching methods, through questionnaires, actual observations, etc. Compare the preferences reflected in the data generated by the model with them and calculate the matching ratio. The formula is the number of matching preference samples / the total number of samples, which is used to intuitively evaluate the degree of fit between the data generated by the model and students' actual preferences.
[0058] Preference Consistency: Divide the student data into a training set and a test set, or collect student preference data at different time points, and calculate the consistency of the model's preference predictions for the same student on these different data sets or time points. Methods such as Kendall's Coefficient of Concordance can be used to measure, which is used to ensure the stability of the model's judgment of student preferences in different situations.
[0059] Learning Curve: Show the changes in the model's performance (such as accuracy, loss, etc.) during the training process as the amount of training data or the number of training rounds increases. By plotting the learning curve, the learning efficiency of the model can be intuitively judged. If the curve converges quickly, it means the model can effectively use the data for learning; if the curve fluctuates greatly or converges slowly, it may imply problems with the data generation strategy, model architecture, or training method.
[0060] 8. Feedback Loop Result Feedback Collection: Feed back the exploration and learning results of the model in the simulation environment, including the generated data, recommended strategies, decisions made, etc., to professional user evaluators.
[0061] As described above, it is only a preferred specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention, according to the technical solution of the present invention and its inventive concept, makes equivalent substitutions or changes, and should be covered by the protection scope of the present invention.
Claims
1. A method for generating preference perturbation reinforcement learning data for teaching, scientific research and training scenarios, characterized in that: The following steps are involved: Data collection and evaluation, collecting user feedback data and evaluating the current performance status of the model to define the required data type and quality; Preference perturbation design: designing perturbation mechanisms to adjust model decision boundaries based on user feedback; Strategy exploration and learning: the model uses perturbations to explore new data spaces and evaluates the exploration results through reinforcement learning; Feedback loop, which feeds new model findings and learning results back to users for evaluation, forming an iterative optimization loop; Optimization and iteration: optimize the model learning strategy and perturbation strategy based on feedback results to improve model performance and generalization ability.
2. According to the method for generating preference perturbation reinforcement learning data for teaching, scientific research and training scenarios in claim 1, it is characterized in that: The perturbation mechanism generates diversified data samples by simulating different preferences of multiple users, and then uses the diversified data samples to train the model. The perturbation mechanism simulates different preference changes by perturbing the input data, adding noise or randomly transforming the samples.
3. According to the method for generating preference perturbation reinforcement learning data for teaching, scientific research and training scenarios in claim 1, it is characterized in that: The method also includes requirements analysis and preference definition, clarifying research objectives and data requirements, and defining researcher preferences.
4. According to the method for generating preference perturbation reinforcement learning data for teaching, scientific research and training scenarios in claim 1, it is characterized in that: The strategy also includes initial data collection, obtaining basic data from existing databases or field collection based on preliminary needs.
5. According to the method for generating preference perturbation reinforcement learning data for teaching, scientific research and training scenarios in claim 1, it is characterized in that: The method also includes evaluating and optimizing the synthetic data, performing quality evaluation on the generated synthetic data, and optimizing the data generation strategy based on the evaluation results.
6. According to the method for generating preference perturbation reinforcement learning data for teaching, scientific research and training scenarios in claim 1, it is characterized in that: The data sources in the data collection phase include collecting operation data of the online learning platform, collecting offline classroom performance data, and integrating students' learning history data. The data collection phase conducts detailed annotations on the collected data, and the annotation content includes students' emotional state, perceived difficulty of learning tasks, and students' homework status.
7. According to the method for generating preference perturbation reinforcement learning data for teaching, scientific research and training scenarios in claim 1, it is characterized in that: It also includes a model building phase, which is divided into: Model architecture selection: For student data with sequence data characteristics, recurrent neural networks and their variants are selected; for multimodal data, a combination of convolutional neural networks and fully connected layers is used. Convolutional neural networks are used to extract local features of data of different modalities, and then feature fusion and information interaction are performed through fully connected layers, or a model with a Transformer architecture is used to achieve efficient parallel processing and deep fusion of data of different modalities through a multi-head attention mechanism. Personalized feature embedding: A special embedding layer is designed in the model to process the personalized features of students, converting features such as students' learning style, interests, hobbies, and knowledge level into low-dimensional vectors, and embedding them into the input layer or hidden layer of the model, so that the model can directly use this personalized information during the learning process; For categorical data, such as learning style categories, one-hot encoding or embedding lookup tables are used for embedding to ensure accurate representation of category information; For continuous data, such as learning time and grade scores, they are directly used as input after normalization or embedded through a linear layer.
8. According to claim 7, a method for generating preference perturbation reinforcement learning data for teaching, scientific research and training scenarios is characterized in that: After the model architecture is selected, the generated training data is used for model training, and pre-training is performed in combination with the labeled data in supervised learning. Reinforcement learning is then used to optimize the strategy in a simulated environment, and the training direction is continuously adjusted according to real-time feedback from the task.
9. A model trained using the method according to any one of claims 1 to 8, characterized in that: The model is able to generate customized learning materials based on students' individual needs.
10. A computer-readable storage medium storing a computer program of the method according to any one of claims 1 to 8.