Education large model tuning method and device based on dynamic optimization
Through the dynamic optimization method of multimodal data fusion and adaptive reward mechanism, the problems of single data, rigid rewards and architecture limitations in the educational model are solved, and personalized learning and real-time response are improved.
Patent Information
- Application Number
- CN202510428043.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-07
- Publication Date
- 2025-07-25
AI Technical Summary
The existing educational models have single data dimensions, rigid reward mechanisms, limited model architecture and insufficient real-time performance, resulting in the inability to effectively identify the learning efficiency caused by students' fatigue or mood swings, and the calculation overhead is high and the response delay is significant.
Multimodal data fusion technology is used to process physiological data, learning behavior data and emotional feedback data using heterogeneous models (GLM-4 model, graph neural network, convolutional neural network), and determine the Q value through reinforcement learning models and adaptive reward functions, and dynamically optimize the recommendation strategy.
It improves the personalized ability and real-time response efficiency of educational models, can better adapt to individual differences between different students, and improves learning effect and computing efficiency.
Smart Images

Figure CN120373624A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology, and in particular, to an optimization method and device for an educational large model based on dynamic optimization. Background Art
[0002] With the rapid development of artificial intelligence technology, educational large models have shown significant potential in fields such as teaching assistance, personalized recommendation, and learning assessment. However, the following key problems still exist in the existing technology for the dynamic optimization method of educational large models:
[0003] Single data dimension: Existing systems mainly rely on students' learning behavior data (such as progress, answering accuracy) and emotional feedback (such as comment sentiment analysis), lacking real-time monitoring and analysis of students' physiological states (such as concentration, fatigue), and thus, unable to identify problems of decreased learning efficiency caused by fatigue or mood swings;
[0004] Rigid reward mechanism: Traditional dynamic optimization methods mostly adopt linear reward functions with fixed weights (such as weighted learning progress and answering performance), which are difficult to adapt to the individual differences of different students and affect the learning effect;
[0005] Limitations of model architecture: Existing educational large models are mostly based on a single architecture (such as Transformer), and there are problems of low information fusion efficiency when dealing with multi-modal data (such as text, images, time-series physiological signals);
[0006] Insufficient real-time performance and high resource consumption: Traditional optimization methods rely on periodic updates of global model parameters, with high computational overhead and significant response delays. Summary of the Invention
[0007] The purpose of the present invention is to provide an optimization method and device for an educational large model based on dynamic optimization, in order to improve the personalized ability and real-time response efficiency of the educational large model, aiming at the deficiencies in the above-mentioned existing technology.
[0008] To achieve the above purpose, the technical solutions adopted in the embodiments of the present application are as follows:
[0009] First aspect, an embodiment of the present application provides an optimization method for an educational large model based on dynamic optimization, including: obtaining multimodal data; the multimodal data includes physiological data, learning behavior data, and emotional feedback data; inputting the multimodal data into a heterogeneous model in the educational large model, using the heterogeneous model to process the multimodal data to obtain multimodal feature vectors, and fusing the multimodal feature vectors to obtain fused feature vectors; the heterogeneous model includes a GLM-4 model, a graph neural network, and a convolutional neural network; inputting the fused feature vectors into a reinforcement learning model in the educational large model, determining a reward through a reward function, determining a Q value according to the reward, and determining a recommendation strategy based on the Q value.
[0010] In one implementation, after obtaining the multimodal data, the method further includes: performing normalization processing on the multimodal data.
[0011] In one implementation, the performing normalization processing on the multimodal data includes: performing Fourier transform on the physiological data to obtain a frequency-domain signal, extracting features based on the frequency-domain signal, and normalizing the features; normalizing the learning behavior data and the emotional feedback data.
[0012] In one implementation, the using the heterogeneous model to process the multimodal data to obtain multimodal feature vectors includes: using the GLM-4 model to process the learning behavior data and the emotional feedback data, and outputting semantic feature vectors; using the graph neural network to process the interaction data in the emotional feedback data, and outputting a learning pattern; using the convolutional neural network to process the physiological data, and outputting an attention distribution feature; the multimodal feature vectors include the semantic feature vectors, the learning pattern, and the attention distribution feature.
[0013] In one implementation, the fusing the multimodal feature vectors to obtain fused feature vectors includes: using an attention mechanism to dynamically assign weights to the multimodal feature vectors to generate the fused feature vectors.
[0014] In one implementation, the reward function is as follows:
[0015] R t = w1ΔP t + w2ΔE t + w3ΔB t
[0016] where R tas a reward; w1, w2, and w3 are the weights of physiological data, learning behavior data, and emotional feedback data respectively, and w1, w2, and w3 are calculated in real time using a lightweight neural network based on historical physiological data, historical learning behavior data, and historical emotional feedback data; ΔP t is the change in learning behavior data; ΔE t is the change in emotional feedback data; B t is the change in physiological data.
[0017] In one implementation, the Q-value calculation formula is as follows:
[0018] Q(s t , a t ) ← Q(s t , a t ) + α(R final + γ max a Q(s t+1 , a) - Q(s t , a t ))
[0019] where Q(s t , a t ) is the Q-value of executing action a in state s, that is, the expected cumulative reward; α is the learning rate, γ is the discount factor, and α and γ are dynamically adjusted according to student adaptability; R final = tanh(R t ); s t is the state at time t, s t = [P norm , E norm , B norm ; a t is the action at time t.
[0020] In one implementation, after determining the recommendation strategy based on the Q-value, the method further includes: collecting the latest physiological data, the latest learning behavior data, and the latest emotional feedback data; inputting the latest physiological data, the latest learning behavior data, and the latest emotional feedback data into an evaluation model, using the evaluation model to evaluate the learning effect of the student, and outputting an evaluation result; updating the model parameters of the reinforcement learning model and the weights of the reward function according to the evaluation result.
[0021] In one implementation, the evaluation model is as follows:
[0022] E t = w4ΔP′ t + w5ΔE′ t + w6ΔB′ t
[0023] Among them, E t is the evaluation result; w4, w5, and w6 are the weights of the latest physiological data, the latest learning behavior data, and the latest emotional feedback data respectively; P t ′ is the change amount of the latest learning behavior data; ΔE t ′ is the change amount of the latest emotional feedback data; ΔB t ′ is the change amount of the latest physiological data.
[0024] In a second aspect, the embodiments of the present application further provide an educational large model tuning device based on dynamic optimization, including: an acquisition module configured to acquire multimodal data; the multimodal data includes physiological data, learning behavior data, and emotional feedback data; a processing module configured to input the multimodal data into a heterogeneous model in the educational large model, process the multimodal data using the heterogeneous model to obtain a multimodal feature vector, and fuse the multimodal feature vector to obtain a fused feature vector; the heterogeneous model includes a GLM-4 model, a graph neural network, and a convolutional neural network; a determination module configured to input the fused feature vector into a reinforcement learning model in the educational large model, determine a reward through a reward function, determine a Q value according to the reward, and determine a recommendation strategy based on the Q value.
[0025] In one implementation manner, the acquisition module is configured to: perform normalization processing on the multimodal data.
[0026] In one implementation manner, the acquisition module is configured to: perform Fourier transform on the physiological data to obtain a frequency domain signal, extract features based on the frequency domain signal, and normalize the features; perform normalization on the learning behavior data and the emotional feedback data.
[0027] In one implementation manner, the processing module is configured to: process the learning behavior data and the emotional feedback data using the GLM-4 model and output a semantic feature vector; process the interaction data in the emotional feedback data using the graph neural network and output a learning pattern; process the physiological data using the convolutional neural network and output an attention distribution feature; the multimodal feature vector includes the semantic feature vector, the learning pattern, and the attention distribution feature.
[0028] In one implementation manner, the processing module is configured to: dynamically allocate weights to the multimodal feature vector using an attention mechanism to generate the fused feature vector.
[0029] In one implementation manner, the determination module is configured that the reward function is as follows:
[0030] R t= w1ΔP t + w2ΔE t + w3ΔB t
[0031] where R t is the reward; w1, w2, and w3 are the weights of physiological data, learning behavior data, and emotional feedback data respectively. w1, w2, and w3 are calculated in real time using a lightweight neural network based on historical physiological data, historical learning behavior data, and historical emotional feedback data; ΔP t is the change in learning behavior data; ΔE t is the change in emotional feedback data; B t is the change in physiological data.
[0032] In one implementation, the determination module is configured such that the Q-value calculation formula is as follows:
[0033] Q(s t , a t ) ← Q(s t , a t ) + α(R final + γmax a Q(s t+1 , a) - Q(s t , a t ))
[0034] where Q(s t , a t ) is the Q-value of performing action a in state s, i.e., the expected cumulative reward; α is the learning rate, γ is the discount factor, and α and γ are dynamically adjusted according to student adaptability; R final = tanh(R t ); s t is the state at time t, s t = [P norm , E norm , B norm ; a t is the action at time t.
[0035] In one implementation, the educational large model tuning device based on dynamic optimization further includes a feedback module, and the feedback module is configured to: collect the latest physiological data, the latest learning behavior data, and the latest emotional feedback data; input the latest physiological data, the latest learning behavior data, and the latest emotional feedback data into an evaluation model, use the evaluation model to evaluate the learning effect of the student, and output an evaluation result; update the model parameters of the reinforcement learning model and the weights of the reward function according to the evaluation result.
[0036] In one embodiment, the feedback module is configured such that the evaluation model is as follows:
[0037] E t = w4ΔP′ t + w5ΔE′ t + w6ΔB′ t
[0038] where E t is the evaluation result; w4, w5, and w6 are the weights of the latest physiological data, the latest learning behavior data, and the latest emotional feedback data respectively; P t ′ is the change amount of the latest learning behavior data; ΔE t ′ is the change amount of the latest emotional feedback data; ΔB t ′ is the change amount of the latest physiological data.
[0039] In a third aspect, an embodiment of the present application provides a computer device, including: a processor, a storage medium, and a bus. The storage medium stores program instructions executable by the processor. When the computer device runs, the processor communicates with the storage medium through the bus, and the processor executes the program instructions to perform the steps of any of the above methods.
[0040] In a fourth aspect, an embodiment of the present application provides a non-volatile computer-readable storage medium, on which a computer program is stored. When the computer program is run by a processor, it performs the steps of any of the above methods.
[0041] The beneficial effects of the present application are as follows: First, multi-modal data is acquired; the multi-modal data includes physiological data, learning behavior data, and emotional feedback data; Second, the multi-modal data is input into a heterogeneous model in an educational large model, and the heterogeneous model is used to process the multi-modal data to obtain multi-modal feature vectors, and the multi-modal feature vectors are fused to obtain fused feature vectors; the heterogeneous model includes a GLM-4 model, a graph neural network, and a convolutional neural network; Finally, the fused feature vectors are input into a reinforcement learning model in the educational large model, a reward is determined through a reward function, a Q value is determined according to the reward, and a recommendation strategy is determined based on the Q value. In this way, through a dynamic optimization method combining physiological data, an adaptive reward mechanism, and heterogeneous model fusion, the personalized ability and real-time response efficiency of the educational large model can be comprehensively improved. Description of the Drawings
[0042] To more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings required for the embodiments will be briefly introduced below. It should be understood that the following drawings only show some embodiments of the present invention and should not be regarded as limiting the scope. For those of ordinary skill in the art, without creative efforts, other related drawings can be obtained based on these drawings.
[0043] Figure 1 It is a schematic flowchart of a method for optimizing an educational large model based on dynamic optimization provided by an embodiment of the present application;
[0044] Figure 2 It is a schematic flowchart of a method for optimizing an educational large model based on dynamic optimization provided by an embodiment of the present application;
[0045] Figure 3 It is a schematic flowchart of a method for optimizing an educational large model based on dynamic optimization provided by an embodiment of the present application;
[0046] Figure 4 It is a schematic flowchart of a method for optimizing an educational large model based on dynamic optimization provided by an embodiment of the present application;
[0047] Figure 5 It is a schematic structural diagram of a device for optimizing an educational large model based on dynamic optimization provided by an embodiment of the present application;
[0048] Figure 6 It is a schematic structural diagram of a computer device provided by an embodiment of the present application. Detailed implementation manners
[0049] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention.
[0050] Therefore, the following detailed description of the embodiments of the present application provided in the drawings is not intended to limit the scope of the present application claimed, but merely represents selected embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the scope of protection of the present application.
[0051] In the description of the present application, it should be noted that if terms such as "upper", "lower", etc. are used to indicate the orientation or positional relationship, it is based on the orientation or positional relationship shown in the drawings, or the orientation or positional relationship in which the product of this application is usually placed during use. This is only for the convenience of describing the present application and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation. Therefore, it should not be construed as a limitation to the present application.
[0052] In addition, terms such as "first", "second", etc. in the description and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects, and do not necessarily need to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances, so that the embodiments of the present invention described here can be implemented in an order other than those illustrated or described here. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device comprising a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.
[0053] It should be noted that, without conflict, the features in the embodiments of the present application can be combined with each other.
[0054] Figure 1 The flowchart of a method for optimizing an educational large model based on dynamic optimization provided by an embodiment of the present application; as Figure 1 shown, the method includes the following steps 110 to 130:
[0055] Step 110, obtain multimodal data.
[0056] Among them, the multimodal data includes physiological data, learning behavior data, and emotional feedback data.
[0057] Physiological data can be used to collect heart rate variability (HRV), eye movement trajectory, and galvanic skin response (GSR) in real time through wearable devices (smart bracelets, eye trackers).
[0058] Learning behavior data can obtain learning progress, answer accuracy rate, interaction frequency, etc. from an online education platform.
[0059] Emotional feedback data can analyze student comments and speech content through natural language processing tools (such as BERT) to generate an emotional score (-1, 1).
[0060] In actual operation, due to the different dimensions of multimodal data, the multimodal data can be standardized; specifically, after the above step 110, the method further includes:
[0061] Normalize multi-modal data.
[0062] Among them, the normalization of physiological data, learning behavior data, and emotional feedback data is different; specifically, as Figure 2 shown, the above steps may further include the following steps 210 and 220:
[0063] Step 210: Perform Fourier transform on physiological data to obtain a frequency-domain signal, extract features based on the frequency-domain signal, and normalize the features.
[0064] Among them, the Fourier transform is to transform the signal in the time domain (i.e., the time domain) into the signal in the frequency domain (i.e., the frequency domain). As the domain is different, the understanding angle of the same thing will change accordingly. Therefore, some places that are difficult to process in the time domain can be processed more simply in the frequency domain.
[0065] In actual operation, the Fourier transform in this step can adopt the fast Fourier transform. The fast Fourier transform (FFT) is an optimized algorithm of the discrete Fourier transform (DFT) and is used to convert the time-domain signal into the frequency-domain signal. The DFT analyzes the frequency components of the signal by decomposing the signal into sine waves and cosine waves of different frequencies. The FFT reduces the computational complexity from O(N 2 ) to O(NlogN) by using the symmetry and periodicity of the DFT, significantly improving the computational efficiency.
[0066] The feature extraction in this step can be achieved through the following process: First, divide the frequency band according to the physiological meaning; second, extract the features of each frequency band; here, the features include: low frequency (LF, 0.04 - 0.15Hz): reflecting the mixed activities of the sympathetic and parasympathetic nerves, related to stress and cognitive load; high frequency (HF, 0.15 - 0.4Hz): reflecting the activity of the parasympathetic nerve (vagus nerve), synchronized with the respiratory rhythm; LF / HF ratio: characterizing the balance state of the autonomic nervous system (a high ratio indicates active sympathetic nerve).
[0067] The normalization in this step can map features with different dimensions to a unified scale (such as: 0, 1).
[0068] Step 220: Normalize the learning behavior data and emotional feedback data.
[0069] Among them, the learning behavior data can be normalized to 0, 1 using the following formula (1):
[0070]
[0071] where P t represents the current learning progress value, indicating the completion degree of the student in a certain course or knowledge point; Pmin Represents the minimum value of the learning progress, usually set to the starting state of the course; P max Represents the maximum value of the learning progress, indicating the state of complete mastery of the course or knowledge point.
[0072] Map the sentiment feedback data to 0, 1 using the following formula (2):
[0073]
[0074] where, E t Represents the original sentiment score, calculated by a sentiment analysis tool (such as BERT, sentiment dictionary).
[0075] Step 120: Input the multimodal data into the heterogeneous model in the education large model, process the multimodal data using the heterogeneous model to obtain a multimodal feature vector, and fuse the multimodal feature vectors to obtain a fused feature vector.
[0076] Among them, the heterogeneous model includes the GLM-4 model, graph neural network, and convolutional neural network.
[0077] The GLM-4 model, graph neural network, and convolutional neural network process different data; specifically, as Figure 3 shown, the above step 120 may further include the following steps 310 to 330:
[0078] Step 310: Process the learning behavior data and sentiment feedback data using the GLM-4 model and output a semantic feature vector.
[0079] Among them, the GLM-4 model is based on deep learning and aims to learn and analyze a large amount of text data, so as to be able to generate high-quality text responses. It uses the Transformer architecture, which enables the model to handle long sequence data well. The GLM-4 model continuously absorbs knowledge from a large amount of text corpora. It learns the relationships between words, the structures of sentences, and semantic information, etc.; for example: it can understand the meaning of the word "apple" in different contexts, which can be either a kind of fruit or the brand of a technology company. Through learning a large amount of text, the GLM-4 model can predict the next possible word or generate logical and grammatically correct sentences.
[0080] In this step, the GLM-4 model can process text data, such as learning behavior data and sentiment feedback data, and then generate a semantic feature vector.
[0081] Step 320: Process the interaction data in the sentiment feedback data using the graph neural network and output a learning pattern.
[0082] Among them, a graph neural network (GNN) refers to an algorithm that uses a neural network to learn graph-structured data, extract and discover features and patterns in the graph-structured data, and meet the requirements of graph learning tasks such as clustering, classification, prediction, segmentation, and generation.
[0083] In actual operation, this step can be achieved through the following process: First, construct a student social interaction graph, where the nodes are defined as follows: Each student corresponds to a node, and the node features include: Static attributes: Subject preferences, historical grades, activity levels (such as high / medium / low); Dynamic behaviors: Posting frequency in the last week, contribution to collaborative tasks (such as the number of code submissions, the amount of document editing); The edges are defined as follows: Edge definition: Interaction types: Discussion replies, collaborative editing, @ mentions, likes, etc.; Edge weights: Basic weight: Interaction frequency (such as the number of replies); Enhanced weight: Combining sentiment analysis (positive interaction weight +1, negative interaction weight -1); Second, each node aggregates the information of its direct neighbors, combines the aggregated neighbor information with its own features, and updates the node; Third, through multiple cascaded GNN layers (for example: 2-3 layers), gradually capture high-order social relationships; High-order social relationships include: The first layer: Learn the local patterns of direct neighbors (for example: groups of students with frequent interactions); The second layer: Learn the global patterns across communities (for example: interdisciplinary collaboration groups); Then, through the output of the last layer of the GNN, obtain the low-dimensional embedding vector of each student (such as 64 dimensions); Finally, use unsupervised clustering algorithms (such as K-means, DBSCAN) to group the embedding vectors and identify group learning patterns; Group learning patterns include: Collaborative groups: High collaboration weights, frequent interdisciplinary interactions; Discussion-dominated groups: High posting volume, high centrality (such as a large PageRank value); Isolated students: Low connectivity, and the sum of edge weights is close to 0.
[0084] Step 330: Process the physiological data using a convolutional neural network and output the attention distribution features.
[0085] Among them, the multi-modal feature vector includes the above semantic feature vector, learning pattern, and attention distribution feature.
[0086] Convolutional Neural Networks (CNNs) are a type of feedforward neural network with convolutional computations and a deep structure. They are one of the representative algorithms in deep learning. Convolutional neural networks have the ability of representation learning and can perform shift-invariant classification on input information according to their hierarchical structure. Therefore, they are also known as "Shift-Invariant Artificial Neural Networks (SIANN)".
[0087] In this step, the convolutional neural network can analyze physiological data, such as eye movement trajectory images, and then identify the attention distribution characteristics.
[0088] Furthermore, the step of "fusing multi-modal feature vectors to obtain a fused feature vector" in the above step 120 can further include the following steps:
[0089] Use the attention mechanism to dynamically assign weights to the multi-modal feature vectors to generate a fused feature vector.
[0090] Among them, the attention mechanism enables the neural network to have the ability to focus on a subset of its inputs (or features): select specific inputs. Attention can be applied to any type of input regardless of its shape. In the case of limited computing power, the attention mechanism is a resource allocation scheme that is one of the main means to solve the problem of information overload, and allocates computing resources to more important tasks.
[0091] In actual operation, this step can dynamically assign weights to the semantic feature vectors, learning patterns, and attention distribution characteristics in the multi-modal feature vectors to generate a fused feature vector.
[0092] Step 130: Input the fused feature vector into the reinforcement learning model in the education large model, determine the reward through the reward function, determine the Q value according to the reward, and determine the recommendation strategy based on the Q value.
[0093] Among them, Reinforcement Learning (RL) is a machine learning method. The basic framework of reinforcement learning is the Markov decision process, which allows an agent to learn the optimal policy through trial and error in the interaction with the environment. The agent executes actions in the environment and receives feedback, i.e., rewards, based on the results of the actions. These reward signals guide the agent to adjust its policy to maximize the long-term cumulative reward.
[0094] Reward calculation is the core mechanism of Reinforcement Learning (RL), and the relationship between the two is inseparable. It is specifically reflected in the following aspects:
[0095] 1. Rewards are the driving signals of reinforcement learning
[0096] Feedback mechanism: Rewards are the immediate evaluation of the agent's behavior by the environment, telling the agent whether the current action is good or bad. For example, in an educational large model, positive rewards are given when students answer questions correctly, and negative rewards are given when they are wrong.
[0097] Goal orientation: The ultimate goal of reinforcement learning is to maximize the cumulative reward (i.e., long-term benefits), and reward calculation directly defines the goal that the agent needs to optimize.
[0098] 2. The design of the reward function determines the learning direction
[0099] Task definition: The design of the reward function implies the specific requirements of the task. For example, if the reward function only focuses on the question answering accuracy rate, the agent may ignore the emotional state (such as when students are under too much pressure); if the reward function incorporates learning progress, emotional feedback, and physiological state, the agent needs to balance multiple goals and generate a more user-friendly strategy.
[0100] Dynamic adaptability: Adaptive reward functions (such as dynamic weight adjustment) can adapt to the individual differences of different students. For example, increase the weight of emotional feedback for anxious students and increase the weight of learning progress for students who are lagging behind.
[0101] 3. Reward calculation affects the stability and efficiency of the algorithm
[0102] Sparsity and density:
[0103] Sparse rewards (such as only giving rewards when the task is completed) may lead to slow learning and require a large amount of exploration.
[0104] Dense rewards (such as giving small rewards at each step) can accelerate learning, but it is necessary to avoid misleading signals (such as when the reward function is not designed reasonably, the agent may "cheat").
[0105] Nonlinear optimization:
[0106] Introducing a non - linear function (e.g., Tanh) to compress the reward value range (e.g., to (-1, 1)) can suppress the interference of extreme values and improve the stability of policy updates. For example, when the student's answer correct rate suddenly drops sharply (extreme negative reward), the Tanh function can smooth the process and avoid drastic policy fluctuations.
[0107] 4. Reinforcement learning algorithms rely on rewards for policy optimization
[0108] Q - learning: Select the optimal action by updating the Q - value (expected cumulative reward). The reward is directly used to calculate the update of the Q - value;
[0109] Policy gradient: Directly optimize the policy parameters through gradient ascent to maximize the expected cumulative reward. The reward value is used to weight the gradient direction of policy updates;
[0110] Deep reinforcement learning (e.g., DQN): Combine a deep neural network to fit the Q - function, and the reward signal guides the update of network parameters.
[0111] 5. Reward Shaping improves learning efficiency
[0112] Intermediate reward design: Accelerate the learning process of the agent by adding auxiliary rewards (such as encouraging exploration of unknown states);
[0113] Adaptive mechanism: In an educational large - model, dynamically adjusting the reward weight (such as reducing the learning progress weight according to the student's fatigue) is a form of reward shaping, making the policy more in line with actual needs.
[0114] In this step, the reward function is as shown in formula (3) below:
[0115] R t = w1ΔP t + w2ΔE t + w3ΔB t (3)
[0116] Where R t is the reward; w1, w2, and w3 are the weights of physiological data, learning behavior data, and emotional feedback data respectively. w1, w2, and w3 are calculated in real - time using a lightweight neural network based on historical physiological data, historical learning behavior data, and historical emotional feedback data; ΔP t is the change in learning behavior data; ΔE t is the change in emotional feedback data; B t is the change in physiological data.
[0117] The Q - value calculation formula is as shown in formula (4) below:
[0118] Q(st , a t ) ← Q(s t , a t ) + α(R final + γmax a Q(s t+1 , a) - Q(s t , a t )) (4)
[0119] Among them, Q(s t , a t ) is the Q - value of executing action a in state s, that is, the expected cumulative reward; α is the learning rate, γ is the discount factor, and α and γ are dynamically adjusted according to student adaptability; R final = tanh(R t ); s t is the state at time t, s t = [P norm , E norm , B norm ; a t is the action at time t.
[0120] In actual operation, this step can dynamically optimize the educational large - model through a reinforcement learning model; based on student feedback, the reinforcement learning model uses the Q - learning algorithm (Q - value calculation formula) to continuously optimize the recommendation strategy.
[0121] In actual operation, after each learning cycle ends, the reinforcement learning model can also collect the latest learning data, evaluate the learning effect of students, and update the recommendation strategy according to the evaluation results; specifically, as Figure 4 shown, a method for tuning an educational large - model based on dynamic optimization provided by an embodiment of the present application may further include the following steps 410 to step 430:
[0122] Step 410: Collect the latest physiological data, the latest learning behavior data, and the latest emotional feedback data.
[0123] Among them, the latest physiological data is the most recent physiological data; the same applies to the latest learning behavior data and the latest emotional feedback data.
[0124] This step is the same as step 110 above and will not be elaborated here.
[0125] Step 420: Input the latest physiological data, the latest learning behavior data, and the latest emotional feedback data into the evaluation model, use the evaluation model to evaluate the learning effect of students, and output the evaluation result.
[0126] Among them, the evaluation model is shown in the following formula (5):
[0127] Et = w4ΔP′ t + w5ΔE′ t + w6ΔB′ t (5)
[0128] E t is the evaluation result; w4, w5, and w6 are the weights of the latest physiological data, the latest learning behavior data, and the latest emotional feedback data respectively; P t ′ is the change amount of the latest learning behavior data; ΔE t ′ is the change amount of the latest emotional feedback data; ΔB t ′ is the change amount of the latest physiological data.
[0129] Step 430: Update the model parameters of the reinforcement learning model and the weights of the reward function according to the evaluation result.
[0130] Among them, through the Figures 1 to 4 steps, a closed loop of "acquisition → training → recommendation → feedback" can be formed.
[0131] The method for optimizing an educational large model based on dynamic optimization provided by the embodiments of the present application first obtains multimodal data; the multimodal data includes physiological data, learning behavior data, and emotional feedback data; secondly, inputs the multimodal data into a heterogeneous model in the educational large model, processes the multimodal data by using the heterogeneous model to obtain a multimodal feature vector, and fuses the multimodal feature vectors to obtain a fused feature vector; the heterogeneous model includes a GLM-4 model, a graph neural network, and a convolutional neural network; finally, inputs the fused feature vector into the reinforcement learning model in the educational large model, determines the reward through the reward function, determines the Q value according to the reward, and determines the recommendation strategy based on the Q value. In this way, through a dynamic optimization method combining physiological data, an adaptive reward mechanism, and heterogeneous model fusion, the personalized ability and real-time response efficiency of the educational large model can be comprehensively improved.
[0132] After introducing the method for optimizing an educational large model based on dynamic optimization of the exemplary embodiments of the present disclosure, next, refer to Figure 5 to describe the device 500 for optimizing an educational large model based on dynamic optimization of the exemplary embodiments of the present disclosure.
[0133] Refer to Figure 5, an educational large model tuning device 500 based on dynamic optimization, includes: an acquisition module 510 configured to acquire multimodal data; the multimodal data includes physiological data, learning behavior data, and emotional feedback data; a processing module 520 configured to input the multimodal data into heterogeneous models in the educational large model, process the multimodal data using the heterogeneous models to obtain multimodal feature vectors, and fuse the multimodal feature vectors to obtain fused feature vectors; the heterogeneous models include a GLM-4 model, a graph neural network, and a convolutional neural network; a determination module 530 configured to input the fused feature vectors into a reinforcement learning model in the educational large model, determine a reward through a reward function, determine a Q value based on the reward, and determine a recommendation strategy based on the Q value.
[0134] In one implementation, the acquisition module 510 is configured to: perform normalization processing on the multimodal data.
[0135] In one implementation, the acquisition module 510 is configured to: perform a Fourier transform on the physiological data to obtain a frequency-domain signal, extract features based on the frequency-domain signal, and normalize the features; perform normalization on the learning behavior data and the emotional feedback data.
[0136] In one implementation, the processing module 520 is configured to: process the learning behavior data and the emotional feedback data using the GLM-4 model and output semantic feature vectors; process the interaction data in the emotional feedback data using the graph neural network and output a learning pattern; process the physiological data using the convolutional neural network and output an attention distribution feature; the multimodal feature vectors include semantic feature vectors, a learning pattern, and an attention distribution feature.
[0137] In one implementation, the processing module 520 is configured to: dynamically assign weights to the multimodal feature vectors using an attention mechanism to generate fused feature vectors.
[0138] In one implementation, the determination module 530 is configured that the reward function is as follows:
[0139] R t = w1ΔP t + w2ΔE t + w3ΔB t
[0140] where, R t is the reward; w1, w2, and w3 are the weights of the physiological data, the learning behavior data, and the emotional feedback data respectively, and w1, w2, and w3 are calculated in real time using a lightweight neural network based on historical physiological data, historical learning behavior data, and historical emotional feedback data; ΔP t is the change amount of the learning behavior data; ΔE t is the change amount of the emotional feedback data; Bt is the change in physiological data.
[0141] In one embodiment, the determination module 530 is configured such that the Q-value calculation formula is as follows:
[0142] Q(s t , a t ) ← Q(s t , a t ) + α(R final + γmax a Q(s t+1 , a) - Q(s t , a t ))
[0143] where Q(s t , a t ) is the Q-value of performing action a in state s, i.e., the expected cumulative reward; α is the learning rate, γ is the discount factor, and α and γ are dynamically adjusted according to student adaptability; R final = tanh(R t ); s t is the state at time t, s t = [P norm , E norm , B norm ; a t is the action at time t.
[0144] In one embodiment, the educational large model tuning device 500 based on dynamic optimization further includes a feedback module, and the feedback module is configured to: collect the latest physiological data, the latest learning behavior data, and the latest emotional feedback data; input the latest physiological data, the latest learning behavior data, and the latest emotional feedback data into the evaluation model, use the evaluation model to evaluate the learning effect of the student, and output the evaluation result; update the model parameters of the reinforcement learning model and the weights of the reward function according to the evaluation result.
[0145] In one embodiment, the feedback module is configured such that the evaluation model is as follows:
[0146] E t = w4ΔP′ t + w5ΔE′ t + w6ΔB t ′ where E t is the evaluation result; w4, w5, and w6 are the weights of the latest physiological data, the latest learning behavior data, and the latest emotional feedback data respectively; P t ′ is the change in the latest learning behavior data; ΔE t ′ is the change in the latest emotional feedback data; ΔB t ′ is the change in the latest physiological data.
[0147] The above device is used to execute the method provided in the foregoing embodiment, and its implementation principle and technical effects are similar, which will not be elaborated here.
[0148] The above modules may be one or more integrated circuits configured to implement the above method. For example: one or more Application Specific Integrated Circuits (ASICs), or, one or more microprocessors, or, one or more Field Programmable Gate Arrays (FPGAs), etc. Again, when a certain module above is implemented in the form of a processing element dispatching program code, the processing element may be a general-purpose processor, such as a Central Processing Unit (CPU) or other processors that can call program code. Again, these modules may be integrated together and implemented in the form of a system-on-a-chip (SOC).
[0149] Figure 6 It is a schematic diagram of the computer device provided by the embodiment of the present application. This device may be integrated into the terminal device or the chip of the terminal device, and the terminal may be a computing device with data processing capabilities.
[0150] This device includes: a processor 601, a storage medium 602, and a bus 603.
[0151] The storage medium 602 stores program instructions executable by the processor 601. When the computer device 600 runs, the processor 601 communicates with the storage medium 602 through the bus 603, and the processor 601 executes the program instructions to execute the above method embodiment. The specific implementation manners and technical effects are similar and will not be elaborated here.
[0152] Optionally, the present invention further provides a program product, such as a computer-readable storage medium, including a program, which is used to execute the above method embodiment when executed by a processor.
[0153] In several embodiments provided by the present invention, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces. The indirect couplings or communication connections of the devices or units can be in electrical, mechanical or other forms.
[0154] The units described as separate components may or may not be physically separated. The components shown as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0155] In addition, in each embodiment of the present invention, the functional units can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above integrated units can be implemented in the form of hardware or in the form of hardware plus software functional units.
[0156] The above integrated units implemented in the form of software functional units can be stored in a computer-readable storage medium. The above software functional units stored in a storage medium include several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) or a processor (English: processor) to execute some steps of the methods described in various embodiments of the present invention. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (English: Read-Only Memory, abbreviated as: ROM), random access memories (English: Random Access Memory, abbreviated as: RAM), magnetic disks or optical discs that can store program codes.
[0157] The above is only the specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention can easily think of changes or substitutions, which should all be covered by the protection scope of the present invention. Therefore, the protection scope of the present invention should be subject to the protection scope of the claims.
Claims
1. A method for optimizing an educational large model based on dynamic optimization, characterized in that Including: Obtain multimodal data; the multimodal data includes physiological data, learning behavior data, and emotional feedback data; Input the multimodal data into a heterogeneous model in an educational large model, use the heterogeneous model to process the multimodal data to obtain multimodal feature vectors, and fuse the multimodal feature vectors to obtain fused feature vectors; the heterogeneous model includes a GLM-4 model, a graph neural network, and a convolutional neural network; Input the fused feature vectors into a reinforcement learning model in the educational large model, determine a reward through a reward function, determine a Q value based on the reward, and determine a recommendation strategy based on the Q value.
2. The method according to claim 1, wherein After obtaining the multimodal data, the method further includes: Perform normalization processing on the multimodal data.
3. The method according to claim 2, wherein The normalization processing of the multimodal data includes: Perform Fourier transform on the physiological data to obtain a frequency-domain signal, extract features based on the frequency-domain signal, and normalize the features; Normalize the learning behavior data and the emotional feedback data.
4. The method according to claim 1, wherein The processing of the multimodal data by the heterogeneous model to obtain multimodal feature vectors includes: Use the GLM-4 model to process the learning behavior data and the emotional feedback data and output semantic feature vectors; Use the graph neural network to process the interaction data in the emotional feedback data and output a learning pattern; Use the convolutional neural network to process the physiological data and output an attention distribution feature; the multimodal feature vectors include the semantic feature vectors, the learning pattern, and the attention distribution feature.
5. The method according to claim 1, wherein The fusion of the multimodal feature vectors to obtain fused feature vectors includes: Use an attention mechanism to dynamically assign weights to the multimodal feature vectors to generate the fused feature vectors.
6. The method according to claim 1, wherein The reward function is as follows: R t = w1ΔP t + w2ΔE t + w3ΔB t Among them, R t is the reward; w1, w2, and w3 are the weights of physiological data, learning behavior data, and emotional feedback data respectively, and w1, w2, and w3 are calculated in real time using a lightweight neural network based on historical physiological data, historical learning behavior data, and historical emotional feedback data; ΔP t is the change in learning behavior data; ΔE t is the change in emotional feedback data; B t is the change in physiological data.
7. The method according to claim 1, wherein The calculation formula of the Q value is as follows: Q(s t , a t ) ← Q(s t , a t ) + α(R final + γmax a Q(s t+1 , a) - Q(s t , a t )) Among them, Q(s t , a t ) is the Q-value of executing action a in state s, that is, the expected cumulative reward; α is the learning rate, γ is the discount factor, and α and γ are dynamically adjusted according to student adaptability; R final = tanh(R t ); s t is the state at time t, s t = [P norm , E norm , B norm ; a t is the action at time t.
8. The method according to claim 1, characterized in that, After determining the recommendation strategy based on the Q value, the method further includes: Collect the latest physiological data, the latest learning behavior data, and the latest emotional feedback data; Input the latest physiological data, the latest learning behavior data, and the latest emotional feedback data into an evaluation model, use the evaluation model to evaluate the learning effect of the student, and output an evaluation result; Update the model parameters of the reinforcement learning model and the weights of the reward function according to the evaluation result.
9. The method according to claim 8, wherein The evaluation model is as follows: E t = w4ΔP′ t + w5ΔE′ t + w6ΔB t ′ where E t is the evaluation result; w4, w5, and w6 are the weights of the latest physiological data, the latest learning behavior data, and the latest emotional feedback data respectively; P′ t is the change in the latest learning behavior data; ΔE t ′ is the change in the latest emotional feedback data; ΔB t ′ is the change in the latest physiological data.
10. An educational large model tuning device based on dynamic optimization, characterized in that, Including: An acquisition module configured to acquire multimodal data; the multimodal data includes physiological data, learning behavior data, and emotional feedback data; A processing module configured to input the multimodal data into a heterogeneous model in an educational large model, use the heterogeneous model to process the multimodal data to obtain multimodal feature vectors, and fuse the multimodal feature vectors to obtain fused feature vectors; the heterogeneous model includes a GLM-4 model, a graph neural network, and a convolutional neural network; A determination module, configured to input the fusion feature vector into a reinforcement learning model in the large education model, determine a reward through a reward function, determine a Q-value according to the reward, and determine a recommendation strategy based on the Q-value.
Citation Information
Cited By
Smart classroom adaptability regulation and control method based on multi-mode and large language model
CN121637403A