Online training education informatization teaching method based on big data

Through the information-based teaching methods of online training education based on big data, and the use of technical means such as deep Q networks, the problem that traditional online education platforms cannot achieve personalized recommendations and dynamic teaching strategies adjustments have been solved, and efficient and personalized teaching resource allocation and learning effect have been achieved.

CN120218508AActive Publication Date: 2025-06-27BEIJING LANHAI DAXIN TECH CO LTD

Patent Information

Application Number
CN202510286610.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-12
Publication Date
2025-06-27
Estimated Expiration
2045-03-12

AI Technical Summary

Technical Problem

Traditional online education platforms cannot fully utilize students' real-time data for personalized recommendations, the teaching content is not accurate and lagging, and the lack of intelligent teaching strategy adjustment mechanisms, resulting in limited improvement in learning effects.

Method used

The information-based teaching method of online training education based on big data is adopted. By collecting multiple data sources, constructing student state vectors, establishing state action models, modeling teaching recommendation problems into Markov decision-making processes, and using deep Q networks for reinforcement learning, dynamically adjusting teaching strategies, and real-time recommendations are achieved.

Benefits of technology

It has achieved personalized teaching recommendations that accurately match students' needs, dynamically adjust teaching strategies to improve learning effects, enhance teaching adaptability and flexibility, and improve students' participation and emotional motivation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120218508A_ABST
    Figure CN120218508A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of online education, and discloses an online training education informatization teaching method based on big data, which comprises the following steps: S1, collecting interaction log data of an online teaching platform, a real-time evaluation system, emotion monitoring equipment and a user, and performing cleaning, standardization, noise filtering and feature extraction on the data; s2, constructing a student state vector according to the data in the step S1 so as to implement dimension reduction processing; and S3, establishing a state action model. The personalized teaching recommendation system based on big data is adopted, learning data of students are collected in real time, a learning path is analyzed and optimized through the big data technology, the technical effect of accurately matching student requirements is achieved, compared with a technical scheme in which the states of the students cannot be responded in time in a traditional education platform, recommendation content can be adjusted in real time, and the recommendation efficiency is improved. The problem that personalized teaching resource configuration is not accurate is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of online education, and specifically to an information-based teaching method for online training education based on big data. Background Art

[0002] With the rapid development of online education, especially the rise of teaching systems based on big data, there are many problems in traditional online education platforms. Traditional online education platforms usually cannot make full use of students' real-time data for personalized recommendations, and there are problems of inaccurate and lagging teaching content push. Traditional teaching methods often rely on fixed curriculum designs and teaching processes, and it is difficult to make immediate adjustments according to students' learning progress, emotional states, and feedback. Therefore, the system cannot configure personalized teaching resources according to students' actual needs, which will affect the improvement of learning effects.

[0003] Existing online education platforms, especially traditional platforms, usually rely on static teaching plans. Their methods cannot flexibly meet the diverse needs of students, making it difficult for teaching content to accurately match the actual situation of students, resulting in waste of resources and decline in learning effects.

[0004] Many existing teaching platforms lack effective feedback mechanisms and cannot respond to students' needs in a timely manner. Students' feedback information cannot directly affect the adjustment of teaching content, resulting in problems encountered by students during the learning process not being solved in a timely manner, making the recommendation of personalized teaching resources unable to keep up with students' actual needs.

[0005] Most traditional education platforms lack an intelligent teaching strategy adjustment mechanism. Although students' learning status and behavior data can be collected and analyzed to a certain extent, in actual applications, how to dynamically adjust teaching strategies according to the data remains a difficult problem. Many platforms still use simple rule systems and static models when processing data, unable to make sufficiently intelligent recommendations and dynamic optimizations, and unable to effectively improve teaching quality and students' learning effects.

[0006] Therefore, those skilled in the art provide an information-based teaching method for online training education based on big data to solve the above-mentioned problems. Summary of the Invention

[0007] Aiming at the deficiencies of the prior art, the present invention provides an information-based teaching method for online training education based on big data to solve the problems raised in the above background art.

[0008] To achieve the above purposes, the present invention is realized through the following technical solutions: An information-based teaching method for online training education based on big data, including the following steps:

[0009] Step S1: Collect the online teaching platform, real-time assessment system, emotion monitoring device, and user interaction log data, and perform cleaning, standardization, noise filtering, and feature extraction on the data;

[0010] Step S2: Construct a student state vector based on the data in Step S1 to perform dimensionality reduction processing;

[0011] Step S3: Establish a state-action model, define the recommended actions corresponding to video resources, question resources, and discussion modules, and design a reward function to calculate the reward score;

[0012] Step S4: Abstract the teaching recommendation problem into a Markov decision process, construct a state set, an action set, a state transition probability function, and an immediate reward function, and use the Bellman optimality principle to solve the optimal state value function;

[0013] Step S5: Use a deep Q-network to perform reinforcement learning training on the Markov decision process in Step S4 to obtain an optimal recommendation strategy by iteratively updating the state-action value function;

[0014] Step S6: Construct a dynamic game model to describe the interaction between the system and the students. The system reward and the student reward are expressed in the form of cumulative discounted rewards and satisfy the Nash equilibrium condition;

[0015] Step S7: Integrate Steps S1 to S6 to form a closed-loop feedback system, and collect feedback data in real time to adjust each parameter to verify the system performance.

[0016] Preferably, the data collected in Step S1 includes the student login data, page click data, video viewing duration data, and answer correct rate data in the online teaching platform, the stage performance data in the real-time assessment system, the emotion score data in the emotion monitoring device, and the discussion record data in the user interaction log.

[0017] Preferably, Step S2 further includes:

[0018] Step 2.1: Obtain the original feature data of each student from Step S1 to form a state vector X. Each row corresponds to a student, and the columns are the learning progress component, the assessment score component, the emotion score component, and the interaction frequency component in sequence;

[0019] Step 2.2: Perform standardization processing on each feature in the state vector X according to the following formula to obtain a normalized data matrix x': x' (ij) =(x (ij) -μ j ) / σ j ,

[0020] where x (ij) is the original value of the i-th student on the j-th feature, and μ jis the mean of the j-th feature among all students, σ j is the standard deviation of the j-th feature among all students,

[0021] i = 1, 2, …, n and j = 1, 2, …, d, where n is the total number of students and d is the dimension of the initial state vector.

[0022] Preferably, the step S2 further includes:

[0023] Step 2.3, calculate the covariance matrix S using the normalized data matrix x′, and the calculation formula is:

[0024]

[0025] where, is the transpose of x′, and n is the total number of students;

[0026] Step 2.4, perform eigenvalue decomposition on the covariance matrix S to solve the eigenvalues λ1, λ2, …, λ d and the corresponding eigenvectors v1, v2, …, v d , satisfying:

[0027] S·v j = λ j ·v j , j = 1, 2,..., d,

[0028] where λ j is the j-th eigenvalue of the covariance matrix S, and v j is the j-th eigenvector of the covariance matrix S;

[0029] Step 2.4, select the eigenvectors corresponding to the top k largest eigenvalues according to the eigenvalue magnitudes to form the projection matrix W;

[0030] Step 2.5, perform dimensionality reduction processing on the normalized data matrix x′ using the projection matrix W to obtain the dimensionality-reduced state vector matrix Z, and the calculation formula is: Z = x′·W,

[0031] where Z is the dimensionality-reduced state vector matrix.

[0032] Preferably, the step S3 further includes:

[0033] Step 3.1, establish a state-action model. According to the student state vector constructed in step S2, define the state vector to describe the learning progress, assessment scores, emotional states, and interaction behaviors of students; define the recommended action set U, and this set contains various recommended actions;

[0034] Step 3.2: Design a reward function, which includes learning progress reward, knowledge mastery reward, and emotional incentive reward. The specific calculation formula is as follows:

[0035] r (p) (ξ q ,v q ) = f (p) (ξ q ,v q ),

[0036] r (k) (ξ q ,v q ) = f (k) (ξ q ,v q ),

[0037] r (e) (ξ q ,v q ) = f (e) (ξ q ,v q ),

[0038] Among them, r (p) (ξ q ,v q ) is the learning progress reward, r (k) (ξ q ,v q ) is the knowledge mastery reward, r (e) (ξ q ,v q ) is the emotional incentive reward, ξ q is the state vector of the student, v q is the recommended action, and f (p) , f (k) , and f (e) are functions for calculating each reward;

[0039] Step 3.3: Calculate the cumulative reward. To evaluate the long-term effect of the recommended action, a discount factor δ is used to perform a decreasing accumulation on the immediate reward to obtain the overall reward. The overall reward function is as follows:

[0040]

[0041] Among them, ω p is the learning progress reward, ω k is the knowledge mastery reward, ω e is the weight coefficient of the emotional incentive reward, δ is the discount factor, T is the time step, and J(π) is the overall reward;

[0042] Step 3.4: The system determines the optimal recommended action v qTo maximize the reward value.

[0043] Preferably, step S4 further includes:

[0044] Step 4.1, model the teaching recommendation problem as a Markov decision process, define the state set S, the action set A, the state transition probability function P, and the immediate reward function R;

[0045] The state set S represents all states of the student;

[0046] The action set A contains all actions recommended by the system for the student;

[0047] The state transition probability function P(s′|s,a) represents the probability of transitioning to state s′ after taking action a in state s;

[0048] The immediate reward function R(s,a) represents the immediate reward obtained after taking action a in state s.

[0049] Preferably, step S4 further includes:

[0050] Step 4.2, use the Bellman optimality principle to solve the optimal state value function V * (s), which represents the maximum cumulative reward for taking the optimal policy in state s. The recursive form of the Bellman equation is:

[0051]

[0052] where V * (s) is the optimal value function of state s, a is the currently available action,

[0053] R(s,a) is the immediate reward obtained after taking action a in state s, s′ is the next state,

[0054] P(s′|s,a) is the probability of transitioning to state s′ after taking action a in state s.

[0055] γ is the discount factor;

[0056] Step 4.3, iteratively update the Bellman equation until the optimal state value function V * (s) converges. During each iterative update, calculate the value of the next state according to the combination of the current state s and action a, and finally obtain the optimal policy and the optimal state value function.

[0057] Preferably, step S5 further includes:

[0058] Step 5.1, after obtaining the optimal state value function through the Markov decision process in step S4, the deep Q-network is used to approximate this optimal value function. Define the Q-network as a neural network, whose input is the state vector s q , and the output is the Q value Q(s q ,a q ) corresponding to each action a q ;

[0059] The Q value represents the expected cumulative reward after executing action a q under the given state s q ;

[0060] Step 5.2, during the training process, by interacting with the environment, update the Q-value function based on the current Q-network. At each update, calculate the actually obtained reward and the Q value of the next state to update the current Q value, and use the Bellman equation to update the Q value. The formula is:

[0061]

[0062] where Q(s q ,a q ) is the Q value of executing action a q under the current state s q ,

[0063] α is the learning rate, r q is the immediate reward obtained after executing action a q under the current state s q ,

[0064] γ is the discount factor, is the maximum Q value under the next state s q+1 .

[0065] Preferably, the step S5 further includes:

[0066] Step 5.3, to increase the stability of the training process, the deep Q-network usually uses a target network to calculate the target Q value. The parameters θ - of the target network are regularly copied from the parameters θ of the main Q-network. The update formula is:

[0067] θ - ←θ. The update of the target network can avoid large fluctuations during the training process;

[0068] Step 5.4, the deep Q-network adjusts the network parameters θ through multiple sets of training cycles to reduce the error between the Q value prediction and the actual reward, and finally enables the Q-network to accurately predict the Q values of each state-action pair, and finally obtains the optimal strategy.

[0069] Preferably, the step S6 further includes:

[0070] Step 6.1: Model the interaction between the system and the student as a dynamic game model to describe the interaction process between the system and the student. The dynamic game model includes the rewards of the system and the rewards of the student.

[0071] Step 6.2: Define the reward functions of the system and the student. The reward functions of the system and the student are represented by the following formulas:

[0072] System reward r s : Measure the effectiveness of the system's recommended actions;

[0073] Student reward r st : Reflect the student's feedback on the recommended actions;

[0074] Step 6.3: Nash equilibrium condition. To ensure the stability of the strategies in the game, the strategies of the system and the student need to satisfy the Nash equilibrium condition. Given the strategies of the other party, neither party has the incentive to unilaterally change its own strategy. The Nash equilibrium condition is expressed as:

[0075]

[0076] where r s is the reward function of the system, r st is the reward function of the student, a s is the recommended action of the system, a st is the reaction action of the student, E[r s (a s ,a st )] is the expected value of the system reward, and E[r st (a s ,a st )] is the expected value of the student reward.

[0077] Step 6.4: In the dynamic game, the system and the student optimize their respective behaviors according to their own reward functions through dynamic game strategies. The strategy π s of the system and the strategy π st of the student can both be optimized through the iterative game process until the Nash equilibrium is reached.

[0078] The present invention provides an information-based teaching method for online training education based on big data, having the following

[0079] Beneficial effects:

[0080] 1. The present invention adopts a personalized teaching recommendation system based on big data. By collecting students' learning data in real time and using big data technology to analyze and optimize the learning path, it achieves the technical effect of accurately matching students' needs. Compared with the technical solutions in traditional educational platforms that cannot respond to students' status in a timely manner, the present invention can adjust the recommended content in real time and solve the problem of inaccurate allocation of personalized teaching resources.

[0081] 2. The present invention utilizes deep reinforcement learning and multi-objective optimization to dynamically adjust teaching strategies, ensuring that the learning resources in the teaching process can be automatically optimized according to students' status and feedback, improving the learning effect. Compared with the static teaching methods in the prior art, the present invention solves the deficiency that fixed teaching strategies cannot cope with the diverse needs of students, thereby greatly enhancing the teaching adaptability and flexibility.

[0082] 3. The present invention realizes the interactive optimization between the system and students through the Markov decision process and dynamic game model, ensuring a high degree of fit between the recommended actions and students' behaviors. Different from the technical solutions of single evaluation systems in traditional teaching platforms, the present invention effectively improves students' participation and emotional motivation through the optimization of the reward function and game strategy, and solves the problems of lagging student feedback and insufficient interactivity. BRIEF DESCRIPTION OF THE DRAWINGS

[0083] Figure 1 is a flowchart of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0084] To enable those skilled in the art to understand the solution of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. Based on the embodiments of the present invention, other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0085] The present invention will be described in detail below with reference to the accompanying drawings:

[0086] Embodiment:

[0087] Please refer to the attached Figure 1 , the embodiment of the present invention provides an information-based teaching method for online training education based on big data, including the following steps:

[0088] Step S1: Collect data from the online teaching platform, real-time assessment system, emotion monitoring device, and user interaction logs, and perform cleaning, standardization, noise filtering, and feature extraction on the data. The data collected in Step S1 includes student login data, page click data, video viewing duration data, and answer accuracy data on the online teaching platform, periodic performance data in the real-time assessment system, emotion score data in the emotion monitoring device, and discussion record data in the user interaction logs;

[0089] Step S2: Construct a student state vector based on the data in Step S1 to perform dimensionality reduction processing;

[0090] Step S3: Establish a state-action model, define the recommended actions corresponding to video resources, question resources, and discussion modules, and design a reward function to calculate the reward score;

[0091] Step S4: Abstract the teaching recommendation problem into a Markov decision process, construct a state set, an action set, a state transition probability function, and an immediate reward function, and use the Bellman optimality principle to solve the optimal state value function;

[0092] Step S5: Use a deep Q-network to perform reinforcement learning training on the Markov decision process in Step S4 to obtain an optimal recommendation strategy by iteratively updating the state-action value function;

[0093] Step S6: Construct a dynamic game model to describe the interaction between the system and the students. The system reward and the student reward are expressed in the form of cumulative discounted rewards and satisfy the Nash equilibrium condition;

[0094] Step S7: Integrate Steps S1 to S6 to form a closed-loop feedback system, and collect feedback data in real time to adjust each parameter to verify the system performance.

[0095] The benefits of data collection and preprocessing in Step S1 are to remove noise and redundant information, ensuring that subsequent analysis can accurately reflect the real learning situation of students. Compared with the single data source of traditional systems, the present invention provides a comprehensive student behavior portrait through multi-source data integration, laying a foundation for subsequent personalized recommendations.

[0096] The benefits of constructing and reducing the dimension of the student state vector in Step S2 are to help the system efficiently process large-scale student data, improve computational performance, and at the same time ensure the response speed and accuracy of the recommendation system. Compared with the traditional method of directly using high-dimensional data for processing, the present invention can balance performance and effect.

[0097] The benefits of designing the state-action model and the reward function in Step S3 are to encourage students' active participation and improve learning effects. At the same time, dynamically adjusting the recommended actions can meet the personalized needs of students and overcome the rigidity problem brought by the fixed teaching strategies in traditional educational platforms.

[0098] Step S4: The application benefits of Markov decision process and Bellman optimality principle achieve dynamic optimization of the student learning process. Compared with the traditional static teaching model, the present invention can dynamically adjust the recommendation strategy according to the students' instant feedback and learning status, significantly improving the adaptability and flexibility of teaching.

[0099] Step S5: The benefits of deep Q-network reinforcement learning training enable the recommendation system to gradually adapt to the personalized learning needs of students and maximize the learning benefits of students. Compared with the traditional recommendation methods based on rules and simple models, the deep Q-network can perform intelligent learning and decision-making based on a large amount of data to provide accurate teaching resource recommendations.

[0100] Step S6: The benefits of dynamic game model and Nash equilibrium optimization enable the system to effectively promote the interaction between the system and students, enhancing the feedback and participation in the learning process. Different from the traditional one-way evaluation system, the present invention adjusts the teaching strategy by considering the students' feedback and behavior, enhancing the interactivity and incentive effect of teaching, and helping students maintain high learning interest and emotional enthusiasm.

[0101] Step S7: The benefits of the integration of the closed-loop feedback system enable the system to update the strategy according to the changes of students, improving the long-term effectiveness of the system. Compared with the mode lacking dynamic feedback in traditional educational platforms, the present invention ensures the continuous fit between the recommended content and the students' needs.

[0102] Step S2 further includes:

[0103] Step 2.1, obtaining the original feature data of each student from Step S1 to form a state vector X, where each row corresponds to a student, and the columns are successively the learning progress component, the assessment score component, the emotion score component, and the interaction frequency component;

[0104] Step 2.2, performing standardization processing on each feature in the state vector X according to the following formula to obtain a normalized data matrix x': x' (ij) =(x (ij) -μ j ) / σ j ,

[0105] where, x (ij) is the original value of the i-th student on the j-th feature, μ j is the mean value of the j-th feature among all students, σ j is the standard deviation of the j-th feature among all students,

[0106] i = 1, 2,..., n and j = 1, 2,..., d, here, n is the total number of students, and d is the dimension of the initial state vector;

[0107] Step 2.3, calculate the covariance matrix S using the normalized data matrix x′, and the calculation formula is:

[0108]

[0109] where, is the transpose of x′, and n is the total number of students;

[0110] Step 2.4, perform eigenvalue decomposition on the covariance matrix S to solve the eigenvalues λ1, λ2, …, λ d and the corresponding eigenvectors v1, v2, …, v d , satisfying:

[0111] S·v j =λ j ·v j , j = 1, 2, ..., d,

[0112] where, λ j is the j-th eigenvalue of the covariance matrix S, and v j is the j-th eigenvector of the covariance matrix S;

[0113] Step 2.4, select the eigenvectors corresponding to the top k largest eigenvalues according to the eigenvalue magnitudes to form the projection matrix W;

[0114] Step 2.5, use the projection matrix W to perform dimensionality reduction on the normalized data matrix x′ to obtain the dimensionality-reduced state vector matrix Z, and the calculation formula is: Z = x′·W,

[0115] where, Z is the dimensionality-reduced state vector matrix.

[0116] The advantage of Step 2.1 is that the data is comprehensive, avoiding missing key information in subsequent processing.

[0117] The benefit of Step 2.2 is to improve data consistency and reduce the influence of dimensions.

[0118] Step 2.3 helps to reveal the internal structure of the data, facilitating subsequent dimensionality reduction processing and enhancing the data expression ability.

[0119] The advantage of Step 2.4 is to highlight key information and reduce noise interference.

[0120] The advantage of Step 2.5 is to reduce the data dimension, maintain the main information, reduce the calculation amount, and facilitate subsequent efficient analysis.

[0121] Step S3 further includes:

[0122] Step 3.1, establish a state-action model, and according to the student state vector constructed in Step S2, define the state vector Describe the learning progress, assessment scores, emotional states, and interaction behaviors of students; define a set of recommended actions U, which includes various recommended actions;

[0123] Step 3.2, design a reward function, which includes learning progress rewards, knowledge mastery rewards, and emotional incentive rewards. The specific calculation formula is:

[0124] r (p) (ξ q ,v q )=f (p) (ξ q ,v q ),

[0125] r (k) (ξ q ,v q )=f (k) (ξ q ,v q ),

[0126] r (e) (ξ q ,v q )=f (e) (ξ q ,v q ),

[0127] Among them, r (p) (ξ q ,v q ) is the learning progress reward, r (k) (ξ q ,v q ) is the knowledge mastery reward, r (e) (ξ q ,v q ) is the emotional incentive reward, ξ q is the state vector of the student, v q is the recommended action, and f (p) , f (k) and f (e) are functions for calculating each reward;

[0128] Step 3.3, calculate the cumulative reward. To evaluate the long-term effect of the recommended action, a discount factor δ is used to perform a decreasing accumulation of the immediate reward to obtain the overall reward. The overall reward function is:

[0129]

[0130] Among them, ω p is the learning progress reward, ω k is the knowledge mastery reward, ω eis the weight coefficient of emotional incentive rewards, δ is the discount factor, T is the time step, and J(π) is the overall reward;

[0131] Step 3.4, the system determines the optimal recommended action v according to the cumulative reward J(π) q to maximize the reward value.

[0132] Benefits of Step 3.1: Establishing a state-action model. Compared with the single-dimensional learning progress tracking in traditional teaching platforms, the present invention can comprehensively consider various aspects of information, comprehensively capture the learning status of students, and provide accurate basis for the subsequent recommendation system.

[0133] Benefits of designing the reward function: Designing learning progress rewards, knowledge mastery rewards, and emotional incentive rewards enables the recommendation strategy to be targeted and effectively improve learning enthusiasm. Compared with traditional reward mechanisms, this step fully considers the multi-dimensional feedback of students to ensure the adaptability of the recommendation strategy.

[0134] Benefits of calculating the cumulative reward: By using the discount factor to perform a decreasing accumulation of the immediate reward to calculate the cumulative reward, the importance of the current reward and future rewards can be balanced. Compared with the maximization of short-term benefits in traditional systems, this step can guide the system to optimize the long-term learning effect and solve the problem of insufficient short-term incentives.

[0135] Benefits of determining the optimal recommended action: It can ensure that students obtain resource recommendations that can best improve their learning effects. Compared with traditional fixed recommendation strategies, this step enables the recommended content to be dynamically optimized according to the immediate needs and long-term development goals of students, improving the learning effect.

[0136] Step S4 further includes:

[0137] Step 4.1, model the teaching recommendation problem as a Markov decision process, define the state set S, the action set A, the state transition probability function P, and the immediate reward function R;

[0138] The state set S represents all the states of the students;

[0139] The action set A contains all the actions recommended by the system for the students;

[0140] The state transition probability function P(s′|s,a) represents the probability of transitioning to state s′ after taking action a in state s;

[0141] The immediate reward function R(s,a) represents the immediate reward obtained after taking action a in state s;

[0142] Step 4.2, use the Bellman optimality principle to solve the optimal state value function V *(s), this function represents the maximum cumulative reward when taking the optimal strategy in state s. The recursive form of the Bellman equation is as follows:

[0143]

[0144] where V * (s) is the optimal value function of state s, a is the currently available action,

[0145] R(s,a) is the immediate reward obtained after taking action a in state s, s′ is the next state,

[0146] P(s′|s,a) is the probability of transitioning to state s′ after taking action a in state s.

[0147] γ is the discount factor;

[0148] Step 4.3, update the Bellman equation iteratively until the optimal state value function V * (s) converges. During each iterative update, calculate the value of the next state according to the combination of the current state s and action a, and finally obtain the optimal strategy and the optimal state value function.

[0149] Step 4.1: The advantage of constructing a Markov decision process model is that the model truly restores the dynamic interaction between the student state and the system behavior. It comprehensively captures the changes in the student state and avoids the limitations of single-point data decision-making. Compared with traditional static models, its description has timeliness and adaptability.

[0150] Step 4.2: The advantage of using the Bellman optimality principle to solve the optimal state value function is that it takes into account both immediate rewards and future benefits, ensuring the long-term decision-making of the system. It helps the system evaluate the comprehensive effects of various actions and improve the overall recommendation quality.

[0151] Step 4.3: The advantage of iterative updating until the optimal state value function converges is that the iterative process can continuously correct errors, ensuring that the final optimal strategy obtained is stable and accurate. The dynamic update mechanism can effectively adapt to the changes in the student state and respond in a timely manner. Compared with fixed strategies, iterative solution is more flexible and conforms to the actual teaching situation.

[0152] Step S5 further includes:

[0153] Step 5.1, after obtaining the optimal state value function through the Markov decision process in Step S4, the deep Q-network is used to approximate this optimal value function. Define the Q-network as a neural network, whose input is the state vector s q , and the output is the Q value Q(s q ,a q ,a q ) corresponding to each action a;

[0154] The Q-value represents the expected cumulative reward after performing action a in the given state s q and q performing action a

[0155] Step 5.2: During the training process, by interacting with the environment, update the Q-value function based on the current Q-network. At each update, update the current Q-value by calculating the actually obtained reward and the Q-value of the next state, and use the Bellman equation to update the Q-value. The formula is:

[0156]

[0157] where Q(s q , a q ) is the Q-value of performing action a in the current state s q , q α is the learning rate, r

[0158] is the immediate reward obtained after performing action a in the current state s q , q γ is the discount factor, q is the maximum Q-value in the next state s

[0159] ; q+1 q+1 Step 5.3: To increase the stability of the training process, the deep Q-network usually uses a target network to calculate the target Q-value. The parameters θ

[0160] of the target network are periodically copied from the parameters θ of the main Q-network. The update formula is: -

[0161]

[0161] θ - ← θ. The update of the target network can avoid large fluctuations during the training process;

[0162] Step 5.4: The deep Q-network adjusts the network parameters θ through multiple training cycles to reduce the error between the Q-value prediction and the actual reward, and finally enables the Q-network to accurately predict the Q-values of each state-action pair, and finally obtains the optimal strategy.

[0163] Step 5.1: The benefit of using the deep Q-network to approximate the optimal state value function is that it can capture the complex non-linear relationship between states and actions, accurately predict the long-term benefits of each action, is more flexible than traditional methods, and can adapt to changing environments;

[0164] Step 5.2: The benefit of updating the Q-value function based on environmental interaction is that it can correct prediction errors in real time, quickly adapt to feedback, effectively balance current and future benefits, and has a faster convergence speed compared to fixed update strategies;

[0165] Step 5.3: The benefits of using the target network to improve training stability are to smooth the training process, reduce the drastic fluctuations of parameters, avoid oscillations caused by unstable estimation, and improve the overall robustness of training;

[0166] Step 5.4: The benefits of multi-cycle training optimization to obtain the optimal strategy are that continuous optimization gradually reduces the prediction error, and finally outputs the optimal recommendation strategy, and the dynamic adjustment ability far exceeds that of traditional fixed strategies.

[0167] Step S6 further includes:

[0168] Step 6.1: Model the interaction between the system and the student as a dynamic game model to describe the interaction process between the system and the student. The dynamic game model includes the rewards of the system and the rewards of the student;

[0169] Step 6.2: Define the reward functions of the system and the student. The reward functions of the system and the student are represented by the following formulas:

[0170] System reward r s : Measure the effectiveness of the system's recommended actions;

[0171] Student reward r st : Reflect the student's feedback on the recommended actions;

[0172] Step 6.3: Nash equilibrium condition. To ensure the stability of the strategies in the game, the strategies of the system and the student need to satisfy the Nash equilibrium condition. Given the strategies of the other party, neither party has the incentive to unilaterally change its own strategy. The Nash equilibrium condition is expressed as:

[0173]

[0174] where r s is the reward function of the system, r st is the reward function of the student, a s is the recommended action of the system, a st is the reaction action of the student, E[r s (a s ,a st )] is the expected value of the system reward, and E[r st (a s ,a st )] is the expected value of the student reward.

[0175] Step 6.4: In the dynamic game, the system and the student optimize their respective behaviors according to their own reward functions through dynamic game strategies. The strategy π s of the system and the strategy π st of the student can both be optimized through the iterative game process until the Nash equilibrium is reached.

[0176] Step S6 greatly improves the interaction quality between the system and students by introducing a dynamic game model and Nash equilibrium conditions. Through dynamic adjustment and game strategy optimization, the system can dynamically adjust the recommended content according to students' feedback and behaviors, ensuring the stability and effectiveness of the recommendation strategy. This process enables the recommendation system to flexibly respond to complex and changing student needs and provide accurate and personalized learning resource recommendations. Compared with traditional teaching methods with fixed strategies, the game optimization strategy of the present invention is outstanding in improving student participation and interactivity, thereby enhancing the teaching effect and the adaptability of the system.

[0177] Although the embodiments of the present invention have been shown and described, it will be understood by those of ordinary skill in the art that various changes, modifications, substitutions and variations can be made in these embodiments without departing from the principles and spirit of the present invention, and the scope of the present invention is defined by the appended claims and their equivalents.

Claims

1. An online training education information-based teaching method based on big data, characterized in that it includes the following steps: Step S1: Collect data from the online teaching platform, real-time evaluation system, emotion monitoring equipment, and user interaction logs, and perform cleaning, standardization, noise filtering, and feature extraction on the data; Step S2: constructing a student state vector based on the data in step S1 to implement dimensionality reduction processing; Step S3: Establish a state-action model, define the video resources, topic resources and discussion modules corresponding to the recommended actions, and design a reward function to calculate the reward score; Step S4: Abstract the teaching recommendation problem into a Markov decision process, forming a state set, an action set, a state transition probability function and an immediate reward function, and use the Bellman optimality principle to solve the optimal state value function; Step S5: using the deep Q network to implement reinforcement learning training on the Markov decision process in step S4, so as to obtain the optimal recommendation strategy by iteratively updating the state-action-value function; Step S6: Construct a dynamic game model to describe the interaction between the system and students. The system rewards and student rewards are expressed in the form of cumulative discount rewards and meet the Nash equilibrium conditions. Step S7: Integrate steps S1 to S6 to form a closed-loop feedback system, collect feedback data in real time to adjust various parameters to verify system performance.

2. According to a big data-based online training education informationization teaching method as recorded in claim 1, the data collected in step S1 includes student login data, page click data, video viewing time data and answer accuracy data in the online teaching platform, interim performance data in the real-time evaluation system, emotion scoring data in the emotion monitoring device, and discussion record data in the user interaction log.

3. According to the online training education information-based teaching method based on big data as described in claim 1, it is characterized in that the step S2 further comprises: Step 2.1, the original feature data of each student is obtained from step S1 to form a state vector X, where each row corresponds to a student, and the columns are learning progress component, assessment score component, emotion score component, and interaction frequency component; Step 2.2: Standardize each feature in the state vector X according to the following formula to obtain the normalized data matrix x′: x′( ij )=(x( ij )-μ j ) / σ j , Among them, x (ij) is the original value of the i-th student on the j-th feature, μ j is the mean of the jth feature among all students, σ j is the standard deviation of the jth characteristic among all students, i=1,2,…,n and j=1,2,…,d, where n is the total number of students and d is the dimension of the initial state vector.

4. According to claim 1, the online training education informationization teaching method based on big data is characterized in that: The step S2 further comprises: Step 2.3, use the normalized data matrix x′ to calculate the covariance matrix S, the calculation formula is: in, is the transpose of x′, n is the total number of students; Step 2.4, perform eigendecomposition on the covariance matrix S and solve for the eigenvalues ​​λ1,λ2,…,λ d and the corresponding eigenvectors v1,v2,…,v d ,satisfy: S·v j =λ j ·v j ,j=1,2,...,d, Among them, λ j is the jth eigenvalue of the covariance matrix S, v j is the j-th eigenvector of the covariance matrix S; Step 2.4, select the eigenvectors corresponding to the first k largest eigenvalues ​​according to the eigenvalue size to form the projection matrix W; Step 2.5, use the projection matrix W to reduce the dimension of the normalized data matrix x′ to obtain the reduced-dimensional state vector matrix Z, which is calculated as: Z = x′·W, Among them, Z is the state vector matrix after dimensionality reduction.

5. According to the big data-based online training education information teaching method of claim 1, it is characterized in that: The step S3 further comprises: Step 3.1, establish a state-action model, and define the state vector according to the student state vector constructed in step S2 To describe students' learning progress, assessment scores, emotional states and interactive behaviors; define a recommended action set U, which contains multiple recommended actions; Step 3.2, design a reward function, which includes learning progress rewards, knowledge mastery rewards and emotional motivation rewards. The specific calculation formula is: r (p) (x) q ,v q )=f (p) (x) q ,v q ), r (k) (x) q ,v q )=f (k) (x) q ,v q ), r (e) (x) q ,v q (=f (e) (x) q ,v q ), Among them, r (p) (ξ q ,v q ) is the learning progress reward, r (k) (ξ q ,v q ) is the knowledge mastery reward, r (e) (ξ q ,v q ) is the emotional incentive reward, ξ q is the student's state vector, v q is the recommended action, f (p) 、f (k) and f (e) A function for calculating various rewards; Step 3.3, calculate the cumulative reward. To evaluate the long-term effect of the recommended action, use the discount factor δ to accumulate the immediate reward in descending order to obtain the overall reward. The overall reward function is: Among them, ω p is the learning progress reward, ω k Is the knowledge mastery reward, ω e is the weight coefficient of the emotional incentive reward, δ is the discount factor, T is the number of time steps, and J(π) is the total reward; Step 3.4: The system determines the optimal recommended action v based on the cumulative reward J(π) q to maximize the reward value.

6. According to the big data-based online training education information teaching method of claim 1, it is characterized by: The step S4 further comprises: Step 4.1, model the teaching recommendation problem as a Markov decision process, define the state set S, action set A, state transition probability function P, and immediate reward function R; The state set S represents all states of students; The action set A contains all the actions recommended by the system for students; The state transition probability function P(s′|s,a) represents the probability of transitioning to state s′ after taking action a in state s; The immediate reward function R(s,a) represents the immediate reward obtained after taking action a in state s.

7. The method of online training education informationization based on big data according to claim 1 is characterized in that: The step S4 further comprises: Step 4.2, use Bellman's optimality principle to solve the optimal state value function V * (s), which represents the maximum cumulative reward for taking the optimal strategy in state s. The recursive form of the Bellman equation is: Among them, V * (s) is the optimal value function of state s, a is the currently available action, R(s,a) is the immediate reward after taking action a in state s, s′ is the next state, P(s′|s,a) is the probability of transitioning to state s′ after taking action a in state s. γ is the discount factor; Step 4.3, update the Bellman equation iteratively until the optimal state value function V * (s) converges. During each iterative update, the value of the next state is calculated based on the combination of the current state s and action a, and finally the optimal strategy and optimal state value function are obtained.

8. The method of online training education informationization based on big data according to claim 1 is characterized in that: The step S5 further comprises: Step 5.1: After obtaining the optimal state value function through the Markov decision process in step S4, the deep Q network is used to approximate the optimal value function. The Q network is defined as a neural network whose input is the state vector s q , the output is each action a q The corresponding Q value Q(s q ,a q ); The Q value represents the q Next, perform action a q Expected cumulative rewards after Step 5.2, during the training process, the Q value function is updated based on the current Q network by interacting with the environment. At each update, the current Q value is updated by calculating the actual reward and the Q value of the next state. The Bellman equation is used to update the Q value. The formula is: Among them, Q(s q ,a q ) is the current state s q Next, perform action a q The Q value, α is the learning rate, r q is the current state q Next, perform action a q After receiving the instant reward, γ is the discount factor, is the next state s q+1 The maximum Q value under .

9. The online training education informationization teaching method based on big data according to claim 1 is characterized in that: The step S5 further comprises: Step 5.3, to increase the stability of the training process, the deep Q network usually uses the target network to calculate the target Q value. The parameter θ of the target network - It is periodically copied from the parameters θ of the main Q network, and the update formula is: θ - ←θ, the update of the target network can avoid large fluctuations during training; In step 5.4, the deep Q network adjusts the network parameters θ through multiple training cycles to reduce the error between the Q value prediction and the actual reward, so that the Q network can accurately predict the Q value of each state-action pair and finally obtain the optimal strategy.

10. The online training education informationization teaching method based on big data according to claim 1 is characterized in that: The step S6 further comprises: Step 6.1, model the interaction between the system and students as a dynamic game model to describe the interaction process between the system and students. The dynamic game model includes the system's rewards and the students' rewards. Step 6.2, define the reward functions of the system and the student. The reward functions of the system and the student are expressed by the following formula: System Rewards s : Measures the effectiveness of the actions recommended by the system; Student Rewards st : reflects students’ feedback on recommended actions; Step 6.3, Nash equilibrium condition. To ensure the stability of the strategy in the game, the strategies of the system and the students must meet the Nash equilibrium condition. Given the strategy of the other party, neither party has the motivation to unilaterally change its strategy. The Nash equilibrium condition is expressed as: Among them, r s is the reward function of the system, r st is the student’s reward function, a s is the recommended action of the system, a st is the student's reaction action, E[r s (a s ,a st )] is the expected value of the system reward, E[r st (a s ,a st )] is the expected value of the student’s reward. Step 6.4, in the dynamic game, the system and the student optimize their respective behaviors through dynamic game strategies according to their own reward functions. The system's strategy π s Strategies for students st All of them can be optimized through the iterative game process until the Nash equilibrium is reached.

Citation Information

Patent Citations

  • Adaptive learning content recommendation method and system based on deep reinforcement learning

    CN117009668A

  • Teaching path planning method and recommendation system based on reinforcement learning

    CN117151602A

  • Intelligent learning guiding method based on GPT and multi-agent reinforcement learning

    CN117808637A

  • Personalized music teaching content pushing system

    CN119417665A

  • Method, device and system for providing pre and post competency assessment solution for step by step competency evaluation and customized course recommendation

    KR102775769B1

Cited By

  • Teaching decision-making method, system and equipment based on artificial intelligence and medium

    CN121883211A