Big data online education system and method for theoretical learning

Through multi-dimensional learning state modeling and real-time feedback mechanism, combined with reinforcement learning and nonlinear dynamics systems, the learning path is dynamically adjusted, which solves the problem that the existing system cannot adapt to students' real-time state changes, and achieves a personalized and efficient learning experience.

CN120069291APending Publication Date: 2025-05-30INSPECTION & QUARANTINE TECH CENT SHANDONG ENTRY EXIT INSPECTION & QUARANTINE BUREAU
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510079898.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-18
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

The existing online education system cannot dynamically adjust the learning path, neglecting students' emotional state and cognitive load, resulting in insufficient recommendation effects.

Method used

Through multi-dimensional learning state modeling and real-time feedback mechanism, combined with reinforcement learning and other methods, the learning path is dynamically adjusted, students' emotional and cognitive changes are considered, and the learning process is simulated through a nonlinear dynamic system to optimize the learning path.

Benefits of technology

It realizes personalized learning path recommendations, which can dynamically adapt to students' real-time changes, and improves the adaptability and accuracy of learning effects and paths.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120069291A_ABST
    Figure CN120069291A_ABST
Patent Text Reader

Abstract

The invention relates to the field of online education, and discloses a big data online education system and method for theoretical learning, and the system comprises a data collection and preprocessing module which is used for collecting learning behavior data, emotional attitude data and social interaction data of students, carrying out the cleaning, normalization and feature extraction of the collected data, and carrying out the feature extraction of the collected data; outputting the preprocessed learning state data; and the state space modeling and action space defining module is used for constructing a multi-dimensional student learning state space including knowledge mastery degree, cognitive load, emotional state and social interaction state, and defining an action space recommended by a learning path. According to the method, the learning path is dynamically adjusted through multi-dimensional learning state modeling and a reinforcement learning algorithm, recommendation is optimized according to real-time changes of emotion and cognitive load of students, the problem that in the prior art, the learning path cannot be adjusted in a personalized mode is solved, and more efficient learning experience is provided.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of online education, and specifically to a big data online education system and method for theoretical learning. Background Art

[0002] In existing online education systems, learning path recommendations rely on students' basic data, such as academic performance, homework completion, etc. Most of these systems generate recommended paths through static rules or simple algorithms, aiming to provide students with the most suitable learning tasks. However, the recommendation mechanisms of these systems mostly ignore students' personalized needs, especially the emotional fluctuations, excessive cognitive load, and social interactions that occur during the learning process, resulting in the recommendation effect of learning paths being insufficiently refined and efficient.

[0003] The learning path recommendation systems in the prior art rely on single-dimensional indicators. For example, the difficulty of learning tasks is set based on students' historical grades, without considering the influence of factors such as students' emotional states and cognitive loads. Such a recommendation mechanism often cannot dynamically adapt to the changes of students during the learning process. For example, when students feel anxious or have an excessive cognitive load, traditional systems do not adjust the task difficulty or recommend more interactive learning content. In addition, traditional systems lack the ability of self-optimization and cannot dynamically adjust the recommendation strategy according to students' feedback, thus failing to maximize students' learning effects. Compared with the prior art, the present invention, through multi-dimensional modeling of learning states and a real-time feedback mechanism, can not only adjust the learning path according to students' emotional and cognitive changes, but also achieve continuous optimization of the learning path through methods such as reinforcement learning, thereby providing a more personalized and efficient learning experience. Summary of the Invention

[0004] Aiming at the deficiencies of the prior art, the present invention provides a big data online education system and method for theoretical learning, which solves the problems in the prior art of lacking dynamic adjustment of learning paths and being unable to perform personalized recommendations according to changes in students' emotions and cognitive loads.

[0005] To achieve the above objectives, the present invention is realized through the following technical solutions: A big data online education system and method for theoretical learning, including: A data collection and preprocessing module, which is used to collect students' learning behavior data, emotional attitude data, and social interaction data, and perform cleaning, normalization, and feature extraction on the collected data, and output preprocessed learning state data; A state space modeling and action space definition module, which is used to construct a multi-dimensional student learning state space, including knowledge mastery, cognitive load, emotional state, and social interaction state, and define the action space for learning path recommendations; Reward function design and feedback module, which is used to design a reward function based on the student's learning progress, emotional feedback, and cognitive load factors, calculate the reward after the student completes the learning activity, and generate feedback information; Markov decision process and reinforcement learning module, which is used to calculate the optimal learning path recommendation based on the student's state data and reward information through the reinforcement learning algorithm, and realize optimal path learning by updating the Q value; Nonlinear dynamics system modeling and feedback module, which is used to simulate the student's learning progress and state changes based on the nonlinear dynamics system, and adjust the learning path recommendation strategy according to the output of the dynamic model; Learning path recommendation and real-time adjustment module, which is used to generate a personalized learning path according to the student's real-time learning state and feedback information, and continuously adjust and optimize the learning path; User interaction and feedback module, which is used to provide a user-friendly interface, display the student's learning progress, recommended learning path, and receive the student's learning feedback to realize the optimization of the personalized learning path.

[0006] Preferably, the data acquisition and preprocessing module generates the student's multi-dimensional learning state according to the student's learning duration, homework completion situation, test scores, and emotional analysis result data.

[0007] Preferably, the reward function design and feedback module conducts a weighted evaluation of each learning path by designing a reward function. The rewards include improved knowledge mastery, improved emotional state, reduced cognitive load, and increased social interaction, and it has a personalized feedback function.

[0008] Preferably, the Markov decision process and reinforcement learning module adopts the Q-learning algorithm, updates the Q value according to the student's real-time feedback information, and selects the optimal learning path recommendation.

[0009] Preferably, the nonlinear dynamics system modeling and feedback module adopts the Lorenz system model, and based on the student's cognitive load, emotional state, and social interaction state variables, simulates the dynamic evolution of the student's learning process and generates adjustment strategies.

[0010] Preferably, the learning path recommendation and real-time adjustment module combines reinforcement learning with the nonlinear dynamics system model, and dynamically adjusts the learning path according to the student's learning progress and feedback information in real time to ensure the optimization of the learning path.

[0011] Preferably, the user interaction and feedback module provides the student's learning progress, recommended learning path, and emotional analysis result information, and adjusts the recommendation strategy according to the student's feedback to realize personalized learning optimization.

[0012] A big data online education method for theoretical learning includes the following steps: Step 1: Collect the learning behavior data, emotional attitude data, and social interaction data of students, and preprocess the data to generate the multi-dimensional learning status of students; Step 2: Construct a multi-dimensional learning status space and action space, and define a reward function to evaluate each learning path; Step 3: Use the reinforcement learning algorithm to update the Q value through Q learning and calculate the optimal learning path; Step 4: Simulate the dynamic changes of the learning process according to the non-linear dynamics system model and adjust the learning path recommendation; Step 5: According to the real-time feedback information of students, generate a personalized learning path through the learning path recommendation and real-time adjustment module and optimize it.

[0013] Preferably, the reward function in Step 2 includes: designing rewards according to the improvement of knowledge mastery, the improvement of emotional state, the reduction of cognitive load, and the increase of social interaction to ensure the balance of each dimension.

[0014] Preferably, the non-linear dynamics system model in Step 4 adopts the Lorenz system model to simulate the dynamic changes of each dimension state of students during the learning process and generate a feedback strategy for adjusting the learning path.

[0015] The present invention provides a big data online education system and method for theoretical learning. It has the following beneficial effects: 1. By adopting a dynamic adjustment mechanism based on the multi-dimensional learning status of students and combining information from multiple dimensions such as emotional analysis, knowledge mastery, cognitive load, and social interaction, the present invention achieves the technical effect of personalized learning path recommendation. Compared with the single-dimensional static recommendation scheme in the prior art, it solves the deficiency that the traditional learning recommendation system cannot adapt to the real-time state changes of students, ensuring that each student can obtain the most suitable learning path during the learning process.

[0016] 2. By accurately simulating the learning status changes of students through the non-linear dynamics model, the present invention can optimize and adjust the learning path according to the real-time feedback of students, achieving the technical effects of real-time, intelligent feedback, and learning path optimization. Compared with the system lacking a deep state prediction and adjustment mechanism in the prior art, it solves the deficiency that the traditional system cannot accurately predict the long-term learning status and needs of students, improving the adaptability and accuracy of the learning path.

[0017] 3. By combining the reinforcement learning algorithm and using methods such as Q learning to continuously optimize the learning path recommendation, the present invention achieves the technical effects of self-optimization and intelligent recommendation. Compared with the system with a fixed rule recommendation path in the prior art, it solves the deficiency of being unable to self-adjust through feedback, enabling the system to optimize the learning strategy in real time according to the behavior and feedback of students and continuously improve the learning effect.

[0018] 4. The present invention models the learning process of students through Markov decision process and optimizes and adjusts learning tasks in combination with a reward function, achieving the technical effect of optimal learning path selection. Compared with the existing learning path recommendation solutions lacking a decision-making model and personalized reward mechanism, it solves the deficiency that the existing system fails to dynamically adjust according to the long-term performance and needs of students, making the learning path more flexible and efficient. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] Figure 1 It is the system architecture diagram of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0020] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0021] Please refer to the attached Figure 1 , the embodiments of the present invention provide a big data online education system for theoretical learning, including: A data collection and preprocessing module, configured to collect students' learning behavior data, emotional attitude data, and social interaction data, and perform cleaning, normalization, and feature extraction on the collected data, and output preprocessed learning status data; In this embodiment, the data collection and preprocessing module includes the following key steps: data collection, data cleaning, data normalization, and feature extraction. Through these steps, the system can obtain the learning status of students and provide necessary inputs for subsequent learning path recommendation.

[0022] Data Collection Data collection is the primary task of the system. The system monitors students' learning behaviors in real time through the integrated online learning platform and obtains the following types of data: Learning behavior data: including students' learning duration, homework completion, test scores, video viewing duration, learning progress, etc. These data can be directly extracted from the interaction records of the learning platform.

[0023] Emotional attitude data: Emotional analysis is obtained by analyzing students' text feedback, discussion participation, etc. Natural language processing techniques (such as emotional analysis models) can be used to extract emotional information from students' comments, questions, or discussion content. At the same time, facial expression recognition can also be used as a supplementary data source for emotional attitudes to capture students' emotional fluctuations in real time.

[0024] Social interaction data: Social interaction data refers to the frequency and quality of interactions between students and their classmates or teachers in places such as classroom discussion areas and social platforms. These data reflect the students' participation and cooperation in the learning process.

[0025] During the data collection process, the system continuously receives data streams from different channels and platforms and ensures the timeliness and real-time nature of the data.

[0026] Data cleaning Since the data collected from multiple different sources may contain noise or missing values, data cleaning is an important step. The purpose of data cleaning is to ensure the quality of the input data and avoid affecting subsequent learning path recommendations due to data errors or inconsistencies.

[0027] In some embodiments, the data cleaning step includes removing redundant data, filling in missing values, and removing outliers. For example, if a student's learning duration record is extremely short or long, the system will mark it as abnormal data and use reasonable filling strategies (such as mean filling or interpolation) to fill in the missing data. For inconsistent data, the system will use rule verification or manual intervention for processing.

[0028] Specifically, during the data cleaning process, the system will also unify the formats of data from different platforms. For example, learning behavior data may be in minutes, while emotional data may be quantified by emotional intensity. The system needs to unify them into the same dimension for processing in order to perform further analysis in subsequent processing steps.

[0029] Data normalization After data cleaning, the system will perform normalization processing on all the collected data. Since different types of data may have different dimensions or data ranges, normalization processing is a necessary step to ensure that various data can be effectively compared and integrated.

[0030] In one possible implementation, data normalization uses the following formula:

[0031] Where: is the original data value; and are the minimum and maximum values in this dataset respectively; is the normalized data value, and the range is limited between [0, 1].

[0032] For example, there may be great differences in the learning durations of students. To process the learning durations on the same scale as other data (such as emotional attitudes), the system normalizes all the learning duration data to ensure that the contributions of each data dimension are equal.

[0033] Feature Extraction After data normalization, the system extracts features from various learning data of students for subsequent state modeling and learning path recommendation. The core of the feature extraction process lies in extracting the key features from the students' learning data that can best reflect their current learning state. The feature extraction step generates a multi-dimensional learning state vector as the input for subsequent modules.

[0034] For example, the system converts data such as learning duration, test scores, homework completion status, and emotional state into a single learning state value through weighted average or other methods. By calculating the multi-dimensional feature vector of a student at a certain time point, the system can provide clear basic information for subsequent learning path recommendation.

[0035] In this embodiment, the following calculation method is adopted for feature extraction:

[0036] where: is the learning state vector of the student at time step and contains data in multiple dimensions; represents the student's knowledge mastery at time step ; represents the student's cognitive load at time step ; represents the student's emotional state at time step ; represents the student's social interaction state at time step .

[0037] These features will be input into the state space modeling module for subsequent learning path modeling and recommendation.

[0038] Connection with Other Modules The data collection and preprocessing module is the basic module of the system and completes the data processing required for subsequent operations. First, the collected data is cleaned and normalized and then converted into structured data suitable for subsequent modeling. This data will enter the state modeling module to help the system construct the learning state of students and provide input for the reinforcement learning module. In the reinforcement learning module, based on the state data of students, the system uses the Q-learning algorithm to continuously optimize the learning path recommendation. The non-linear dynamics system further adjusts the learning path based on these state data to produce a more accurate optimization effect during the students' learning process.

[0039] In this embodiment, the data collection and preprocessing module ensures that the system can obtain high-quality data support in subsequent state modeling and path recommendation through the comprehensive collection and processing of students' learning behaviors, emotional attitudes, and social interactions. Steps such as data cleaning, normalization, and feature extraction not only improve the quality of the data but also enhance the system's ability to process data in different dimensions, providing a solid data foundation for the final personalized learning path recommendation.

[0040] The state space modeling and action space definition module is used to construct a multi-dimensional student learning state space, including knowledge mastery, cognitive load, emotional state, and social interaction state, and define the action space for learning path recommendation; In this embodiment, the state space modeling and action space definition module adopts the following method to implement the modeling of students' learning states and the design of the action space for learning path recommendation.

[0041] State space modeling The construction of the state space is one of the core tasks of this module. Generally, a student's learning state cannot be measured by a single-dimensional indicator, so multiple dimensions need to be considered comprehensively. Specifically, a student's learning state can be composed of the following main dimensions: Knowledge mastery : Represents the degree of a student's mastery of the current learning content, usually quantified by data such as test scores and homework completion.

[0042] The value range of is usually between [0, 1], where 0 means no mastery at all and 1 means complete mastery. Specifically, in this embodiment,

[0043] The calculation of is based on the student's homework scores and test scores, and a weighted average method can be used to synthesize the results of multiple different learning tasks into an overall knowledge mastery score.

[0044] Emotional state : The emotional state of students reflects their emotional fluctuations, learning interests, and engagement levels during the learning process. The measurement of emotional state can be extracted from students' written feedback through sentiment analysis techniques or obtain the emotional changes of students during the learning process through facial expression recognition techniques. The value of the emotional state also ranges between [0, 1], where 0 represents extremely negative and 1 represents highly positive.

[0045] Social interaction state : The social interaction state reflects the engagement of students during the learning process and their interactions with classmates. The interactions of students in classroom discussions, on social platforms, and the frequency of asking questions can all be used as measurement criteria for the social interaction state. The value range of this state is also [0, 1], and the larger the value, the more active the social interaction of the student.

[0046] The comprehensive reflection of these dimensions represents the multiple states of students during the learning process. Specifically, the learning state of a student at a certain point in time can be represented as a vector , where: is the learning state of the student at time step ; represents the knowledge mastery of the student at time step ; represents the cognitive load of the student at time step ; represents the emotional state of the student at time step ; represents the social interaction state of the student at time step .

[0047] These state variables will form the basis for the system's decision-making, and subsequent learning path recommendations will be dynamically adjusted according to the changes in these states.

[0048] Definition of the action space The definition of the action space is based on all possible learning paths or learning activities that a student can take based on their current state . The action space is a set that contains all the learning activities that the system can recommend to students. Generally speaking, these actions include but are not limited to: Watching videos: Recommending students to watch specific educational videos, usually instructional videos related to the knowledge points that the student currently masters.

[0049] Doing practice questions: Recommending students to do relevant practice questions or mock tests to help students consolidate the current knowledge points.

[0050] Participating in discussions: Encouraging students to participate in the interactions in the discussion area or online learning groups to improve the social engagement of students.

[0051] Reading materials: Recommend that students read articles, textbooks, or case studies related to the current learning topic.

[0052] These actions are designed as diverse learning activities aimed at providing students with different learning methods to meet their personalized learning needs.

[0053] In this embodiment, the action space includes multiple learning path options, such as:

[0054] Among them: represents recommended video watching; represents recommended exercise doing; represents recommended participation in discussions; represents recommended reading materials; and so on.

[0055] At each time step, the system will, based on the student's current state select an optimal action from the action space and update the student's learning state according to the selected action.

[0056] The mapping relationship between states and actions The design of the state space and the action space ensures that the system can select a suitable learning path according to the student's learning situation and progress. Specifically, the system will, based on the student's current learning state (such as and other dimensions), evaluate and select the optimal action to generate a recommended learning path.

[0057] In this embodiment, the system continuously evaluates the optimal actions in each state through a reinforcement learning algorithm (such as Q-learning). Through Q-value updates, the system can gradually optimize the learning path recommendation to ensure that the student is always in the most suitable learning content and difficulty level.

[0058] Connection with other modules The state space modeling and action space definition module is closely connected to the aforementioned data collection and preprocessing module. The data collection and preprocessing module provides clear and standardized student learning data for this module. These data are transformed into the student's learning state and used as the input for subsequent modules. In addition, the reinforcement learning module in the system depends on the definition of the state space and the action space, continuously updates the Q-value through the Q-learning algorithm, and finally generates an optimal learning path recommendation.

[0059] The state space modeling and action space definition module in this embodiment provides the system with an accurate description of the learning state and a basis for selecting recommended actions. By comprehensively considering multiple dimensions such as students' knowledge mastery, cognitive load, emotional state, and social interaction, the system can comprehensively evaluate the learning state of students and select the most suitable learning path. The close connection between this module and the data collection and preprocessing module and the reinforcement learning module ensures that the system can dynamically adjust the learning path and achieve personalized learning recommendations.

[0060] The reward function design and feedback module is used to design a reward function based on factors such as students' learning progress, emotional feedback, and cognitive load, calculate the reward after students complete learning activities, and generate feedback information. In this embodiment, the reward function design and feedback module adopts the following method to dynamically calculate the reward according to the students' state information and feedback the reward information to the system to optimize the learning path recommendation.

[0061] Reward Function Design The reward function is a key component that guides the system to select the learning path. In the present invention, the design of the reward function takes into account multiple factors, including students' knowledge mastery, emotional state, cognitive load, and social interaction, etc. Through the weighted sum evaluation of these factors, the reward function can quantify the impact of each learning path recommendation on students' learning effects.

[0062] Specifically, the design goal of the reward function is to balance the learning effects of multiple dimensions. The reward function can be expressed as:

[0063] Where: is the immediate reward after taking action at time step .

[0064] is the change in students' knowledge mastery, indicating the improvement in students' knowledge mastery after completing a certain learning activity.

[0065] is the change in students' emotional state, measuring students' learning interest and mood fluctuations.

[0066] is the cognitive load of students. A higher cognitive load will reduce learning efficiency. Therefore, when the cognitive load is high, the reward will be reduced accordingly.

[0067] is the social interaction state of students. The increase in social interaction helps to improve students' learning motivation and efficiency. Therefore, when the interaction frequency is high, the reward will be increased.

[0068] is a weight coefficient used to balance the contributions of each dimension, and the specific weights can be dynamically adjusted according to the individual characteristics of the students.

[0069] Generally, the design of the reward function will be personalized according to the learning stage and goals of the students. For students in the beginner stage, the system may increase the weight of knowledge mastery to encourage students to strengthen the learning of basic knowledge; while for students who already have a certain amount of knowledge accumulation, the weight of emotional state and social interaction may be increased to maintain learning interest and social activity.

[0070] Feedback mechanism The reward signal calculated by the reward function not only guides the selection of the learning path, but also further adjusts the learning strategy through the feedback mechanism. Specifically, after the reward signal is fed back to the reinforcement learning module, the system will update its learning path recommendation strategy to ensure that the next recommendation can maximize the long-term learning benefit.

[0071] In this embodiment, the feedback mechanism is carried out through the Q-learning algorithm. Q-learning is a reinforcement learning method based on value iteration, which optimizes the learning path recommendation by continuously updating the Q-value function. Specifically, in each state of the student and action after that, the system will update the Q-value according to the output of the reward function. The update formula of the Q-value is as follows:

[0072] Where: is the learning rate, which controls the amplitude of each Q-value update; is the discount factor, which is used to balance the importance of the current reward and future rewards; is the immediate reward obtained by executing the action in the current state ; is the maximum Q-value that can be obtained by selecting the optimal action in the state .

[0073] In a possible implementation, the Q-learning algorithm will gradually learn the optimal learning path recommendation strategy through a continuous training process. When the Q-value converges, the system can provide the optimal learning path recommendation according to the learning state of the student .

[0074] Real-time adjustment of the reward function and feedback To improve the personalization and dynamic adjustment capabilities of learning path recommendations, this embodiment also includes real-time adjustment of the reward function and feedback mechanism. Specifically, the learning feedback of students (such as learning satisfaction, emotional state, etc.) can affect the weight coefficient of the reward function in real time. By continuously monitoring the learning state and feedback of students, the system can dynamically adjust the weights in the reward function to ensure that the learning path recommendations are more in line with the current needs of students.

[0075] For example, when the system detects that a student's emotional state is low in a certain learning session, the weight of the emotional state in the reward function can be increased to recommend more interactive and interesting learning content, thereby enhancing the student's learning motivation; if the student's cognitive load is too high, the system may reduce the weight of the cognitive load and recommend easier-to-understand learning tasks.

[0076] Connection with other modules The reward function design and feedback module is closely connected to the aforementioned state space modeling, action space definition, and reinforcement learning modules. Specifically, the state space modeling and action space definition modules provide the learning state of students and the optional learning paths, and the reward function design module generates corresponding reward signals by evaluating the effects of these paths on students. The reinforcement learning module then continuously optimizes the learning path recommendation strategy through methods such as Q-learning based on the reward signals.

[0077] The reward function design and feedback module also has a close relationship with the aforementioned non-linear dynamics system modeling module. The non-linear dynamics model provides more accurate feedback for the reward function design by simulating the dynamic changes in the student's learning process. For example, when the student's emotional state changes, the non-linear dynamics model will adjust the learning path to adapt to this change, and the reward function will evaluate the impact of this adjustment on the student.

[0078] In this embodiment, the reward function design and feedback module accurately calculates the reward value of each learning path by comprehensively considering the multi-dimensional learning state of students (such as knowledge mastery, emotional state, cognitive load, and social interaction state). Through close cooperation with the reinforcement learning module, the reward signal feedback prompts the system to gradually optimize the learning path recommendation strategy, thereby providing students with a personalized and dynamically adjustable learning experience. The design of this module ensures that the system can provide students with the optimal learning path recommendations in a constantly changing learning environment.

[0079] The Markov decision process and reinforcement learning module are used to calculate the optimal learning path recommendation based on the state data and reward information of students through reinforcement learning algorithms, and to achieve optimal path learning by updating the Q value; In this embodiment, the Markov decision process and the reinforcement learning module use algorithms such as Q-learning, combine the state space of the student and the reward function, and update the learning path recommendation strategy in real time, so as to make the student's learning experience more personalized and efficient.

[0080] Markov decision process The Markov decision process (MDP) is a method for describing a decision-making system. It selects actions based on the current state and optimizes future decisions through reward feedback. Generally, MDP is used to simulate dynamic systems with state transition characteristics. In the present invention, the learning process of the student is modeled as a Markov process, where: State space : Represents the learning state of the student at each moment, which has been determined by the aforementioned state space modeling module. Specifically, the state of the student includes the degree of knowledge mastery , cognitive load , emotional state and social interaction state . These states change in each learning step, constituting a multi-dimensional state space.

[0081] Action space : Represents the learning path or learning activity that the system can choose in each state. The action space has been defined by the learning path recommendation in the aforementioned module. The action space includes learning activities such as watching videos, doing exercises, and participating in discussions.

[0082] Reward function ( : The reward function represents the immediate reward given by the system after the student executes a certain learning activity. These rewards will feedback the learning effect of the student, help the system evaluate the value of each learning activity, and have been defined in the reward function design and feedback module.

[0083] State transition : Represents how the system transfers to the next state after taking a certain action . In this embodiment, the state transition can predict the state change of the student after completing a certain learning task by statistically analyzing historical learning data. .

[0084] According to the MDP theory, the learning process of the student can be expressed by the following formula:

[0085] Where: is in the state The optimal value below represents the long-term learning benefit of the student in this state; is the discount factor, indicating the degree of influence of future rewards; is the immediate reward after taking a certain action in the current state; is from state Execute action and then transfer to the next state probability.

[0086] In one possible implementation, the system will dynamically update the recommendation strategy of the learning path by estimating the impact of different learning paths on the long-term learning effect of students.

[0087] Reinforcement learning Reinforcement learning is a machine learning method that learns the optimal strategy by interacting with the environment. In the present invention, a reinforcement learning algorithm (such as Q-learning) is used to continuously optimize the learning path recommendation. Specifically, Q-learning estimates the value (Q-value) of different actions in each state and selects the optimal learning path by iteratively updating these Q-values.

[0088] In Q-learning, the system learns the optimal value of actions in each state through the following update formula:

[0089] where: is the state Execute action Q-value, representing the long-term return of this action; is the learning rate, controlling the amplitude of Q-value update; is the discount factor, weighing the current reward and future rewards; is the immediate reward after executing action in the current state; is the next state the maximum Q-value of all possible actions in.

[0090] By continuously interacting and providing feedback with the learning environment, the system gradually adjusts the recommendation strategy to make the Q-value converge to the optimal state. In one possible implementation, Q-learning will train the Q-value based on the historical data of students (such as learning achievements, homework completion, emotional feedback, etc.) and generate the optimal learning path recommendation according to the training results.

[0091] Combination of reinforcement learning and nonlinear dynamical systems The reinforcement learning module works closely with the non - linear dynamics system modeling module, enabling more precise adjustment of learning path recommendations. In some embodiments, the non - linear dynamics model can provide dynamic feedback by simulating the learning process of students (such as knowledge mastery, emotional fluctuations, etc.), thus affecting the update of Q - learning. Specifically, when there are significant changes in the learning state of students (such as emotional state fluctuations or excessive cognitive load), the non - linear dynamics model will provide feedback to adjust the weights in the reward function, enabling the system to quickly respond to changes in the learning state of students and avoid inefficient learning caused by long - term mismatches in learning paths.

[0092] For example, when the emotional state of the student is low, the non - linear dynamics system will adjust the recommended path, increasing interactive and interesting learning activities to improve the student's emotional state and maintain their learning motivation. At this time, the weight of the emotional state in the reward function may increase, and the reinforcement learning module will adjust the recommendation strategy based on this.

[0093] Connection with other modules The Markov decision process and the reinforcement learning module are closely related to the aforementioned state - space modeling, reward function design, and feedback modules. The aforementioned modules provide the current learning state of the student, the optional paths of the learning task, and the corresponding reward signals. The reinforcement learning module uses this information to update the Q - value and selects the optimal learning path through Q - learning.

[0094] Specifically, the state - space modeling module constructs a multi - dimensional learning state vector for the student, and these state variables (such as knowledge mastery, emotional state, cognitive load, etc.) are the inputs for reinforcement learning; the reward function design and feedback module provides guidance for Q - learning by calculating rewards based on the student's learning feedback; ultimately, the reinforcement learning module continuously optimizes the learning path recommendation based on Q - learning.

[0095] In this embodiment, the Markov decision process and the reinforcement learning module use methods such as Q - learning to combine the state space and reward function of the student to optimize the learning path recommendation in real - time. By dynamically adjusting the Q - value and the weights of the reward function, the system can provide personalized and optimal learning paths for each student. The design of this module not only enables the system to handle complex learning tasks but also ensures the optimization of the learning path, providing an efficient and accurate learning experience.

[0096] The non - linear dynamics system modeling and feedback module is used to simulate the learning progress and state changes of students based on the non - linear dynamics system, and adjust the learning path recommendation strategy according to the output of the dynamic model; In this embodiment, the non - linear dynamics system modeling and feedback module simulates the complex changes in the student's state during the learning process by using non - linear models such as the Lorenz system, and combines the feedback of the reinforcement learning module to optimize the learning path recommendation in real - time.

[0097] Non - linear dynamics system modeling Non - linear dynamics system modeling is the core part of simulating the state changes of students during the learning process. Generally, in a dynamic system, various learning states of students (such as knowledge mastery, emotional state, cognitive load, etc.) are interrelated and non - linearly changing. In order to effectively capture these complex dynamic relationships, the non - linear dynamics system describes the interactions between various state variables during the learning process through a mathematical model.

[0098] Generally, the Lorenz system is a classic non - linear dynamics system model, which is widely used to simulate complex non - linear behaviors in dynamic systems. In this embodiment, we model the student's learning state as a three - dimensional non - linear dynamics system, and use the Lorenz system model to describe the mutual influence between knowledge mastery, cognitive load, emotional state, and social interaction state.

[0099] Specifically, the dynamic model of the system describes the student's learning process through the following three sets of ordinary differential equations:

[0100]

[0101]

[0102] Where: represents the student's knowledge mastery; represents the student's cognitive load; represents the student's emotional state; represents the student's social interaction state; are the parameters of the system, which control the mutual relationship between each dimension.

[0103] These equations describe the dynamic changes between various state variables of students during the learning process. Specifically, the change in knowledge mastery is related to the difference in cognitive load, the change in cognitive load is affected by knowledge mastery and emotional state, and the change in emotional state is related to the intensity of social interaction. By solving these differential equations, the system can simulate the state of students at future time steps and adjust the learning path according to the changes in these states.

[0104] Model feedback and dynamic adjustment In a possible implementation, the non-linear dynamics system modeling and feedback module generates feedback signals by calculating the real-time changes in the learning state of the student, in order to adjust the learning path recommendation in real time. For example, if the system detects that the emotional state of the student decreases, indicating that the student may have lost interest in learning or encountered learning difficulties, the system can enhance the emotional state of the student by adding interactive or interesting learning tasks and provide personalized learning path recommendations.

[0105] Specifically, the system monitors the dynamic changes of the student, such as the cognitive load increases, the system will automatically adjust the difficulty of the recommended content to avoid overloading the student and ensure the rationality of the learning tasks. When the social interaction of the student increases, the system may further recommend more tasks based on cooperative learning to improve learning motivation and learning effect.

[0106] During the feedback process, the non-linear dynamics system provides accurate prediction of the learning state for the system, ensuring that each adjustment of the learning path can meet the learning needs of the student to the greatest extent. This process is not only based on the current state of the student, but also includes the prediction and preparation for possible state changes in the future learning process.

[0107] Connection with other modules The non-linear dynamics system modeling and feedback module is closely connected with the previous state space modeling, reward function design and feedback, and Markov decision process and reinforcement learning modules. In the aforementioned modules, the system has obtained multi-dimensional state information of the student through data collection and preprocessing, and calculated the immediate reward of each learning path through the reward function design and feedback module. In the Markov decision process and reinforcement learning module, the system continuously optimizes the learning path recommendation through methods such as Q-learning.

[0108] The non-linear dynamics system then adjusts the recommendation of the learning path through a feedback mechanism according to the dynamic changes in the learning state of the student. Specifically, when the system detects fluctuations in the emotional state of the student or excessive cognitive load, the non-linear dynamics system will provide feedback to the reward function design and feedback module, adjust the reward weight, and thus affect the update of the learning path of the reinforcement learning algorithm.

[0109] For example, when the emotional state of the student is low, the non-linear dynamics model will reduce the weight of the cognitive load and increase the weight of the emotional state, and the system will recommend more interactive tasks that can increase learning motivation. Conversely, when the social interaction of the student is less, the non-linear dynamics system will increase the weight of the social interaction and promote the system to recommend more cooperative learning tasks.

[0110] Real-time adjustment and feedback of the model The advantage of the non - linear dynamics system is that it can make rapid adjustments based on the real - time feedback of students. During the implementation process, the system continuously tracks the learning status of students, automatically generates learning path recommendations, and makes adjustments according to the feedback of students. Specifically, the system dynamically adjusts the recommended learning content and difficulty according to the real - time changes of variables such as the emotional state and cognitive load of students to ensure that the learning path is always adapted to the learning status of students.

[0111] In this embodiment, the non - linear dynamics system modeling and feedback module provides accurate learning status prediction and feedback for the system by simulating the multi - dimensional state changes in the student learning process. By using non - linear models such as the Lorenz system, the system can capture the complex interactions in the learning process and dynamically adjust the learning path recommendations according to factors such as the emotional state and cognitive load of students. This module works closely with the aforementioned reward function design and feedback, Markov decision process and reinforcement learning modules to ensure that the system can provide personalized and optimal learning path recommendations for students.

[0112] The learning path recommendation and real - time adjustment module is used to generate personalized learning paths according to the real - time learning status and feedback information of students, and optimize the learning paths by continuous adjustment; In this embodiment, the learning path recommendation and real - time adjustment module realizes a personalized learning experience by continuously monitoring the changes in the student learning status and dynamically updating the learning path recommendations according to the feedback. This module not only considers the current learning status of students, but also makes predictions according to historical learning data and emotional attitudes and other factors to ensure that students can obtain the most suitable learning tasks at each learning stage.

[0113] Generation of learning path recommendations In this embodiment, the generation of learning path recommendations is based on the current status information of students. First, the learning status of students is generated by the aforementioned state space modeling and action space definition module, which maps various learning data of students (such as knowledge mastery, emotional state, cognitive load, social interaction, etc.) into a multi - dimensional learning status vector. This vector is used as input and is passed to the learning path recommendation module.

[0114] Generally, the learning path recommendation module selects the most suitable learning tasks or learning activities according to the current status of students. Specifically, the system generates the recommended learning path through reinforcement learning algorithms (such as Q - learning) or rule - based recommendation algorithms according to factors such as knowledge mastery, emotional state and cognitive load in the current learning status. For example: If the student's knowledge mastery is low, the system will give priority to recommending content related to basic knowledge.

[0115] If the student's emotional state is relatively low, the system will recommend more interactive and interesting learning tasks to enhance students' learning motivation.

[0116] When the cognitive load of students is relatively high, the system will recommend relatively simple and less burdensome learning content to avoid overlearning.

[0117] In some embodiments, the learning path recommendation can also be adjusted more precisely by combining the output of the non-linear dynamics model. For example, if the cognitive load of a student starts to increase and the emotional state is relatively low, the system will optimize the recommendation by adjusting the difficulty of the learning content and combining the fun and interactivity of the learning tasks. The real-time adjustment of the learning path Real-time adjustment is another core function of this module. As the learning progress of students advances, the learning state of students will change. To ensure the optimization of the learning path, the system must update the recommended learning path in real time to meet the dynamic needs of students.

[0118] Specifically, after a student executes a certain learning task, the system will update their learning state according to the student's feedback (such as learning performance, changes in emotional state, frequency of social interaction, etc.). This feedback information will be transmitted to the aforementioned reward function design and feedback module in a timely manner, thereby affecting the decision-making of the learning path recommendation module. For example: If a student obtains a high score after a certain learning task, the system will consider that this task is helpful for knowledge mastery and may recommend more similar tasks.

[0119] If the emotional state of a student changes, such as being in a low mood, the system will adjust the learning path through the non-linear dynamics model and recommend some relaxing learning activities to help the student regain their learning interest.

[0120] If the cognitive load of a student is too high, the system will adjust the difficulty of the recommended content in real time and recommend more interactive and less stressful tasks to relieve the student's learning burden.

[0121] During the real-time adjustment process, the update of the reward signal is crucial. Specifically, the system will use the feedback reward signal to evaluate the performance of students under different learning paths, and then select the most suitable learning path for the current state. In some embodiments, this adjustment process may be calculated in real time through the following formula:

[0122] Where: represents the current learning state of the student; represents the learning path recommended by the system for the student; represents the reward obtained by the student after executing the learning path; It is the new state of the student after executing the learning path.

[0123] Through this formula, the system can update its learning path recommendation in real time according to the student's immediate feedback, ensuring the continuous optimization of the learning process.

[0124] Connection with other modules The learning path recommendation and real-time adjustment module has a close cooperation relationship with the aforementioned modules (such as state space modeling, reward function design and feedback, Markov decision process and reinforcement learning, non-linear dynamics system modeling and feedback). Specifically: With the state space modeling and action space definition module: The learning state of the student (such as knowledge mastery, emotional state, cognitive load, etc.) is provided by this module and used as the input for learning path recommendation.

[0125] With the reward function design and feedback module: The system evaluates the student's performance during the learning process through the reward function and provides feedback for real-time adjustment of the learning path. The change of the reward signal directly affects the optimization of the learning path.

[0126] With the Markov decision process and reinforcement learning module: In the recommendation strategy based on Q-learning, the learning path recommendation module provides personalized learning tasks for the student according to the optimal path generated by the reinforcement learning model.

[0127] With the non-linear dynamics system modeling and feedback module: The non-linear dynamics model provides predictions of the student's possible future states by simulating the student's learning process and adjusts the learning path recommendation strategy through feedback.

[0128] The learning path recommendation and real-time adjustment module in this embodiment can dynamically adjust the learning path through real-time feedback of multi-dimensional student states while ensuring a personalized learning experience. The system not only generates recommendations based on the student's current learning state, but also combines the feedback signal and the prediction of the non-linear dynamics system to ensure that the learning path recommendation always matches the student's needs and state. The close cooperation of this module with the aforementioned modules such as state space modeling, reward function design and feedback, reinforcement learning, and non-linear dynamics system enables the system to flexibly respond to the changing needs of students in a complex learning environment and provide the optimal learning path.

[0129] The user interaction and feedback module is used to provide a user-friendly interface, display the student's learning progress, recommended learning path, etc., and receive the student's learning feedback to optimize the personalized learning path; In this embodiment, the user interaction and feedback module not only collects the student's learning feedback, but also displays information such as the student's learning progress, grades, and emotional state in real time, so that the student can clearly understand their learning state and stimulate their motivation to further learn.

[0130] User Interface Design In this embodiment, the user interface provides students with an intuitive and easy-to-operate learning control platform. Students can view their learning progress, current tasks, recommended learning paths, and feedback on the system through this interface. Generally, the interface will display the following information: Learning Progress: Displays the learning tasks that students have completed during the entire learning process, as well as the completion status of the current learning task.

[0131] Recommended Path: The next learning task or learning activity recommended by the system based on the student's current learning status and feedback.

[0132] Grades and Feedback: Includes the grades after students complete tasks, the feedback given by the system, and the analysis of emotional states.

[0133] As an option, the user interface can also provide functions for interacting with teachers or classmates. Students can submit questions, participate in discussions, or share knowledge through the platform, enhancing social interaction and improving learning motivation.

[0134] Specifically, the interface design includes a visual learning progress bar, a display window for recommended content, a feedback area for task completion, etc. After each learning task is completed, students will receive corresponding feedback information, which is based on multi-dimensional data such as the student's immediate performance, emotional attitude, and learning grades.

[0135] Student Feedback Collection and Processing The feedback collection methods in this embodiment include various forms, such as: Emotional Feedback: Through emotion analysis technology, the system can capture students' emotional changes in real time through methods such as students' text input, facial expression recognition, and speech emotion analysis. For example, if the system detects that a student has a strong negative emotion (such as anxiety or boredom) when learning a certain module, the system will preferentially select interesting or highly interactive content when recommending the next learning path to reduce the student's emotional burden.

[0136] Learning Progress Feedback: Students can view their current learning progress through the interface and provide feedback on the learning content. If a student finds a certain knowledge point particularly difficult or not well understood, they can mark or leave a message on the platform, and the system will adjust the recommendation of subsequent learning tasks based on these marks.

[0137] Social Interaction Feedback: The system records students' participation in the discussion area or social platform, including the frequency and depth of asking questions, answering questions, and discussing with classmates. If the system detects that a student has less social interaction, the system may recommend more group learning or discussion tasks to increase the student's social participation.

[0138] In some embodiments, students can submit feedback through simple interface operations. The feedback information includes emotional state selection, evaluation of the difficulty of learning tasks, feedback on learning interest, etc. These feedbacks will be processed by the background algorithm, thereby affecting the adjustment of the learning path of the system.

[0139] Real-time Feedback Mechanism and Learning Path Optimization After being processed by the system, the feedback information of the students will be fed back to the learning path recommendation and real-time adjustment module, affecting the decision-making process of the system. Specifically, the feedback information will be combined with the output of the reward function design module to provide an immediate reward signal for the system. Through this feedback mechanism, the system can respond immediately to the personalized needs of students and optimize the learning path recommendation.

[0140] For example, if a student has a low emotional state after learning a certain module and shows a high cognitive load during the task completion process, the system will adjust the learning path through a nonlinear dynamics model, reduce the difficulty of the task, and recommend more interactive and interesting tasks. This feedback process can be expressed by the formula:

[0141] Where: represents the current learning state of the student; represents the currently recommended learning path; represents the immediate reward; represents the change in emotional state, reflecting the emotional feedback of the student.

[0142] Through this formula, the system can adjust the learning path of the student based on real-time feedback, making the recommended learning tasks more in line with the needs of the student.

[0143] Connection with Other Modules The user interaction and feedback module works closely with the aforementioned multiple modules. First, the data collection and preprocessing module provides the behavior data and emotional data of the students, providing basic data support for the user interaction and feedback module. Then, the student state information generated by the state space modeling and action space definition module provides real-time data for the user interaction module, while the reward function design and feedback module adjusts the learning path according to the feedback signal. On this basis, the learning path recommendation and real-time adjustment module will optimize the recommendation strategy according to the feedback of the students.

[0144] For example, when students feedback that their emotional state is poor through the user interaction interface, this information will be fed back to the reward function design module, affecting the calculation of the reward function. At this time, the system will increase the weight of the emotional state and preferentially recommend learning tasks that can improve the emotional state of the students (such as easy and interesting interactive content). The feedback signal further continuously adjusts the recommended learning path through reinforcement learning algorithms (such as Q-learning).

[0145] In this embodiment, the user interaction and feedback module forms a closed-loop system through cooperation with modules such as learning path recommendation, reward function design, reinforcement learning, and non-linear dynamic system modeling, capable of realizing real-time monitoring and adjustment of students' learning status and emotional state. The system continuously optimizes the learning path by obtaining students' real-time feedback information, ensuring a personalized and dynamic learning experience for each student. This interactive feedback mechanism not only enhances students' learning motivation but also improves learning efficiency, ensuring the adaptability and interactivity of learning content.

[0146] A big data online education method for theoretical learning includes the following steps: Step 1: Collect students' learning behavior data, emotional attitude data, and social interaction data, and preprocess the data to generate students' multi-dimensional learning status; In this embodiment, the data collection and preprocessing module collects students' data through multiple channels, including learning behavior data (such as homework completion and test scores), emotional state data (such as emotional analysis scores), and social interaction data (such as discussion participation). After cleaning, normalizing, and feature extraction, these data will form the input of multi-dimensional learning status and provide a data basis for the subsequent learning path recommendation and real-time adjustment of the system.

[0147] Step 2: Construct a multi-dimensional learning state space and action space, and define a reward function to evaluate each learning path; In this embodiment, the construction of the state space and action space is based on multi-dimensional information such as students' learning behavior, emotional feedback, and social interaction. The system transforms it into a structured state space and defines selectable learning paths according to the current state. By constructing the state space and action space, the system can select the optimal learning activities according to the students' current state in subsequent reinforcement learning.

[0148] Step 3: Use the reinforcement learning algorithm to update the Q value through Q-learning and calculate the optimal learning path; In this embodiment, the reward function design and feedback module is closely connected with the aforementioned "data collection and preprocessing" module, state space modeling and action space definition module, and learning path recommendation and real-time adjustment module. The aforementioned modules provide information such as students' learning data, emotional state, and social interaction as input data for the reward function. The design of the reward function will consider multiple factors such as students' knowledge mastery, emotional changes, cognitive load, and social interaction, and recommend personalized learning paths by dynamically adjusting the reward mechanism.

[0149] Step 4: Simulate the dynamic changes of the learning process according to the non-linear dynamic system model and adjust the learning path recommendation; In this embodiment, the Markov decision process and the reinforcement learning module play a core role. It is responsible for optimizing the selection of learning paths according to the learning status and immediate feedback of students. Through the reinforcement learning algorithm, the system can dynamically adjust the recommended learning tasks based on the state space and reward function of students, enabling the system to continuously learn from the interactions of students and self-optimize. This module is closely connected with the aforementioned "data collection and preprocessing" module, the state space modeling and action space definition module, and the reward function design and feedback module to ensure that each student can obtain a customized learning path recommendation.

[0150] In the data collection and preprocessing module, the system first obtains the learning behavior data, emotional state, and social interaction data of students and preprocesses them. In the state space modeling and action space definition module, the system transforms this data into a multi-dimensional learning state vector and defines the corresponding action space. Then, the reward function design and feedback module calculates the immediate reward based on the feedback of students, and these reward signals are fed back to the Markov decision process and the reinforcement learning module. This module optimizes the recommendation of learning paths through the reinforcement learning algorithm (such as Q-learning) and dynamically adjusts according to the state of students, ultimately ensuring that each student can study efficiently on the most suitable path for themselves.

[0151] Step 5: According to the real-time feedback information of students, generate and optimize a personalized learning path through the learning path recommendation and real-time adjustment module.

[0152] In this embodiment, the non-linear dynamics system modeling and feedback module is responsible for adjusting the learning path recommendation in real time by simulating the state changes during the learning process of students. In close connection with the aforementioned steps, based on the data collection and preprocessing, state space modeling and action space definition, and reward function design and feedback modules, this module further optimizes the learning experience of students. By simulating and feeding back the cognitive and emotional fluctuations of students, it ensures the continuous adaptability and efficiency of the learning path.

[0153] Although the embodiments of the present invention have been shown and described, for those of ordinary skill in the art, it can be understood that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and spirit of the present invention. The scope of the present invention is defined by the appended claims and their equivalents.

Claims

1. A big data online education system for theoretical learning, characterized in that: include: The data collection and preprocessing module is used to collect students' learning behavior data, emotional attitude data and social interaction data, and clean, normalize and extract features from the collected data, and output the preprocessed learning status data; The state space modeling and action space definition module is used to construct a multi-dimensional student learning state space, including knowledge mastery, cognitive load, emotional state, and social interaction state, and define the action space for learning path recommendation; The reward function design and feedback module is used to design the reward function based on the students' learning progress, emotional feedback, and cognitive load factors, calculate the rewards after the students complete the learning activities, and generate feedback information; The Markov decision process and reinforcement learning module is used to calculate the optimal learning path recommendation based on the student's state data and reward information through the reinforcement learning algorithm, and to achieve the optimal path learning by updating the Q value; Nonlinear dynamic system modeling and feedback module, which is used to simulate students’ learning progress and state changes based on nonlinear dynamic systems, and adjust the learning path recommendation strategy according to the output of the dynamic model; The learning path recommendation and real-time adjustment module is used to generate personalized learning paths based on students' real-time learning status and feedback information, and optimize the learning paths through continuous adjustment; The user interaction and feedback module is used to provide a user-friendly interface, display students' learning progress, recommend learning paths, and receive students' learning feedback to achieve the optimization of personalized learning paths.

2. According to the big data online education system for theoretical learning in claim 1, it is characterized in that: The data collection and preprocessing module generates the students' multi-dimensional learning status based on the students' study time, homework completion, test scores, and sentiment analysis results.

3. According to claim 1, a big data online education system for theoretical learning is characterized in that: The reward function design and feedback module performs weighted evaluation on each learning path by designing a reward function. The rewards include improved knowledge mastery, improved emotional state, reduced cognitive load and increased social interaction, and has a personalized feedback function.

4. The big data online education system for theoretical learning according to claim 1 is characterized in that: The Markov decision process and reinforcement learning module adopts a Q learning algorithm to update the Q value according to the real-time feedback information of the students and select the optimal learning path recommendation.

5. The big data online education system for theoretical learning according to claim 1 is characterized in that: The nonlinear dynamic system modeling and feedback module adopts the Lorenz system model to simulate the dynamic evolution of students' learning process and generate adjustment strategies based on students' cognitive load, emotional state and social interaction state variables.

6. A big data online education system for theoretical learning according to claim 1, characterized in that: The learning path recommendation and real-time adjustment module combines reinforcement learning with a nonlinear dynamic system model to dynamically adjust the learning path in real time according to the student's learning progress and feedback information to ensure the optimization of the learning path.

7. A big data online education system for theoretical learning according to claim 1, characterized in that: The user interaction and feedback module provides students with learning progress, recommended learning paths, and sentiment analysis result information, and adjusts the recommendation strategy based on student feedback to achieve personalized learning optimization.

8. A big data online education method for theoretical learning, according to a big data online education system for theoretical learning according to any one of claims 1 to 7, characterized in that: The following steps are involved: Step 1: Collect students’ learning behavior data, emotional attitude data, and social interaction data, and pre-process the data to generate students’ multi-dimensional learning status; Step 2: Construct a multi-dimensional learning state space and action space, and define a reward function to evaluate each learning path; Step 3: Use the reinforcement learning algorithm to update the Q value through Q learning and calculate the optimal learning path; Step 4: Simulate the dynamic changes of the learning process according to the nonlinear dynamic system model and adjust the learning path recommendation; Step 5: Based on students’ real-time feedback, generate and optimize personalized learning paths through the learning path recommendation and real-time adjustment module.

9. The big data online education method for theoretical learning according to claim 7 is characterized in that: The reward function in step 2 includes: designing rewards based on improved knowledge mastery, improved emotional state, reduced cognitive load, and increased social interaction to ensure a balance in each dimension.

10. The big data online education method for theoretical learning according to claim 7 is characterized in that: The nonlinear dynamic system model in step 4 adopts the Lorenz system model to simulate the dynamic changes of the states of each dimension of students during the learning process and generate a feedback strategy for adjusting the learning path.

Citation Information

Cited By

  • Personalized online education system and method based on artificial intelligence

    CN120950521A

  • An artificial intelligence-based personalized online education system and method

    CN120950521B