Cloud education management system and method based on big data
By collecting and integrating students' eye movement and voice data in the cloud education management system, a dual-stream Transformer model is built to predict students' cognitive load index and optimize teaching strategies, the problems of insufficient utilization of multimodal data and limited accuracy of cognitive load prediction in the existing system are solved, and more accurate teaching strategies are optimized and improved teaching effects are achieved.
Patent Information
- Application Number
- CN202510356020.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-25
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2045-03-25
AI Technical Summary
The existing cloud education management system lacks multimodal fusion capabilities in the collection and processing of student behavior data, and the cognitive load prediction model relies on a single feature and has limited prediction accuracy.
By collecting and preprocessing students' eye movement data and speech data, MFCC features and comprehensive eye movement features are extracted, and dynamic time regularization algorithm is used to fuse these features by timestamps to generate comprehensive feature vectors. Based on the hierarchical processing mechanism and cross-modal attention mechanism, a dual-stream Transformer model is built to predict students' cognitive load index and optimize teaching strategies based on the predicted results.
It realizes efficient fusion of multimodal data and accurate prediction of cognitive load, providing a basis for dynamic optimization of teaching strategies and improving teaching effectiveness.
Smart Images

Figure CN120217134A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of cloud education technology, and in particular to a cloud education management system and method based on big data. Background Art
[0002] In recent years, as a new education model, the cloud education management system has realized the centralized management of educational resources and the dynamic generation of personalized teaching plans by integrating cloud computing, Internet of Things, and data analysis technologies. In this context, the education management technology based on big data has become a research hotspot, especially remarkable progress has been made in the acquisition and analysis of student behavior data. In addition, the application of reinforcement learning algorithms in virtual teaching environments enables teaching plans to be dynamically adjusted according to students' real-time feedback, further improving the teaching effect.
[0003] However, there are still some deficiencies in the existing cloud education management systems. First, the existing technologies lack the ability of multimodal fusion in the acquisition and processing of student behavior data. For example, eye movement data and speech data are usually analyzed independently, making it difficult to comprehensively reflect students' learning states. Second, most of the existing cognitive load prediction models rely on single features and fail to fully utilize the complementarity of multimodal data, resulting in limited prediction accuracy. Summary of the Invention
[0004] In view of the above existing problems, the present invention is proposed.
[0005] Therefore, the present invention provides a cloud education management method based on big data to solve the problems of insufficient utilization of multimodal data and limited prediction accuracy of cognitive load in the prior art.
[0006] To solve the above technical problems, the present invention provides the following technical solutions: In a first aspect, the present invention provides a cloud education management method based on big data, which includes collecting students' eye movement data and speech data and performing preprocessing, and at the same time extracting MFCC features and comprehensive eye movement features; using the dynamic time warping algorithm to fuse the MFCC feature vectors and comprehensive eye movement feature vectors according to timestamps to generate comprehensive feature vectors; constructing a two-stream Transformer model based on a hierarchical processing mechanism and a cross-modal attention mechanism, and predicting students' cognitive load index according to the comprehensive feature vectors; constructing a virtual teaching environment, and optimizing three strategy parameters of the explanation speed, interaction frequency, and content difficulty in the virtual teaching environment based on an adaptive strategy optimization algorithm combined with cognitive load data to dynamically generate teaching plans.
[0007] As a preferred solution of the cloud education management method based on big data according to the present invention, wherein: the extraction of MFCC features and comprehensive eye movement features includes the following steps, Frame the speech data, perform fast Fourier transform on each frame of the signal to generate a spectrogram, and extract MFCC feature MFCC feature vectors through discrete cosine transform; Generate a fixation heat map by performing Gaussian kernel density estimation on the fixation point coordinates, identify the distribution statistics of the fixation duration, extract the pupil diameter change rate and fluctuation frequency, and segment the saccade path to calculate the path complexity to generate a comprehensive eye movement feature vector.
[0008] As a preferred solution of the cloud-based education management method based on big data according to the present invention, wherein: the dynamic time warping algorithm is used to fuse the MFCC feature vectors and the comprehensive eye movement feature vectors according to timestamps to generate comprehensive feature vectors, and the specific steps are as follows. Calculate the Euclidean distance between the MFCC feature vector and the eye movement feature vector frame by frame; Based on the Euclidean distance between the MFCC feature vector and the eye movement feature vector, use the global sequence alignment method to construct a cost matrix; Based on the cost matrix, use the dynamic time programming algorithm to find the optimal alignment path according to timestamps, and fuse the MFCC features and the eye movement features according to timestamps along the optimal alignment path to generate comprehensive feature vectors.
[0009] As a preferred solution of the cloud-based education management method based on big data according to the present invention, wherein: based on a hierarchical processing mechanism and a cross-modal attention mechanism, a two-stream Transformer model is constructed, and according to the comprehensive feature vector, the cognitive load index of the student is predicted, and the specific steps are as follows. The temporal embedding layer performs position encoding on the time series of the comprehensive feature vector through Transformer position encoding to generate temporal embedding features; The spatial embedding layer performs linear transformation on the comprehensive feature vector through MLP respectively to generate spatial embedding features; The temporal embedding and the spatial embedding are spliced through the cross-modal attention mechanism, and the weights of the temporal features and the spatial features are adjusted in combination with the Sigmoid gate control mechanism. The temporal embedding layer, the spatial embedding layer and the cross-modal attention mechanism are used to construct a two-stream Transformer model, and the comprehensive feature vector is input to predict the cognitive load index of the student.
[0010] As a preferred solution of the cloud-based education management method based on big data according to the present invention, wherein: based on historical student behavior data and the cognitive load data of the student, a virtual teaching environment is constructed through the Unity3D engine and in combination with the TensorFlow reinforcement learning framework.
[0011] As a preferred solution of the cloud education management method based on big data according to the present invention, wherein: the construction process of the adaptive policy optimization algorithm is as follows. Define the state space, action space, and reward function. Construct a policy network based on the state space, action space, and the cognitive load data of students. Construct a value network based on the state space and the reward function. Construct an adaptive policy optimization algorithm based on the policy network and the value network.
[0012] As a preferred solution of the cloud education management method based on big data according to the present invention, wherein: the specific steps of the dynamically generated teaching plan are as follows. Combine the adaptive policy optimization algorithm with the cognitive load data of the trainees to optimize the policy parameters. Dynamically generate a teaching plan through RLDP according to the optimized policy parameters.
[0013] In the second aspect, the present invention provides a cloud education management system based on big data, including a feature extraction module, a feature fusion module, a cognitive load prediction module, and a teaching plan generation module; the feature extraction module is used to collect the eye movement data and voice data of students and perform preprocessing, and at the same time extract MFCC features and comprehensive eye movement features; the feature fusion module is used to fuse the MFCC feature vectors and the comprehensive eye movement feature vectors according to the time stamps by using the dynamic time warping algorithm to generate a comprehensive feature vector; the cognitive load prediction module is used to construct a two-stream Transformer model based on a hierarchical processing mechanism and a cross-modal attention mechanism, and predict the cognitive load index of students according to the comprehensive feature vector; the teaching plan generation module is used to construct a virtual teaching environment, and optimize three policy parameters of the explanation speed, interaction frequency, and content difficulty in the virtual teaching environment based on the adaptive policy optimization algorithm combined with the cognitive load data, and dynamically generate a teaching plan.
[0014] In the third aspect, the present invention provides a computer device, including a memory and a processor, wherein: when the computer program stored in the memory is executed by the processor, any step of the cloud education management method based on big data as described in the first aspect of the present invention is implemented.
[0015] In the fourth aspect, the present invention provides a computer-readable storage medium, on which a computer program is stored, wherein: when the computer program is executed by the processor, any step of the cloud education management method based on big data as described in the first aspect of the present invention is implemented.
[0016] The beneficial effects of the present invention are as follows: By using the dynamic time warping algorithm to fuse the MFCC feature vectors and the comprehensive eye movement feature vectors according to the timestamps, a comprehensive feature vector is generated, realizing the efficient fusion of multi-modal data; By constructing a two-stream Transformer model, combining the temporal embedding layer and the spatial embedding layer, and dynamically fusing the temporal features and the spatial features through the cross-modal attention mechanism, the accurate prediction of cognitive load is realized, providing a basis for the dynamic optimization of teaching strategies. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0018] Figure 1 It is a flowchart of the cloud education management method based on big data in Embodiment 1.
[0019] Figure 2 It is a schematic diagram of the cloud education management system based on big data in Embodiment 1.
[0020] Figure 3 It is a flowchart of data collection and preprocessing in the cloud education management method based on big data in Embodiment 1.
[0021] Figure 4 It is a flowchart of predicting the cognitive load index of students in the cloud education management method based on big data in Embodiment 1. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0022] In order to make the above objects, features, and advantages of the present invention more obvious and understandable, the following will make a detailed description of the specific embodiments of the present invention with reference to the accompanying drawings of the specification.
[0023] In the following description, many specific details are set forth in order to fully understand the present invention. However, the present invention can also be implemented in other ways different from those described herein. Those skilled in the art can make similar extensions without departing from the connotation of the present invention. Therefore, the present invention is not limited by the specific embodiments disclosed below.
[0024] Secondly, the so-called "one embodiment" or "embodiment" herein refers to a specific feature, structure, or characteristic that can be included in at least one implementation manner of the present invention. The phrase "in one embodiment" appearing in different places in this specification does not necessarily refer to the same embodiment, nor is it a separate or alternative embodiment that excludes other embodiments.
[0025] Embodiment 1, refer to Figures 1 to 4, which is the first embodiment of the present invention. This embodiment provides a cloud education management method based on big data, including the following steps: S1. Collect the eye movement data and voice data of students and perform preprocessing, and at the same time extract MFCC features and comprehensive eye movement features.
[0026] The eye movement data includes fixation point coordinates, fixation duration, and pupil diameter changes, and the voice data includes fundamental frequency, speech rate, silent interval, and formants.
[0027] It should be noted that the eye movement data is obtained by capturing the eye movement trajectory with an infrared camera and calculating the changes in fixation points and pupil diameters, and the voice data is obtained by a microphone array.
[0028] The preprocessing includes data cleaning, denoising, fixation point segmentation, time alignment, and normalization, specifically as follows: First, start with data cleaning to ensure the integrity and accuracy of the data set by removing outliers in the eye movement and voice data and filling in missing parts; then perform the denoising step, use a filter to smooth the eye movement data and reduce high-frequency noise, and use audio processing tools to remove background noise in the voice data to improve clarity; subsequently, perform fixation point segmentation on the eye movement data, and use an algorithm to distinguish different fixation behaviors for subsequent analysis; time alignment ensures that the time axes between different samples of both eye movement and voice data are consistent, laying a foundation for comparative analysis; the last step is normalization, by normalizing the data to ensure the consistency of data from different individuals or acquisition conditions.
[0029] Frame the voice data, perform a fast Fourier transform on each frame of the signal to generate a spectrogram, and extract MFCC feature MFCC feature vectors through discrete cosine transform; Furthermore, first divide the voice signal into several frames at a fixed duration (such as 25 milliseconds) and perform windowing processing (such as Hamming window) on each frame of the signal to reduce boundary effects; then perform a fast Fourier transform on each frame of the signal to convert the time-domain signal into a frequency-domain signal and generate a spectrogram; then take the logarithm of the spectrogram and calculate the Mel filter bank energy to obtain the Mel spectrogram; finally, perform a discrete cosine transform on the Mel spectrogram to extract MFCC feature vectors. For example, a 1-second voice signal is divided into 40 frames, each frame of the signal generates a spectrogram through a fast Fourier transform, and then a 13-dimensional MFCC feature vector is obtained through a discrete cosine transform.
[0030] Generate a fixation heat map by performing Gaussian kernel density estimation on the fixation point coordinates, identify the distribution statistics (mean, variance, peak) of the fixation duration, extract the pupil diameter change rate and fluctuation frequency, and segment the saccade path to calculate the path complexity (curvature change) to generate a comprehensive eye movement feature vector.
[0031] Further, first, perform Gaussian kernel smoothing on the fixation point coordinates to generate a fixation heat map to reflect the concentration degree of the fixation area; then calculate the distribution statistics of the fixation duration, including the mean, variance, and peak value, which are used to describe the concentration and discreteness of the fixation duration; then extract the change rate and fluctuation frequency of the pupil diameter to reflect the contraction and dilation dynamics of the pupil; finally, segment the saccade path and calculate the curvature change of each segment of the path to obtain the path complexity, which is used to describe the smoothness of the eye movement trajectory. For example, in a 5-second eye movement data, the fixation point coordinates are used to generate a fixation heat map through Gaussian kernel density estimation. The mean of the fixation duration is 200 milliseconds, the variance is 50 milliseconds, the peak value is 300 milliseconds, the change rate of the pupil diameter is 0.5, the fluctuation frequency is 2 Hz, and the curvature change of the saccade path is 0.2, and finally a comprehensive eye movement feature vector is generated.
[0032] S2. Use the dynamic time warping algorithm to fuse the MFCC feature vector and the comprehensive eye movement feature vector according to the time stamp to generate a comprehensive feature vector.
[0033] Calculate the Euclidean distance between the MFCC feature vector and the eye movement feature vector frame by frame. The expression is: ; Among them, represents the Euclidean distance between the MFCC feature vector of the th frame and the comprehensive eye movement feature vector of the th frame. represents the th dimension value of the MFCC feature vector of the th frame. represents the th dimension value of the comprehensive eye movement feature vector of the th frame. represents the dimension of the MFCC feature vector and the eye movement feature vector; It should be noted that for the MFCC feature vector and the comprehensive eye movement feature vector of each frame, their dimension values are respectively extracted, the difference between the corresponding dimension values is calculated and squared, and after accumulating the squared differences of all dimensions and taking the square root, the Euclidean distance between the two-frame feature vectors is obtained.
[0034] Based on the Euclidean distance between the MFCC feature vector and the eye movement feature vector, use the global sequence alignment method to construct a cost matrix; Further, first initialize an initial cost matrix. The number of rows of the initial cost matrix is the number of frames of the MFCC feature vector, and the number of columns is the number of frames of the comprehensive eye movement feature vector. Then, initialize the first row and the first column of the matrix as the cumulative Euclidean distance. For example, in the first row, the Euclidean distance of each column is accumulated from the first column to the last column, and in the first column, the Euclidean distance of each row is accumulated from the first row to the last row. Then, starting from the second row and the second column, calculate the cost of the current position element by element. The formula is the current Euclidean distance plus the minimum value of the three adjacent positions on the left, above, and upper left. Finally, fill in the complete matrix to obtain the cost matrix. For example, if the MFCC feature vector has 3 frames and the eye movement feature vector has 4 frames, initialize a 3-row and 4-column matrix. The first row is filled with the cumulative Euclidean distance from the first column to the fourth column, the first column is filled with the cumulative Euclidean distance from the first row to the third row, and the second row and the second column are filled with the current Euclidean distance plus the minimum value of the three positions on the left, above, and upper left, and so on to generate the cost matrix.
[0035] Based on the cost matrix, use the dynamic time warping algorithm to find the optimal alignment path according to the time stamps, and fuse the MFCC feature vector and the comprehensive eye movement feature vector according to the optimal alignment path and the time stamps to generate a comprehensive feature vector.
[0036] Further, starting from the lower right corner of the cost matrix, gradually backtrack to the upper left corner, and select the path with the minimum cost among the three adjacent positions on the left, above, and upper left as the optimal alignment path. Then, align the MFCC feature vector and the neutral eye movement feature vector according to the optimal alignment path and the time stamps. For example, the MFCC feature of the i-th frame is aligned with the eye movement feature of the j-th frame. Finally, splice the aligned MFCC feature and the eye movement feature according to the time stamps to generate a comprehensive feature vector. For example, if the optimal alignment path of the cost matrix is (1, 1), (2, 2), and (3, 3), then align the MFCC feature of the first frame with the eye movement feature of the first frame, the MFCC feature of the second frame with the eye movement feature of the second frame, and the MFCC feature of the third frame with the eye movement feature of the third frame, and splice them to generate a comprehensive feature vector.
[0037] S3. Based on the hierarchical processing mechanism and the cross-modal attention mechanism, construct a two-stream Transformer model, and predict the cognitive load index of the student according to the comprehensive feature vector; The hierarchical processing mechanism includes a temporal embedding layer and a spatial embedding layer; The temporal embedding layer performs position encoding on the time series of the comprehensive feature vector through the Transformer position encoding to generate temporal embedding features. The expression is: ; Among them, is the temporal embedding feature at the time step and is the comprehensive feature vector. is the comprehensive feature vector of the position encoding of the time series; Furthermore, first generate a position encoding for each time step of the comprehensive feature vector. The position encoding is composed of a sine function and a cosine function and is used to represent the order information of the time steps. Then add the comprehensive feature vector to the position encoding of the corresponding time step to generate the temporal embedding feature. For example, add the comprehensive feature vector of the first frame to the position encoding of the first frame, add the comprehensive feature vector of the second frame to the position encoding of the second frame, and so on. Finally, generate the temporal embedding feature containing the time order information.
[0038] The spatial embedding layer performs a linear transformation on the comprehensive feature vector through an MLP to generate the spatial embedding feature, and the expression is: ; where, is the spatial embedding feature at time step , is the weight matrix of the MLP, is the bias term of the MLP; Furthermore, first multiply the comprehensive feature vector by the weight matrix of the MLP, and then add the bias term to obtain the result of the linear transformation. Then apply the ReLU activation function to the result of the linear transformation to generate the spatial embedding feature. For example, the comprehensive feature vector at time step obtains an intermediate result after the linear transformation of the weight matrix and the bias term, and finally generates the spatial embedding feature through the processing of the ReLU activation function.
[0039] Concatenate the temporal embedding and the spatial embedding through the cross-modal attention mechanism, and combine the Sigmoid gate control mechanism to adjust the weights of the temporal features and the spatial features; It should be noted that first concatenate the temporal embedding feature and the spatial embedding feature by dimension to form a fused feature. Then input the fused feature into the Sigmoid gate control mechanism to calculate the weights of the temporal feature and the spatial feature. Then multiply the temporal embedding feature by the temporal feature weight, multiply the spatial embedding feature by the spatial feature weight, and add the results to generate the weighted fused feature.
[0040] Construct a two-stream Transformer model with the temporal embedding layer, the spatial embedding layer, and the cross-modal attention mechanism; Input the comprehensive feature vector into the two-stream Transformer model to predict the cognitive load index of the student, and the expression is: ; where, is the weight matrix that converts the features output by the cross-modal attention mechanism into the cognitive load index of the trainee, is the total number of time steps, is the index of the time step, is the Sigmoid function, is the gating weight matrix, is the element-wise multiplication operator, is the query matrix, is the key matrix, is the value matrix, is the dimension of the feature, is at time step the cognitive load index of the student, is function, is the bias term in the prediction process of the trainee's cognitive load index.
[0041] Furthermore, first add the temporal embedding feature and the spatial embedding feature to obtain ; then calculate the gating weight through the gating weight matrix and the Sigmoid function , then multiply the temporal embedding feature by the query matrix , multiply the spatial embedding feature by the key matrix , calculate the attention score between the two, and normalize through the Softmax function to obtain the attention weight; then multiply the spatial embedding feature by the value matrix , and then multiply element-wise with the attention weight to obtain the weighted feature, then multiply the weighted feature element-wise with the gating weight to obtain the gated weighted feature, and finally take the mean of the gated weighted features for all time steps, and perform a linear transformation through the weight matrix and the bias term to generate the cognitive load index at time step .
[0042] It should be noted that the comprehensive feature vector is calculated in the form of the temporal embedding feature in this expression. The temporal embedding feature performs position encoding on the time series of the comprehensive feature vector through the Transformer position encoding to capture the time dynamics; the spatial embedding feature performs linear transformation on the comprehensive vector through the MLP respectively to extract the spatial semantic information; the two are dynamically fused through feature concatenation and the Sigmoid gating mechanism.
[0043] S4. Construct a virtual teaching environment, and optimize the three policy parameters of the explanation speed, interaction frequency, and content difficulty in the virtual teaching environment based on the adaptive strategy optimization algorithm combined with the cognitive load data, and dynamically generate a teaching plan; Construct a virtual teaching environment based on the historical student behavior data and the students' cognitive load data through the Unity3D engine and combined with the TensorFlow reinforcement learning framework; Furthermore, first design a virtual teaching scenario in the Unity3D engine, including classroom layout, teaching tools, and virtual teacher roles, and at the same time import the historical student behavior data (such as fixation points, voice features) and cognitive load data (such as pupil diameter changes, fundamental frequency fluctuations) into the scenario; then define the state space through the TensorFlow reinforcement learning framework. The state space includes the three policy parameters of the explanation speed, interaction frequency, and content difficulty. For example, the range of the explanation speed parameter is from 1 to 5, where 1 represents the slowest and 5 represents the fastest; then define the action space. The action space includes the adjustment amounts for the explanation speed, interaction frequency, and content difficulty. For example, the adjustment amount for the explanation speed for each action is -1, 0, or 1; finally, based on the historical student behavior data and cognitive load data, train the reinforcement learning model in the virtual teaching environment. For example, when the cognitive load index is lower than the target value, the reinforcement learning model improves the explanation speed and content difficulty by adjusting the action space parameters, generates an optimized teaching strategy, and feeds it back to the virtual teaching scenario in real time to complete the construction and dynamic adjustment of the virtual teaching environment.
[0044] Define the state space based on the comprehensive eye movement features and voice features ; Furthermore, ; is the explanation speed, indicating the speed of teaching content delivery ( ∈[1, 5], 1 is the slowest, 5 is the fastest); is the interaction frequency, indicating the frequency of teacher-student interaction during the teaching process ( ∈[1, 5], 1 is the lowest, 5 is the highest); is the content difficulty, indicating the complexity of the teaching content ( ∈[1, 5], 1 is the simplest, 5 is the most difficult).
[0045] Define the action space based on the adjustment strategies of the explanation speed, interaction frequency, and content difficulty in the historical courses ; Furthermore, ; is the adjustment amount of the explanation speed, indicating the increase or decrease of the explanation speed for each action ; is the adjustment amount of the interaction frequency, representing the increase or decrease of the interaction frequency for each action ; is the adjustment amount of the content difficulty, representing the increase or decrease of the content difficulty for each action .
[0046] Define the reward function based on the deviation between the student's cognitive load index and the target cognitive load index and the stability of the strategy parameters , and the expression is: ; where, is the target cognitive load index, is the weight coefficient of the stability of the strategy parameters; It should be noted that, is used to balance the deviation between the student's cognitive load index and the target cognitive load index and the weight of the stability of the strategy parameters. Its value is based on the degree of influence of the strategy parameter adjustment on the teaching effect. If the strategy parameter adjustment has a greater impact on the teaching effect, then takes a smaller value (such as 0.1) to preferentially optimize the cognitive load index; if the strategy parameter adjustment has a smaller impact on the teaching effect, then λ takes a larger value (such as 0.5) to preferentially maintain the stability of the strategy parameters. The value range of is usually between 0 and 1.
[0047] Combine the state space with the student's cognitive load data as the input, and the probability distribution of the action space as the output to construct the policy network ; Furthermore, first combine the state space s with the student's cognitive load data as the input. The state space includes three strategy parameters: the explanation speed, the interaction frequency, and the content difficulty. For example, the explanation speed is 3, the interaction frequency is 2, the content difficulty is 4, and the cognitive load data is 0.8; then process the input data through a multi-layer neural network to generate the probability distribution of the action space , and the action space includes the adjustment amounts for the explanation speed, the interaction frequency, and the content difficulty. For example, the probabilities of the adjustment amount of the explanation speed being -1, 0, or 1 are 0.2, 0.6, and 0.2 respectively; finally, output the probability distribution of the action space to guide the adjustment of the teaching strategy. For example, the state space is [3, 2, 4, 0.8], and the policy network outputs the probabilities of the adjustment amount of the explanation speed being -1, 0, or 1 as 0.2, 0.6, and 0.2 respectively.
[0048] Take the state space as the input, and through the reward function take the expected cumulative reward of the current state obtained in the state space as the output to construct a value network ; Furthermore, first take the state space as the input. The state space s includes three policy parameters: the explanation speed, the interaction frequency, and the content difficulty. For example, the explanation speed is 3, the interaction frequency is 2, and the content difficulty is 4. Then process the input data through a multi-layer neural network to calculate the expected cumulative reward of the current state. The reward function is based on the deviation between the cognitive load index and the target cognitive load index and the stability of the policy parameters. Finally, output the expected cumulative reward of the current state to evaluate the quality of the teaching strategy. For example, the state space is [3, 2, 4], and the value network outputs the expected cumulative reward of the current state as 0.75.
[0049] Based on the policy network and the value network, construct an adaptive policy optimization algorithm, and combine the cognitive load data of the trainees to optimize the three policy parameters of the explanation speed, the interaction frequency, and the content difficulty in the virtual teaching environment. The expression is: ; wherein, is the optimized th policy parameter at the current time step , is the th policy parameter at time step , means restricting the value of the policy parameter to [1, 5]; Furthermore, first obtain the policy parameter values of the previous time step. For example, the explanation speed is 3, the interaction frequency is 2, and the content difficulty is 4. Then generate the action space probability distribution through the policy network, combine the expected cumulative reward output by the value network, and calculate the adjustment amount of the policy parameter at the current time step. Then add the policy parameter value of the previous time step to the adjustment amount to obtain the policy parameter value at the current time step. Finally, limit the policy parameter value within the range of 1 to 5 to ensure that the parameter value is within a reasonable range. For example, the explanation speed of the previous time step is 3, the probability generated by the policy network for the adjustment amount of 1 is 0.6, and the expected cumulative reward output by the value network is 0.8. Calculate the explanation speed of the current time step as 4, and it remains 4 after limitation, completing the optimization of the policy parameter.
[0050] Based on the optimized policy parameters , dynamically generate a teaching plan through RLDP; Based on the optimized policy parameters, the specific process of dynamically generating teaching plans through RLDP is as follows: First, according to the deviation between the real-time cognitive load index and the target value, the explanation speed is dynamically adjusted through RLDP. If the current cognitive load is low, the explanation rhythm is accelerated. For example, the explanation duration of each knowledge point is shortened from 5 minutes to 4 minutes, and the voice broadcast speed is increased; otherwise, the rhythm is slowed down, extended to 6 minutes, and the speech rate is decreased. Secondly, according to the interaction frequency value in the policy parameters, the teaching interaction nodes are dynamically set. When the frequency parameter increases, two scenario Q&A tasks are inserted within 8 minutes. For example, through virtual scene simulation, students are allowed to solve practical problems; when the frequency parameter decreases, it is adjusted to conduct a basic multiple-choice question test every 15 minutes. Finally, based on the content difficulty parameter, matching content is dynamically selected from the pre-set knowledge base. When the difficulty parameter increases, a comprehensive case analysis involving multiple disciplines is selected (such as the bridge load-bearing calculation combining mathematics and physics), and step-by-step guiding prompts are added; when the parameter decreases, it is switched to basic concept diagrams and single-step practice questions. Throughout the process, through the dynamic programming mechanism of reinforcement learning, cognitive load data is collected every 5 minutes. For example, when students show efficient understanding during case discussions, the difficulty parameter is immediately increased to introduce advanced topics; if abnormal enlargement of pupil diameter fluctuations is detected, the speed parameter is immediately decreased and relaxation guidance voice is inserted.
[0051] This embodiment also provides a cloud education management system based on big data, including: a feature extraction module, a feature fusion module, a cognitive load prediction module, and a teaching plan generation module; the feature extraction module is used to collect and preprocess the eye movement data and voice data of students, and at the same time extract MFCC features and comprehensive eye movement features; the feature fusion module is used to fuse the MFCC feature vectors and comprehensive eye movement feature vectors according to time stamps by using the dynamic time warping algorithm to generate comprehensive feature vectors; the cognitive load prediction module is used to construct a two-stream Transformer model based on a hierarchical processing mechanism and a cross-modal attention mechanism, and predict the cognitive load index of students according to the comprehensive feature vectors; the teaching plan generation module is used to construct a virtual teaching environment, optimize the three policy parameters of explanation speed, interaction frequency, and content difficulty in the virtual teaching environment based on the adaptive policy optimization algorithm combined with cognitive load data, and dynamically generate teaching plans.
[0052] This embodiment also provides a computer device applicable to the situation of the cloud education management method based on big data, including: a memory and a processor; the memory is used to store computer-executable instructions, and the processor is used to execute the computer-executable instructions to implement the cloud education management method based on big data proposed in the above embodiment.
[0053] The computer device can be a terminal, which includes a processor, a memory, a communication interface, a display screen, and an input device connected via a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage medium. The communication interface of the computer device is used to communicate with external terminals in a wired or wireless manner, and the wireless manner can be achieved through WIFI, carrier network, NFC (Near Field Communication), or other technologies. The display screen of the computer device can be a liquid crystal display screen or an electronic ink display screen, and the input device of the computer device can be a touch layer covering the display screen, or buttons, trackballs, or touchpads provided on the housing of the computer device, or an external keyboard, touchpad, or mouse, etc.
[0054] This embodiment also provides a storage medium, on which a computer program is stored. When the program is executed by a processor, it implements the method for realizing cloud education management based on big data proposed in the above embodiment; the storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as Static Random Access Memory (SRAM), Electrically Erasable Programmable Read-Only Memory (EEPROM), Erasable Programmable Read Only Memory (EPROM), Programmable Read-Only Memory (PROM), Read-Only Memory (ROM), magnetic memory, flash memory, a magnetic disk, or an optical disc.
[0055] In summary, in the present invention: the dynamic time warping algorithm fuses the MFCC feature vectors and the comprehensive eye movement feature vectors according to timestamps to generate comprehensive feature vectors, realizing the efficient fusion of multi-modal data; by constructing a two-stream Transformer model, combining a temporal embedding layer and a spatial embedding layer, and dynamically fusing temporal features and spatial features through a cross-modal attention mechanism, the accurate prediction of cognitive load is realized, providing a basis for the dynamic optimization of teaching strategies.
[0056] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical solutions of the present invention can be modified or equivalently replaced without departing from the spirit and scope of the technical solutions of the present invention, and they should all be covered by the scope of the claims of the present invention.
Claims
1. A cloud-based education management method based on big data, characterized by: include, Collecting students' eye movement data and speech data and preprocessing them, and extracting MFCC features and comprehensive eye movement features, the eye movement data including gaze point coordinates, gaze duration and pupil diameter change, and the speech data including fundamental frequency, speech rate, silence interval and formant; The dynamic time warping algorithm is used to fuse the MFCC feature vector and the comprehensive eye movement feature vector according to the timestamp to generate a comprehensive feature vector; Based on the hierarchical processing mechanism and cross-modal attention mechanism, a two-stream Transformer model is constructed, and the students' cognitive load index is predicted according to the comprehensive feature vector; Construct a virtual teaching environment, optimize the three strategic parameters of explanation speed, interaction frequency and content difficulty in the virtual teaching environment based on the adaptive strategy optimization algorithm combined with cognitive load data, and dynamically generate teaching plans.
2. The cloud-based education management method based on big data as claimed in claim 1, characterized in that: The extraction of MFCC features and comprehensive eye movement features comprises the following steps: The speech data is divided into frames, each frame signal is subjected to fast Fourier transform to generate a spectrum diagram, and the MFCC feature vector is extracted through discrete cosine transform; The gaze heat map is generated by performing Gaussian kernel density estimation on the gaze point coordinates, the distribution statistics of the gaze duration are identified, the pupil diameter change rate and fluctuation frequency are extracted, and the path complexity is calculated by segmenting the scanning path to generate a comprehensive eye movement feature vector.
3. The cloud-based education management method based on big data as claimed in claim 1, characterized in that: The dynamic time warping algorithm is used to fuse the MFCC feature vector and the comprehensive eye movement feature vector according to the timestamp to generate a comprehensive feature vector. The specific steps are as follows: Calculate the Euclidean distance between the MFCC feature vector and the eye movement feature vector frame by frame; Based on the Euclidean distance between the MFCC feature vector and the eye movement feature vector, the cost matrix is constructed using the global sequence alignment method; Based on the cost matrix, a dynamic time planning algorithm is used to find the optimal alignment path according to timestamps, and the MFCC features and eye movement features are fused according to the timestamps according to the optimal alignment path to generate a comprehensive feature vector.
4. The cloud-based education management method based on big data as claimed in claim 1, characterized in that: The dual-stream Transformer model is constructed based on the hierarchical processing mechanism and the cross-modal attention mechanism, and the cognitive load index of students is predicted according to the comprehensive feature vector. The specific steps are as follows: The time series embedding layer performs position encoding on the time series of the comprehensive feature vector through Transformer position encoding to generate time series embedding features; The spatial embedding layer uses MLP to perform linear transformation on the comprehensive feature vectors to generate spatial embedding features; The temporal embedding and spatial embedding are spliced through the cross-modal attention mechanism, and the weights of temporal features and spatial features are adjusted in combination with the Sigmoid gate control mechanism. The temporal embedding layer, spatial embedding layer and cross-modal attention mechanism are combined to construct a two-stream Transformer model, and the comprehensive feature vector is input to predict the students' cognitive load index.
5. The cloud-based education management method based on big data as claimed in claim 4, characterized in that: Based on historical student behavior data and students' cognitive load data, a virtual teaching environment is constructed through the Unity3D engine combined with the TensorFlow reinforcement learning framework.
6. The cloud-based education management method based on big data as claimed in claim 4, characterized in that: The adaptive strategy optimization algorithm is constructed as follows: Defining the state space Action space and reward function; According to the state space Action space and students’ cognitive load data, constructing strategy network; Construct a value network based on the state space and reward function; Based on the policy network and value network, an adaptive strategy optimization algorithm is constructed.
7. The cloud-based education management method based on big data as claimed in claim 6, characterized in that: The specific steps of dynamically generating a teaching plan are as follows: Combine the adaptive strategy optimization algorithm with the learner’s cognitive load data to optimize strategy parameters; According to the optimized strategy parameters, the teaching plan is dynamically generated through RLDP.
8. A cloud-based education management system based on big data, based on the cloud-based education management method based on big data according to any one of claims 1 to 7, characterized in that: It includes feature extraction module, feature fusion module, cognitive load prediction module and teaching plan generation module; The feature extraction module is used to collect students' eye movement data and speech data and perform preprocessing, and extract MFCC features and comprehensive eye movement features; A feature fusion module is used to fuse the MFCC feature vector and the comprehensive eye movement feature vector according to the timestamp using a dynamic time warping algorithm to generate a comprehensive feature vector; The cognitive load prediction module is used to build a two-stream Transformer model based on the hierarchical processing mechanism and cross-modal attention mechanism, and predict the students' cognitive load index based on the comprehensive feature vector; The teaching plan generation module is used to build a virtual teaching environment. Based on the adaptive strategy optimization algorithm combined with cognitive load data, it optimizes the three strategy parameters of explanation speed, interaction frequency and content difficulty in the virtual teaching environment and dynamically generates teaching plans.
9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the cloud-based education management method based on big data described in any one of claims 1 to 7 are implemented.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the cloud-based education management method based on big data described in any one of claims 1 to 7 are implemented.
Citation Information
Patent Citations
Teaching method based on electroencephalogram education system, education system, equipment and medium
CN111402643A
Air traffic controller cognitive load assessment method based on multi-feature fusion
CN116595423A
Teaching scheme recommendation method and device based on student behavior analysis
CN116701774A
Student participation degree analysis method in talent social practice teaching based on VR
CN118674168A
Controller working state detection method based on multi-modal cognitive data fusion
CN118761035A
Cited By
User analysis method, device and equipment based on large model and storage medium
CN120524251A
AI teaching interaction method and system based on multi-modal emotion perception and generative strategy optimization
CN121742655A