Intelligent educational administration management system based on digital learning platform
Patent Information
- Application Number
- CN202611182253.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-08-05
- Publication Date
- 2026-09-29
AI Technical Summary
[0003]本发明旨在提供一种基于数字化学习平台的智慧教务管理系统,以解决现有技术中认知状态的实时精准解析不足以及学习资源调度无法动态响应认知状态变化的问题
通过认知状态解析模块执行脑波行为双流分析,对脑电波信号进行多频段功率谱密度分解,提取α波、β波及θ波的相对能量比值,同时将比值与点击间隔时长、页面停留时长等操作轨迹数据进行时间戳对齐融合,生成认知负荷时序特征矩阵,再调用注意力残差网络对特征矩阵进行空间
时间联合编码,得到实时认知负荷指数;对鼠标移动路径及视线聚焦热点图进行曲率方差分析,计算视线偏离焦点中心点的欧氏距离序列并输入卡尔曼滤波器进行状态估计,得到注意力漂移向量。该方案将脑电生理信号的频域特征与操作行为的时序特征在统一时间轴上融合,使认知负荷的解析不再依赖单一行为推断,而是获得与神经活动直接关联的多维度量化指标,同时注意力漂移向量通过空间轨迹的方差分析及滤波估计,实现对视线偏移方向和幅度的连续跟踪,从而提高认知状态辨识的实时性和精细度。内容适配引擎模块以实时认知负荷指数及注意力漂移向量作为状态输入,通过深度强化学习网络的策略网络层计算动作概率分布,生成候选动作集,再利用价值网络层对每个候选动作进行长期累积奖励估计,结合贪心
随机混合策略选取最优动作,输出个性化学习资源调度策略。该动态路径规划方式允许系统在每一次状态变化时即时评估不同资源调整动作的远期收益,而非基于预设规则或静态模型,使得学习资源的呈现顺序和难度级别能够根据认知负荷的高低及注意力偏移程度持续优化,在注意力分散时自动切换至焦点引导内容,在认知负荷过大时降低材料复杂度,维持学习过程的高效性与适应性。
Smart Images

Figure CN122840804A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of digital learning and academic affairs management technology, specifically to a smart academic affairs management system based on a digital learning platform. Background Technology
[0002] Digital learning platforms are widely used in academic affairs management, but existing intelligent academic affairs management systems typically rely solely on static analysis of learning behavior logs and historical grade data, failing to accurately capture changes in students' real-time cognitive states during the learning process. Conventional technical solutions often infer learning engagement through behavioral parameters such as clickstreams and page dwell time, lacking the integration and utilization of deeper cognitive physiological signals such as EEG responses, resulting in a coarse-grained portrayal of student attention shifts and cognitive load fluctuations. In existing technologies, learning resource recommendation strategies generally employ collaborative filtering or rule-based matching methods, which cannot dynamically adjust the order, difficulty, and presentation of learning content based on real-time changes in cognitive states. When students' attention drifts or cognitive load exceeds an appropriate range, the system cannot immediately replan the learning path, leading to a disconnect between the recommended resources and current cognitive needs, affecting the effectiveness of personalized teaching. Joint analysis of real-time cognitive load indices and attention drift vectors requires extracting high-dimensional joint representations from brainwave frequency energy distribution and operational behavior temporal characteristics, and constructing a dynamic scheduling mechanism capable of handling sequential decision-making. Conventional independent behavioral analysis or post-event EEG analysis cannot provide a real-time closed-loop cognitive state-driven scheduling scheme. Summary of the Invention
[0003] The present invention aims to provide a smart academic affairs management system based on a digital learning platform to solve the problems of insufficient real-time and accurate analysis of cognitive state and the inability of learning resource scheduling to dynamically respond to changes in cognitive state in the existing technology.
[0004] To achieve the above objectives, the present invention provides the following technical solution: The present invention provides a smart academic affairs management system based on a digital learning platform. This system acquires user operation trajectory data and biometric feedback data on the learning platform through a learning behavior acquisition module, thereby achieving multi-dimensional perception of the user's learning status. A cognitive state analysis module receives the operation trajectory data and biometric feedback data and executes brainwave-based... Behavioral dual-stream analysis generates a real-time cognitive load index and attention drift vector for each user, accurately depicting their current energy input and attention distribution. The content adaptation engine module utilizes a pre-trained deep reinforcement learning network model to dynamically plan learning paths based on the real-time cognitive load index and attention drift vector, generating personalized learning resource scheduling strategies. This ensures that the timing, presentation order, and difficulty level of learning resources are dynamically matched to the user's immediate cognitive state. The knowledge graph construction module extracts concept nodes and relationships from the learning content based on the personalized learning resource scheduling strategy, constructing a user-specific knowledge mastery graph. This enables visual representation and quantitative evaluation of individual knowledge structures. The academic decision optimization module reallocates teaching resources based on the topological characteristics of the knowledge mastery graph, generating course schedule adjustment tables and tutoring intervention plans. This directly transforms individual learning status analysis results into actionable academic management actions, improving the targeting and efficiency of teaching resource allocation.
[0005] As a technical solution of the present invention, the cognitive state analysis module performs brainwave-based analysis on operation trajectory data and biometric feedback data. The specific implementation of behavioral dual-stream analysis includes: performing multi-band power spectral density decomposition on the EEG signals in the biometric feedback data, extracting the relative energy ratios of alpha, beta, and theta waves, and then aligning and fusing these relative energy ratios with the click interval duration and page dwell duration in the operation trajectory data using timestamps to generate a cognitive load temporal feature matrix. A pre-trained attention residual network is then invoked to spatially analyze this cognitive load temporal feature matrix. Temporal features are jointly encoded to obtain a real-time cognitive load index. Simultaneously, curvature variance analysis is performed on the mouse movement path and gaze focus heatmap in the operation trajectory data to calculate the Euclidean distance sequence of the gaze deviation from the focal point, which is then input into a Kalman filter for state estimation, generating an attention drift vector. This process allows the assessment of cognitive state to rely not only on explicit behavioral indicators but also to integrate implicit electroencephalographic features, significantly improving the accuracy and real-time performance of cognitive load and attentional jitter detection.
[0006] Preferably, the content adaptation engine module calls a deep reinforcement learning network model for dynamic learning path planning as follows: Real-time cognitive load index and attention drift vector are used as state inputs, which are then input into the policy network layer of the deep reinforcement learning network model to calculate the action probability distribution, generating a candidate action set. Each candidate action corresponds to a learning resource node adjustment instruction. The value network layer performs long-term cumulative reward estimation for each candidate action in the candidate action set, generating an action value score set. Based on this score set, a greedy algorithm is then used... A randomized hybrid strategy selects the optimal action in the current state to adjust the presentation order or difficulty level of learning resources. The corresponding learning resource node adjustment instruction for this optimal action is output as a personalized learning resource scheduling strategy. This approach enables the adjustment of the learning path to maximize long-term learning benefits, avoiding frequent content switching due to short-term cognitive fluctuations, thereby achieving a balance between maintaining a stable learning state and improving learning effectiveness.
[0007] Furthermore, the implementation of the knowledge graph construction module in constructing a user-specific knowledge mastery graph based on a personalized learning resource scheduling strategy includes: performing named entity recognition on the learning resource text targeted by the scheduling strategy to extract a set of subject concept entities; performing co-occurrence frequency analysis and semantic dependency parsing on each pair of entities in the set to generate a correlation strength rating matrix between entities. A pre-trained graph attention network model is then invoked, using this correlation strength rating matrix as edge weights, to perform node representation learning on the set of subject concept entities, generating an initial knowledge graph with embedded vector representations. Subsequently, the initial knowledge graph is corrected based on the user's historical answer accuracy, and the corrected edge weights are updated in the graph to obtain the user-specific knowledge mastery graph. The knowledge graph constructed in this way can truly reflect an individual's mastery of the relationships between concepts, overcoming the problem of false associations caused by relying solely on content co-occurrence.
[0008] The specific method by which the academic affairs decision optimization module reallocates teaching resources based on the topological characteristics of the knowledge mastery graph is as follows: It calculates the degree centrality and betweenness centrality of each concept node in the knowledge mastery graph to identify the set of key concept nodes with weak mastery; based on the comparison of the coverage of this set of weak nodes with the preset course syllabus, it calculates the teaching resource gap matrix. A pre-trained swarm intelligence optimization algorithm is then used to allocate teachers and rearrange class hours on the teaching resource gap matrix, generating a course schedule adjustment table; simultaneously, based on the path distance between nodes in the set of weak nodes, a tutoring intervention plan containing recommended tutoring sequences for each weak node is generated. This decision-making process enables the school's academic affairs department to adjust class scheduling and tutoring resources based on objective knowledge mastery data rather than experience-based judgment, reducing resource mismatch and improving the accuracy of tutoring intervention.
[0009] A preferred implementation method for the behavior acquisition module in this system to obtain operation trajectory data and biometric feedback data is as follows: A behavior capture script is embedded in the user terminal to capture the timing data of mouse click events, keyboard input events, and page scrolling events through an event listening interface. After data cleaning, operation trajectory data is generated. Simultaneously, a wearable EEG device is connected via Bluetooth to acquire raw EEG signals from the user's prefrontal cortex at a preset sampling rate. This signal is then subjected to bandpass filtering and independent component analysis (ICA) for noise reduction, extracting clean EEG feature waveforms. A timestamp synchronization unit is used to perform frame alignment processing between the operation trajectory data and the clean EEG feature waveforms, generating structured data packets for subsequent analysis. This scheme ensures the temporal consistency of behavioral and physiological data, providing a basis for EEG data processing. Behavioral dual-stream analysis provides a high-quality data foundation.
[0010] As a preferred method for training deep reinforcement learning network models, the training approach includes: constructing a historical learning resource library containing labels of different difficulty levels and knowledge domains; defining a comprehensive reward function based on learning completion time, test accuracy, and cognitive load fluctuations; generating a training sample set containing state features, action vectors, and reward values using the historical learning resource library; and iteratively training the policy network layer and value network layer using a proximal policy optimization algorithm until the reward values converge. The model trained in this way can fully learn the optimal scheduling strategy under different cognitive states, resulting in highly adaptive and robust decision-making.
[0011] Regarding the application integration of the academic affairs decision optimization module, this module is also used to convert course schedule adjustment tables and tutoring intervention plans into data formats conforming to the academic affairs management protocol, and push them to the academic affairs management system via application programming interfaces (APIs). The APIs include course scheduling interfaces and tutoring task allocation interfaces. The course schedule adjustment table, in JSON structure format, includes course identifier, adjusted time slot, and instructor ID fields. The tutoring intervention plan, in XML structure format, includes student ID, tutoring concept sequence, and recommended tutoring duration fields. This design enables the system to seamlessly integrate with existing academic affairs management platforms, reducing deployment and integration complexity.
[0012] In addition, this system includes a learning performance evaluation module. Within a preset evaluation period, it collects users' test scores and learning time data, calls a pre-trained Bayesian knowledge tracing model to update the knowledge state of this data, and generates a skill mastery probability matrix for the user. The system then compares the node similarity of this skill mastery probability matrix with the knowledge mastery graph, calculates the forgetting curve decay coefficient, and feeds this decay coefficient back to the cognitive state analysis module to update the calculation parameters of the attention drift vector. Through this closed-loop feedback mechanism, the system can continuously calibrate the cognitive state estimation model, making the attention assessment more closely reflect the user's actual knowledge forgetting patterns.
[0013] The system further includes a resource index database for storing learning resource files, including video files, document files, and interactive exercise files, that are targeted by personalized learning resource scheduling strategies. This resource index database is organized using a content-addressed file system, with each resource file corresponding to a hash address. The content adaptation engine module retrieves the corresponding learning resource file by querying this hash address and pushes the file to the user terminal via a streaming media transmission protocol. Utilizing content addressing ensures efficient resource retrieval and tamper-proof characteristics, while streaming media transmission guarantees smooth presentation on the user end, reducing the impact of waiting delays on the learning process.
[0014] The technical effects and advantages provided by the present invention in the above technical solution are as follows: Brainwaves are executed through the cognitive state analysis module. Behavioral dual-stream analysis decomposes EEG signals into multi-band power spectral density, extracting the relative energy ratios of alpha, beta, and theta waves. These ratios are then fused with operation trajectory data such as click interval duration and page dwell time using timestamp alignment to generate a cognitive load temporal feature matrix. Finally, an attention residual network is invoked to spatially process this feature matrix. Temporal joint encoding yields a real-time cognitive load index. Curvature variance analysis is performed on the mouse movement path and gaze focus heatmap to calculate the Euclidean distance sequence of gaze deviation from the focal point, which is then input into a Kalman filter for state estimation, resulting in an attention drift vector. This scheme fuses the frequency domain features of EEG signals with the temporal features of operational behavior on a unified time axis. This allows cognitive load analysis to move beyond relying on single-behavioral inferences and instead obtain multi-dimensional quantitative indicators directly related to neural activity. Simultaneously, the attention drift vector, through variance analysis and filtering estimation of spatial trajectories, enables continuous tracking of the direction and magnitude of gaze deviation, thereby improving the real-time performance and precision of cognitive state identification. The content adaptation engine module uses the real-time cognitive load index and attention drift vector as state inputs. It calculates the action probability distribution through the policy network layer of a deep reinforcement learning network, generating a candidate action set. Then, the value network layer performs long-term cumulative reward estimation for each candidate action, combined with a greedy algorithm. A randomized hybrid strategy selects the optimal action and outputs a personalized learning resource scheduling strategy. This dynamic path planning approach allows the system to evaluate the long-term benefits of different resource adjustment actions in real time at each state change, rather than based on preset rules or static models. This enables the presentation order and difficulty level of learning resources to be continuously optimized according to the level of cognitive load and the degree of attention shift. When attention is distracted, the system automatically switches to focus-guided content, and when the cognitive load is too high, it reduces the complexity of the material, maintaining the efficiency and adaptability of the learning process. Attached Figure Description
[0015] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in this invention. For those skilled in the art, other drawings can be obtained based on these drawings.
[0016] Figure 1 This is a schematic diagram of the intelligent academic affairs management system. Figure 2 This is a flowchart of deep reinforcement learning network model training and personalized learning resource scheduling. Figure 3 It is a curve showing the changes in reward and loss values during the training process of a deep reinforcement learning network; Figure 4 This is a comparison chart of the correlation strength score distribution curves. Detailed Implementation
[0017] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0018] See Figure 1 This invention provides a smart academic affairs management system based on a digital learning platform, including a learning behavior acquisition module, a cognitive state analysis module, a content adaptation engine module, a knowledge graph construction module, and an academic affairs decision optimization module. The learning behavior acquisition module acquires user operation trajectory data and biometric feedback data on the learning platform; the cognitive state analysis module performs brainwave-behavioral dual-stream analysis on the operation trajectory data and biometric feedback data to generate the user's real-time cognitive load index and attention drift vector; the content adaptation engine module uses a pre-trained deep reinforcement learning network model to perform dynamic learning path planning on the real-time cognitive load index and attention drift vector, generating a personalized learning resource scheduling strategy; the knowledge graph construction module extracts concept nodes and relationships from the learning content based on the personalized learning resource scheduling strategy, constructing a user-specific knowledge mastery graph; and the academic affairs decision optimization module performs teaching resource reallocation processing based on the topological characteristics of the knowledge mastery graph, generating a course schedule adjustment table and tutoring intervention plan.
[0019] Example 1: In practice, the learning behavior acquisition module obtains operation trajectory data through a behavior capture script embedded in the user terminal. The behavior capture script utilizes the browser's native event listening interface to capture mouse click events, keyboard input events, and page scrolling events, recording the timestamp, event type, and associated page coordinates or key values for each event, forming a raw event stream. The raw event stream is then cleaned to remove duplicate event records caused by system lag and automatically triggered events from the browser. The cleaned time-series data is organized into operation trajectory data, with each record in the operation trajectory data containing the event type, trigger timestamp, and event parameters.
[0020] While acquiring operational trajectory data, the learning behavior acquisition module establishes a data transmission channel with the user's wearable EEG device via Bluetooth Low Energy protocol, acquiring raw EEG signals from the Fp1 and Fp2 electrode positions in the user's prefrontal cortex at a preset sampling rate of 256 Hz. The raw EEG signals are then bandpass filtered from 0.5 Hz to 30 Hz to remove DC drift components and high-frequency electromyographic interference. Independent component analysis is performed on the bandpass-filtered raw EEG signals, decomposing the multi-channel signal into statistically independent components. Independent components corresponding to eye movement artifacts are identified and removed, reconstructing a clean EEG characteristic waveform.
[0021] The learning behavior acquisition module includes a timestamp synchronization unit. This unit uses the network time protocol server as a reference and aligns the timestamps of each record in the operation trajectory data with the timestamps of the clean EEG feature waveforms to generate a structured data packet. The structured data packet contains timetamp-aligned behavioral feature vectors and EEG feature vectors. The behavioral feature vectors consist of click interval duration, page dwell time, mouse movement speed, and scrolling speed, while the EEG feature vectors consist of the EEG voltage amplitude at the corresponding moment.
[0022] In generating the real-time cognitive load index, the cognitive state analysis module performs multi-band power spectral density decomposition on the EEG waveforms contained in the biometric feedback data. Using the Welch method, the EEG waveforms are segmented into short time segments with a time window length of 2 seconds and 50% overlap. A Hamming window is applied to each short segment, and a periodogram is calculated. The power spectral density estimate is obtained by averaging the periodograms of all short segments. The power spectral density estimate is integrated to extract the power from the α-wave band (8 Hz to 13 Hz), the β-wave band (13 Hz to 30 Hz), and the θ-wave band (4 Hz to 8 Hz). The ratio of α-wave power to θ-wave power is calculated as the first relative energy ratio, and the ratio of β-wave power to α-wave power is calculated as the second relative energy ratio. The first and second relative energy ratios are then used as a set of relative energy ratios.
[0023] The cognitive state analysis module performs timestamp alignment and fusion processing on the set of relative energy ratios and the click interval duration and page dwell time in the operation trajectory data. Using the unified timestamp in the structured data package as an index, the first relative energy ratio, second relative energy ratio, click interval duration, and page dwell time at the corresponding moment are concatenated into a four-dimensional feature vector. All four-dimensional feature vectors are arranged in chronological order to generate a cognitive load temporal feature matrix.
[0024] The cognitive state parsing module calls a pre-trained attention residual network to perform joint spatial-temporal feature encoding on the cognitive load temporal feature matrix. The network architecture of the attention residual network consists of an input mapping layer, three stacked attention residual blocks, and an output prediction layer. The input mapping layer is a fully connected layer that maps the four-dimensional feature vector at each time step in the cognitive load temporal feature matrix to a 64-dimensional embedding vector. Each attention residual block contains a multi-head self-attention layer and a position-wise feedforward network layer. The multi-head self-attention layer has four heads, each with a dimension of 16. Within the attention residual block, the input vector sequence is first normalized, then fed into the multi-head self-attention layer to calculate the attention-weighted representation. This attention-weighted representation is residually concatenated with the input vector sequence, and then normalized again. The result is fed into the position-wise feedforward network layer, which uses two fully connected layers with a hidden layer dimension of 128. The output is then residually concatenated with the vector after the first residual concatenation to obtain the output of the attention residual block. Three attention residual blocks are stacked sequentially, and the output vector sequence of the last attention residual block is fed into the output prediction layer. The output prediction layer applies a fully connected layer with a sigmoid activation function to the 64-dimensional output vector of the last time step of the sequence, mapping it to a scalar. This scalar is the real-time cognitive load index, which ranges from 0 to 1, with higher values indicating greater cognitive load.
[0025] The pre-trained attention residual network is trained using a pre-constructed cognitive load labeled dataset. This dataset contains multiple temporal feature matrix samples of historical cognitive load and corresponding cognitive load index labels. The cognitive load index labels are obtained by averaging the scores given independently by three educational psychology experts based on user behavior playback videos and EEG rhythm maps. During training, the temporal feature matrix of cognitive load is used as input, with mean squared error as the loss function. An adaptive moment estimation optimizer is employed, with an initial learning rate of 0.0005, a batch size of 32, and 60 iterations. Training stops when the validation set loss no longer decreases for five consecutive iterations, and the model parameters with the minimum validation set loss are saved as pre-training weights.
[0026] During the generation of attention drift vectors, the cognitive state analysis module performs curvature variance analysis on the mouse movement path and gaze focus heatmap in the operation trajectory data. The mouse movement path is represented by a series of screen coordinate points ordered by time. For every three consecutive screen coordinate points, the curvature of the circumcircle defined by these three points is calculated. The variance of all curvature values within the sliding window is calculated to obtain the curvature variance sequence of the mouse movement path. The gaze focus heatmap is constructed using the gaze point coordinates collected by the eye-tracking device. The variance of the two-dimensional spatial distribution of the gaze point coordinate set within a fixed time window is calculated. The variance value is used as the dispersion of the gaze distribution. At the same time, the gaze focus center point is calculated, which is the arithmetic mean coordinate of the gaze point coordinates.
[0027] For the center point of the gaze focus, the cognitive state analysis module calculates the Euclidean distance between the user's actual gaze point coordinates and the coordinates of the center point of the gaze focus at each moment, forming a sequence of Euclidean distances between the gaze and the center point of the gaze focus in chronological order.
[0028] The cognitive state analysis module inputs the Euclidean distance sequence between the line of sight and the focal point into the Kalman filter for state estimation. ,in, This represents the x-axis coordinate of the focus of attention at time t in the screen coordinate system. This represents the y-axis coordinate of the focus of attention at time t in the screen coordinate system. This represents the speed at which the focus of attention moves along the x-axis at time t in the screen coordinate system. This represents the velocity of the attention focus along the y-axis at time t in the screen coordinate system. The Kalman filter's time update equation uses a uniform motion model, mapping the state vector from the previous time step to the predicted state vector at the current time step through a state transition matrix. The observation vector of the Kalman filter is set as the Euclidean distance between the line of sight and the center point of the focus. The observation matrix is constructed based on the relationship between the predicted position and the coordinates of the center point of the focus, mapping the predicted state vector to the predicted Euclidean distance value. The process noise covariance matrix is set as a diagonal matrix, with its diagonal elements set to 0.01 based on the physiological inertia characteristics of attention shift in the prefrontal cortex. The observation noise covariance scalar is set to 0.5 based on the nominal measurement error of the eye-tracking device. After recursively calculating prediction and update at each time step, the Kalman filter outputs a posterior state estimation vector. The combination of the two-dimensional position coordinate components and the two-dimensional velocity components extracted from the posterior state estimation vector is the attention drift vector, which represents the instantaneous position, direction of movement, and velocity of the attention focus on the screen.
[0029] Example 2: In specific implementation, please refer to Figure 2The training process of a deep reinforcement learning network model begins with the construction of a historical learning resource repository. This repository contains multiple historical learning records, each corresponding to a learning resource node. These nodes are labeled with difficulty and knowledge domain tags. The difficulty tags are divided into four levels: "memory," "understanding," "application," and "analysis." The knowledge domain tags use subject-specific knowledge classification codes, with each subject's knowledge domain tags stored according to a three-level classification code. The historical learning resource repository stores records associated with each learning resource node, including the user's completion time for that node, test accuracy, cognitive load fluctuations during the learning process, and the node identifier.
[0030] After constructing the historical learning resource repository, a comprehensive reward function is defined for training the deep reinforcement learning network model. The expression for the comprehensive reward function is: in, This represents the overall reward value; This indicates the maximum standard learning time limit specified by the learning resource node. Based on the average grade level, for video resources Set to 1.2 times the total video length for text-based resources. Set to 1.5 times the average reading time of the text; This indicates the actual learning completion time for the user, extracted from the operation trajectory data recorded by the learning behavior collection module, representing the time difference from the start of resource loading to the end of learning; This indicates the user's accuracy rate in the quiz associated with this learning resource node, with the accuracy rate ranging from 0 to 1; This represents the standard deviation of the cognitive load fluctuation during the user's learning of this resource node. The cognitive load fluctuation sequence is composed of the real-time cognitive load index output by the cognitive state analysis module within the corresponding time period. The standard deviation of the real-time cognitive load index sequence is calculated to obtain... ; , , These are the weighting coefficients. The value of 0.3 is chosen because learning efficiency should account for a significant proportion of the overall reward, but not exceed the importance of accuracy. The value is set at 0.5 because test accuracy is a core indicator for measuring learning effectiveness. The value of 0.2 is chosen because cognitive load stability has an auxiliary evaluative effect on the learning experience. , , The sum of them is always 1.
[0031] After defining the comprehensive reward function, a training sample set is generated using the historical learning resource library. Each training sample in the training sample set consists of state features, action vectors, and a reward value. State features are extracted from historical learning records, including the real-time cognitive load index and attention drift vector prior to the learning record. The real-time cognitive load index is a one-dimensional scalar, and the attention drift vector is a four-dimensional vector; therefore, the state feature is a five-dimensional vector. Action vectors are selected from historical learning records. Each dimension of the action vector corresponds to the encoding of a learning resource node adjustment instruction. These instructions include five types: "increase learning resource difficulty," "decrease learning resource difficulty," "present subsequent resources earlier," "delay presenting subsequent resources," and "keep current resources unchanged." The action vector uses a unique encoding method, has a length of 5, and each dimension corresponds to one of the five instructions. The reward value is calculated using the comprehensive reward function based on the learning completion time, test accuracy, and cognitive load fluctuations in the historical learning record. The reward value is associated with the action vector at time steps. The state features of each training sample correspond to a time step, the action vector corresponds to the action performed at that time step, and the reward value is the immediate reward obtained from the learning environment after performing the action.
[0032] The deep reinforcement learning network model consists of a policy network layer and a value network layer. The core architecture of the policy network layer is a multilayer perceptron, composed of an input layer, two hidden layers, and an output layer. The input layer has 5 neurons, corresponding to the dimension of the state features. The first hidden layer contains 128 neurons, using a modified linear unit (MRU) activation function. The second hidden layer contains 64 neurons, also using a MRU activation function. The output layer contains 5 neurons, corresponding to five different learning resource node adjustment instructions, and uses a softmax function to generate the action probability distribution. The core architecture of the value network layer is also a multilayer perceptron, consisting of an input layer, two hidden layers, and an output layer. The input layer of the value network layer has 5 neurons, consistent with the dimension of the state features. The first hidden layer contains 128 neurons, using a MRU activation function. The second hidden layer contains 64 neurons, also using a MRU activation function. The output layer contains 1 neuron with no activation function, used to output the estimated state value.
[0033] The proximal policy optimization algorithm is used to iteratively train the policy network layer and the value network layer. During training, samples with a batch size of 64 are randomly selected from the training sample set, and a pruning parameter is set. It is 0.2. Used to limit the change in the probability ratio of the new and old strategies from exceeding the range. The setting of 0.2 is based on allowing for some exploration while maintaining training stability. The loss function of the value network layer is the mean square value of the temporal difference error, the loss function of the policy network layer uses a cut-off objective function from the proximal policy optimization, the optimizer uses an adaptive moment estimator, and the learning rate of the policy network layer is set to... The learning rate of the value network layer is set to In each iteration, an experience trajectory of a fixed length of 2048 steps is first collected, the generalized advantage estimate is calculated, and then 10 rounds of optimization and updates are performed on the policy network layer and the value network layer. After 50 rounds of iterative training, when the average reward value changes by less than 0.01 for 5 consecutive rounds, the reward value is considered to have converged, training is stopped, and the current policy network layer parameters and value network layer parameters are saved.
[0034] When the content adaptation engine module calls the trained deep reinforcement learning network model, it concatenates the real-time cognitive load index and attention drift vector generated by the cognitive state analysis module into a five-dimensional state vector, which is then input into the policy network layer of the deep reinforcement learning network model to calculate the action probability distribution. The forward propagation of the policy network layer is executed sequentially at different time steps. For the state vector at the current time step, after calculation through the input layer, the first hidden layer, and the second hidden layer, a five-dimensional action probability vector is output at the output layer using the Softmax function. The value of each dimension in the five-dimensional action probability vector represents the probability value of selecting the corresponding learning resource node adjustment instruction. The action probability vector constitutes a candidate action set, and each candidate action in the candidate action set corresponds to a learning resource node adjustment instruction.
[0035] The content adaptation engine module performs long-term cumulative reward estimation for each candidate action in the candidate action set through the value network layer of a deep reinforcement learning network model. For the learning resource node adjustment instruction corresponding to each candidate action in the candidate action set, the instruction is taken as the assumed action to be executed. Combined with the current state vector, the next state vector is predicted through the environment transition model. The next state vector is then input into the value network layer. The forward propagation output of the value network layer is a scalar as the state value estimate. The state value estimate is added to the immediate comprehensive reward value to obtain the long-term cumulative reward estimate. After performing the above calculation on all candidate actions in the candidate action set, an action value score set is obtained, which contains the action value score of each candidate action.
[0036] The content adaptation engine module uses a greedy-random hybrid strategy based on the action value score set to select the optimal action in the current state. In the greedy-random hybrid strategy, the candidate action with the highest action value score is selected as the optimal action with a probability of 0.9, and a candidate action is randomly and evenly selected from the candidate action set with a probability of 0.1. The probability value of 0.9 is set to utilize existing experience while preserving opportunities to try actions that have not been fully explored. The learning resource node adjustment instruction corresponding to the optimal action is used to adjust the presentation order or difficulty level of learning resources. The content adaptation engine module encapsulates the learning resource node adjustment instruction corresponding to the optimal action into a personalized learning resource scheduling strategy. The personalized learning resource scheduling strategy includes the instruction type, the target learning resource node identifier, and adjustment parameters.
[0037] See Figure 3 In the graph, the horizontal axis represents the number of training steps, ranging from 0 to 2000; the left vertical axis represents the average reward value; and the right vertical axis represents the loss value. Three curves are plotted: the solid blue line represents the average reward value curve, corresponding to the left vertical axis; the dashed red line represents the value network loss curve; and the dotted purple line represents the policy network loss curve, all corresponding to the right vertical axis.
[0038] As the number of training steps increases, the average reward value curve generally shows a fluctuating trend of first rising rapidly and then stabilizing. In the early stage of training (from 0 to about 500 steps), the average reward value quickly climbs from a negative value to a positive value close to 0.5, indicating that the deep reinforcement learning network model has a rapid improvement in learning performance in the early training stage. Subsequently, in the middle training stage (about 500 to 1500 steps), the average reward value continues to rise slowly to about 2.5, accompanied by small oscillations, showing that the model strategy is gradually optimized and tends to stabilize. In the later training stage (after about 1500 steps), the average reward value remains at a high level, indicating that the training has reached an optimal state.
[0039] The loss curve of the value network starts to decrease rapidly from a relatively high value of about 2.5 at the initial time, and approaches zero after about 1000 training steps, accompanied by small fluctuations. This indicates that the state value estimation error of the value network layer gradually decreases and the model performance is effectively improved.
[0040] The policy network loss curve generally shows a downward trend, initially dropping rapidly to approximately 0.75, then continuing to decrease at a slower rate, stabilizing around 0.15 towards the end of training. This indicates that the parameters of the policy network layers gradually converge, and the policy probability distribution gradually stabilizes. The loss curve exhibits periodic small fluctuations, reflecting the exploration changes brought about by the greedy-random hybrid policy used during training.
[0041] Example 3: In practice, the knowledge graph construction module obtains the personalized learning resource scheduling strategy output by the content adaptation engine module. The personalized learning resource scheduling strategy contains the identifier of the target learning resource node. Based on the identifier of the target learning resource node, the knowledge graph construction module retrieves the corresponding learning resource text from the resource index database.
[0042] The knowledge graph construction module performs named entity recognition (NAME) processing on the retrieved learning resource text. The NAME recognition process employs a pre-trained NAME recognition model based on a bidirectional long short-term memory (BSSM) network and a conditional random field (CRF). The architecture of this model consists of a word embedding layer, a bidirectional BSSM network layer, and a CRF decoding layer. The word embedding layer transforms each input word in the learning resource text into a word vector, with a 300-dimensional word vector. The bidirectional BSSM network layer is composed of a forward BSSM network and a backward BSSM network connected in series. The forward BSSM network receives the word vector sequence from the beginning of the learning resource text to the current word position and extracts contextual semantic features. The backward BSSM network receives the word vector sequence from the end of the learning resource text to the current word position and extracts contextual semantic features. The hidden layer dimensions of both the forward and backward BSSM networks are set to 256. The hidden states of the forward and backward BSSM networks at each time step are concatenated into a 512-dimensional vector. The Conditional Random Field (CRF) decoding layer receives a 512-dimensional vector from all time steps of the bidirectional Long Short-Term Memory (LSTM) network layer. This vector is mapped to an emission probability matrix in the label space dimension via a fully connected layer. The CRF decoding layer maintains a label transition matrix internally. The globally optimal label sequence is solved using the Viterbi algorithm on the emission probability matrix and the label transition matrix. The label sequence adopts the BIO (Browser-Indexed I / O) labeling pattern, where "B-CONCEPT" represents the starting word of a subject concept entity, "I-CONCEPT" represents the internal word of a subject concept entity, and "O" represents a non-concept word. All starting words labeled "B-CONCEPT" and their immediately following consecutive words labeled "I-CONCEPT" are extracted from the globally optimal label sequence. These starting words and consecutive words are then concatenated sequentially to form a set of subject concept entities.
[0043] In the specific implementation, the pre-training process of the named entity recognition model based on bidirectional long short-term memory network and conditional random field uses a manually annotated subject-specific text corpus. The subject-specific text corpus contains 20,000 sentences, and the subject-specific concept entities in each sentence are labeled with "B-CONCEPT" and "I-CONCEPT" boundary tags. During training, sentences from the subject-specific text corpus are input into the word embedding layer to obtain word vector sequences. After forward propagation, the conditional random field decoding layer outputs the predicted label sequence. The loss function is the negative log-likelihood loss function of conditional random field, and the optimizer is the stochastic gradient descent optimizer. The learning rate of the stochastic gradient descent optimizer is set to 0.005, the momentum is set to 0.9, the batch size is set to 32, and the training is iterated for 40 rounds. After each round, the entity recognition F1 score is calculated on the validation set, and the model parameters with the highest F1 score on the validation set are saved as pre-training weights.
[0044] After extracting the subject concept entity set, the knowledge graph construction module performs co-occurrence frequency analysis and semantic dependency relation parsing on each pair of entities in the subject concept entity set. The co-occurrence frequency analysis uses the full-text corpus consisting of learning resource texts corresponding to all courses in the learning platform as the statistical scope. It traverses any two different concept entities in the subject concept entity set, counts the number of times the two concept entities co-occur in the same natural paragraph, divides the number of co-occurrences by the total number of natural paragraphs for normalization, and obtains the co-occurrence frequency score. The co-occurrence frequency score ranges from 0 to 1.
[0045] Semantic dependency parsing calls a pre-trained dependency parsing model based on dual affine attention. The architecture of this model consists of a deep bidirectional encoder representation transformer layer, a bidirectional long short-term memory network layer, and a dual affine attention layer. The deep bidirectional encoder representation transformer layer uses a pre-trained Chinese deep bidirectional encoder representation transformer model to encode the input sentence into a context-dependent word representation sequence. The bidirectional long short-term memory network layer receives the word representation sequence output by the deep bidirectional encoder representation transformer layer and outputs the center word representation vector and dependency word representation vector for each word position. The dual affine attention layer calculates the dependency arc score and dependency relation label score for each pair of word positions and outputs a dependency parsing tree. For any two concept entities in the subject concept entity set, the shortest path between the nodes containing the two concept entities is determined in the dependency parsing tree. The number of edges contained in the shortest path is extracted as the dependency path length, and the dependency relation type sequence on the shortest path is extracted. Based on a pre-defined dependency relationship weight mapping table, each dependency relationship type in the dependency relationship type sequence is converted into a base score. The sum of the base scores is divided by the dependency path length plus 1 to obtain the semantic dependency score, which is normalized to a range of 0 to 1. In the pre-defined dependency relationship weight mapping table, the base score for "Agent Relationship" and "Patient Relationship" is set to 0.9, the base score for "Limited Relationship" is set to 0.8, the base score for "Parallel Relationship" is set to 0.7, the base score for "Supplementary Relationship" is set to 0.5, and the base score for "Punctuation Relationship" is set to 0.1.
[0046] After obtaining the co-occurrence frequency score and semantic dependency score for each pair of concept entities, the knowledge graph construction module performs a weighted sum of the co-occurrence frequency score and the semantic dependency score. The weight of the co-occurrence frequency score is set to 0.6, and the weight of the semantic dependency score is set to 0.4. The weighted sum is the association strength score for that pair of concept entities. The association strength scores of all concept entity pairs are arranged in entity order to form an association strength score matrix. The i-th element of the association strength score matrix... Line 1 Column elements represent the first The first conceptual entity and the first A score for the strength of association between conceptual entities.
[0047] The knowledge graph construction module calls a pre-trained graph attention network model, using the association strength rating matrix as edge weights, to perform node representation learning on the subject concept entity set. The graph attention network model architecture consists of a first graph attention layer and a second graph attention layer stacked together. The first graph attention layer receives the initial node feature vector and association strength rating matrix for each concept entity in the subject concept entity set. The initial node feature vector is obtained by querying a pre-trained 300-dimensional word vector matrix, and the dimension of the initial node feature vector is 300.
[0048] In the first graph attention layer, for any conceptual entity and conceptual entities Any concept entity in the set of neighboring concept entities with a non-zero association strength score Calculate the attention coefficient. The calculation process for the attention coefficient is as follows: [The text abruptly shifts to a different topic] ...the concept entity... Node feature vectors and concept entities Each node feature vector is multiplied by a learnable linear transformation matrix, the dimension of which is... This process yields two 128-dimensional linear transformation vectors. These vectors are then concatenated to obtain a 256-dimensional concatenated vector. This concatenated vector is then used as a dot product with the learnable attention weight vector (which has a dimension of 256). The dot product is then applied to the LeakyReLU activation function, which outputs a scalar attention value. The negative slope of the LeakyReLU activation function is set to 0.2. The scalar attention value is then applied to the concept entity. and conceptual entities The weighted attention value is obtained by multiplying the correlation strength scores between concepts and entities. The weighted attention values of all neighboring conceptual entities are normalized using the Softmax function to obtain the attention coefficient for each neighboring conceptual entity. The first attention layer uses four attention heads, each independently performing the attention coefficient calculation process. Each attention head performs a weighted summation of the node feature vectors of neighboring conceptual entities based on the attention coefficient, resulting in a 128-dimensional node feature representation output by that attention head. These 128-dimensional node feature representations from the four attention heads are concatenated into a 512-dimensional vector, which is then transformed to 128 dimensions using a learnable linear transformation matrix. The dimension of the linear transformation matrix is... To obtain the conceptual entity Embed the vector of the intermediate node output by the attention layer of the first graph.
[0049] The second attention layer receives the intermediate node embedding vectors of all concept entities output by the first attention layer. These intermediate node embedding vectors have a dimension of 128. The computational structure of the second attention layer is the same as the first, with four attention heads, each outputting a 128-dimensional vector. The outputs of the four attention heads are concatenated and linearly transformed to output the final node embedding vectors of the concept entities, also with a 128-dimensional vector. All the final node embedding vectors of the concept entities are stacked in entity index order to form an initial knowledge graph represented by embedding vectors. This initial knowledge graph contains a set of nodes, a 128-dimensional embedding vector for each node, and an association strength scoring matrix as the edge weight matrix.
[0050] In practice, the pre-training process of the graph attention network model is performed offline. The training data uses subgraph samples extracted from a public subject knowledge graph database. Each subgraph sample contains a set of entity nodes and their real edge connections. The training task is set as a link prediction task. For each real edge in a subgraph sample, it is used as a positive sample, and two node pairs without edges are randomly sampled as negative samples. During training, the initial node features of the subgraph samples and the pre-computed association strength score matrix are input into the graph attention network model. Forward propagation yields node embedding vectors. The dot product of the node embedding vectors of positive and negative sample node pairs is calculated and mapped to the link existence probability using the sigmoid function. The loss function used is the binary cross-entropy loss function, and the optimizer is the adaptive moment estimation optimizer. The learning rate of the adaptive moment estimation optimizer is set to 0.001, the weight decay coefficient is set to 0.0005, and the training epochs are set to 200 epochs. The average loss value on the validation set is calculated in each epoch. Training is terminated when the average loss value on the validation set does not decrease for 10 consecutive epochs. The model parameters with the minimum average loss value on the validation set are saved as the pre-training weights of the graph attention network model.
[0051] After generating the initial knowledge graph, the knowledge graph construction module performs confidence correction based on the user's historical answer accuracy. The module extracts user historical practice records from the learning platform database. Each record contains the practice question number, the concept entity pair involved in the question, and whether the user's answer to the question was correct. For any pair of concept entities in the initial knowledge graph... and conceptual entities Filter out all instances involving conceptual entities from the user's historical practice records. and conceptual entities The practice question records are used to calculate the ratio of the number of correct answers on these practice question records to the total number of answers. This ratio is used as the accuracy rate of the concept entity pair association, denoted as . A modified formula is used for conceptual entities. and conceptual entities The edge weights between them are adjusted, and the adjustment formula is expressed as: in, Represents the concept entity before the revision With conceptual entities The edge weights between them, i.e., the weights in the association strength score matrix. Line 1 The column's score value; This indicates that the user has views on the conceptual entities involved. With conceptual entities The average accuracy rate of related practice questions. The value range is from 0 to 1; Indicates the corrected strength coefficient. The value is 0.8. The reason for choosing 0.8 is that when the user accuracy changes from 0 to 1, the correction factor... The range of change is The adjustment range of edge weights is limited to 0.6 to 1.4 times the original edge weights, which is sufficient to distinguish between the two extreme states of complete lack of mastery and complete mastery. Represents the revised conceptual entity With conceptual entities The edge weights between concept entity pairs are calculated. The knowledge graph construction module performs the above correction calculation on all concept entity pairs, and then calculates the corrected edge weights. Replace the original edge weights in the association strength score matrix By keeping the node set and node embedding vector unchanged, a user-specific knowledge mastery graph is obtained.
[0052] See Figure 4 In the graph, the horizontal axis represents the association strength score, ranging from 0 to 1, and the vertical axis represents the probability density of the corresponding score. The graph contains three curves, corresponding to the "original association strength score," the "user-specific corrected weight," and the "average weight of the global knowledge graph," respectively.
[0053] The "Original Association Strength Score" curve is represented by a solid line, exhibiting a unimodal distribution. The peak value is located in the range of approximately 0.20 to 0.25, with a maximum probability density close to 2.7. The association strength scores are relatively concentrated within the peak region, indicating that the association strength between most concept entities in the initially constructed knowledge graph is concentrated in the low to medium range. Within the range of 0.25 to 0.6, the probability density gradually decreases, approaching zero after reaching 0.6, indicating that there are relatively few associations with high initial scores.
[0054] The “User-Specific Corrected Edge Weights” curve is represented by a dashed line. Its shape is basically the same as the “Original Association Strength Score”, but its overall curve height is slightly lower than the original score. The probability density near the peak, in the range of 0.20 to 0.25, is about 2.4, and the probability density in the range of 0.3 to 0.5 is slightly higher than the original score. This reflects that after the edge weights are corrected based on the user’s historical answer accuracy, the weights of some edges with lower association strength have been improved, enhancing the personalized expression of the user’s mastery.
[0055] The "average weight of the global knowledge graph" curve, represented by a dotted line, exhibits a significantly different distribution compared to the previous two, displaying a multi-peaked shape. The peak values are primarily concentrated in the 0.4 to 0.6 range, with the highest probability density at approximately 1.7. The curve falls below the previous two curves in the low association strength score range (0 to 0.15), indicating a more balanced and moderately strong distribution of edge weights in the global knowledge graph. This suggests that the edge weights of the global knowledge graph reflect the general association degree of the overall knowledge network, rather than the personalized performance of a single user.
[0056] Example 4: In its implementation, the academic decision optimization module obtains a user-specific knowledge mastery graph from the knowledge graph construction module. This graph contains a set of concept nodes, a 128-dimensional embedding vector for each concept node, and a corrected edge weight matrix. The module then calculates the degree centrality and betweenness centrality of each concept node in the knowledge mastery graph. For any given concept node, the module counts the number of neighboring concept nodes directly connected to it via non-zero edge weights. This number is then divided by the total number of concept nodes in the knowledge mastery graph minus one to obtain the degree centrality of that concept node. The degree centrality value ranges from 0 to 1. For any given concept node, the academic decision optimization module performs the betweenness centrality calculation: iterates through all concept node pairs in the knowledge mastery graph except for the given concept node, counts the total number of shortest paths between any two concept nodes, counts the number of times the concept node appears in these shortest paths, divides the number of times the concept node appears by the total number of shortest paths, sums the statistical results for all concept node pairs, and then divides the sum by the total number of concept node pairs to normalize, thus obtaining the betweenness centrality of the given concept node.
[0057] The academic decision optimization module performs a weighted summation of the degree centrality and betweenness centrality of each concept node. The weight of degree centrality is set to 0.4, and the weight of betweenness centrality is set to 0.6. The weighting is based on the fact that betweenness centrality has a stronger representational ability than degree centrality in reflecting the bridging role of concept nodes in the knowledge network. The weighted summation result constitutes the comprehensive weakness score of the concept node. Concept nodes with a comprehensive weakness score below a preset weakness threshold of 0.25 are identified as key concept nodes with weak mastery and added to the set of key concept nodes with weak mastery. The preset weakness threshold of 0.25 is based on the inflection point value of the cumulative proportion of low scores in the comprehensive weakness score distribution; the number of concept nodes with scores less than 0.25 accounts for approximately 20% of all concept nodes.
[0058] The academic decision optimization module compares the coverage of the set of key concept nodes where students have weak grasps with the pre-set course syllabus. The pre-set course syllabus is stored in structured data format, including course unit identifiers, course unit names, and a list of concept nodes covered by the course unit. For each course unit in the pre-set course syllabus, the academic decision optimization module queries the number of concept nodes in the concept node list covered by the course unit that belong to the set of key concept nodes where students have weak grasps, and calculates the proportion of this number to the total number of concept nodes in the list covered by the course unit, denoted as the weak concept coverage ratio. Based on the weak concept coverage ratio, the academic decision optimization module constructs a teaching resource gap matrix. The row index of the teaching resource gap matrix represents the course unit, and the column index represents the teaching resource type, including instructors, class hours, and tutoring materials. The teaching resource gap matrix... Line 1 The elements of a column are calculated using the following formula: in, Indicates the first The course unit in the The gap value in various types of teaching resources The value of is a real number between 0 and 1. The higher the value, the more severe the shortage of teaching resources. Indicates the first The percentage of weak concepts covered in each course unit. The value is obtained directly from the comparison between the set of key concept nodes where the students have weak grasp and the coverage of the pre-set course syllabus, and the value ranges from 0 to 1. Indicates the first The first course unit The upper limit of utilization rate of various types of teaching resources Extracted from historical resource allocation records in the academic affairs management system database. Extraction method: Retrieve the current semester's [number of records]. The total number of class hours for each course unit and the number of allocated teachers are used to calculate the ratio of the maximum number of class hours that the allocated teachers can handle to the total number of class hours, which serves as the upper limit for the utilization rate of teacher resources. The following steps are taken to obtain the first... The proportion of scheduled class hours for each course unit to the total available class hours is used as the upper limit for the utilization rate of class hour resources; obtain the first The ratio of the number of supplementary materials provided to the number of standard supplementary materials for each course unit serves as the upper limit for the utilization rate of the supplementary material resource type. The value is selected from the set {1,2,3}. The representative type of teaching resource is the teaching staff. The representative type of teaching resource is class hour allocation. The representative type of teaching resource is supplementary materials.
[0059] The academic affairs decision optimization module calls a pre-trained swarm intelligence optimization algorithm to process teacher allocation and class time rearrangement for the teaching resource gap matrix. The pre-trained swarm intelligence optimization algorithm adopts the discrete particle swarm optimization algorithm. The core architecture of the discrete particle swarm optimization algorithm consists of a particle encoding module, a fitness calculation module, an individual optimal update module, and a swarm optimal update module. The particle encoding module encodes the teacher allocation scheme and class time rearrangement scheme into an integer vector. The dimension of the integer vector is equal to the product of the number of course units and the number of teaching resource types. Each element of the integer vector represents the number of incremental resource units invested in a teaching resource type in a course unit. The value of each element is an integer between -3 and 3. The fitness calculation module receives an integer vector, decodes it into an adjustment matrix of the teaching resource gap matrix, adds the corresponding elements of the adjustment matrix to the teaching resource gap matrix to obtain the adjusted gap matrix, and calculates the sum of squares of all elements in the adjusted gap matrix as the fitness value. The smaller the fitness value, the more balanced the resource allocation.
[0060] In the training process of the Discrete Particle Swarm Optimization (DPO) algorithm, the initial particle swarm consists of 40 particles. The initial positions of these 40 particles are randomly generated within the range of integer vector values, and the initial velocity vector is randomly initialized in all dimensions between -2 and 2. The inertia weight is set to 0.8, the cognitive acceleration coefficient is set to 1.5, and the social acceleration coefficient is set to 1.5. The inertia weight of 0.8 is chosen to maintain a balance between global search and local convergence. During velocity updates, the individual's historical best position and the group's historical best position are selected by comparing fitness values. In continuous iterations, each particle updates its position based on its velocity in each round, and the updated position is rounded to ensure integer constraints. The maximum number of iterations is set to 100 rounds, and iteration is terminated early when the change in the group's optimal fitness value is less than 0.001 for 10 consecutive rounds. After the iteration terminates, the integer vector corresponding to the group's historical best position is the optimal resource allocation scheme. The optimal resource allocation scheme is decoded into the number of teacher adjustments and class hours adjustments for each course unit, generating a course schedule adjustment table. Each record in the course schedule adjustment table includes the course unit identifier, the original class hour time slot, the adjusted class hour time slot, the original teacher identifier, the adjusted teacher identifier, and the reason for the adjustment.
[0061] The academic decision optimization module generates tutoring intervention plans based on the path distances between concept nodes in the set of key concept nodes where mastery is weak. In the knowledge mastery graph, using concept nodes in the set of key concept nodes where mastery is weak as starting and target nodes, the module uses Dijkstra's algorithm to calculate the shortest path between any two concept nodes. The length of the shortest path is the sum of the reciprocals of the weights of all edges on the path, with the reciprocals of the edge weights reflecting the mastery distance between concepts. The shortest path length is directly used as the path distance between two concept nodes; a larger path distance value indicates a weaker mastery link between concept nodes. For each key concept node where mastery is weak, the module calculates the path distance from that concept node to all other key concept nodes where mastery is weak, and sums all path distance values to obtain the overall isolation degree of that concept node. All key concept nodes where mastery is weak are sorted from high to low according to their overall isolation degree, generating a recommended tutoring sequence, which specifies the tutoring order. For each concept node in the recommended tutoring sequence, the academic decision optimization module calculates the product of the degree centrality of that concept node and the comprehensive weakness score, and multiplies the product by the standard tutoring unit duration of 45 minutes to obtain the recommended tutoring duration. The tutoring intervention plan includes student identification, a tutoring concept sequence, and the recommended tutoring duration for each tutoring concept. The tutoring concept sequence is arranged in the order of the recommended tutoring sequence.
[0062] The academic affairs decision optimization module converts the course schedule adjustment table and tutoring intervention plan into a data format that conforms to the academic affairs management agreement. The course schedule adjustment table is organized in the form of a JSON structure. The top level of the JSON structure contains a "Course Adjustment List" field, which is an array. Each element of the array is a JSON object, and each JSON object contains a "Course Identifier" field, an "Adjusted Time Slot" field, and an "Instructor ID" field. The value of the "Course Identifier" field is a string type, corresponding to the unique number of the course unit; the value of the "Adjusted Time Slot" field is a string type, in the format of "Day of the Week - Period Segment", where the day of the week is 1 to 7, and the period segment is one of "1-2", "3-4", "5-6", or "7-8"; the value of the "Instructor ID" field is a string type, corresponding to the instructor's employee number in the academic affairs management system. The tutoring intervention plan is organized in the form of an XML structure. The root element of the XML is "Tutoring Intervention Plan". Under the root element, there are "Student ID" elements and "Tutoring Concept List" elements. The content of the "Student ID" element is the student's student ID as a string. Under the "Tutoring Concept List" element, there are multiple "Tutoring Concept" child elements. Each "Tutoring Concept" child element has a "Concept Name" attribute and a "Recommended Tutoring Duration" attribute. The value of the "Concept Name" attribute is the text name of the concept node, and the value of the "Recommended Tutoring Duration" attribute is the integer duration in minutes.
[0063] The academic affairs decision optimization module pushes course schedule adjustment tables and tutoring intervention plans to the academic affairs management system via application programming interfaces (APIs). The APIs include a course scheduling interface and a tutoring task assignment interface. The course scheduling interface uses a POST request method, with the interface address "apischeduleupdate", the request body payload being a JSON string containing the course schedule adjustment table, and the content type field in the request header set to "applicationjson". The tutoring task assignment interface uses a POST request method, with the interface address "apitutoringassign", the request body payload being an XML string containing the tutoring intervention plan, and the content type field in the request header set to "applicationxml". Upon receiving the request, the academic affairs management system's API gateway returns an HTTP status code of 200 indicating successful push, and the academic affairs decision optimization module records the push log.
[0064] Example 5: In its implementation, the learning effectiveness evaluation module collects users' test scores and learning time data within a preset evaluation period. The preset evaluation period is set to seven calendar days, synchronized with the weekly teaching plan of the academic management system, with data collection triggered at midnight on the first day of each week. The learning effectiveness evaluation module uses an application programming interface (API) to batch extract all users' test records from the learning platform database within the current evaluation period. Test score data includes user ID, test number, a list of knowledge point identifiers involved in the test, and the score for each question. The score for each question is either 0 or 1, where 0 indicates an incorrect answer and 1 indicates a correct answer. Learning time data is aggregated and extracted from structured data packets recorded by the learning behavior collection module. The active periods for the same user ID across all sessions are summed. Active periods are defined as the time intervals during which the user remains on the learning page and the mouse moves or the keyboard inputs.
[0065] The learning effectiveness evaluation module uses a pre-trained Bayesian knowledge tracing model to update the knowledge state of test scores and learning time data. The core architecture of the pre-trained Bayesian knowledge tracing model is a standard Hidden Markov Model, comprising a knowledge state latent variable layer and an observation output layer. In the knowledge state latent variable layer, each knowledge point corresponds to a binary latent variable, with values of 0 or 1, where 0 indicates the user has not mastered the knowledge point and 1 indicates the user has mastered it. In the observation output layer, the observation variable corresponding to each knowledge point represents the user's performance on questions related to that knowledge point, with values of 0 or 1, where 0 indicates an incorrect answer and 1 indicates a correct answer.
[0066] The Bayesian knowledge tracing model maintains four parameters for each knowledge point: the initial mastery probability is symbolically represented as... This represents the probability that a user has mastered the knowledge point before engaging in any learning activities; the learning probability symbol is denoted as . , representing the probability that a user transitions from an uncontrolled state to a controlled state; the probability of guessing is represented by the symbol . This represents the probability that a user answers a question correctly without having mastered the knowledge points; the probability of error is represented by the symbol . This represents the probability that a user answers a question incorrectly when they have already mastered the knowledge points. The value range is from 0.15 to 0.35. The value range is from 0.05 to 0.25. The value range is from 0.10 to 0.30. The value range is from 0.05 to 0.15.
[0067] The pre-training process of the Bayesian knowledge tracing model uses a historical answer sequence dataset. This dataset contains multiple sets of user responses to knowledge point sequences. Each set of responses is an ordered sequence, with each response corresponding to a knowledge point and an observation value of 0 or 1. During training, the expectation-maximization algorithm is used to estimate the four parameters of the Bayesian knowledge tracing model. In the expectation step of the expectation-maximization algorithm, the posterior probability distribution of the user's mastery status of each knowledge point at each time step is calculated based on the current parameter values. In the maximization step, the four parameter values are updated by maximizing the log-likelihood function based on the posterior probability distribution obtained in the expectation step. The expectation and maximization steps are executed iteratively. The iteration terminates when the change in the log-likelihood value is less than 0.0001, and the converged result is saved. Parameter values Parameter values Parameter values and The parameter values are used as pre-training parameters.
[0068] When using the pre-trained Bayesian knowledge tracing model to update knowledge states, the learning performance evaluation module uses test score data within a preset evaluation period as the observation sequence and executes the forward inference algorithm of Bayesian knowledge tracing for each knowledge point involved in the test. The input to the forward inference algorithm is the observation sequence, the four pre-trained parameters, and the prior mastery probability at the start of the current preset evaluation period. The prior mastery probability is used in the first preset evaluation period. The parameter values are used in subsequent preset evaluation periods based on the posterior mastery probability at the end of the previous preset evaluation period. The forward inference algorithm, at each time step, uses the posterior mastery probability from the previous time step... Parameter values Parameter values and The parameter values, combined with the current observations, are used to calculate the posterior mastery probability at the current time step using Bayes' theorem. The posterior mastery probabilities of each knowledge point across all relevant time steps are arranged by knowledge point to form a skill mastery probability matrix. The row indices of the skill mastery probability matrix correspond to the knowledge point identifiers, the column indices correspond to the time step sequence numbers, and the matrix element values are the posterior mastery probabilities, which range from 0 to 1.
[0069] The learning effectiveness evaluation module performs node similarity comparison between the skill mastery probability matrix and the knowledge mastery graph. The module obtains the user-specific knowledge mastery graph from the knowledge graph construction module. This graph contains a set of concept nodes and a 128-dimensional embedding vector for each node. For each concept node in the knowledge mastery graph, the module locates the corresponding knowledge point in the skill mastery probability matrix. If a knowledge point row in the skill mastery probability matrix directly maps to the concept node, the module extracts the posterior mastery probability of that row at the last time step as the mastery strength value of the concept node in the knowledge mastery graph. If no direct mapping exists, the module uses the 128-dimensional embedding vector of the concept node as the query vector and searches for the matching knowledge point with the highest cosine similarity among all the embedding vectors corresponding to knowledge points in the skill mastery probability matrix. The module then extracts the mastery strength value of this matching knowledge point as the mastery strength value of the concept node in the knowledge mastery graph. The embedding vectors corresponding to all knowledge points in the skill mastery probability matrix are obtained by inputting the knowledge point text into a pre-trained word embedding model. The pre-trained word embedding model and the graph attention network model share the same pre-trained 300-dimensional word vector matrix. After principal component analysis to reduce the dimension to 128, it is aligned with the embedding vector dimension of the concept nodes in the knowledge mastery graph.
[0070] After obtaining the mastery strength values of all concept nodes in the knowledge mastery map, the learning effect evaluation module calculates the forgetting curve decay coefficient. The calculation of the forgetting curve decay coefficient is based on the Ebbinghaus forgetting curve mechanism. For each concept node in the knowledge mastery map, the following value is calculated: the difference between the mastery strength value of the previous preset evaluation period and the mastery strength value of the current preset evaluation period, divided by the mastery strength value of the previous preset evaluation period, and then divided by the time span of the preset evaluation period, with the time span taken as 7 days, to obtain the forgetting rate of a single concept node. The arithmetic mean of the forgetting rates of all concept nodes is then calculated to obtain the forgetting curve decay coefficient. The formula for calculating the forgetting curve decay coefficient is expressed as follows: in, This represents the decay coefficient of the forgetting curve. The range of values for is real numbers greater than or equal to 0; This represents the total number of concept nodes in the knowledge mastery graph. Indicates the first The mastery level of each concept node in the previous preset evaluation period. The value range is from 0 to 1; Indicates the first The mastery level of each concept node in the current preset evaluation period. The value range is from 0 to 1; Indicates the time span of the preset evaluation period. The fixed value is 7.
[0071] The learning performance evaluation module feeds back the forgetting curve decay coefficient to the cognitive state analysis module to update the calculation parameters of the attention drift vector. Within the cognitive state analysis module, the diagonal elements of the process noise covariance matrix of the Kalman filter are initially set to 0.01. Upon receiving the forgetting curve decay coefficient, the diagonal elements of the process noise covariance matrix are updated to... An increase in the decay coefficient of the forgetting curve indicates that forgetting is intensified, the process noise covariance increases accordingly, and the Kalman filter increases the dependence weight of the observations on the state estimation.
[0072] In practice, the resource index database stores the learning resource files pointed to by personalized learning resource scheduling strategies. These learning resource files include video files, document files, and interactive exercise files. Video files are encapsulated in H.264 encoded MP4 format, document files are encapsulated in a portable document format, and interactive exercise files are encapsulated in a data packet format conforming to the SCORM standard. The resource index database is organized using a content-addressable file system. This file system uses the SHA-256 hash algorithm to calculate the hash value of each resource file, which serves as the hash address of the resource file. Internally, the resource index database maintains a key-value storage mapping table, where the key is the hash address represented by 64 hexadecimal characters, and the value is the physical storage path of the resource file.
[0073] When storing learning resource files into the resource index database, the content-addressable file system reads the complete binary stream of the learning resource file, calculates a 256-bit hash digest of the binary stream using the SHA-256 hash algorithm, encodes the hash digest into a 64-bit hexadecimal string, and checks whether the same hash address already exists in the key-value storage mapping table. If it does not exist, the binary stream of the learning resource file is persistently stored in the distributed storage node, and the mapping relationship between the hash address and the physical storage path is recorded in the key-value storage mapping table. If the same hash address exists, only a reference count is added without duplicate storage of the file content.
[0074] The content adaptation engine module retrieves the corresponding learning resource file by querying its hash address. When generating a personalized learning resource scheduling strategy, the module retrieves the hash address corresponding to the target learning resource node from the resource index database and sends a query request to the database. The query request format is a GET command from a key-value storage system, with the command parameter being a 64-bit hexadecimal hash address string. Upon receiving the query request, the resource index database looks up the physical storage path corresponding to the hash address in the key-value storage mapping table, reads the binary stream of the learning resource file through the physical storage path, and returns the binary stream of the learning resource file to the content adaptation engine module.
[0075] The content adaptation engine module pushes the acquired learning resource files to the user terminal via a streaming media transmission protocol. For video files, a dynamic adaptive streaming media protocol based on HTTP is used for transmission. The content adaptation engine module pre-segments the video file into a sequence of media segments with a duration of 2 seconds, generating a media segment index file. The media segment index file contains the URL address and bitrate parameters of each media segment. The user terminal player adaptively selects the media segment with an appropriate bitrate for download and playback based on the current network bandwidth. For document files, HTTP chunked transmission encoding is used. The content adaptation engine module divides the document file into a sequence of data blocks of 256 kilobytes each. Each data block is independently appended with an HTTP response header and sent sequentially to the user terminal via a persistent connection. After receiving the data, the user terminal reassembles it into a complete document file locally. For interactive exercise files, the content adaptation engine module pushes the entire SCORM data package of the interactive exercise file to the user terminal in ZIP compressed format via HTTP file download. After decompression, the user terminal loads and executes the data using the learning platform's runtime environment.
[0076] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any changes or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application.
Claims
1. A smart academic affairs management system based on a digital learning platform, characterized in that: include: The learning behavior collection module is used to acquire user operation trajectory data and biometric feedback data on the learning platform; The cognitive state analysis module is used to perform brainwave-behavioral dual-stream analysis on the operation trajectory data and biometric feedback data to generate the user's real-time cognitive load index and attention drift vector. The content adaptation engine module is used to call a pre-trained deep reinforcement learning network model to perform dynamic learning path planning on the real-time cognitive load index and attention drift vector, and generate a personalized learning resource scheduling strategy. The knowledge graph construction module is used to extract concept nodes and relationships in the learning content based on the personalized learning resource scheduling strategy, and to construct a knowledge mastery graph exclusive to the user. The teaching affairs decision optimization module is used to redistribute teaching resources based on the topological characteristics of the knowledge mastery graph, and generate course arrangement adjustment tables and tutoring intervention plans.
2. The intelligent academic affairs management system based on a digital learning platform according to claim 1, characterized in that, The cognitive state analysis module performs brainwave-behavioral dual-stream analysis on the operation trajectory data and biometric feedback data to generate the user's real-time cognitive load index and attention drift vector. The implementation methods include: The EEG signals in the biometric feedback data are subjected to multi-band power spectral density decomposition processing to extract the relative energy ratios of alpha waves, beta waves, and theta waves. The relative energy ratio is timestamped and fused with the click interval duration and page dwell duration in the operation trajectory data to generate a cognitive load time series feature matrix. The pre-trained attention residual network is invoked to perform spatial-temporal joint encoding on the cognitive load temporal feature matrix to generate the real-time cognitive load index. The mouse movement path and gaze focus heatmap in the operation trajectory data are subjected to curvature variance analysis to calculate the Euclidean distance sequence of the gaze deviation from the center point of focus. The Euclidean distance sequence is then input into a Kalman filter for state estimation to generate the attention drift vector.
3. The intelligent academic affairs management system based on a digital learning platform according to claim 2, characterized in that, The content adaptation engine module calls a pre-trained deep reinforcement learning network model to perform dynamic learning path planning on the real-time cognitive load index and attention drift vector, generating a personalized learning resource scheduling strategy. The implementation methods include: The real-time cognitive load index and attention drift vector are used as state inputs and input to the policy network layer of the deep reinforcement learning network model to calculate the action probability distribution and generate a candidate action set. Each candidate action in the candidate action set corresponds to a learning resource node adjustment instruction. The value network layer of the deep reinforcement learning network model is used to perform long-term cumulative reward estimation on each candidate action in the candidate action set to generate an action value score set. Based on the action value score set, a greedy-random hybrid strategy is used to select the optimal action in the current state. The optimal action is used to adjust the presentation order or difficulty level of learning resources. The instruction to adjust the learning resource node corresponding to the optimal action is output as the personalized learning resource scheduling strategy.
4. The intelligent academic affairs management system based on a digital learning platform according to claim 1, characterized in that, The knowledge graph construction module extracts concept nodes and relationships from the learning content based on the personalized learning resource scheduling strategy, and constructs a user-specific knowledge mastery graph in the following ways: Named entity recognition processing is performed on the learning resource text pointed to by the personalized learning resource scheduling strategy to extract the subject concept entity set; For each pair of entities in the subject concept entity set, co-occurrence frequency analysis and semantic dependency relation parsing are performed to generate an entity association strength score matrix; The pre-trained graph attention network model is invoked to perform node representation learning on the subject concept entity set using the association strength score matrix as edge weights, thereby generating an initial knowledge graph with embedded vector representations. The initial knowledge graph is subjected to confidence correction based on the user's historical answer accuracy rate, and the corrected edge weights are updated in the initial knowledge graph to obtain the user's exclusive knowledge mastery graph.
5. The intelligent academic affairs management system based on a digital learning platform according to claim 1, characterized in that, The teaching affairs decision optimization module performs teaching resource reallocation processing based on the topological characteristics of the knowledge mastery graph, and generates course schedule adjustment tables and tutoring intervention plans in the following ways: Calculate the degree centrality and betweenness centrality of each concept node in the knowledge mastery graph, and identify the set of key concept nodes with weak mastery. Based on the comparison results of the set of key concept nodes with weak mastery and the coverage of the preset course syllabus, the teaching resource gap matrix is calculated; The pre-trained swarm intelligence optimization algorithm is invoked to perform teacher allocation and class hour rearrangement on the teaching resource gap matrix, generating the course schedule adjustment table; Based on the path distances between nodes in the set of weak key concept nodes, the tutoring intervention plan is generated, which includes a recommended tutoring sequence for each weak node.
6. The intelligent academic affairs management system based on a digital learning platform according to claim 1, characterized in that, The learning behavior acquisition module acquires operation trajectory data and biometric feedback data in the following ways: A behavior capture script is embedded in the user terminal to capture the timing data of mouse click events, keyboard input events, and page scrolling events through an event listening interface. After data cleaning, the operation trajectory data is generated. The device connects to a wearable EEG device via Bluetooth and collects raw EEG signals from the user's prefrontal cortex at a preset sampling rate. The raw EEG signal was subjected to bandpass filtering and independent component analysis for noise reduction, and pure EEG feature waveforms were extracted. The operation trajectory data and the purified EEG feature waveform are frame-aligned using a timestamp synchronization unit to generate a structured data packet.
7. The intelligent academic affairs management system based on a digital learning platform according to claim 1, characterized in that, The training methods for the deep reinforcement learning network model include: Construct a historical learning resource library containing tags of different difficulty levels and knowledge domains; Define a comprehensive reward function based on learning completion time, test accuracy, and cognitive load fluctuations; The historical learning resource library is used to generate a training sample set, and each training sample contains state features, action vectors and reward values; The policy network layer and value network layer of the deep reinforcement learning network model are iteratively trained using a proximal policy optimization algorithm until the reward value converges.
8. The intelligent academic affairs management system based on a digital learning platform according to claim 1, characterized in that, The academic affairs decision optimization module is also used to convert the course schedule adjustment table and tutoring intervention plan into a data format that conforms to the academic affairs management agreement, and push it to the academic affairs management system through the application programming interface; The application programming interface includes a course scheduling interface and a tutoring task allocation interface; The course schedule adjustment table is in JSON structure format and includes the course identifier, adjusted time slot, and instructor ID fields. The tutoring intervention plan is in the form of an XML structure, which includes fields for student ID, tutoring concept sequence, and recommended tutoring duration.
9. The intelligent academic affairs management system based on a digital learning platform according to claim 1, characterized in that, Also includes: The learning effectiveness evaluation module is used to collect users' test scores and learning duration data within a preset evaluation period. The pre-trained Bayesian knowledge tracing model is invoked to update the knowledge status of the test score data and learning time data, generating a probability matrix of the user's skill mastery. The node similarity of the skill mastery probability matrix and the knowledge mastery graph is compared, and the forgetting curve decay coefficient is calculated. The forgetting curve decay coefficient is fed back to the cognitive state analysis module to update the calculation parameters of the attention drift vector.
10. The intelligent academic affairs management system based on a digital learning platform according to claim 1, characterized in that, Also includes: A resource index database is used to store the learning resource files pointed to by the personalized learning resource scheduling strategy, including video files, document files, and interactive exercise files; The resource index database is organized in the manner of a content-addressable file system, with each resource file corresponding to a hash address; The content adaptation engine module retrieves the corresponding learning resource file by querying the hash address and pushes the learning resource file to the user terminal via a streaming media transmission protocol.