Education resource recommendation method and system based on artificial intelligence
By collecting and analyzing audio, interactive, and task data from online education systems, and using artificial intelligence models to identify the learning psychology and cognitive load of rural students, course recommendations are dynamically adjusted. This solves the problem of course recommendations being out of sync with student needs in existing systems, thereby improving the learning experience and the quality of education.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-01-30
- Publication Date
- 2026-04-07
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Existing online education systems are unable to accurately identify the learning psychology and cognitive load of rural students, resulting in a disconnect between course recommendations and students' actual needs, leading to a vicious cycle of "the more they learn, the harder it becomes" and "the more they learn, the more tired they become."
By collecting audio, interaction, and task data from user terminals, and using artificial intelligence models for multimodal fusion analysis, we can extract emotional features, learning pace, and content features, calculate cognitive load, and adjust course recommendation strategies based on psychological state, including adjusting difficulty levels and content order.
It enables accurate identification and dynamic adjustment of the learning status of rural students, improves the interactivity and teaching quality of online education, provides personalized learning support, and promotes educational equity and digital development.
Smart Images

Figure CN121808153A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of intelligent education, in particular to an education resource recommendation method and system based on artificial intelligence. BACKGROUND
[0002] In recent years, with the promotion of the national "Internet + education" strategy, the popularization rate of online education platforms in urban and rural areas has gradually increased. However, there are still obvious shortcomings in the education resources in rural and remote areas. On the one hand, high-quality course content and professional teacher guidance resources are concentrated in urban areas, resulting in a disconnection between rural students and the mainstream curriculum system in terms of learning content, teaching methods and learning pace; on the other hand, the self-learning ability of students in rural areas is generally weak, and there is a lack of systematic learning guidance and personalized recommendation mechanism, and problems such as lack of learning confidence, large psychological fluctuations and weak motivation for continuous learning are common.
[0003] The existing online education system can recommend courses based on answer data or learning time, but it mainly stays at the level of performance orientation or static behavior statistics. For example, some platforms judge the learning level according to the correct rate of students' answers, completion rate and other indicators and recommend the next learning module, but they do not consider the differences in psychological state, emotional changes and cognitive carrying capacity of learners. For students in rural areas who have learning difficulties, this recommendation method that relies only on surface data often has the following shortcomings: Students with learning difficulties often have understanding lag, attention dispersion and other phenomena in the learning process. Traditional platforms only judge concentration by learning time or page clicks, which cannot reflect the actual cognitive burden, and even less can identify potential "active abandonment" and "cognitive overload" and other learning state differences. When students' learning motivation decreases or cognitive load is too high, the system still pushes high-difficulty courses in the established order, leading to a vicious cycle of "the more you learn, the more difficult it is" and "the more you learn, the more tired you are". SUMMARY
[0004] In view of the problem that the existing online education system lacks a more gentle, flexible and intelligent learning guide for education course recommendation, the present application provides an education resource recommendation method and system based on artificial intelligence.
[0005] In order to achieve the above purpose, the technical scheme of the present application is as follows: In a first aspect, the present application discloses an education resource recommendation method based on artificial intelligence, comprising the following steps: obtaining audio data, interaction data and task data collected by a user terminal; wherein the audio data includes user voice questions or reading tasks; the interaction data includes learning time, click frequency, page dwell time, fast-forward times and pause times; the task data includes course content tags and question structure information; The audio data, the interaction data and the task data are respectively subjected to feature extraction to obtain a sentiment feature vector, a learning rhythm vector and a content feature vector; The sentiment feature vector and the learning rhythm vector are input into a pre-constructed sentiment recognition model to obtain a psychological state vector representing motivation degree, tension degree and burnout degree; The cognitive load is calculated based on the learning rhythm vector and the content feature vector, and then the cognitive load is corrected by the psychological state vector to obtain a final cognitive load; The final cognitive load is matched with a state interval according to a pre-set threshold interval, and the state interval includes a comfortable interval, an overload interval and a low maneuvering interval; The difficulty coefficient of a recommended course, the sequence of course content and a learning aid are adjusted according to the state interval, and structured data is generated and output as an intelligent learning recommendation result.
[0006] In a second aspect, the application discloses an education resource recommendation system based on artificial intelligence, which adopts the education resource recommendation method based on artificial intelligence as described above; the system comprises a data acquisition module, a feature extraction module, a cognitive load calculation module, a state interval identification module and a result output module.
[0007] The data acquisition module is used for acquiring audio data, interaction data and task data collected by a user terminal; wherein the audio data includes user voice questions or reading tasks; the interaction data includes learning duration, click frequency, page dwell time, fast-forward times and pause times; the task data includes course content labels and question structure information; The feature extraction module is used for respectively extracting features of the audio data, the interaction data and the task data to obtain a sentiment feature vector, a learning rhythm vector and a content feature vector; and is also used for inputting the sentiment feature vector and the learning rhythm vector into a pre-constructed sentiment recognition model to obtain a psychological state vector representing motivation degree, tension degree and burnout degree; The cognitive load calculation module is used for calculating a cognitive load based on the learning rhythm vector and the content feature vector, and then correcting the cognitive load by the psychological state vector to obtain a final cognitive load; The state interval identification module is used for matching the final cognitive load with a state interval according to a pre-set threshold interval, and the state interval includes a comfortable interval, an overload interval and a low maneuvering interval; The result output module is used for adjusting the difficulty coefficient of a recommended course, rearranging the sequence of course content and a learning aid according to the state interval, generating structured data and outputting the structured data as an intelligent learning recommendation result.
[0008] Compared with the prior art, the application has the following beneficial effects: The application obtains cognitive load through psychological-cognitive-behavioral three-modal fusion analysis, and realizes the state of students according to the cognitive load, and realizes the adaptive adjustment of teaching content, rhythm and emotional incentive according to the state of students, effectively improves the interactivity and teaching quality of online learning in rural and remote areas, so as to provide intelligent auxiliary learning support of "tailor-made, emotion adjustment and time optimization" for students in need in the environment of relatively weak education resources, and promote the in-depth development of education equity and education digitalization. BRIEF DESCRIPTION OF DRAWINGS
[0009] The disclosure of the present application will be described with reference to the accompanying drawings. It should be understood that the drawings are only for illustrative purposes, and are not intended to limit the scope of protection of the present application. In the drawings, the same reference numerals are used to refer to the same parts. Among them: Figure 1 The flowchart of the education resource recommendation method based on artificial intelligence introduced in embodiment 1 of the present application; Figure 2 The overall logic flowchart of the education resource recommendation method based on artificial intelligence; Figure 1 Figure 3 The feedback adjustment flowchart when the intelligent auxiliary learning recommendation result is executed based on Figure 1 The feedback adjustment flowchart after the intelligent auxiliary learning recommendation result is executed based on Figure 4 Figure 1 Figure 5 The scene description diagram of the education resource recommendation method based on artificial intelligence; Figure 1 Figure 6 The block diagram of the education resource recommendation system based on artificial intelligence introduced in embodiment 2. DETAILED DESCRIPTION
[0010] It is easy to understand that according to the technical scheme of the present application, those skilled in the art can propose a plurality of structure modes and implementation modes which can be replaced with each other without changing the essential spirit of the present application. Therefore, the following specific embodiments and drawings are only exemplary descriptions of the technical scheme of the present application, and should not be regarded as the whole or as the limitation or restriction of the technical scheme of the present application.
[0011] In the prior art, online education platforms usually only rely on student behavior data (such as learning time, page clicks, correct answer rate, etc.) to infer learning state, and make course recommendations based on fixed rules. This kind of method fails to identify the psychological state of students, only reflects learning ability with behavior statistics, and is prone to misjudgment of "extreme surface, actual overload" or "low activity but psychological tension", resulting in a disconnection between the recommended results and the psychological and cognitive state of students, poor learning experience and delayed intervention.
[0012] To solve the above problems, research and development found that by fusing multi-modal data analysis and psychological-cognitive joint means can effectively improve the recommendation accuracy. Specifically, by constructing a psychological and cognitive model reflecting the comprehensive state of learners through audio data, interaction data and task data, the coupling relationship between the psychological state and cognitive load of students can be accurately described. On this basis, by setting multiple threshold intervals and dividing the state interval based on the psychological state corrected cognitive load, this mechanism enables learning recommendation to shift from "result-oriented" to "process regulation", realizing dynamic and closed-loop learning guidance.
[0013] After introducing the basic idea of the present application, the embodiments of the present application will be specifically introduced with reference to the drawings.
[0014] Embodiment 1: As Figures 1-2 shown, the education resource recommendation method based on artificial intelligence is introduced, including the following steps: S100. Acquire audio data, interaction data and task data collected by a user terminal; wherein the audio data includes user voice questions or reading tasks; the interaction data includes learning duration, click frequency, page dwell time, fast-forward times and pause times; the task data includes course content labels and question structure information.
[0015] When the user enters the learning interface or initiates voice interaction in the learning terminal (such as a learning APP, a learning applet or a web terminal), the system automatically starts the data collection module; The data collection module monitors user learning behavior, voice input and learning task calling information in real time through local sensors, microphones and operation log interfaces, and starts the corresponding data channel according to the task type.
[0016] The system calls the audio input device composed of local sensors, microphones, etc. to collect the voice signal of students in the learning process, and the collection content includes: voice question and answer data when the student asks questions; voice reading data when the student performs reading tasks; oral reaction, tone and pause information of the student in the task execution.
[0017] The collected original audio signal is saved in the form of waveform data or time domain signal, which contains features such as tone, speed, energy intensity and phoneme distribution, providing data basis for subsequent emotion feature extraction. In order to ensure data quality, the system monitors the audio sampling rate (such as 16kHz) and signal-to-noise ratio, and triggers re-recording or signal enhancement processing when it is lower than the set threshold.
[0018] The system records the operation behavior of students in the learning interface through the operation log interface, including: learning duration (the duration of a single session and the cumulative time); click frequency (the number of interactions such as page buttons, answer options, and play controls); page dwell time (the length of time spent on each content unit or video); fast-forward frequency and amplitude (reflecting content skipping or repeated learning behavior); pause frequency and interval (reflecting learning pace and cognitive stagnation characteristics).
[0019] All interaction data is stored in timestamp form, forming a sequence of behavior events, which is used for subsequent calculation of the learning pace vector. The system filters or labels abnormal data (such as leaving the page, network interruption, etc.) in the background to ensure the continuity of the pace characteristics.
[0020] The system synchronously reads teaching content and task structure data during student learning, including: course content label information (discipline, knowledge point, chapter level, difficulty level, etc.); question structure information (question type, answer steps, problem solving logic, time estimation, etc.); task context information (task start and end time, corresponding learning module number, etc.). Task data is automatically called from the background course database and corresponds to interaction behavior and audio data one by one to ensure synchronization and alignment of the three types of data in the time dimension.
[0021] The system generates a unique identification ID for each learning session and binds the audio data, interaction data, and task data according to the timestamp alignment. This binding structure ensures that the subsequent feature extraction process can achieve cross-modal correlation based on the same time window.
[0022] To facilitate feature extraction in step S200, the system performs appropriate preprocessing operations on the collected raw data, such as removing silent segments or background noise; eliminating non-learning interaction events (such as page switching, returning to the home page, and irrelevant operations); filling in missing time slices and standardizing sampling; and caching and segmenting data according to learning units or question dimensions. The cleaned data is temporarily stored in the local buffer or cloud database, providing input data for subsequent step S200.
[0023] S200. Extract features from audio data, interaction data, and task data to obtain emotion feature vectors, learning pace vectors, and content feature vectors. Specifically: Input the audio data obtained in S100 into a pre-trained audio encoding model such as CLAP or Wav2Vec2 model. Perform time-frequency feature encoding on the original speech signal to extract a 512-dimensional static semantic embedding vector. At the same time, input the original audio signal into the OpenSmile toolkit and extract 88-dimensional acoustic features (energy, fundamental frequency, speech rate, intonation, acoustic texture, etc.) based on the eGeMAPS scheme to form an acoustic feature matrix , where each row represents the acoustic features of a time slice. Perform frame-level statistical operations (mean, variance, skewness, kurtosis, etc.) to convert the time series into a fixed-length statistical feature vector. Then, the static semantic embedding vector is... With statistical eigenvectors By concatenating and merging the data, we obtain the sentiment feature vector: .
[0024] The training process of the audio coding model is as follows: Samples are collected and processed beforehand to obtain training samples. A contrastive loss is used to train the CLAP, which employs a dual-tower structure, including an audio encoder and a text encoder. Training uses the AdamW optimizer, combined with a learning rate decay strategy to optimize parameters. Specifically, the training samples are input into the audio encoder and text encoder respectively. The audio encoder and text encoder extract feature vectors from the training samples, respectively. The contrastive loss between audio and text features is calculated. Backpropagation is used to update the parameters of the audio encoder and text encoder. The optimization objective is to minimize the contrastive loss. The batch size during training is 16, the initial learning rate is 1e-5, and the learning rate is dynamically adjusted using a Cosine AnnealingWarm Restarts strategy. The training epochs are 50, and the early stopping patience value is set to 5. After fine-tuning and evaluation, the trained CLAP is obtained.
[0025] Extracting key metrics from interaction data: Video view ratio Pause frequency Page Switching Rate ,in: ; ; ; The actual duration of video viewing by students (in seconds); The total video duration (in seconds) is when A higher score indicates that students have watched the entire video and are highly engaged in learning. This represents the number of times a student triggers a pause operation during video playback. A higher score indicates that the learner has difficulty understanding the current content or has reduced concentration. This represents the actual number of page switches students made while watching the video. The higher the score, the more likely the learner is to be distracted, skipping sections, or multitasking, indicating poor learning stability.
[0026] The above metrics are subjected to min-max normalization to eliminate numerical differences between different tasks or devices. After thresholding abnormal behaviors (such as frequent pausing or extremely short viewing), the features are concatenated in a fixed order to form a learning rhythm vector. .
[0027] Retrieves structured tags of the current learning content from the task data, including the number of knowledge points. Content duration And the number of visual or multimodal elements Combine structured tags into content feature vectors .
[0028] Sentiment feature vector Learning rhythm vectors Content feature vector All three correspond to the same learning segment in a synchronous manner over time, serving as input for subsequent steps.
[0029] S300. Input the emotional feature vector and the learning rhythm vector into the pre-constructed emotion recognition model to obtain psychological state vectors representing the degree of motivation, tension, and burnout.
[0030] Emotion recognition models include: Convolutional layers (CNNs): extract local temporal features and high-dimensional semantic information related to sentiment; Bidirectional Long Short-Term Memory Network (BiLSTM): Captures the dynamic dependence of learners' mental states over time; Attention Layer: Adaptively assigns importance weights to different features, highlighting feature segments that are more sensitive to psychological state judgments; Fully connected layer (Dense Layer): Maps high-dimensional features to the mental state output space.
[0031] The specific steps are as follows: To ensure consistency of units, the sentiment feature vectors were analyzed separately. Learning rhythm vectors Perform linear mapping and standardization, then fuse the processed features proportionally in a weighted manner to form a joint input sequence: ; in, , These represent the processed sentiment feature vector and the learning rhythm vector, respectively. This indicates vector concatenation; Indicates the modal weighting coefficient; This represents the fused feature sequence at time step t.
[0032] The fused feature sequence is input into a convolutional layer (CNN) to extract emotional change patterns and rhythmic features from local time segments. After convolution, ReLU activation and Dropout are applied to prevent overfitting, resulting in intermediate feature maps. .
[0033] Mapping intermediate features The input is a Bidirectional Long Short-Term Memory (BiLSTM) network, which uses sequential dependencies to capture the dynamic trend of students' psychological states over time and outputs time series features. .
[0034] The attention layer performs attention-weighted calculations on the time-series features to highlight time segments that significantly influence psychological states, calculating the attention weight for each time step: ; This represents the attention energy score at time step t; This represents the trainable weight vector of the attention layer; This represents the weight matrix in the attention layer, used to... Perform a linear transformation to project it onto a unified attention space; This represents the hidden state vector output by the previous layer (BiLSTM layer) at time step t; This represents the bias term vector of the attention layer; This represents the hyperbolic tangent nonlinear activation function; This represents the attention weight coefficient at time step t; This represents the total number of steps in the time series, i.e., the length of the input feature sequence; This represents the attention energy score at time step k.
[0035] The final global contextual sentiment representation is obtained as follows: .
[0036] The aggregated sentiment context vector C is input into a fully connected layer (MLP), which outputs intensity scores for the three psychological states: ; Indicates the degree of motivation; Indicates the level of tension; Indicates the degree of burnout; This is the Sigmoid function, used to normalize the result to [0,1].
[0037] The model outputs the student's mental state vector at the current moment: .
[0038] S400. Cognitive load is calculated based on the learning rhythm vector and content feature vector, and then corrected using the mental state vector to obtain the final cognitive load. The specific steps are as follows: Content complexity is obtained by weighting and standardizing the intrinsic features in the content feature vector; The intrinsic feature of content feature vectors is the number of knowledge points. Content duration And the number of visual or multimodal elements Normalization and weighting are performed on each feature, and weight coefficients are assigned according to the degree of influence of different features on learning complexity. Therefore, the content complexity... The calculation formula is: ; , , For the weight parameters, satisfying .
[0039] Extracting pause frequencies from the learning rhythm vector and page switching rate and content complexity Cognitive load is obtained through weighted calculation; the calculation formula is: ; in, , , This is an empirical weight used to balance the impact of behavioral and content factors on cognitive load.
[0040] The motivation level is extracted from the mental state vector, and the cognitive load is adjusted to obtain the final cognitive load. The adjustment formula is as follows: ; in, Indicates the final cognitive load; Indicates cognitive load; Indicates the degree of motivation; This represents the motivation adjustment coefficient.
[0041] This step, by weighting and integrating behavioral indicators with content complexity, effectively distinguishes the essential difference between "voluntary abandonment" and "cognitive overload." It effectively identifies low-interaction behaviors caused by a lack of confidence and differentiates them from genuine cognitive overload, thus avoiding pushing overly difficult courses that exacerbate the learning burden. Simultaneously, it can promptly detect high-load states and trigger appropriate supplementary learning prompts to help learners alleviate stress and boost their learning confidence.
[0042] S500. Match the final cognitive load state range according to the pre-set threshold range. The state range includes the comfort zone, overload zone, and low mobility zone. The specific steps are as follows: A threshold range is preset, which includes a first threshold whose values gradually increase. Second threshold and the third threshold . < < It is important to emphasize that the initial value of the threshold interval is based on the mean of the population sample. with standard deviation Settings, namely: .
[0043] For example, suppose the final cognitive load mean and standard deviation of the population sample data are as follows: (The final load value range has been standardized to [0,1]). Therefore, After the system is in operation, the platform will retrain and dynamically update the threshold every semester or a set period (e.g., every 2 months). This process ensures that the threshold can be dynamically adapted to changes in the group's learning behavior and psychological characteristics.
[0044] Based on the final cognitive load of the current cycle The position within the threshold range combined with the degree of motivation Tension The matching logic is as follows: (1) If ≤ and A value ≥0.6 indicates the patient is in the comfort zone. (2) If ≤ and <0.4 indicates the area is in a low-mobility zone; (3) If ≥ and ≥0.7 indicates the area is in the overload zone; If none of them meet the requirements, the state interval of the historical nearest cycle will be used for matching; the historical nearest cycle is the historical cycle with the shortest time interval between the current cycle and that is in the state interval.
[0045] By matching state intervals with multi-level threshold intervals, the problem of inaccurate state recognition caused by relying solely on cognitive load thresholds is solved. In particular, it addresses the common problem of fluctuating learning confidence among rural students, effectively distinguishing between low mobility zones caused by insufficient motivation and overload zones caused by cognitive overload, thereby improving the accuracy and adaptability of educational resource recommendations.
[0046] Before using the state interval for matching the historical nearest cycle, determine whether the current cycle is in the transition zone. If it is, mark it as transition and perform nearest neighbor calculation. If the nearest neighbor calculation results of K consecutive cycles are all in the same state interval, then determine that it is in that state interval. The transition zone refers to the intermediate stage where a learner's cognitive state gradually changes from one stable interval to another. It is determined using a multi-dimensional parameter joint judgment method, aiming to accurately capture the critical characteristics of the learner's state evolution and avoid misjudging normal transitions as stable states. In practical applications, nearest neighbor calculation uses data from neighboring periods to predict short-term trends, thus providing a more reliable intermediate basis for state determination. Nearest neighbor calculation involves the system searching for state interval labels within the nearest neighbor period set. If multiple different state labels exist in the nearest neighbor period, the system uses a majority voting strategy.
[0047] Among them, the conditions for being in the transition zone are: < < And 0.4≤ <0.6 and <0.7, or cross , or The value is less than the boundary threshold The value range is 0.03–0.05.
[0048] Where K is an integer less than 5, used to maintain the system's responsiveness and avoid excessively long transition periods that could delay learning intervention. A smaller K (e.g., 2–3) is suitable for real-time learning scenarios with rapid responses; a larger K (e.g., 4) is suitable for learning scenarios with stable emotions and a slow pace.
[0049] By introducing transition zone judgment and continuous consistency verification mechanisms, we can not only effectively deal with typical transitional scenarios where cognitive load is in the middle range and learning motivation is moderate and tension is low, but also handle edge cases where load changes suddenly but has not reached a stable state, solve the technical problems caused by the complexity of state evolution, and improve the stability and adaptability of educational resource recommendations.
[0050] After marking a course as transitional, a transition strategy will be implemented regarding the difficulty level of recommended courses, the order of course content rearrangement, and supplementary learning tips. The transition strategy is as follows: Obtain historical adjustment strategies for the difficulty coefficient of recommended courses, the order of course content reordering, and supplementary learning tips in the near future; Based on historical adjustment strategies, a minor adjustment is made according to a preset adjustment mechanism, and the adjusted historical adjustment strategy is used as a transition strategy for the current cycle.
[0051] Specifically, the adjustment mechanism involves making minor adjustments to the difficulty level of recommended courses, following the principle of "light adjustments without disruption": ; in, This indicates the recommended difficulty level of the course after adjustments during the transition period. This indicates the standard difficulty value of the content recommended by the system within the comfort zone. This represents the adjustment rate coefficient, which ranges from [0.05, 0.1].
[0052] Furthermore, based on the historical adjustment strategy's course content sequencing, content with high information density or strong emotional load is displayed one level lower (e.g., complex explanatory videos or high-cognitive tasks are postponed to the next unit). Light content of about 1-2 minutes is inserted in the adjusted position, such as "knowledge consolidation Q&A" or "positive emotion video clips," to alleviate cognitive fatigue and maintain learning motivation, while triggering motivational cues.
[0053] S600. Adjust the difficulty coefficient of recommended courses, rearrange the order of course content, and add supplementary learning tips based on the state interval, generating structured data and outputting it as the intelligent supplementary learning recommendation result. The specific steps are as follows: The specific steps for adjusting the difficulty level of recommended courses, rearranging the order of course content, and adding supplementary learning tips based on the student's condition range are as follows: Based on the identified state intervals and results, the following decisions are made: (1) If the recognition result is the comfort zone, the multimodal synchronization of the recommended courses will be gradually increased; that is, the synchronous display of text, images and audio will be gradually increased.
[0054] (2) If the identification result is a low mobility zone, the difficulty coefficient of the recommended course is adjusted based on the preset first amplitude, and the course content is rearranged and incentive prompts are triggered with positive emotion as the guide.
[0055] Based on the first amplitude modulation Adjustments will be made downwards: ; , These are the difficulty levels of the recommended courses before and after the adjustment. The value range is [0.1, 0.2].
[0056] To rearrange the order of course content, the first step is to retrieve the set of teaching resources corresponding to the current learning task from the content database. Each teaching resource contains multimodal material information such as videos, images, text, or audio. For each teaching material, the system extracts and calculates its multimodal content features, including information density, duration, content complexity, and emotional elements. These features are then normalized to ensure that all feature values fall within a uniform range. Next, a comprehensive weight score is calculated for each material based on preset weights, and the materials are arranged from highest to lowest based on their comprehensive weight scores. Lower emotional feature values indicate a more positive sentiment.
[0057] In low mobility zones, materials with low emotional value are given higher priority when reordering course content.
[0058] Trigger "incentive-based support prompts" using positive emotional expressions and short task guidance, such as "Try one more question." Increase the frequency of prompts appropriately (reducing cooldown time by 20%).
[0059] (3) If the identification result is an overloaded area, the difficulty coefficient of the recommended courses is adjusted based on the preset second amplitude, and the course content order is rearranged and a rest prompt is triggered with the guidance of reducing content complexity and duration.
[0060] According to the second amplitude Reduce the difficulty of recommended courses: ; The value range is [0.15, 0.25].
[0061] Because it is in the overload zone, when rearranging the order of course content, give priority to materials with low information density and low content complexity.
[0062] Trigger "rest prompts" including phrases such as "take a proper rest" and "it is recommended to pause your review," and embed psychological adjustment guidance (such as deep breathing exercises and links to soft music) in the prompts.
[0063] In detail, this scheme ensures targeted adjustments by identifying state intervals as the basis for decision-making. When learners are in their comfort zone, the system gradually increases the multimodal synchronization of recommended courses. This process can be triggered by analyzing learners' moderate cognitive load, thereby enhancing learning interest and promoting deep learning. For learners in the low-mobility zone, the system adjusts the difficulty coefficient of recommended courses based on a preset first amplitude, while prioritizing positive content to rebuild learning confidence and providing immediate reinforcement of positive behaviors through incentive prompts. These measures work together to enhance learning motivation. When learners are in the overload zone, the system significantly reduces course difficulty based on a preset second amplitude, directly alleviating cognitive burden by reducing content complexity and information density, and forcibly interrupting learning to restore state and prevent fatigue accumulation. This differentiated adjustment strategy achieves refined and adaptive educational resource recommendations, effectively alleviating the problems of cognitive overload and insufficient motivation during the learning process.
[0064] The system can provide personalized learning support based on the learner's specific situation, avoiding the phenomenon of "the more you learn, the harder it becomes" or "the more you learn, the more tired you become" caused by the one-size-fits-all recommendation in traditional methods. It can better meet the personalized learning needs of rural students with learning difficulties.
[0065] like Figure 3 As shown, in order to achieve feedback correction, when executing the intelligent learning recommendation results, the system receives feedback audio data, interaction data, and task data in real time, calculates the final cognitive load after feedback, and calculates the difference between the final cognitive load before feedback to obtain the load change. ; Determine the load change If the value is less than zero, continue to execute the intelligent tutoring recommendation results; Otherwise, the first adjustment to the difficulty level Second amplitude Adjustments were made and the intelligent tutoring recommendation results were regenerated.
[0066] Adjustment methods can be based on motivation level. Tension Dynamic adjustment, firstly, the load change Normalization process yields ; This indicates the final cognitive load before feedback. This represents the smallest positive number used to prevent the denominator from being zero, and its value is 10. -3 .
[0067] Therefore, the adjustment formula for the first amplitude modulation is: ; in, This indicates the first amplitude adjustment after the change; , In this embodiment, the gain coefficient is... =0.1, =0.05.
[0068] Similarly, the adjustment formula for the second amplitude modulation is: ; in, This indicates the adjusted second amplitude; , In this embodiment, the gain coefficient is... =0.12, =0.06.
[0069] The system calculates the difference between the final cognitive load after feedback and the initial value to determine the load change, focusing on the net load change resulting from the recommended implementation. Based on the logic of whether the load change is less than zero, the system automates decision-making: when the change is less than zero, it indicates that the recommended strategy has effectively reduced the cognitive load, and therefore implementation continues to maintain learning continuity; otherwise, an adjustment mechanism is triggered, refining the first and second amplitudes of the difficulty coefficient to generate more suitable recommended content. This process not only solves the problem of lacking real-time monitoring of the recommended implementation effect but also ensures a high degree of match between the recommended content and the student's current cognitive state through dynamic adjustment, thereby effectively preventing the risk of continuous deterioration of cognitive load. Rapidly adjusting the recommendation strategy based on the student's immediate behavioral performance and psychological fluctuations avoids the recommendation bias caused by static analysis in traditional platforms, thus enhancing learning confidence and sustained motivation.
[0070] like Figure 4 As shown, after the intelligent learning recommendation results of the current stage are completed, the performance value of the current stage and the performance value of the previous stage are obtained, and the difference between the two is used to calculate the performance difference. If the difference in effectiveness is greater than zero, continue to execute the intelligent tutoring recommendation results; Otherwise, the first and second adjustments to the difficulty level will be made and used for the next stage of intelligent tutoring recommendations. The adjustment methods are the same as those described above.
[0071] The performance value is calculated by weighting the learning outcome accuracy, load adjustment amount, and effective data rate in the interactive data for each stage.
[0072] Learning outcome accuracy The calculation formula is: ; This indicates the number of correct answers a student received on a test or task. This indicates the total number of questions.
[0073] Load regulation The calculation formula is: ; This indicates the maximum load threshold.
[0074] The formula for calculating the effective data rate is: ; , and Let be the weighting coefficients, and let sum to 1.
[0075] After normalizing all the above parameters, the performance value is calculated using the following formula: ; The weighting coefficients are preferably 0.5, 0.3, or 0.2.
[0076] Simultaneously, by filtering out invalid interactions through effective data rate, learning effectiveness can be quantified more objectively. This not only solves the problem of strategy rigidity caused by the lack of effectiveness evaluation in recommendation systems, but also improves the adaptability of recommendation strategies to students in rural areas with weak learning motivation and fluctuating cognitive states.
[0077] Therefore, the above steps can adjust the recommended parameters in a timely manner according to the dynamic changes in learning outcomes, avoiding the occurrence of the "recommendation-implementation-ineffectiveness" cycle, thereby effectively alleviating the learning dilemmas faced by rural students due to insufficient learning motivation or cognitive overload.
[0078] The main solution of this application has been introduced above. In practical applications, the extension of this application can also acquire audio data, interaction data and task data collected by user terminals, as well as acquire user visual data, extract features from the visual data to obtain facial expression vectors, and input them together with emotion feature vectors and learning rhythm vectors into a pre-built emotion recognition model to assist in the recognition of psychological states.
[0079] Visual data is acquired in real-time via the user's front-facing camera during the learning process, using RGB cameras, infrared cameras, or depth cameras. After preprocessing operations such as face detection, illumination normalization, and face alignment, the user's facial images or videos are input into a facial expression encoding model (e.g., FER-ResNet). The model extracts emotion-related local and global facial texture features, outputting fixed-dimensional expression vectors. The model outputs confidence scores for seven or eight basic emotions (e.g., happiness, anger, fear, boredom, calmness) through a softmax layer. The emotion categories and confidence scores are encoded into a set of emotion distribution vectors, which are then dimensionality-reduced and integrated into an expression emotion vector. This vector is then concatenated with the emotion feature vector and the learning rhythm vector to form the input to the emotion recognition model.
[0080] This application addresses the problem of existing systems relying solely on surface-level behavioral data and failing to accurately identify learners' psychological states and cognitive load by integrating multi-dimensional learning data and constructing a dynamic perception and feedback mechanism. Specifically, it achieves comprehensive perception of the learning process by collecting multi-source data from user terminals; quantifies psychological states and cognitive load through feature extraction and model recognition; and dynamically adjusts recommendation strategies through state interval matching to ensure accurate matching of course recommendations with learning states. This solution avoids the problem of recommendation strategies being disconnected from learning states caused by traditional recommendation systems ignoring psychological states and cognitive load, thereby improving the personalized recommendation capabilities of online education systems.
[0081] To facilitate understanding of the above embodiments, a specific application scenario of the above embodiments will be used as an example for illustration below: like Figure 5 As shown, taking a sixth-grade student in a rural primary school in a poverty-stricken county as an example, this student, who struggles with math, logged into a mini-program on a tablet provided by the school after class for a 15-minute micro-lesson on fraction word problems followed by 15 minutes of online practice. This analysis focuses only on the first 15 minutes of the video learning segment. During these 15 minutes, the mini-program automatically collected the following data: Audio data: Reading aloud key formulas explained by the teacher + verbal answers to two simple questions. Actual effective audio duration: approximately 90 seconds (the rest was silent or ambient sound, which has been removed through preprocessing). The processed data used for sentiment analysis includes: overall voice volume is moderately low, speech rate is slightly slow, but noticeably faster in the second answer, and intonation is more varied than in his usual classroom responses (with an upward inflection when asking questions).
[0082] Interactive data: Total video duration =900 seconds (15 minutes), actual viewing time =780 seconds (system automatically deducts downtime), number of pauses =5 (5-15 seconds each time, mostly when the teacher gives examples), fast forward count is 2 times, page switching count is 5. =3 (Switch to the "Drafts" module, return to the video). The system calculates relevant key indicators, such as the video view ratio. Pause frequency Page Switching Rate For example: ; times / second; / Second; The system then performs min-max normalization on these values, mapping them to the 0-1 range, forming part of the learning rhythm features. Assuming a large sample of statistics for this grade level: Common ranges for video watch rate: [0.4, 1.0]; common ranges for pause frequency: [0, 0.02]; common ranges for page switching rate: [0, 0.01].
[0083] The normalized video watch ratio is 0.78, the normalized pause frequency is 0.32, and the normalized page switch rate is 0.38.
[0084] Task Data: Structured Tags for this "Fraction Word Problems Micro-Lesson": Number of Knowledge Points =7 (2 sections on the meaning of fractions, 3 types of word problems, and 2 sections on breaking down problem-solving steps). Content duration At 15 minutes, this is considered "medium to long" among all micro-courses on the platform. (Number of multimodal elements) =4 (blackboard writing video + animated blackboard writing + illustrated handouts + a small interactive exercise).
[0085] When normalizing content features, the platform has internal statistics: the number of knowledge points in a single lesson is usually between 1 and 10; the duration is concentrated between 5 and 20 minutes; and the number of multimodal elements is generally between 1 and 8. Therefore, simple linear normalization yields the normalized result. It is 0.7, normalized. 0.5, normalized The value is 0.4. The system feeds the collected speech into the pre-trained CLAP audio coding model and combines it with the eGeMAPS acoustic features extracted by OpenSmile to form an emotion feature vector. After inference, the model gives a simplified psychological state estimate (value 0–1): Motivation level =0.80 (Xiao Jun was willing to cooperate and read along, and his tone was noticeably more energetic when answering questions correctly). Tension level =0.35 (Speaking speed is slightly slow, but there is no obvious stuttering or prolonged silence). Fatigue level =0.20 (Mental state was acceptable at the beginning of evening self-study).
[0086] Content complexity is obtained by weighting the number of knowledge points, duration, and multimodal elements, with the weights set as follows: weight of the number of knowledge points. =0.4, content duration weight =0.3, weight of the number of multimodal elements =0.3. Computational complexity: The calculation results indicate that this lesson falls into the "moderately complex" category within the overall resource pool, presenting a challenge for struggling sixth-grade students but not an extremely difficult one.
[0087] Cognitive load is calculated by weighting behavioral indicators such as content complexity, pause frequency, and page switching rate, with pause frequency weighted. =0.3, page switch rate weight =0.2, content complexity weight =0.5. Calculate cognitive load: The unmodified cognitive load was approximately 0.573 (standardized to 0–1), slightly above the median of 0.5.
[0088] The formula for adjusting cognitive load through motivation can be understood as follows: the higher the motivation, the less difficult the perceived objective load will seem; conversely, the lower the motivation, the more the perceived load will be amplified. Let's define a motivation moderating coefficient. =0.4, motivation level =0.80, therefore: .
[0089] For Xiaojun, who has high motivation, the subjective load of the same slightly complex lesson was "reduced" from 0.573 to about 0.504, which is closer to the comfort range.
[0090] The system will use the mean and standard deviation of the group data to set three thresholds, assuming the final cognitive load statistics for all students in the same grade this semester are as follows: Therefore, the threshold is: .because ≤ and ≥0.6, therefore, it is determined to be in the comfort zone. The adjustment strategy is: "Gradually increase the multimodal synchronization of recommended courses," such as increasing the synchronized display of text, images, and audio. The next recommended course will rearrange the order of materials, first showing situational animations + life-like examples (pictures + narration), anchoring to real-life situations such as "cutting a watermelon," then showing a whiteboard derivation video (with step-by-step annotations) to improve multimodal synchronization. Finally, three small practice questions of increasing difficulty will be arranged, and "praise voice + animated badge" will be given when the answer is correct. Because Xiaojun is in the comfort zone and has high motivation, the system selects "advanced encouragement prompts": after answering two questions correctly, a pop-up will appear: "Your mastery of fractional questions is slowly catching up with the class average. You can try more challenging questions!" If you answer incorrectly for a series of times, you will be gently guided to review the corresponding segments in the video, rather than simply being told "wrong."
[0091] After the online learning session concludes, the system will automatically generate a brief learning summary and send it to the volunteer teacher on duty that day: "Student xx completed the micro-lesson on 'Fraction Application Problems' during the xx time period, with a final cognitive load of 0.504 (comfort zone) and a motivation score of 0.80. Pauses were mainly concentrated in the example explanation stage. It is recommended to design 1-2 more real-life scenario problems for him during offline tutoring." The volunteer can then follow up on this topic more effectively during offline tutoring the next day, achieving a true "online-offline closed loop."
[0092] Example 2: like Figure 6 As shown in the figure, this embodiment introduces an artificial intelligence-based educational resource recommendation system, which adopts the aforementioned artificial intelligence-based educational resource recommendation method; the system includes a data acquisition module, a feature extraction module, a cognitive load calculation module, a state interval identification module, and a result output module.
[0093] The data acquisition module is used to acquire audio data, interaction data, and task data collected from the user terminal; among which, audio data includes user voice questions or reading tasks; interaction data includes learning duration, click frequency, page dwell time, fast forward times, and pause times; task data includes course content tags and question structure information; The feature extraction module is used to extract features from audio data, interaction data, and task data respectively, to obtain emotional feature vectors, learning rhythm vectors, and content feature vectors; it is also used to input the emotional feature vectors and learning rhythm vectors into a pre-built emotion recognition model to obtain psychological state vectors representing motivation, tension, and fatigue. The cognitive load calculation module is used to calculate the cognitive load based on the learning rhythm vector and content feature vector, and then correct the cognitive load through the mental state vector to obtain the final cognitive load; The state interval identification module is used to match the state interval of the final cognitive load according to a pre-set threshold interval. The state intervals include the comfort zone, overload zone, and low mobility zone. The results output module is used to adjust the difficulty coefficient of recommended courses, rearrange the order of course content and supplementary learning tips according to the state interval, generate structured data and output it as intelligent supplementary learning recommendation results.
[0094] In practical applications, at least one application server is needed to deploy core programs such as feature extraction, emotion recognition models, cognitive load calculation, state interval recognition, and result output. A separate database server can be set up, or it can be combined with the application server to store user historical behavior data, course content tags, question structure information, and system configurations such as model parameters and threshold ranges. If a separate database server is set up, its configuration can be slightly lower than that of the application server, focusing on large capacity and stable I / O. Network attached storage (NAS) or an external hard drive is used for periodically backing up user data and model files to ensure system reliability and recoverability. This corresponds to the feature extraction module, cognitive load calculation module, state interval recognition module, and result output module.
[0095] The client can be any of a smartphone, tablet, or desktop computer, used to run the learning client or access the web-based learning platform. Basic configuration includes: a microphone for collecting audio data (voice questions, reading tasks); a display screen and touch / mouse input for showing course content, recording click frequency, page dwell time, fast-forward times, and pause times; internal storage and caching for temporarily caching task data and logs locally. It corresponds to the data acquisition module (front-end acquisition end) and the interface presentation end in the result output module (displaying recommended results and supplementary learning tips). Multiple student terminals are connected to the campus LAN or the Internet, reporting the collected audio data, interactive data, and task data to the server. Wired / wireless access is supported, such as Wi-Fi AP + wired router. In areas with poor network conditions, a small edge server or interactive whiteboard can be deployed to locally cache course content and temporarily store interactive logs, which are then batch-synchronized to the main server when the network becomes available.
[0096] Regarding the operating system, a Linux system is used on the server side to deploy web services, model inference services, and the database. The client side uses a common operating system such as Android, iOS, or Windows to run the learning app or browser. In addition, a relational database (such as MySQL or PostgreSQL) or a NoSQL database is used to store user information, course resources, feature data, and recommendation results. A web server / application server (such as Nginx + a backend framework) handles client requests and API calls. An AI / ML framework is used on the application server side to deploy the sentiment recognition model and potential recommendation models.
[0097] The client component calls the terminal microphone interface to collect user voice questions and reading audio at a preset sampling rate, and uploads them in real time or caches them locally. In the course playback interface, it automatically records learning duration, number of clicks, page dwell time, fast-forwarding, and pause operations. When a user enters a course or question, it retrieves the course's tag information and question structure information from the server and associates them with the local behavior log.
[0098] The server-side component receives audio streams and interaction logs from the client, performs data verification, cleaning, and storage; it also manages course content tags and question structure information, providing query and distribution interfaces. After completing data processing and obtaining structured recommendation results, the structured data is encapsulated in JSON or other standard formats and returned to the user terminal or learning platform frontend via HTTP / HTTPS interfaces.
[0099] The client-side component parses the structured recommendation results in the learning app or webpage, adjusts the sorting and difficulty tags of the subsequent recommended course list, controls the content order of the front-end playback page (e.g., animations first, then example questions), and pops up supplementary learning tips, rest suggestions, or motivational copy at appropriate times.
[0100] This embodiment has the same beneficial effects as Embodiment 1.
[0101] The technical scope of this invention is not limited to the content described above. Those skilled in the art can make various modifications and variations to the above embodiments without departing from the technical concept of this invention, and all such modifications and variations should fall within the protection scope of this invention.
Claims
1. An artificial intelligence-based method for recommending educational resources, characterized in that: Includes the following steps: Acquire audio data, interaction data, and task data collected from user terminals; among which, audio data includes user voice questions or reading tasks; interaction data includes learning duration, click frequency, page dwell time, fast forward times, and pause times; task data includes course content tags and question structure information; Feature extraction was performed on audio data, interaction data, and task data respectively to obtain emotional feature vectors, learning rhythm vectors, and content feature vectors; By inputting the emotional feature vector and the learning rhythm vector into the pre-built emotion recognition model, psychological state vectors representing the degree of motivation, tension, and fatigue are obtained. Cognitive load is calculated based on the learning rhythm vector and content feature vector, and then the cognitive load is corrected by the mental state vector to obtain the final cognitive load. The final cognitive load is matched to the state range of the pre-set threshold range, which includes the comfort zone, overload zone and low mobility zone. The difficulty level of recommended courses is adjusted based on the state interval, the order of course content is rearranged, and supplementary learning tips are added to generate structured data, which is then output as the intelligent supplementary learning recommendation result.
2. The method for recommending educational resources based on artificial intelligence according to claim 1, characterized in that, The specific steps for calculating cognitive load based on learning rhythm vector and content feature vector, and then correcting the cognitive load using mental state vector to obtain the final cognitive load, are as follows: Content complexity is obtained by weighting and standardizing the intrinsic features in the content feature vector; Pause frequency and page switching rate are extracted from the learning rhythm vector and weighted with content complexity to obtain cognitive load; where pause frequency is the ratio of the number of pauses to the learning duration, and page switching rate is the ratio of the number of page switches to the learning duration. The motivation level is extracted from the mental state vector, and the cognitive load is adjusted to obtain the final cognitive load. The adjustment formula is as follows: ; in, Indicates the final cognitive load; Indicates cognitive load; Indicates the degree of motivation; This represents the motivation adjustment coefficient.
3. The method for recommending educational resources based on artificial intelligence according to claim 1, characterized in that, The specific steps for matching the final cognitive load state range according to the pre-set threshold range, which includes the comfort zone, overload zone, and low mobility zone, are as follows: A threshold range is preset, which includes a first threshold whose values gradually increase. Second threshold and the third threshold ; Based on the final cognitive load of the current cycle The position within the threshold range combined with the degree of motivation Tension The matching logic is as follows: (1) If ≤ and A value ≥0.6 indicates the patient is in the comfort zone. (2) If ≤ and <0.4 indicates the area is in a low-mobility zone; (3) If ≥ and ≥0.7 indicates the area is in the overload zone; If none of them meet the requirements, the state interval of the historical nearest cycle will be used for matching; the historical nearest cycle is the historical cycle with the shortest time interval between the current cycle and that is in the state interval.
4. The method for recommending educational resources based on artificial intelligence according to claim 3, characterized in that, Before using the state interval for matching the historical nearest cycle, determine whether the current cycle is in the transition zone. If it is, mark it as transition and perform nearest neighbor calculation. If the nearest neighbor calculation results of K consecutive cycles are all in the same state interval, then determine that it is in that state interval. Among them, the conditions for being in the transition zone are: < < And 0.4≤ <0.6 and <0.7, or cross , or The value is less than the boundary threshold; Where K is an integer less than 5.
5. The method for recommending educational resources based on artificial intelligence according to claim 4, characterized in that, Once marked as transitional, a transition strategy is implemented based on the difficulty level of recommended courses, the order of course content rearrangement, and supplementary learning tips. The specific steps are as follows: Obtain historical adjustment strategies for the difficulty coefficient of recommended courses, the order of course content reordering, and supplementary learning tips in the near future; Based on historical adjustment strategies, a minor adjustment is made according to a preset adjustment mechanism, and the adjusted historical adjustment strategy is used as a transition strategy for the current cycle.
6. The method for recommending educational resources based on artificial intelligence according to claim 1, characterized in that, The specific steps for adjusting the difficulty level of recommended courses, rearranging the order of course content, and adding supplementary learning tips based on the student's condition range are as follows: Identify the state intervals and make the following decisions based on the identification results: (1) If the identification result is the comfort zone, the multimodal synchronization of the recommended courses will be gradually increased; (2) If the identification result is a low mobility zone, the difficulty coefficient of the recommended course is adjusted based on the preset first amplitude, and the course content order is rearranged and incentive prompts are triggered with positive emotion as the guide. (3) If the identification result is an overloaded area, the difficulty coefficient of the recommended courses is adjusted based on the preset second amplitude, and the course content order is rearranged and a rest prompt is triggered with the guidance of reducing content complexity and duration.
7. The method for recommending educational resources based on artificial intelligence according to claim 1, characterized in that, When acquiring audio data, interaction data, and task data collected from user terminals, the process also includes acquiring the user's visual data, extracting features from the visual data to obtain facial expression vectors, and inputting them together with emotion feature vectors and learning rhythm vectors into a pre-built emotion recognition model to assist in the recognition of psychological states.
8. The method for recommending educational resources based on artificial intelligence according to claim 1, characterized in that, When implementing intelligent tutoring recommendation results, the following are also included: The system receives audio data, interaction data, and task data in real time, calculates the final cognitive load after feedback, and calculates the difference between the final cognitive load before feedback to obtain the load change. Determine if the load change is less than zero; if so, continue to execute the intelligent learning recommendation results. Otherwise, adjust the first and second amplitudes of the difficulty coefficient and regenerate the intelligent tutoring recommendation results.
9. The method for recommending educational resources based on artificial intelligence according to claim 1, characterized in that, After implementing the intelligent learning recommendation results, adjustments are also made based on the performance differences between the current stage and the previous stage, specifically including: After the intelligent learning recommendation results of the current stage are completed, the performance value of the current stage and the performance value of the previous stage are obtained, and the difference between the two is used to calculate the performance difference. If the difference in effectiveness is greater than zero, continue to execute the intelligent tutoring recommendation results; Otherwise, the first and second amplitudes of the difficulty coefficient will be adjusted and used for the next stage of intelligent tutoring recommendations; The performance value is calculated by weighting the learning outcome accuracy, load adjustment amount, and effective data rate in the interactive data for each stage.
10. An artificial intelligence-based educational resource recommendation system, which employs the artificial intelligence-based educational resource recommendation method as described in any one of claims 1-9, characterized in that, It includes: The data acquisition module is used to acquire audio data, interaction data, and task data collected from the user terminal. The audio data includes user voice questions or reading tasks; the interaction data includes learning duration, click frequency, page dwell time, fast forward times, and pause times; and the task data includes course content tags and question structure information. The feature extraction module is used to extract features from audio data, interaction data, and task data respectively, to obtain emotional feature vectors, learning rhythm vectors, and content feature vectors; it is also used to input the emotional feature vectors and learning rhythm vectors into a pre-built emotion recognition model to obtain psychological state vectors representing motivation, tension, and fatigue. The cognitive load calculation module is used to calculate the cognitive load based on the learning rhythm vector and the content feature vector, and then correct the cognitive load through the mental state vector to obtain the final cognitive load. The state interval identification module is used to match the state interval of the final cognitive load according to a pre-set threshold interval. The state intervals include the comfort zone, overload zone, and low mobility zone. The results output module is used to adjust the difficulty coefficient of recommended courses, rearrange the order of course content and supplementary learning tips according to the state interval, generate structured data and output it as intelligent supplementary learning recommendation results.
Citation Information
Cited By
Multi-terminal collaborative remote academic tutoring resource scheduling method and system
CN122027818A
Course structure adaptive optimization method and system, electronic equipment and storage medium
CN122089532A