Cognitive load real-time intervention method and device based on multi-modal behavior fingerprints
By acquiring multimodal micro-behavioral data for time series alignment and feature fusion, generating behavioral fingerprint vectors, and combining them with macro indicators for cognitive load assessment, the problem of delayed monitoring results in existing technologies is solved, fine-grained perception and real-time intervention of cognitive load are achieved, and learning effects are improved.
Patent Information
- Application Number
- CN202511128748.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-13
- Publication Date
- 2025-10-17
AI Technical Summary
Existing cognitive load monitoring technologies rely on single or macro indicators and are unable to capture early micro-behavioral signals of cognitive overload, resulting in delayed monitoring results and high misjudgment rates, making it difficult to meet the needs of dynamic learning scenarios.
By acquiring multimodal micro-behavioral data, including the frequency of draft corrections, formula skipping, and writing pauses, temporal alignment and feature fusion are performed to generate behavioral fingerprint vectors. Cognitive load assessment is then conducted in combination with macro indicators to dynamically adjust the difficulty of the question bank.
It achieves fine-grained perception and real-time intervention of cognitive load status, improves the accuracy of monitoring results and learning effects, avoids cognitive overload or slackness, and promotes the balanced development of students' abilities.
Smart Images

Figure CN120807241A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of educational informatization, and in particular to a cognitive load real-time intervention method and device based on multi-modal behavior fingerprints. BACKGROUND
[0002] Existing cognitive load monitoring technologies generally rely on macro indicators such as answer time, accuracy, expression recognition, or physiological signals. These macro indicators can provide rough cognitive state evaluation, but they have fundamental defects. First, they cannot capture early micro behavior signals of cognitive overload. Second, they lack fine-grained perception and real-time intervention capabilities. These two defects result in lagging monitoring results and high misjudgment rates, making it difficult to meet the needs of dynamic learning scenarios. SUMMARY
[0003] The present application provides a cognitive load real-time intervention method and device based on multi-modal behavior fingerprints to solve the problem of low accuracy of monitoring results caused by relying on single or macro indicators in existing cognitive load monitoring technologies.
[0004] In a first aspect, the present application provides a cognitive load real-time intervention method based on multi-modal behavior fingerprints, comprising:
[0005] Obtaining multi-modal micro behavior data of a student to be tested during answering, the multi-modal micro behavior data including draft revision frequency, formula jump frequency, and writing pause frequency;
[0006] Performing time series alignment and feature fusion on the multi-modal micro behavior data, extracting cognitive load features corresponding to each modality, and generating a behavior fingerprint vector from all cognitive load features;
[0007] Inputting the behavior fingerprint vector and macro indicators into a dynamic cognitive load evaluation model to output a cognitive load level of the student to be tested, the macro indicators including answer time, accuracy, expression, and physiological signals;
[0008] According to the cognitive load level and the portrait of the student to be tested, dynamically updating the difficulty level of the question bank of the student to be tested.
[0009] In a second aspect, the present application provides a cognitive load real-time intervention device based on multi-modal behavior fingerprints, comprising:
[0010] A behavior data acquisition module for obtaining multi-modal micro behavior data of a student to be tested during answering, the multi-modal micro behavior data including draft revision frequency, formula jump frequency, and writing pause frequency;
[0011] a fingerprint vector generation module configured to perform time sequence alignment and feature fusion on the multi-modal micro-behavior data, extract cognitive load features corresponding to each modality, and generate a behavior fingerprint vector from all the cognitive load features;
[0012] a level determination module configured to input the behavior fingerprint vector and macro-indicators into a dynamic cognitive load evaluation model, and output a cognitive load level of the student under test, wherein the macro-indicators include answering time, accuracy, expression, and physiological signals;
[0013] a question bank updating module configured to dynamically update the difficulty level of the question bank of the student under test according to the cognitive load level and the student profile.
[0014] The present application provides a cognitive load real-time intervention method and device based on multi-modal behavior fingerprint. The method includes the following steps: obtaining multi-modal micro-behavior data of a student under test during answering, wherein the multi-modal micro-behavior data includes draft revision frequency, formula skip frequency, and writing pause frequency; performing time sequence alignment and feature fusion on the multi-modal micro-behavior data, extracting cognitive load features corresponding to each modality, and generating a behavior fingerprint vector from all the cognitive load features; inputting the behavior fingerprint vector and macro-indicators into a dynamic cognitive load evaluation model, and outputting a cognitive load level of the student under test, wherein the macro-indicators include answering time, accuracy, expression, and physiological signals; and dynamically updating the difficulty level of the question bank of the student under test according to the cognitive load level and the student profile. The present application combines micro-behavior data such as draft revision frequency, formula skip frequency, and writing pause frequency, and macro-indicators such as answering time, accuracy, expression, and physiological signals. The micro-behavior data can capture the immediate reaction and detailed operation of the student during problem solving, while the macro-indicators provide a reference for overall performance and physiological state. The combination of the two can more comprehensively reflect the cognitive load state of the student, avoiding the limitations of a single data source. In addition, by performing time sequence alignment on the multi-modal micro-behavior data, the consistency of different modal data on the time axis is ensured, and then feature fusion is performed, which can preserve the time sequence relationship and associated information between data, extract more representative cognitive load features, and improve the accuracy of evaluation. At the same time, by dynamically adjusting the difficulty level of the question bank, the method can guide the student to balance between the comfort zone and the challenge zone, avoiding cognitive overload caused by excessive pressure and preventing cognitive slack caused by lack of challenge. This balance helps to promote the balanced development of the student's ability and improve the learning effect. BRIEF DESCRIPTION OF DRAWINGS
[0015] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings needed to be used in the embodiments or prior art description. Obviously, the drawings in the following description only some of the embodiments of the present application, and for those skilled in the art, other drawings can also be obtained without creative labor.
[0016] Figure 1 is a flowchart of the cognitive load real-time intervention method based on multi-modal behavior fingerprint provided by the embodiments of the present application;
[0017] Figure 2 is a structural schematic diagram of the cognitive load real-time intervention device based on multi-modal behavior fingerprint provided by the embodiments of the present application. DETAILED DESCRIPTION
[0018] In the following description, specific details such as specific system structures, techniques, etc. are presented in order to thoroughly understand the embodiments of the present application. However, it should be clear to those skilled in the art that the present application can also be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits and methods are omitted to avoid unnecessary details that hinder the description of the present application.
[0019] In order to make the purpose, technical solutions and advantages of the present application more clear, the following will be described by specific embodiments in conjunction with the drawings.
[0020] Figure 1 The implementation flowchart of the cognitive load real-time intervention method based on multi-modal behavior fingerprint provided by the embodiments of the present application is described in detail as follows:
[0021] In step 101, multi-modal micro-behavior data of the student to be tested when answering the question is obtained, and the multi-modal micro-behavior data includes the frequency of draft correction, the frequency of formula skip and the frequency of writing pause.
[0022] In the embodiments of the present application, in order to view the acceptance of the student to be tested to the difficulty of the questions in the current question bank, the multi-modal micro-behavior data of the student to be tested when answering the question including the frequency of draft correction, the frequency of formula skip and the frequency of writing pause is obtained through the corresponding device.
[0023] The embodiments of the present application use micro-behavior data to capture the immediate reaction and detailed operation of the student in the problem solving process, and provide a basis for reflecting the cognitive load state of the student subsequently.
[0024] In one possible implementation, obtaining the multi-modal micro-behavior data of the student to be tested when answering the question can include:
[0025] The interactive recording device is used to capture the touch behavior sequence data of the student to be tested when answering the questions, and the frequency of the student to be tested correcting the draft when answering the questions is determined based on the touch behavior sequence data;
[0026] An image of the test paper of the student under test is obtained through an image sensor, and the answer steps written by the student under test in the answer area of each question are extracted from the answer paper image. For each question, the answer steps written in the answer area of the question are compared with the standard answer steps of the question to determine the frequency of formula step skipping by the student under test when answering the question;
[0027] The note track, writing pressure and writing speed of the student to be tested when answering the questions are obtained through the handwriting board, and the writing pause frequency of the student to be tested when answering the questions is determined based on the note track, writing pressure and writing speed of the student to be tested when answering the questions.
[0028] Optionally, the multimodal micro-behavior data in the embodiment of the present application may include the frequency of draft corrections, the frequency of formula skipping, and the frequency of writing pauses. The process of determining the multimodal micro-behavior data is as follows:
[0029] (1) The process of determining the frequency of draft revisions:
[0030] For each student, as they answer questions on a tablet using a stylus, the interactive recording device (i.e., the interaction between the tablet and the stylus) captures real-time data on the student's touch behavior, such as smudge displacement and writing pressure. The smudge displacement and writing pressure are then combined to determine the frequency of corrections made by the student during their draft.
[0031] That is, the erasing action time is extracted from the touch behavior sequence data, when the erasing position Δ satisfies:
[0032] Δ>λW and Counted as a valid modification.
[0033] Where W is the width of the writing area, λ is the smear displacement coefficient, P is the writing pressure, P th is the writing pressure change rate threshold, and dt is the differential time.
[0034] Then, the draft modification frequency is determined by using the ratio of the number of valid modifications to the number of all modifications in the acquired touch behavior sequence data.
[0035] For example, the smear displacement and writing pressure of the stylus on the writing board during the student A's answering process are obtained to determine the frequency of student A's draft corrections during the answering process.
[0036] (2) The process of determining the frequency of the formula jump:
[0037] For each student to be tested, when answering questions on the handwriting board, the image sensor above the handwriting board is used to obtain the answer sheet image of the student to be tested when answering questions. Then the answer steps written by the student to be tested in the answer area are identified from the answer sheet image. Finally, the answer steps of each question in the corresponding answer area are compared with the standard answer steps of the corresponding question one by one to determine the formula jump frequency of each question.
[0038] In a possible implementation, extracting the answer steps written by the student to be tested in the answer area of each question from the answer sheet image can include:
[0039] The answer sheet image is input into a text extraction model, and the answer steps written in each answer area in the answer sheet image are output. The text extraction model is constructed based on a long short-term memory network.
[0040] Optionally, the obtained answer sheet image is input into the trained text extraction model, and the answer steps written in each answer area in the answer sheet image are output.
[0041] The construction and training process of the text extraction model is as follows:
[0042] The historical answer sheet image and the answer steps of each answer area in the historical answer sheet image are obtained.
[0043] The text extraction model is constructed using a long short-term memory network.
[0044] The historical answer sheet image is input, and the answer steps of each answer area in the historical answer sheet image are output to train the text extraction model.
[0045] (3) The determination process of the writing pause frequency is as follows:
[0046] For each student to be tested, when the student to be tested answers questions on the handwriting board, the pen trajectory, writing pressure, and writing speed are collected in real time. Then, according to the fluent and unfluent features of the pen trajectory, the uniform and non-uniform features of the writing pressure, and the fast and slow features of the writing speed, the writing pause frequency of the student to be tested during the current answering is determined.
[0047] For example, when the pen trajectory is fluent, the writing pressure is uniform, and the writing speed is fast, it indicates that the student has a low writing pause frequency during the current answering process.
[0048] For example, when the pen trajectory is not fluent, the writing pressure is not uniform, and the writing speed is too slow, it indicates that the student has too many writing pause frequencies during the answering process.
[0049] In a possible implementation, after obtaining the multi-modal micro-behavior data of the student to be tested during the answering, the method can further include:
[0050] performing noise reduction processing on the multi-modal micro-behavior data;
[0051] performing normalization processing on the multi-modal micro-behavior data after the noise reduction processing;
[0052] correspondingly, the time sequence alignment and feature fusion on the multi-modal micro-behavior data, the extraction of the cognitive load feature, and the generation of the behavior fingerprint vector from the cognitive load feature can include:
[0053] performing time sequence alignment and feature fusion on the multi-modal micro-behavior data after the normalization processing, extracting a cognitive load feature, and generating a behavior fingerprint vector from the cognitive load feature.
[0054] Optionally, after obtaining the multi-modal behavior data of the student to be tested during the answering, in the process of time sequence alignment and feature fusion, the quality impairment of any one of the modal data caused by noise pollution will affect the extraction of the final cognitive load feature, therefore, the processing of the noisy multi-modal micro-behavior data is an indispensable link.
[0055] In a possible implementation, the noise reduction processing on the multi-modal micro-behavior data can include:
[0056] applying a low-rank and sparse penalty decomposition method to impose a sparse constraint on the multi-modal micro-behavior data;
[0057] According to the local correlation inside the multi-modal micro-behavior data, the global structural information is mined by using the internal correlation of the multi-modal micro-behavior data, so as to remove the noise of the multi-modal micro-behavior data.
[0058] Optionally, when processing the multi-modal micro-behavior data, the data source is extensive and complex, and often contains a large amount of redundant information and noise, which brings great challenges to the accurate analysis and effective use of the data. As an advanced data processing technology, the low-rank and sparse penalty decomposition method can effectively deal with this problem.
[0059] The core idea of the low-rank and sparse penalty decomposition method is based on two important assumptions: one is that the multi-modal micro-behavior data has a low-rank characteristic, which means that there is a potential low-rank structure in the data, and the main features can be described by a few key factors; the other is that there is sparsity in the data, that is, most data elements contribute less to the overall information, and only a few elements contain important information.
[0060] Specifically, multi-modal micro-behavior data is represented as a high-dimensional matrix or tensor. By introducing a low-rank constraint, we assume that the matrix or tensor can be decomposed into the product of multiple low-rank matrices or tensors. For example, for a matrix A ∈ R m×n , we can decompose it as A ≈ U∑V T , where U ∈ R m×k , ∑ ∈ R k×k , V ∈ R n×k , and k ≤ min(m, n). Here, k represents the rank of the low-rank matrix, which is much smaller than the number of rows and columns of the original matrix, thus achieving dimensionality reduction and low-rank representation of data.
[0061] At the same time, in order to further mine the sparse information in the data, we impose a sparsity constraint during the decomposition process. This is usually achieved by introducing a sparsity penalty term, and common sparsity penalty terms include the L1 norm (i.e., the sum of absolute values) and the L0 norm (i.e., the number of non-zero elements). Since the L0 norm is difficult to optimize, we usually use the L1 norm as an approximation. By adding an L1 norm penalty term to the objective function, we encourage some elements in the decomposed matrix or tensor to be zero or close to zero, thus highlighting the key information in the data and removing redundancy and noise.
[0062] In actual operation, optimization algorithms such as Alternating Least Squares (ALS) can be used to solve the low-rank and sparse penalty decomposition problem. The ALS algorithm optimizes other variables by fixing part of the variables alternately, gradually approaching the optimal solution. The specific steps are as follows:
[0063] (1) Initialize parameters: randomly initialize the decomposed matrices U, ∑, and V.
[0064] (2) Alternating optimization:
[0065] 1) Fix ∑ and V, update U: update U by solving a least squares problem to minimize the error A ≈ U∑V T .
[0066] 2) Fix U and ∑, update V: also update V by solving a least squares problem.
[0067] 3) Update ∑: according to the updated results of U and V, recalculate ∑ to maintain the accuracy of the decomposition.
[0068] (3) Introduce sparsity constraint: in the process of updating U, V, and ∑ each time, add an L1 norm penalty term to constrain the sparsity of the elements in the decomposed matrix.
[0069] (4) Iterative convergence: repeat the above alternating optimization steps until the value of the objective function converges or reaches the preset number of iterations.
[0070] In addition, since the multi-modal micro-behavior data not only has low-rank and sparse characteristics, but also has internal local correlation. This local correlation reflects the similarity and correlation of data at different modalities or different time and spatial scales. By mining this local correlation, the internal structure of the data can be further revealed, thereby achieving more effective noise removal.
[0071] First, a local correlation analysis method, such as Locally Linear Embedding (LLE) or Locality Preserving Projections (LPP), is used to identify the local neighborhood structure in the data. These methods connect each data point to its nearest neighbor points by calculating the similarity or distance between data points, forming a local neighborhood graph. In the local neighborhood graph, adjacent data points have similar features and behavior patterns, thereby reflecting the local correlation of the data.
[0072] Next, the information in the local neighborhood graph is used to mine global structure information.
[0073] Specifically, by optimizing a global objective function, the data maintains the local neighborhood structure in the global range while revealing the low-rank and sparse characteristics of the data as much as possible. This objective function usually includes two parts: one part is the local preserving term, which is used to maintain the local correlation of the data; the other part is the low-rank and sparse penalty term, which is used to promote the low-rank representation and sparsity of the data.
[0074] In the process of optimizing the global objective function, the embodiments of the present application use an iterative optimization method to gradually adjust the representation of the data so that the value of the objective function continuously decreases. The specific steps are as follows:
[0075] (1) Construct a local neighborhood graph: according to the similarity or distance between data points, construct a local neighborhood graph to determine the neighbor points of each data point.
[0076] (2) Initialize the data representation: randomly initialize the low-rank representation matrix X of the data.
[0077] (3) Iterative optimization:
[0078] 1) Calculate the local preserving term: according to the local neighborhood graph, calculate the similarity between each data point and its neighbor points, and construct the local preserving matrix W.
[0079] 2) Calculate the low-rank and sparse penalty term: according to the value of X, calculate the low-rank penalty term (such as the nuclear norm) and the sparse penalty term (such as the L1 norm).
[0080] 3) Update X: update the value of X by solving an optimization problem to minimize the value of the global objective function.
[0081] 4) Noise removal: after multiple iterations of optimization, the matrix of X obtained is the data representation after removing noise. Since the X matrix has low rank and sparse characteristics, and maintains the local correlation of the data, it can more accurately reflect the essential characteristics of the data.
[0082] Based on the above, the embodiments of the present application not only can remove the redundancy and noise in the multi-modal micro-behavior data using the low-rank and sparse penalty decomposition method, but also can mine the global structure information according to the internal local correlation of the data, further improving the quality and usability of the data. This has important significance for subsequent data analysis, pattern recognition and decision making tasks.
[0083] Since the multi-modal micro-behavior data obtained by the embodiments of the present application exist in different unit levels, the embodiments of the present application also need to normalize the multi-modal micro-behavior data after noise reduction, so that the multi-modal micro-behavior data are all within a range.
[0084] In step 102, the multi-modal micro-behavior data is time-aligned and feature-fused, the cognitive load features corresponding to each modality are extracted, and all cognitive load features are generated into a behavior fingerprint vector.
[0085] Wherein, the cognitive load refers to the degree of consumption of cognitive resources borne by an individual when completing a task.
[0086] The behavior fingerprint vector is a comprehensive representation of the individual's cognitive load state, which can integrate the cognitive load features of each modality together to form a unique and representative vector.
[0087] In the embodiments of the present application, since the multi-modal micro-behavior data usually covers multiple different data modalities, such as draft revision frequency, formula skip frequency and writing pause frequency, etc. Due to the differences in their own data acquisition frequency, starting time and time resolution, they are not strictly aligned on the time axis during the acquisition process. Therefore, the embodiments of the present application need to first time-align and feature-fuse the multi-modal micro-behavior data, which is a key prerequisite for accurately extracting cognitive load features and generating behavior fingerprint vectors. Then, after fusion, the cognitive load features corresponding to each modality are extracted, and all cognitive load features are generated into a behavior fingerprint vector.
[0088] The embodiment of the application ensures the consistency of different modal data on the time axis by time sequence alignment of multi-modal micro-behavior data, and then performs feature fusion. This technology can preserve the time sequence relationship and correlation information between data, extract more representative cognitive load features, and thus improve the accuracy of evaluation.
[0089] In a possible implementation, the time sequence alignment and feature fusion of multi-modal micro-behavior data, and the extraction of cognitive load features corresponding to each modality can include:
[0090] The time sequence alignment of multi-modal micro-behavior data is performed by using a time warping dynamic alignment algorithm.
[0091] The feature fusion of the time sequence aligned multi-modal micro-behavior data is performed by using a multi-modal feature fusion technology.
[0092] The features related to the preset cognitive load of the corresponding modality are extracted from the multi-modal micro-behavior data after feature fusion by using a correlation analysis algorithm, to obtain the cognitive load features of the corresponding modality.
[0093] Optionally, in the process of collecting multi-modal micro-behavior data, due to the differences in the collection mechanism and triggering conditions of different data sources, the three kinds of data of draft revision frequency, formula skip frequency and writing pause frequency obtained are not strictly synchronized on the time axis. For example, the record of draft revision frequency may depend on the real-time monitoring of writing content by a specific image recognition algorithm, and the time accuracy is affected by the image processing speed; the formula skip frequency may be determined by analyzing the logical structure of the writing content, which involves complex semantic understanding and relatively long processing time; and the record of writing pause frequency may be based on real-time sampling of writing device pressure or motion sensor, and the time accuracy is high but there may be noise interference. This different synchronization in time will seriously affect the subsequent comprehensive analysis and processing of multi-modal micro-behavior data, so the time warping dynamic alignment algorithm is used to realize the time sequence alignment of multi-modal micro-behavior data.
[0094] The time warping dynamic alignment algorithm is a time sequence alignment method based on the dynamic programming idea, which can automatically find the optimal matching path between two time sequences, and thus align them on the time axis. In the context of multi-modal micro-behavior data, the operation steps are as follows:
[0095] Data preprocessing: First, the collected draft revision frequency, formula jump frequency and writing pause frequency data are preprocessed. For the draft revision frequency data, remove abnormal revision records caused by image recognition errors, such as frequent and meaningless revision points in a short time; for the formula jump frequency data, correct the misjudgment jump situation caused by semantic understanding errors, which can be combined with manual inspection and algorithm optimization; for the writing pause frequency data, use filtering algorithm to remove false pause signals caused by writing device jitter or environmental interference.
[0096] Constructing distance matrix: Select two time series that need to be aligned, such as the draft revision frequency time series and the formula jump frequency time series. Calculate the distance between each pair of corresponding points in the two sequences to construct a distance matrix. The distance calculation can use a weighted distance measurement method, considering that different behavior characteristics have different effects on cognitive load, and assigning different weights to draft revision, formula jump and writing pause. For example, draft revision may better reflect the uncertainty and revision process of thinking, and is assigned a higher weight; formula jump may imply the jumping nature of thinking, and is assigned a medium weight; while writing pause may reflect the temporary stagnation of thinking, and is assigned a lower weight. The distance calculated in this way can more accurately reflect the difference between the two time series.
[0097] Finding the optimal path: Use dynamic programming algorithm to find an optimal path from the top left corner to the bottom right corner in the distance matrix, so that the sum of distances on the path is minimized. This optimal path represents the best matching relationship between the two time series. In the process of finding the optimal path, set the limit conditions of time distortion, such as limiting the maximum stretching or compression ratio of the time series, to avoid excessive distortion and ensure that the aligned data still has practical significance. At the same time, considering the correlation between multi-modal micro-behavior data, introduce constraint conditions in the path search process, so that the alignment result is more in line with the logic of cognitive load change.
[0098] Data alignment: According to the found optimal path, align the two time series. For each point on the path, match the data points in the corresponding time series, and for the data points that do not match, use interpolation method to fill in. For example, use cubic spline interpolation method to insert new data points between adjacent data points, so that the inserted data points can better fit the trend of the original data, and ensure that the two time series are consistent on the time axis. Repeat the above steps to align the writing pause frequency time series with the already aligned draft revision frequency and formula jump frequency time series, and finally realize the time series alignment of the three multi-modal micro-behavior data.
[0099] The calculation formula of the time distortion dynamic alignment algorithm is:
[0100]
[0101] where L algn is the alignment loss, W is the time warp matrix, x i is the i-th data point in the first time series, r(i) is the mapping function, y r(i) is the data point in the second time series corresponding to x i , λ is the regularization parameter, and TV(W) is the total variation of the time warp matrix W.
[0102] The advantage of the time warp dynamic alignment algorithm is that it can handle the alignment problem of time series with different lengths and different sampling rates, and can automatically adapt to local changes in the time series, with high alignment accuracy and robustness. Through time series alignment, the three multi-modal micro-behavior data of draft correction frequency, formula skip frequency and writing pause frequency can be integrated into the same time frame, laying a foundation for subsequent feature fusion and cognitive load feature extraction.
[0103] In addition, although the draft correction frequency, formula skip frequency and writing pause frequency data after time series alignment have the same time reference, the data of different modalities still exist in independent forms. In order to fully utilize the complementary information between multi-modal micro-behavior data and improve the representation ability of cognitive load, the multi-modal feature fusion technology is adopted to fuse data of different modalities in the embodiment of the present application.
[0104] The multi-modal feature fusion technology mainly includes two types of model-based methods and statistical-based methods, and in this scenario, a deep learning-based model fusion method can be used, and the specific operation steps are as follows:
[0105] Feature extraction: features are extracted from the micro-behavior data of each modality after time series alignment. For the draft correction frequency data, the number of corrections, the duration of corrections, the size of the correction area, etc. can be extracted; for the formula skip frequency data, the number of skips, the interval time of skips, the complexity change of formulas before and after skips, etc. can be extracted; for the writing pause frequency data, the number of pauses, the duration distribution of pauses, the writing speed change before and after pauses, etc. can be extracted. These features can reflect the individual's behavior patterns and cognitive state from different angles.
[0106] Building deep learning model: A multilayer perceptron (MLP) is used as the base model for multi-modal feature fusion. MLP is a feedforward artificial neural network composed of an input layer, multiple hidden layers, and an output layer. The extracted features of different modalities are input into the input layer of the MLP. The hidden layers transform and combine the input features through nonlinear activation functions to extract higher-level abstract features. In the design of hidden layers, the number of neurons can be decreased layer by layer to achieve gradual compression and abstraction of features. The output layer outputs the fused feature vector, which contains information from multiple modalities and can more comprehensively reflect the individual's cognitive load state.
[0107] Model training and optimization: The constructed MLP model is trained using the annotated data set. The annotated data set contains multi-modal micro-behavior data under different cognitive load levels and their corresponding cognitive load labels. During training, the backpropagation algorithm and stochastic gradient descent optimization algorithm are used to adjust the model parameters to minimize the error between the model output and the true label. At the same time, cross-validation is used to evaluate and optimize the model, and the optimal model structure and parameter settings are selected to improve the model's generalization ability and fusion effect.
[0108] Feature fusion result evaluation: The effectiveness of feature fusion is evaluated by calculating the correlation between the fused features and the cognitive load labels, classification accuracy, and other indicators. If the evaluation results show that the fused features can better reflect the changes in cognitive load, it means that the feature fusion is successful; otherwise, the feature extraction method, model structure, or training parameters need to be adjusted and optimized until a satisfactory fusion effect is achieved.
[0109] Multi-modal feature fusion technology can fully utilize the complementary information between different modal data, improve the representation ability and classification accuracy of cognitive load. Through feature fusion, the three multi-modal data of draft revision frequency, formula skip frequency, and writing pause frequency can be integrated into a unified feature representation, providing more rich and accurate information for subsequent cognitive load feature extraction.
[0110] Then, the multi-modal micro-behavior data after feature fusion contains rich information from multiple modalities, but not all features are related to cognitive load. In order to extract features that can accurately reflect the cognitive load state, the correlation analysis algorithm is used in the present application to filter out features related to the pre-set cognitive load of the corresponding modal from the fused features.
[0111] In the present application, a correlation analysis method combining Pearson correlation coefficient and partial correlation coefficient can be used, and the specific operation steps are as follows:
[0112] Calculate the Pearson correlation coefficient: First, calculate the Pearson correlation coefficient between each feature after feature fusion and the preset cognitive load. The Pearson correlation coefficient is used to measure the degree of linear correlation between two variables, and its value ranges between [-1, 1]. When the correlation coefficient is 1, it means that the two variables are completely positively correlated; when the correlation coefficient is -1, it means that the two variables are completely negatively correlated; when the correlation coefficient is 0, it means that there is no linear correlation between the two variables. By calculating the Pearson correlation coefficient, features with certain linear correlation with cognitive load can be preliminarily screened out.
[0113] Calculate the partial correlation coefficient: Considering that there may be complex relationships between multi-modal micro-behavior data, using only the Pearson correlation coefficient may be disturbed by other features. Therefore, further calculate the partial correlation coefficient between each feature and the cognitive load. The partial correlation coefficient measures the correlation between two variables after controlling the influence of other features. By calculating the partial correlation coefficient, the independent correlation between each feature and the cognitive load can be more accurately determined, excluding the interference of other features.
[0114] Feature selection: According to the calculated Pearson correlation coefficient and partial correlation coefficient, set the corresponding threshold for feature selection. For example, select the features with an absolute value of the Pearson correlation coefficient greater than 0.5 and an absolute value of the partial correlation coefficient greater than 0.3 as the features related to the preset cognitive load of the corresponding modality. At the same time, combined with domain knowledge and practical experience, the selected features are further analyzed and verified to ensure that these features can truly reflect the changes of cognitive load.
[0115] Get cognitive load features corresponding to the modality: classify the selected features related to the corresponding modality of the draft revision frequency, formula jump frequency and writing pause frequency, and get the cognitive load features corresponding to the modality. These cognitive load features can more accurately reflect the cognitive load state of individuals in different behavior patterns, providing strong support for subsequent cognitive load evaluation, learning effect analysis and personalized teaching and other applications.
[0116] The embodiments of the present application use correlation analysis algorithm to extract features related to the preset cognitive load of the corresponding modality from the multi-modal micro-behavior data after feature fusion, which provides an important basis for in-depth understanding of individual cognitive process and behavior patterns.
[0117] In one possible implementation, generating a behavior fingerprint vector from all cognitive load features can include:
[0118] Splicing all cognitive load features into a target vector according to a preset order;
[0119] Inputting the target vector into the feature encoder to generate a fingerprint vector.
[0120] Optionally, after obtaining the cognitive load features of each modality, they cannot be simply combined at will, but need to be spliced in a certain logical order to form a complete and representative target vector. The determination of this preset order is crucial, which directly affects the generation effect of the subsequent fingerprint vector and the representation ability of cognitive load.
[0121] In one possible implementation, splicing all cognitive load features into a target vector according to a preset order can include:
[0122] Using XGBoost to calculate the contribution degree of each cognitive load feature to the corresponding cognitive load, and sorting each cognitive load feature according to the contribution degree from strong to weak, and taking the sorting as the preset order;
[0123] Splicing the cognitive load features into a target vector according to the preset order.
[0124] Optionally, the embodiments of the present application can use the XGBoost algorithm to calculate the contribution degree of each cognitive load feature to the corresponding cognitive load. Wherein, XGBoost is a high-efficiency integrated learning algorithm, which can accurately evaluate the importance of each feature in the prediction task by integrating multiple decision trees.
[0125] The specific splicing process into a target vector is:
[0126] (1) Data preparation and model training: first, all the extracted cognitive load features and corresponding cognitive load labels are combined to form a training data set. The label can be the cognitive load level defined according to actual needs, such as high, medium and low three levels. Then, use this data set to train the XGBoost model. During the training process, XGBoost will automatically adjust the parameters of the model according to the relationship between the features and the labels to minimize the prediction error.
[0127] (2) Feature contribution degree calculation: after training, the XGBoost model will calculate a contribution degree index for each cognitive load feature. This index reflects the importance of the feature in predicting cognitive load. The calculation method of contribution degree is based on the split point selection of the feature in the decision tree and the contribution of information gain. For example, if a feature is frequently used in the split points of multiple decision trees, and each use can bring a larger information gain, then the contribution degree of this feature will be higher.
[0128] (3) Feature ranking and preset order determination: According to the calculated contribution degree of each cognitive load feature, it is ranked in order from strong to weak. This ranking result will be used as the preset order of the subsequent splicing target vector. The feature with high contribution degree means that it has a greater impact on cognitive load and should be placed in a more important position in the target vector so that its information can be better preserved in the subsequent feature encoding process.
[0129] (4) Target vector splicing: According to the above determined preset order, all cognitive load features are spliced in turn to form a target vector. For example, assuming that there are three cognitive load features A, B, and C, and their contribution degree ranking is A > B > C, then the spliced target vector is in the form of [A, B, C].
[0130] The spliced target vector contains all the cognitive load feature information, but may have problems such as high dimension and information redundancy. Therefore, the embodiment of the present application needs to input it into the feature encoder to further extract and compress the information and generate a more representative and discriminative fingerprint vector.
[0131] Among them, the feature encoder can adopt various forms, such as auto-encoder (Auto-encoder), principal components analysis (PCA), etc. Taking the auto-encoder as an example:
[0132] (1) Auto-encoder structure: Auto-encoder is an unsupervised neural network model, which consists of an encoder and a decoder. The encoder is responsible for compressing the input target vector into a low-dimensional latent representation, i.e. the fingerprint vector; the decoder tries to reconstruct the original target vector from the latent representation.
[0133] (2) Model training: Use the spliced target vector as training data to train the auto-encoder. The goal of training is to minimize the reconstruction error, i.e. to make the reconstructed vector output by the decoder as close as possible to the original target vector. During the training process, the auto-encoder will learn the main features and patterns in the target vector, thereby achieving effective compression and extraction of information.
[0134] (3) Fingerprint vector generation: After training, input the target vector into the encoder part of the trained auto-encoder, and the encoder will convert it into a low-dimensional fingerprint vector. This fingerprint vector retains the key information in the target vector while removing redundancy and noise, and can more accurately represent the individual's cognitive load state.
[0135] The embodiment of the application splices all cognitive load features in a preset order into a target vector, and inputs the same into a feature encoder to generate a fingerprint vector, thereby providing more concise and effective data representation for subsequent cognitive load evaluation, behavior analysis and other tasks.
[0136] In step 103, the behavior fingerprint vector and the macroscopic index are input into a dynamic cognitive load evaluation model, and the cognitive load level of the student to be tested is output. The macroscopic index includes the answering time, accuracy, expression and physiological signal.
[0137] In the embodiment of the application, the behavior fingerprint vector generated according to step 102 and the macroscopic index including the answering time, accuracy, expression and physiological signal are input into the dynamic cognitive load evaluation model which has been trained, and the cognitive load level of the student to be tested is output.
[0138] The dynamic cognitive load evaluation model is constructed by using a long short-term memory network combined with a double-flow attention mechanism, and the construction process is as follows:
[0139] The historical behavior fingerprint vector and the historical macroscopic index corresponding to the historical behavior fingerprint vector are obtained, and the historical cognitive load level corresponding to the historical behavior fingerprint vector and the historical macroscopic index is obtained.
[0140] The dynamic cognitive load evaluation model is constructed by using a long short-term memory network combined with a double-flow attention mechanism.
[0141] The historical behavior fingerprint vector and the historical macroscopic index are used as input, and the historical cognitive load level is used as output, and the dynamic cognitive load evaluation model is trained.
[0142] The embodiment of the application introduces the historical behavior fingerprint vector and the historical macroscopic index at the same time, and the model constructs a “micro-macro” double-dimensional feature system. The time series modeling capability of the LSTM can capture the relevance between the instantaneous operation mode in the behavior fingerprint and the long-term change trend of the macroscopic index, and the double-flow attention mechanism automatically selects the feature combination with higher contribution to load evaluation in different scenes by dynamically allocating weights. The synergistic fusion of data and features enables the model to maintain stable evaluation performance when facing data noise, modal missing or scene migration, and significantly improves the robustness.
[0143] For example, the behavior fingerprint vector and the macroscopic index of student A are obtained, and both of them are input into the dynamic cognitive load evaluation model, and the cognitive load level of student A is output, including three cases respectively.
[0144] In the first case, the cognitive load level is low, indicating that student A can easily cope with the difficulty of the current question when coping with the current question.
[0145] The second case is that the cognitive load level is medium, indicating that student A can accept the difficulty of the current question when dealing with it.
[0146] The third case is that the cognitive load level is high, indicating that student A has difficulty dealing with the difficulty of the current question.
[0147] In the answering process, the method can monitor the cognitive load level of the student in real time and provide immediate feedback. This feedback can help the student understand his cognitive state in time and adjust his learning strategy, such as speeding up or slowing down the problem-solving speed, seeking help, etc.
[0148] In step 104, the difficulty level of the question bank of the student to be tested is dynamically updated according to the cognitive load level and the portrait of the student to be tested.
[0149] In the embodiments of the present application, after accurately obtaining the cognitive load level of the student to be tested and combining the detailed student portrait, the difficulty level of the student's question bank can be dynamically updated to achieve more personalized and efficient learning support.
[0150] The cognitive load level intuitively reflects the cognitive pressure that the student bears in the current learning task. If the cognitive load level is high, it indicates that the student has great difficulty in dealing with the existing question, which may have exceeded his current knowledge and cognitive ability. For example, in the mathematics subject, if the cognitive load level of the student is high when solving a series of complex function synthesis questions, it means that these questions are too difficult for the student, and continuing to provide the same level of questions may cause the student to feel frustrated, reduce learning enthusiasm, and even affect the overall interest in the mathematics subject.
[0151] The portrait of the student to be tested comprehensively describes the student from multiple dimensions, including but not limited to the student's stress resistance, learning style, knowledge base, learning habits, interests and hobbies, and past academic performance, etc. For example, some students are good at logical reasoning and understand geometric problems in mathematics quickly, but are relatively weak in algebraic operations; some students are used to consolidating knowledge through a lot of practice, while some students prefer to learn through understanding concepts and principles. These information is crucial for determining the appropriate difficulty level of the question bank.
[0152] In the embodiments of the present application, the student portrait to be tested is directly extracted from an existing portrait database. The portrait database exists in the form of a table, including the student's name and the scores of the student's stress resistance ability, learning style, knowledge base, learning habit, interest and hobby, etc., and then the total score calculated according to the corresponding weight is obtained as the student portrait of each student to be tested. For example, if the total score is between 80 and 100, it indicates that the student's ability in stress resistance, knowledge acquisition, etc. is relatively high; if the total score is between 60 and 80, it indicates that the student's ability in stress resistance, knowledge acquisition, etc. is moderate; and if the total score is less than 60, it indicates that the student's ability in stress resistance, knowledge acquisition, etc. is weak.
[0153] Based on the cognitive load level and the student portrait, the specific operation of dynamically updating the difficulty level of the question bank is as follows:
[0154] When the cognitive load level is high and the student portrait score is less than 60, the difficulty of the related type of questions in the question bank is appropriately reduced.
[0155] For example, if the portrait score of the student portrait is less than 60, i.e. the student lacks knowledge of English grammar, and the current cognitive load level of the English reading comprehension question is high, then the proportion of long and difficult sentences and complex grammatical structures in the reading comprehension question can be appropriately reduced, some articles with relatively simple grammatical structures and moderate vocabulary can be increased, and some special exercise questions for grammar knowledge points can be matched to help the student gradually consolidate the foundation and reduce the cognitive burden.
[0156] For another example, if the portrait score of the student portrait is less than 60, i.e. the student's stress resistance ability is poor, and the current cognitive load level of the question is high, then the difficulty of the question can be appropriately reduced to reduce the cognitive burden of the student.
[0157] If the cognitive load level is moderate, it means that the difficulty of the current question bank is matched with the student's ability. At this time, according to the portrait score of the student portrait, the difficulty distribution of the question bank can be appropriately adjusted.
[0158] For example, if the portrait score of the student portrait is greater than 80, i.e. the student is preparing for the upcoming physics competition, and the current cognitive load is at a moderate level, then some questions with similar difficulty to the competition can be added to the question bank, while a part of the basic questions are retained for consolidation, so as to maintain the student's learning motivation and challenge spirit.
[0159] If the cognitive load level is low, it means that the student can easily cope with the existing questions, at this time, according to the portrait score of the student portrait, the difficulty of the question bank can be appropriately increased.
[0160] For example, if the image score of the student image is greater than 80, that is, the student has a good foundation in both experimental operation and theoretical knowledge of chemistry, and the cognitive load level is low, some chemistry questions with strong integration and involving cutting-edge knowledge, such as research and application of new materials and analysis of complex chemical reaction mechanisms, can be added to stimulate the students' desire to explore and further improve their chemical literacy.
[0161] The embodiments of the present application can dynamically update the difficulty level of the question bank according to the cognitive load level and the student image, so that the question bank can always be dynamically adapted to the learning state and ability level of the student, provide more accurate and effective learning resources for the student, promote the personalized development of the student, and improve the learning effect.
[0162] The present application provides a cognitive load real-time intervention method based on multi-modal behavior fingerprint, by acquiring multi-modal micro-behavior data of the student to be tested when answering questions, the multi-modal micro-behavior data includes the frequency of draft revision, the frequency of formula jump and the frequency of writing pause; the multi-modal micro-behavior data is time-aligned and feature-fused, the cognitive load features corresponding to each mode are extracted, and all cognitive load features are generated into a behavior fingerprint vector; the behavior fingerprint vector and the macro index are input into a dynamic cognitive load evaluation model, and the cognitive load level of the student to be tested is output, the macro index includes the time spent on answering, the accuracy, the expression and the physiological signal; the difficulty level of the question bank of the student to be tested is dynamically updated according to the cognitive load level and the student image. The present application combines micro-behavior data such as draft revision frequency, formula jump frequency and writing pause frequency, and macro indicators such as answering time, accuracy, expression and physiological signal, wherein the micro-behavior data can capture the immediate reaction and detailed operation of the student in the problem-solving process, and the macro indicators provide a reference for the overall performance and physiological state, the combination of the two can more comprehensively reflect the cognitive load state of the student, avoiding the limitations of a single data source; in addition, by time-aligning the multi-modal micro-behavior data, the consistency of different modal data on the time axis is ensured, and then feature fusion is performed, which can preserve the time sequence relationship and associated information between data, extract more representative cognitive load features, and thus improve the accuracy of evaluation; at the same time, by dynamically adjusting the difficulty level of the question bank, the method can guide the student to balance between the comfort zone and the challenge zone, avoiding cognitive overload caused by excessive pressure and preventing cognitive slack caused by lack of challenge, and this balance helps to promote the balanced development of the student's ability and improve the learning effect.
[0163] It should be understood that the size of the serial number of each step in the above embodiments does not mean the order of execution, and the execution order of each process should be determined according to its function and inherent logic, and should not constitute any limitation on the implementation process of the embodiments of the present application.
[0164] The following is an apparatus embodiment of the present application. For details not described in detail, reference can be made to the corresponding method embodiments described above.
[0165] Figure 2 A structural schematic diagram of the cognitive load real-time intervention apparatus based on multi-modal behavior fingerprints provided by the embodiments of the present application is shown. For ease of illustration, only the parts related to the embodiments of the present application are shown, and the details are described as follows:
[0166] As shown in Figure 2 The cognitive load real-time intervention apparatus 4 based on multi-modal behavior fingerprints includes:
[0167] The behavior data acquisition module 21 is configured to acquire multi-modal micro-behavior data of the student to be tested when answering questions, and the multi-modal micro-behavior data includes draft revision frequency, formula skip frequency and writing pause frequency.
[0168] The fingerprint vector generation module 22 is configured to perform time sequence alignment and feature fusion on the multi-modal micro-behavior data, extract cognitive load features corresponding to each mode, and generate a behavior fingerprint vector from all the cognitive load features.
[0169] The grade determination module 23 is configured to input the behavior fingerprint vector and macro indicators into a dynamic cognitive load evaluation model, and output a cognitive load grade of the student to be tested, wherein the macro indicators include answering time, accuracy, expression and physiological signals.
[0170] The question bank updating module 24 is configured to dynamically update the difficulty level of the question bank of the student to be tested according to the cognitive load grade and the portrait of the student to be tested.
[0171] The application provides a cognitive load real-time intervention device based on a multi-modal behavior fingerprint, acquires multi-modal micro-behavior data of a student to be tested when the student is answering a question, the multi-modal micro-behavior data including a draft revision frequency, a formula jump frequency and a writing pause frequency; performs time sequence alignment and feature fusion on the multi-modal micro-behavior data, extracts cognitive load features corresponding to each mode, and generates a behavior fingerprint vector from all the cognitive load features; inputs the behavior fingerprint vector and macroscopic indexes into a dynamic cognitive load evaluation model, and outputs a cognitive load grade of the student to be tested, the macroscopic indexes including an answering time, an accuracy, an expression and a physiological signal; and dynamically updates a difficulty level of a question bank of the student to be tested according to the cognitive load grade and a portrait of the student to be tested. The application fuses micro-behavior data such as the draft revision frequency, the formula jump frequency and the writing pause frequency, and macroscopic indexes such as the answering time, the accuracy, the expression and the physiological signal, wherein the micro-behavior data can capture the instant reaction and the detail operation of the student in the problem solving process, and the macroscopic indexes provide a reference for the overall performance and the physiological state, the combination of the two can more comprehensively reflect the cognitive load state of the student, and avoid the limitation of a single data source; in addition, the time sequence alignment is performed on the multi-modal micro-behavior data, the consistency of the different modal data on the time axis is ensured, and then the feature fusion is performed, the time sequence relationship and the associated information between the data can be retained, more representative cognitive load features can be extracted, and the evaluation accuracy is improved; meanwhile, the difficulty level of the question bank is dynamically adjusted, the method can guide the student to balance between the comfort zone and the challenge zone, avoid cognitive overload caused by excessive pressure, and prevent cognitive slack caused by lack of challenge, and the balance is helpful to promote the balanced development of the ability of the student and improve the learning effect.
[0172] In a possible implementation, the fingerprint vector generation module can be configured to:
[0173] perform time sequence alignment on the multi-modal micro-behavior data by using a time warping dynamic alignment algorithm;
[0174] perform feature fusion on the time sequence aligned multi-modal micro-behavior data by using a multi-modal feature fusion technology;
[0175] extract features related to preset cognitive load of a corresponding mode from the feature fused multi-modal micro-behavior data by using a correlation analysis algorithm, and obtain cognitive load features of the corresponding mode.
[0176] In a possible implementation, the fingerprint vector generation module can be further configured to:
[0177] splice all the cognitive load features into a target vector according to a preset order;
[0178] input the target vector into a feature encoder to generate a fingerprint vector.
[0179] In a possible implementation, the fingerprint vector generation module can also be configured to:
[0180] The XGBoost is used to calculate the contribution degree of each cognitive load feature to the corresponding cognitive load, and the contribution degrees of each cognitive load feature are sorted in order of contribution degree, and the sorting is taken as the preset order;
[0181] The cognitive load features are spliced into the target vector according to the preset order.
[0182] In a possible implementation, the grade determination module can be configured to:
[0183] The behavior fingerprint vector and the macroscopic index are input into the dynamic cognitive load evaluation model, and the cognitive load grade of the student to be tested is output, and the dynamic cognitive load evaluation model is constructed by using a double-flow attention mechanism.
[0184] In a possible implementation, the behavior data acquisition module can be configured to:
[0185] The touch behavior sequence data of the student to be tested when answering the questions is captured through the interaction recording device, and the draft revision frequency of the student to be tested when answering the questions is determined according to the touch behavior sequence data.
[0186] The image sensor is used to obtain the image of the answer sheet of the student to be tested when answering the questions, and the answer steps written in the answer area of each question by the student to be tested are extracted from the image of the answer sheet, and for each question, the answer steps written in the answer area of the question are compared with the standard answer steps of the question to determine the formula skip frequency of the student to be tested when answering the question.
[0187] The handwriting board is used to obtain the note trajectory, writing pressure and writing speed of the student to be tested when answering the questions, and the writing pause frequency of the student to be tested when answering the questions is determined according to the note trajectory, writing pressure and writing speed of the student to be tested when answering the questions.
[0188] In a possible implementation, the behavior data acquisition module can also be configured to:
[0189] The image of the answer sheet is input into the text extraction model, and the answer steps written in each answer area of the image of the answer sheet are output, and the text extraction model is constructed based on a long short-term memory network.
[0190] In a possible implementation, the device can further include a preprocessing module, which can be configured to:
[0191] The multi-modal micro-behavior data is denoised;
[0192] The multi-modal micro-behavior data after denoising is normalized;
[0193] Correspondingly, the multi-modal micro-behavior data is time-aligned and feature-fused, cognitive load features are extracted, and the cognitive load features are generated into behavior fingerprint vectors, including:
[0194] The multi-modal micro-behavior data after normalization is time-aligned and feature-fused, cognitive load features are extracted, and the cognitive load features are generated into behavior fingerprint vectors.
[0195] In a possible implementation, the preprocessing module can be specifically used for:
[0196] A low-rank and sparse penalty decomposition method is used to impose a sparse constraint on the multi-modal micro-behavior data.
[0197] According to the local correlation inside the multi-modal micro-behavior data, the global structural information is mined by using the internal correlation of the multi-modal micro-behavior data, so as to remove the noise of the multi-modal micro-behavior data.
[0198] In the above embodiments, the description of each embodiment has its own emphasis, and the parts not described or recorded in a certain embodiment can be referred to the related description of other embodiments.
[0199] Those skilled in the art can realize that the templates, units and algorithm steps of each example described in combination with the embodiments disclosed herein can be realized by electronic hardware or a combination of computer software and electronic hardware. Whether the functions are realized in hardware or software depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of the present application.
[0200] The modules / units, if realized in the form of software function units and sold or used as independent products, can be stored in a computer readable storage medium. Based on such understanding, all or part of the processes in the above-mentioned embodiment methods can also be completed by a computer program instructing related hardware, and the computer program can be stored in a computer readable storage medium. When the processor executes the computer program, the steps of each multi-modal behavior fingerprint-based cognitive load real-time intervention method embodiment described above can be implemented. The computer program includes computer program code, which can be in the form of source code, object code, executable files or some intermediate forms, etc. The computer readable medium can include any entity or device, recording medium, U disk, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory, random access memory, electrical carrier signal, telecommunication signal and software distribution medium, etc. that can carry the computer program code.
[0201] The above-described embodiments are only used to illustrate the technical solutions of the present application, and are not intended to limit the present application; although the present application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that: it can still modify the technical solutions recorded in the foregoing embodiments, or make equivalent replacement for part of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application, and should be included in the protection scope of the present application.
Claims
1. A real-time cognitive load intervention method based on multimodal behavioral fingerprints, characterized in that: include: Acquiring multimodal micro-behavioral data of the students to be tested when answering questions, wherein the multimodal micro-behavioral data includes the frequency of draft corrections, the frequency of formula skipping, and the frequency of writing pauses; Performing time series alignment and feature fusion on the multimodal micro-behavioral data, extracting cognitive load features corresponding to each modality, and generating a behavioral fingerprint vector from all cognitive load features; Inputting the behavioral fingerprint vector and macro indicators into a dynamic cognitive load assessment model to output the cognitive load level of the student to be tested, wherein the macro indicators include answering time, accuracy, facial expression and physiological signals; The difficulty level of the question bank of the student to be tested is dynamically updated according to the cognitive load level and the portrait of the student to be tested.
2. The real-time cognitive load intervention method based on multimodal behavioral fingerprint according to claim 1 is characterized in that: The step of performing time series alignment and feature fusion on the multimodal micro-behavior data to extract cognitive load features corresponding to each modality includes: Using a time warp dynamic alignment algorithm to perform temporal alignment on the multimodal microscopic behavior data; Use multimodal feature fusion technology to perform feature fusion on multimodal micro-behavior data after time series alignment; Using the correlation analysis algorithm, features related to the preset cognitive load of the corresponding modality are extracted from the multimodal micro-behavioral data after feature fusion to obtain the cognitive load characteristics of the corresponding modality.
3. The real-time cognitive load intervention method based on multimodal behavioral fingerprint according to claim 1 is characterized in that: The method generates a behavioral fingerprint vector from all cognitive load features, including: All cognitive load features are concatenated into a target vector in a preset order; The target vector is input into a feature encoder to generate the fingerprint vector.
4. The real-time cognitive load intervention method based on multimodal behavioral fingerprint according to claim 3 is characterized in that: The step of concatenating all cognitive load features into a target vector in a preset order includes: Calculating the contribution of each cognitive load feature to the corresponding cognitive load using XGBoost, and sorting the contribution of each cognitive load feature in order of strength, and using the sorting as the preset order; The cognitive load features are concatenated into the target vector according to the preset order.
5. The real-time cognitive load intervention method based on multimodal behavioral fingerprint according to claim 1 is characterized in that: The step of inputting the behavioral fingerprint vector and the macro-indicator into a dynamic cognitive load assessment model and outputting the cognitive load level of the student to be tested includes: The behavioral fingerprint vector and the macro-indicator are input into a dynamic cognitive load assessment model to output the cognitive load level of the student to be tested. The dynamic cognitive load assessment model is constructed based on a long short-term memory network combined with a dual-stream attention mechanism.
6. The real-time cognitive load intervention method based on multimodal behavioral fingerprint according to claim 1 is characterized in that: The obtaining of multimodal micro-behavioral data of the students to be tested when answering questions includes: capturing touch behavior sequence data of the student to be tested when answering questions through an interactive recording device, and determining the frequency of the student to be tested correcting the draft when answering questions based on the touch behavior sequence data; Acquiring an image of the answer sheet of the student to be tested when answering questions through an image sensor, extracting the answer steps written by the student to be tested in the answer area of each question from the answer sheet image, and comparing the answer steps written in the answer area of each question with the standard answer steps of the question to determine the frequency of formula step skipping by the student to be tested when answering the question; The note track, writing pressure and writing speed of the student to be tested when answering the questions are obtained through a handwriting board, and the writing pause frequency of the student to be tested when answering the questions is determined based on the note track, writing pressure and writing speed of the student to be tested when answering the questions.
7. The real-time cognitive load intervention method based on multimodal behavioral fingerprint according to claim 6 is characterized in that: The step of extracting the answer written by the student to be tested in the answer area of each question from the answer sheet image includes: The answer sheet image is input into the text extraction model, and the answer steps written in each answer area in the answer sheet image are output. The text extraction model is constructed based on a long short-term memory network.
8. The real-time cognitive load intervention method based on multimodal behavioral fingerprint according to claim 1 is characterized in that: After obtaining the multimodal micro-behavioral data of the student to be tested when answering questions, the method further includes: performing noise reduction processing on the multimodal micro-behavior data; Normalize the multimodal micro-behavioral data after noise reduction; Accordingly, performing time series alignment and feature fusion on the multimodal micro-behavioral data, extracting cognitive load features, and generating a behavioral fingerprint vector from the cognitive load features includes: The normalized multimodal micro-behavioral data are subjected to time series alignment and feature fusion to extract cognitive load features, which are then used to generate behavioral fingerprint vectors.
9. The real-time cognitive load intervention method based on multimodal behavioral fingerprint according to claim 8 is characterized in that: The performing noise reduction processing on the multimodal micro-behavior data includes: A low-rank and sparse penalty decomposition method is used to impose sparsity constraints on the multimodal micro-behavioral data; According to the local correlation within the multimodal micro-behavior data, the inherent correlation relationship of the multimodal micro-behavior data is utilized to mine global structural information to achieve noise removal of the multimodal micro-behavior data.
10. A real-time cognitive load intervention device based on multimodal behavioral fingerprints, characterized in that: include: A behavior data acquisition module is used to acquire multimodal micro-behavioral data of the students when answering questions, wherein the multimodal micro-behavioral data includes the frequency of draft corrections, the frequency of formula skipping, and the frequency of writing pauses; A fingerprint vector generation module is used to perform time series alignment and feature fusion on the multimodal micro-behavioral data, extract cognitive load features corresponding to each modality, and generate a behavioral fingerprint vector from all cognitive load features; a level determination module, configured to input the behavioral fingerprint vector and macro indicators into a dynamic cognitive load assessment model and output the cognitive load level of the student to be tested, wherein the macro indicators include answering time, accuracy, facial expression, and physiological signals; The question bank updating module is used to dynamically update the difficulty level of the question bank of the student to be tested based on the cognitive load level and the portrait of the student to be tested.