A neuro-physiological behavior-driven multi-modal intelligent tutoring system and method

By using a multimodal signal fusion system, which utilizes EEG, cerebral blood oxygen metabolism, eye movement, and facial expression data, cognitive and emotional state characteristics are constructed, and learning tasks are dynamically adjusted. This solves the problems of lag and lack of emotional consideration in adaptive learning systems, and achieves precise teaching intervention and improved learning efficiency.

CN122176993APending Publication Date: 2026-06-09HUAZHONG NORMAL UNIV

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
HUAZHONG NORMAL UNIV
Filing Date
2026-02-06
Publication Date
2026-06-09

AI Technical Summary

Technical Problem

Existing adaptive learning systems suffer from lag in perceiving learners' states, ambiguous signals, a lack of systematic consideration of emotional and motivational dimensions, and a lack of hierarchical targeting in intervention strategies, resulting in an inability to achieve precise teaching interventions.

Method used

By employing a multimodal signal fusion system, and through parallel decoding of EEG, cerebral blood oxygen metabolism, eye movement, and facial expression data, cognitive and emotional state characteristics are constructed, and the difficulty and interaction of learning tasks are dynamically adjusted to achieve precise adaptive control.

Benefits of technology

It achieves precise pathological localization of the source of learning difficulties, prevents negative emotions from encroaching on cognitive resources, provides tiered and refined interventions, and improves learning efficiency and user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122176993A_ABST
    Figure CN122176993A_ABST
Patent Text Reader

Abstract

This invention discloses a neurophysiological behavior-driven multimodal intelligent learning guidance system and method. The system achieves closed-loop control by constructing a cognitive-emotional-task adaptive adjustment space: 1) In the cognitive dimension, it integrates EEG neural oscillations and frontoparietal network connections, fNIRS cerebral oxygen metabolism indicators, and eye-tracking data to construct a multimodal cognitive load assessment model; 2) In the emotional dimension, it combines the nonlinear characteristics of skin conductance signals with facial expression recognition to quantify learners' motivation levels and emotional valence in real time; 3) In the task dimension, it parameterizes task difficulty and interaction form as moderating variables. During system operation, the system locates the learner's position on the cognitive-emotional state plane in real time and dynamically adjusts task dimension parameters through a nonlinear mapping strategy. This invention effectively utilizes the decoupling ability of multimodal signals and the interaction of intention to achieve precise adaptive delivery of teaching content.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of intelligent education technology, and more specifically, relates to a multimodal intelligent learning guidance system and method driven by neurophysiological behavior. Background Technology

[0002] With the deep integration of artificial intelligence and educational technology, Intelligent Tutoring Systems (ITS) have gradually shifted from simply digitizing teaching content to personalized, adaptive learning centered on learners. How to accurately and in real-time perceive learners' learning status and provide appropriate instructional interventions is currently the core technological challenge in this field.

[0003] Existing adaptive learning technologies mainly suffer from the following technical bottlenecks: First, status assessment based on explicit behavioral data suffers from lag and attribution ambiguity. Traditional online learning systems primarily rely on behavioral data such as learners' clickstream, answer accuracy, and reaction time to infer their learning status. This outcome-based assessment model is inherently lagging; by the time the system detects a significant decline in performance, learners are often already in a state of severe cognitive overload or frustration, missing the optimal intervention window. Furthermore, single behavioral data points cannot effectively analyze the deeper mechanisms leading to learning failure: is the learner failing due to insufficient cognitive ability (e.g., working memory overflow), language decoding difficulties (e.g., syntactic complexity), or a lack of learning motivation (e.g., learned helplessness)? Existing "black box" systems lack the pathological diagnostic capabilities to address these issues.

[0004] Second, single-modal neurophysiological monitoring suffers from signal ambiguity and blind spots. To overcome the limitations of behavioral data, some existing technologies attempt to introduce physiological signal monitoring. However, single-modal approaches often have limitations: for example, while EEG alone has high temporal resolution, it is highly susceptible to eye artifacts in natural reading scenarios and struggles to accurately locate specific brain regions involved in language processing (such as syntax / semantics); eye tracking alone can only capture the gaze point location and cannot distinguish between "deep thinking" and ineffective "ignoring" gazes; while near-infrared spectroscopy (fNIRS) alone offers good spatial localization, it suffers from significant temporal delays. The lack of cross-validation with multimodal data makes it difficult for the system to meet the accuracy requirements of complex cognitive states in practical applications.

[0005] Third, there is a lack of systematic consideration of the "emotional and motivational" dimension. Most existing brain-computer interface-assisted education systems overemphasize the cold, hard measurement of "cognitive load" (i.e., whether the brain is working hard enough), while systematically neglecting the protection of the learner's "emotional and motivational" dimension (i.e., monitoring their willingness to learn). According to the Yerkes-Dodson Law and the Zone of Proximal Development (ZPD) theory, negative emotions such as anxiety, boredom, or inefficiency are often the primary factors hindering cognitive performance. The lack of real-time monitoring of "motivational boundaries" makes it easy for the system to force-feed highly challenging content even when the learner is emotionally overwhelmed, which is counterproductive.

[0006] Fourth, intervention strategies lack a tiered, targeted closed-loop approach. Existing technologies often employ a linear "difficulty increase / decrease" strategy in feedback control, lacking refined intervention methods based on specific cognitive bottlenecks. For example, when learning difficulties are detected, the system often cannot distinguish whether to provide visual guidance (for strategy deficiencies), syntactic scaffolding (for language barriers), or content breakdown (for memory overload). This crude feedback model struggles to maintain the learner's "flow experience" and fails to achieve dual protection of both learning motivation and learning ability.

[0007] In summary, there is an urgent need to develop a multimodal neurophysiological signal fusion system that can decouple cognitive abilities and emotional states in real time and perform hierarchical and precise adaptive control based on multidimensional states. Summary of the Invention

[0008] To address the aforementioned deficiencies or improvement needs of existing technologies, this invention provides a neurophysiological behavior-driven multimodal intelligent learning guidance system and method, which effectively utilizes the decoupling capability of multimodal signals and the interaction of intentions to achieve precise adaptive delivery of teaching content.

[0009] To achieve the above objectives, according to one aspect of the present invention, a neurophysiological behavior-driven multimodal intelligent learning guidance system is provided, comprising: The multimodal data acquisition module is used to collect multimodal data during the user's learning process. The multimodal data includes multiple types of data such as electroencephalogram (EEG) data, cerebral blood oxygen metabolism data, eye movement data, skin conductance data, and facial expression data. The multi-dimensional feature decoding and three-dimensional state space feature construction module is used to perform parallel decoding on the multimodal data to obtain multiple features, and divide the decoded features into cognitive state dimension features and emotional state dimension features. The cognitive state dimension features are used to represent the cognitive state of the user during the learning process, and the emotional state dimension features are used to represent the motivation level and emotional state of the user during the learning process. The adaptive control module is used to dynamically adjust the task difficulty parameters and interaction form parameters in the learning task based on the cognitive state dimension features and the emotional state dimension features. The teaching content presentation module is used to call up teaching content that meets the requirements based on the adjusted task difficulty parameters, and to present the called teaching content in an adjusted interactive format.

[0010] Furthermore, the cognitive state dimension features include: working memory load for characterizing the brain's information storage pressure during the learning process, semantic integration load for characterizing the brain's semantic integration pressure, syntactic metabolism ratio for characterizing the brain's syntactic decoding pressure, cognitive locking coefficient for characterizing whether the brain network is in an overloaded locked state, objective cognitive load comprehensive features for characterizing cognitive load pressure, and visual strategy state for characterizing whether it is in a high-frequency sampling mode or an inefficient gaze mode.

[0011] Furthermore, decoding to obtain working memory load includes the following steps: extracting the power spectral density of the 4-8 Hz band of the parietal lobe channel of the EEG data, calculating the standard score of the power spectral density, and using the standard score of the power spectral density as the working memory load; Decoding to obtain semantic integration load includes the following steps: extracting the power spectral density of the 13-30Hz band of the parietal and midline channels in the EEG data, calculating the standard score of the power spectral density, and using the standard score of the power spectral density as the semantic integration load; Decoding to obtain the syntactic metabolic ratio includes the following steps: calculating the ratio of the concentration changes of oxyhemoglobin to deoxyhemoglobin in the left inferior frontal gyrus, and using this ratio as the syntactic metabolic ratio; Decoding to obtain the cognitive lock-in coefficient includes the following steps: calculating the Pearson correlation coefficient between frontal and parietal EEG data as the cognitive lock-in coefficient; Decoding to obtain the comprehensive objective cognitive load features includes the following steps: calculating the comprehensive objective cognitive load features based on the number of fixations, blinks, and pupil size in the eye movement data. The calculation formula is as follows: ,in In order to objectively understand the comprehensive characteristics of the load, For the number of fixations, The number of blinks, The size of the pupil; Decoding to obtain the visual strategy state includes the following steps: K-Means clustering based on eye-tracking data to identify it as a "high-frequency sampling pattern" or an "inefficient gaze pattern".

[0012] Furthermore, the step of identifying K-Means clustering based on eye-tracking data as a "high-frequency sampling pattern" or an "inefficient gaze pattern" includes the following steps: Each time new eye movement data is collected, the eye movement data is standardized so that the horizontal axis represents the number of gazes and the vertical axis represents the comprehensive characteristics of objective cognitive load. The data points are dynamically clustered and divided into one of two categories: the first category is "high-frequency sampling pattern" data and the second category is "inefficient gaze pattern" data. The silhouette coefficient of the cluster is calculated after each clustering. Value, if S If the data points for n consecutive time windows all fall into the second category, it indicates that the visual strategy state has been switched to "inefficient gaze mode". The preset contour coefficient threshold is n, where n is the preset data point count threshold. ,and > If '' is detected, it means that the visual strategy state has switched to "inefficient gaze mode". 'For default settings' Threshold.

[0013] Furthermore, the emotional state dimension features include: skin conductance arousal level, which characterizes the degree of physiological activation, and negative emotion confidence level, which characterizes the probability of frustration or boredom.

[0014] Furthermore, the interaction form parameters include visual guide line parameters and syntactic highlighting parameters, while the task difficulty parameters include text decomposition parameters, annotation parameters, and motivational repair task parameters.

[0015] Furthermore, the step of dynamically adjusting the task difficulty parameters and interaction form parameters in the learning task based on the cognitive state dimension features and the emotional state dimension features includes the following steps: If the skin arousal level is continuously rising and reading performance declines beyond the user's zone of proximal development motivation boundary, or if the confidence level of negative emotions continuously exceeds the preset negative emotion confidence level threshold, then the motivation repair task parameters will be adjusted to repair the user's motivation. If the visual strategy state is detected to switch to "inefficient gaze mode", adjust the visual guide line parameters to generate a visual guide line. If the syntactic metabolism ratio is consistently lower than the preset syntactic metabolism ratio threshold, or the semantic integration load is consistently higher than the preset semantic integration load threshold, the syntactic highlighting parameters and annotation parameters will be adjusted to highlight and provide real-time annotations. If the working memory load exceeds the preset working memory load threshold or the cognitive lockout coefficient exceeds the preset cognitive lockout coefficient threshold, adjust the text decomposition parameters to reduce the difficulty of text comprehension.

[0016] Furthermore, the monitoring of continuously increasing skin arousal levels and declining reading performance beyond the user's zone of proximal development motivation boundary includes the following steps: We continuously collect reading performance data for each user, using the reading performance at each sampling point as the vertical axis and the corresponding skin conductance arousal level as the horizontal axis. The skin conductance arousal level when the reading performance changes from rising to falling, which is the baseline skin conductance arousal threshold for each user, is used as the skin conductance arousal level when the skin conductance arousal level increases. Determine the motivation resilience coefficient for each user; Each user's skin charge arousal threshold is calculated based on each user's motivation resilience coefficient and each user's baseline skin charge arousal threshold. When a user's real-time skin charge arousal exceeds the skin charge arousal threshold, it is determined that the user has broken through the motivation boundary of the user's zone of proximal development.

[0017] Furthermore, adjusting the visual guide line setting parameters to generate the visual guide line includes the following steps: The guide line's movement speed is set to ,in The speed at which the guide line moves. The current reading speed is captured in real time by the eye tracker. The pre-set traction coefficient is used; the color and / or transparency of the guide line dynamically change with the current eye movement state. When the fixation point is detected to be lagging behind the guide line, the line changes to a high-contrast color to enhance the cues. When the fixation point and the guide line are in good synchronization, it turns to a low-interference light color.

[0018] According to another aspect of the present invention, a multimodal intelligent teaching method driven by neurophysiological behavior is provided, comprising the steps of: Collect multimodal data during the user's learning process, including multiple types of data such as electroencephalogram (EEG) data, cerebral blood oxygen metabolism data, eye movement data, electrodermal conductance data, and facial expression data; The multimodal data is decoded in parallel to obtain multiple features, and the decoded features are divided into cognitive state dimension features and emotional state dimension features. The cognitive state dimension features are used to represent the cognitive state of the user during the learning process, and the emotional state dimension features are used to represent the motivation level and emotional state of the user during the learning process. The task difficulty parameters and interaction form parameters in the learning task are dynamically adjusted based on the cognitive state dimension features and the emotional state dimension features. The system calls up teaching content that meets the requirements based on the adjusted task difficulty parameters, and presents the called teaching content using the adjusted interactive format parameters.

[0019] Overall, the technical solutions conceived in this invention have beneficial effects compared with the prior art: (1) By effectively utilizing the interaction between the decoupling ability of multimodal signals and willingness, the system achieves precise adaptive delivery of teaching content. First, it achieves precise pathological localization of the source of learning difficulties. Through cross-validation of multimodal data, the system can clearly distinguish whether the learning disorder is due to cognitive overload, failure of information acquisition strategy, or lack of learning motivation. Second, it establishes a unique motivation protection mechanism, which effectively prevents negative emotions from encroaching on cognitive resources and avoids the occurrence of learned helplessness.

[0020] (2) In constructing the emotion and motivation dimension (A dimension), this invention introduces a dynamic motivation protection mechanism based on physiological benchmarks. The system combines the nonlinear characteristics of electrodermal signaling (EDA) with facial micro-expression recognition technology to construct an emotion vector that includes physiological arousal and emotional valence. In particular, this invention defines and quantifies the “ZPD (Zone of Proximal Development Motivation Boundary),” establishing an individualized psychological tolerance threshold through initial benchmark calibration of the system. The system can track the trajectory of the emotion vector in the state space in real time, identifying the nonlinear motivational collapse inflection point before the learner is about to experience learned helplessness or severe frustration, thereby elevating simple emotion recognition to predictive psychodynamic monitoring.

[0021] (3) In constructing the cognitive dimension (C dimension), this invention abandons the traditional single-indicator monitoring and proposes a deep perception scheme that integrates multi-source information. The system simultaneously collects and processes EEG, near-infrared brain functional imaging (fNIRS), and eye-tracking signals. By extracting the power spectral density of the Theta and Beta bands in the parietal lobe region of EEG, the working memory load and semantic integration load of learners are quantified respectively. Combined with fNIRS monitoring of the blood oxygen metabolism ratio in the left inferior frontal gyrus (Broca area), the bioenergy cost of syntactic processing is accurately located. More importantly, this invention creatively incorporates eye-tracking data into the cognitive evaluation system. Through cluster analysis, it identifies the learner's visual information acquisition strategy, thereby effectively distinguishing between "effective cognition through high-frequency sampling" and "strategic failure through inefficient gazing," achieving decoupling analysis of subjective cognitive input and objective brain processing state.

[0022] (4) This invention establishes a hierarchical adaptive decision-making logic in the task dimension (T dimension). This logic breaks the limitations of traditional linear feedback and adopts a blocking control strategy with strict priorities: the system prioritizes the highest priority motivation circuit breaker protection. When the A dimension is detected to reach the ZPD motivation boundary, the current high-load task is forcibly interrupted and switched to the comfort zone of low arousal to prioritize the repair of the learner's psychomotivation; secondly, behavioral strategy guidance is implemented, generating visual guidance lines for inefficient gaze patterns detected by eye-tracking; thirdly, semantic or syntactic scaffolding is provided, highlighting keywords or annotating terms for specific language processing obstacles detected by EEG or fNIRS; finally, cognitive unloading strategy is implemented, breaking down long and difficult sentences or adjusting speech rate for working memory overflow. Through hierarchical adaptive intervention, the system can provide "targeted therapy" teaching assistance according to specific physiological and psychological bottlenecks, ensuring that learners are always in the optimal learning range that matches their abilities and challenges, significantly improving the efficiency of online learning and user experience. Attached Figure Description

[0023] Figure 1 This is a schematic diagram of the principle of a neurophysiological behavior-driven multimodal intelligent learning system according to an embodiment of the present invention; Figure 2 This is a schematic diagram of the overall architecture of the neurophysiological behavior-driven multimodal intelligent tutoring system according to an embodiment of the present invention; Figure 3 This is a schematic diagram of the spatial complementary layout of the lightweight EEG and near-infrared sensor in an embodiment of the present invention; Figure 4 This is a schematic diagram of the visual strategy state based on eye-tracking clustering according to an embodiment of the present invention; Figure 5 This is a schematic diagram of the Zone of Proximal Development (ZPD) motor boundary detection model based on electrodermal signals according to an embodiment of the present invention; Figure 6 This is a flowchart of the four-level adaptive feedback control strategy for multimodal feature mapping according to an embodiment of the present invention. Detailed Implementation

[0024] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the invention. Furthermore, the technical features involved in the various embodiments of this invention described below can be combined with each other as long as they do not conflict with each other.

[0025] In the description of the embodiments of this application, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion, for example, a process, method, apparatus, product or device that includes a series of steps or modules is not necessarily limited to those steps or modules that are explicitly listed, but may include other steps or modules that are not explicitly listed or that are inherent to such process, method, product or device.

[0026] The naming or numbering of steps in the embodiments of the present invention does not mean that the steps in the method flow must be executed in the time / logical order indicated by the naming or numbering. The execution order of the named or numbered process steps can be changed according to the technical purpose to be achieved, as long as the same or similar technical effect can be achieved.

[0027] In this document, the term "embodiment" means that a particular feature, structure, or characteristic described in connection with an embodiment may be included in at least one embodiment of the invention. The appearance of this phrase in various places throughout the specification does not necessarily refer to the same embodiment, nor is it a separate or alternative embodiment mutually exclusive with other embodiments. It will be explicitly and implicitly understood by those skilled in the art that the embodiments described herein can be combined with other embodiments.

[0028] This invention provides a multimodal intelligent learning system and method driven by neurophysiological behavior, which will be described below.

[0029] An embodiment of the present invention provides a neurophysiological behavior-driven multimodal intelligent tutoring system, comprising: The multimodal data acquisition module is used to collect multimodal data during the user's learning process. The multimodal data includes multiple types of data such as electroencephalogram (EEG) data, cerebral blood oxygen metabolism data, eye movement data, skin conductance data, and facial expression data. The multi-dimensional feature decoding and three-dimensional state space feature construction module is used to perform parallel decoding on the multimodal data to obtain multiple features, and divide the decoded features into cognitive state dimension features and emotional state dimension features. The cognitive state dimension features are used to represent the cognitive state of the user during the learning process, and the emotional state dimension features are used to represent the motivation level and emotional state of the user during the learning process. The adaptive control module is used to dynamically adjust the task difficulty parameters and interaction form parameters in the learning task based on the cognitive state dimension features and the emotional state dimension features. The teaching content presentation module is used to call up teaching content that meets the requirements based on the adjusted task difficulty parameters, and to present the called teaching content in an adjusted interactive format.

[0030] The following is a detailed combination Figures 1-6 Detailed explanation.

[0031] I. System Hardware Architecture and Multimodal Sensor Layout To achieve non-intrusive monitoring of core brain networks during reading tasks, this embodiment constructs the system architecture shown in Figure 2 and adopts the "spatial complementary" sensor layout shown in Figure 3. 1. EEG Acquisition Module: A portable dry electrode system with a sampling rate of 500Hz is used. To accommodate multimodal acquisition, the electrode layout (specially configured) avoids the left inferior frontal gyrus region, focusing on covering the midline frontal region (Fz, FCz) in the international 10-20 system to monitor the executive control network, and the parietal region (Pz, P3, P4) to monitor the semantic integration network.

[0032] 2. Near-infrared (fNIRS) acquisition module: A dual-source, four-detector fNIRS optical probe (sampling rate 10Hz) is positioned in the left inferior frontal gyrus (BA44 / 45 region), where the EEG electrodes are intentionally left empty. This region corresponds to Broca's area and is used to specifically monitor changes in blood oxygen metabolism caused by syntactic processing during reading.

[0033] 3. Peripheral modal acquisition module: integrates a desktop infrared eye tracker (60Hz) for acquiring reading gaze trajectory; a skin conductance sensor worn on the non-dominant wrist for acquiring EDA signals; and a front-facing RGB camera for capturing facial micro-expressions.

[0034] 4. Timing synchronization module: To address the sampling rate differences between heterogeneous signals such as EEG and fNIRS, the system has a built-in data synchronization unit based on the Lab Streaming Layer (LSL) protocol to ensure that all modes are strictly aligned on the time axis of L2 reading.

[0035] II. Multidimensional Feature Decoding and CAT State Space Construction Based on the aforementioned hardware environment, the system performs parallel decoding on the acquired multimodal data to construct an orthogonal CAT three-dimensional state space describing the L2 reading state. To eliminate individual differences, all metrics are standardized using Z-score relative to the baseline. 1. C-dimensional (cognitive state) decoding: Working memory load ( The Theta band (4-8Hz) power spectral density of the EEG apical channels (Pz, P3, P4) is extracted, and its standard score is calculated. This standard score is used as the working memory load. This metric is used to assess the information retention pressure during reading.

[0036] Semantic integration load ( Power spectral density (PSD) in the Beta band (13-30 Hz) of the parietal and midline channels (Pz, CPz) of the EEG is extracted, and the standard score of the PSD is calculated. This ratio is used as the syntactic metabolism ratio. This metric is used to assess semantic integration pressure during reading.

[0037] Syntactic metabolism ratio ( ): Calculate the ratio of changes in oxyhemoglobin to deoxyhemoglobin concentrations in the left inferior frontal gyrus according to fNIRS ( ).when A significant decrease indicates that the brain is performing high-cost syntactic decoding.

[0038] Cognitive lock-in coefficient ( ): Calculate the Pearson correlation coefficient between EEG frontal lobe (Fz) data and parietal lobe (Pz) data as the cognitive lock-in coefficient, used to characterize the functional connectivity strength of brain networks. The specific determination logic is as follows: Set a cognitive lock-in threshold ( For example, in this embodiment, we take 0.85), when the calculated value is... When the value remains above this threshold, the brain network is considered to be in an "overload-locked" state. This indicates a high degree of over-coupling between the frontal lobe executive control network and the parietal semantic network, with neural resources already saturated by the current task, making it difficult to flexibly allocate resources to process new input information.

[0039] The Objective Cognitive Load Index (OCLI) is a composite index calculated from data collected by the eye tracker in the multimodal signal acquisition layer. Unlike the single 'number of fixations' on the horizontal axis, it is a comprehensive model integrating information sampling, fatigue level, and alertness. The system extracts fixation, blink, and pupil features from the eye-tracking data, standardizes them using Z-scores, and then calculates the index using the following linear combination model: This formula integrates three dimensions of physiological signals to quantify the true cognitive cost: (Positive gain): Characterizes the information sampling requirement. The more frequent the fixation, the higher the demand for visual information processing. (Positive gain): Characterizes the accumulation of visual fatigue. An increase in blinking frequency is usually positively correlated with the depletion of cognitive resources and feelings of fatigue. (Negative Modulation): Represents alertness resources. Pupil diameter is usually related to arousal and cognitive resource mobilization levels. In this model, it is used as a negative term to correct for the resource compensation effect caused by high alertness.

[0040] Visual strategy state ( As shown in Figure 4, K-Means clustering was performed based on eye-tracking data to identify "high-frequency sampling mode (Mode_A)" (effective reading) and "inefficient gaze mode (Mode_B)" (strategy failure).

[0041] 2. Decoding A-Dimensional (Emotional State): Skin conductance arousal level ( ): Calculate the standard score of skin conductance level (SCL) to characterize the degree of physiological activation.

[0042] Confidence level of negative sentiment ( This system identifies the probability of frustration or boredom based on facial action units (AUs). Specific AU features are extracted, including AU4 (frowning, associated with high stress), AU15 (drooping corners of the mouth, associated with disappointment), and AU43 (closed eyes, associated with boredom). The intensity values ​​and combination patterns of these AUs are input into a pre-trained emotion classifier (such as a support vector machine or convolutional neural network), outputting a probability value in the interval [0, 1]. When this value is close to 1, the system has a very high confidence level in indicating that the user is in a negative emotional state.

[0043] 3. T-dimensional (task control) parameters: The system has a pre-set intervention library for L2 reading, and parameterizes the intervention methods into the following two sub-dimensions: (1) Interaction form parameters (corresponding to auxiliary strategies): including visual guide parameters (Reading Guide, changing visual presentation) and syntactic highlight parameters (Highlight, changing attention cues); (2) Task difficulty parameters (corresponding to content adjustment): including text decomposition parameters (Chunking, reducing syntactic complexity), annotation parameters (providing real-time annotations) and motivation repair task parameters. The motivation repair task parameters are used to repair user motivation by adjusting the difficulty of teaching content, such as switching to a low-difficulty game mode.

[0044] III. Cascaded Blocking Adaptive Feedback Control Logic To achieve precise intervention, the system operates a four-level priority adaptive feedback control logic as shown in Figure 6. The system can be set to operate with a period of 1000ms, prioritizing (Level 1)... Level 4) Perform the following blocking judgment: Level 1 judgment: Engine circuit breaker protection Triggering condition: detected A sustained increase in reading performance coupled with a non-linear decline (exceeding the ZPD motivation boundary), or It consistently exceeds the preset confidence threshold for negative emotions.

[0045] Among them, monitored The process of continuously increasing reading performance while experiencing a non-linear decline (breaking through the ZPD motivation boundary) includes the following steps: Continuously collecting reading performance data for each user, using the reading performance at each sampling point as the ordinate and the corresponding skin conductance arousal level as the abscissa; using the skin conductance arousal level at which reading performance changes from increasing to decreasing as the baseline skin conductance arousal threshold for each user; determining the motivation resilience coefficient for each user; calculating the skin conductance arousal threshold for each user based on the motivation resilience coefficient and the baseline skin conductance arousal threshold; and determining that a user's real-time skin conductance arousal level exceeds the skin conductance arousal threshold, indicating a breakthrough to the user's ZPD motivation boundary.

[0046] The "ZPD motive boundary" in the triggering condition is not a fixed value, but is based on the motive resilience coefficient. The dynamic threshold for dynamic calibration. Motivation resilience coefficient ( This can be generated by mapping pre-task motivation questionnaire scores (such as the MSLQ scale). The system uses a formula... Individualized adjustments were made to the general physiological arousal threshold, among which... The baseline skin conductance arousal threshold, i.e. Figure 5 The skin conductance arousal level corresponding to the mid-inflection point For skin conductance arousal threshold: for learners with high achievement motivation ( The system automatically increases the circuit breaker threshold to relax the tolerance for physiological arousal; for learners with low self-efficacy ( The system lowers the threshold to trigger the protection mechanism in advance, thereby achieving the initial configuration of objective control parameters based on subjective psychological characteristics. Therefore, "benchmark calibration" refers to the individualized initialization of objective physiological control thresholds using subjective psychological scale data. When a user's real-time skin conductance arousal exceeds the user's skin conductance arousal threshold, it is determined that the user has crossed the motivational boundary of the user's zone of proximal development.

[0047] Corresponding strategy (Strategy D): Adjust the motivation repair task parameters to repair user motivation, such as reducing the difficulty of teaching content and activating the "Dynamic Difficulty Adjustment (DDA)" mechanism for blocking intervention.

[0048] (1) Task level labeling: Each task unit in the system's pre-set teaching task library is labeled with a difficulty level (Level 1~Level 5). For example, reading long and difficult sentences is Level 5, reading short sentences is Level 3, and recognizing words from pictures is Level 1.

[0049] (2) Stepped Degradation (Downward Logic): When the ZPD (Zero-Point Damping) boundary is detected to be exceeded, the system calculates the target degradation step size based on the extent of the skin conductance arousal. For severe overload (such as...), For moderate overload, a "circuit breaker" degradation is executed, switching directly from the current level to Level 1 (comfort zone) to ensure survival; for moderate overload, a "step-by-step degradation" is executed (e.g., Level 5). Level 3) is used to try to maintain a state of flow.

[0050] (3) Gradual Recovery (Upward Logic): When the system detects that physiological indicators (skin conductance and facial expression) have returned to baseline and maintained a stable time window (e.g., 30 seconds), the system will not immediately jump back to the original difficulty level, but will instead enter "Recovery Climbing Mode". That is, press Level 1. Level 2 ...The difficulty of the task is increased step by step in sequence, and the detection stops at each level until the user's current best ZPD area is repositioned.

[0051] Second-level judgment: behavioral strategy guidance Triggering condition: Detection of "inefficient prolonged fixation" in eye movement state. ).

[0052] Corresponding strategy (Strategy C): Generate dynamic visual guide lines between lines to physically guide eye movement. Specific implementation includes: (1) Shape and position: Render a semi-transparent gradient underline or a rectangular cursor that covers the text below the currently read text line.

[0053] (2) Adaptive speed control: the moving speed of the guide line It's not fixed, but set to .in The current reading speed is captured in real time by the eye tracker. This is the traction factor (e.g., 10%~15%). By maintaining a movement speed slightly faster than the user's, the visual tracing mechanism (Visual Pursuit) forces the gaze forward, preventing back-looking and stopping.

[0054] (3) Visual attribute feedback: The color / transparency of the guide line can change dynamically according to the state. For example, when the gaze point is detected to be lagging behind the guide line (not keeping up), the line changes to a high-contrast color (such as dark blue) to enhance the cue; when the gaze point is in good synchronization with the guide line, it turns to a low-interference light color (such as light green). Third-level judgment: Semantic / syntactic aids Triggering conditions: Any of the following language processing indicators exceeding the threshold is detected: (1) Abnormal fNIRS syntactic-metabolic ratio (low metabolic ratio): The oxygen / deoxyhemoglobin ratio in the left inferior frontal gyrus is detected ( ) consistently below a preset syntactic threshold (e.g. ), characterizing the exhaustion of syntactic processing resources; (2) EEG Beta wave anomaly (high semantic load): the standard score of the power in the parietal Beta band (13-30Hz) was detected (i.e. semantic integration load). ) consistently higher than a preset semantic threshold (e.g. This indicates that the semantic integration process exhibits high-intensity cognitive blockage.

[0055] Corresponding strategy (Strategy B): Highlight the subject-predicate structure of long and complex sentences or provide real-time annotations.

[0056] Level 4 judgment: Memory unloading Triggering condition: Detection of excessive EEG Theta wave ( ) or cognitive lock-in ).

[0057] Corresponding strategy (Strategy A): Adjust the text decomposition parameters to reduce the difficulty of text understanding, such as starting the NLP algorithm to decompose long and difficult sentences into short sentences (Chunking) and reduce the speech rate.

[0058] IV. Specific Operational Procedures in Second Language Reading Scenarios To verify the effectiveness of this system in a real-world environment, this embodiment of the invention recorded the entire process of a 25-minute English long and complex sentence reading task. Based on the real-time decoded CAT status, the system triggered the following continuous interventions according to the logic described above: 1. Phase One (Benchmark Calibration): Before the task begins (0-2 minutes), the system collects resting-state data and initializes individualized benchmark values. ).

[0059] 2. Phase Two (Visual Strategy Correction): At the 5th minute, the system detected that the learner exhibited single-point fixation >800ms. The system detects that a second-level judgment has been triggered, and strategy C is executed, generating a light blue guide line at a speed of 200 words / minute, successfully inducing the user to resume high-frequency sampling.

[0060] 3. Phase Three (Syntactic Barrier Breakthrough): At the 12th minute, the system detected fNIRS in the left inferior frontal gyrus. A value of 0.75 indicates that syntactic processing resources are exhausted. This triggers a third-level judgment, executing strategy B, which highlights the current complex clause in yellow.

[0061] 4. Phase Four (Cognitive Overload Relief): At the 18th minute, the system detected a surge in the parietal Theta wave on the EEG ( And network connection strength The reading reached 0.88. This triggered the fourth-level judgment, executing strategy A, which automatically breaks down the subsequent text into short sentences and slows down the playback speed.

[0062] 5. Phase Five (Motor Fuse Repair): At the 24th minute, the system detected skin conductance... The student's performance spiked to 200% of baseline and declined significantly, while their facial expression displayed high confidence and frustration. This indicated that the student had reached the ZPD (Zero-Probability Development) motivation boundary, triggering a Level 1 decision (highest priority). The system prioritized Strategy D, forcibly interrupting the challenging reading session and seamlessly switching to a "picture-based word recognition" game for 3 minutes to restore learning motivation.

[0063] V. Extended Explanation Those skilled in the art will understand that the specific implementations of the feature extraction algorithm, CAT state space construction method, and cascaded adaptive control logic in the above embodiments can exist in the form of computer software products, or can be embodied as a sequence of instructions stored in a computer-readable storage medium, or can be embedded in hardware chips such as DSP (Digital Signal Processor), FPGA (Field Programmable Gate Array), or ASIC (Application-Specific Integrated Circuit).

[0064] Furthermore, although this embodiment uses "second language (L2) reading" as an example, the application scenarios of this invention are not limited to this. The cognitive-emotional decoupling and hierarchical intervention ideas based on multimodal fusion proposed in this invention are also applicable to other complex learning scenarios with high cognitive load and the need to consider learning motivation, such as mathematical logic reasoning training, programming skills teaching, and vocational skills training. Any non-substantial modifications or equivalent substitutions made to specific signal modality combinations, feature calculation formulas, or parameter thresholds based on the core design ideas of this invention should be included within the protection scope of this invention.

[0065] The implementation of the key steps in this invention will be explained below.

[0066] Figure 2 shows a schematic diagram of the overall architecture of the multimodal neurophysiological signal fusion adaptive learning system according to an embodiment of the present invention. The system mainly consists of a four-layer structure: Acquisition layer: Includes spatially complementary EEG and fNIRS sensors, infrared eye trackers, skin conductance sensors, and facial expression acquisition cameras for synchronously acquiring raw physiological data.

[0067] Processing layer: Built-in multimodal fusion engine, containing four parallel decoding units: (1) EEG / fNIRS neural decoding unit, used to extract three-process oscillation features and syntactic metabolism ratio; (2) Eye movement behavior decoding unit, used to decouple subjective and objective load based on cluster analysis; (3) Electrodermal psychological decoding unit, used to detect ZPD motivation boundary; (4) Visual emotion decoding unit, used to identify emotional valence based on facial expression to distinguish emotional polarity under high arousal.

[0068] Control layer: Receives multi-dimensional state vectors from the processing layer, generates feedback instructions through an adaptive controller, and drives the user terminal (screen) to adjust the difficulty gradient and interaction format of the teaching content in real time.

[0069] Interaction layer: The user terminal (screen) presents the difficulty level and interactive format of the teaching content.

[0070] Figure 3 shows a schematic diagram of the spatially complementary layout of the lightweight EEG and near-infrared sensors in this invention. This is a top-down plan view of the scalp.

[0071] Solid dots represent EEG electrodes, which are concentrated in the midline frontal lobe (Fz, FCz), right frontal lobe (F4, F8), and parietal lobe (Pz, P3, P4) regions, and are used to capture high-frequency oscillating signals of the general execution control network and semantic integration network.

[0072] The dashed box shows the left inferior frontal gyrus region, which, instead of having EEG electrodes, is equipped with an fNIRS optical probe (represented by a square) for specific monitoring of the blood oxygen metabolic cost of syntactic processing.

[0073] This layout design achieves effective coverage of the core cognitive function network of the whole brain while simplifying the number of hardware channels.

[0074] Figure 4 illustrates the logic of decoupling subjective and objective workloads and identifying strategies based on eye-tracking clustering. The figure is a two-dimensional scatter plot.

[0075] The horizontal axis (number of gazes) represents only the frequency of the behavior, while the vertical axis (OCLI) represents the overall physiological cost. By comparing these two dimensions, the system can identify "inefficient gaze patterns" (i.e., although the number of gazes is high, it is combined with the OCLI characteristics of high blinking and low arousal, indicating ineffective fatigue-induced gaze), thereby achieving accurate strategy recognition.

[0076] The following three-tiered mechanism of "cleaning-verification-hysteresis" significantly improves the robustness of distinguishing between high-frequency sampling and inefficient gaze patterns. An outlier cleaning mechanism (pre-clustering cleaning) can be added to perform cleaning based on physiological thresholds before inputting the data into the clustering algorithm: short fixation removal; removal duration. Micro-fixation (often noise during saccades and does not represent actual cognitive processing). Pupil artifact removal: removing artifacts with a pupil diameter change rate exceeding [a certain threshold]. The data segment (standard deviation) is used to exclude measurement errors caused by eyelash occlusion or rapid blinking. The core function of this mechanism is to significantly improve the performance of the input to the OCLI model. , , The signal-to-noise ratio of the feature vectors. By filtering out physiological artifacts, it is ensured that the calculated OCL1 index can truly reflect the learner's cognitive load and fatigue state, preventing system misjudgments caused by measurement noise.

[0077] To prevent the K-Means algorithm from forcibly classifying data when the data distribution is not obvious, the system introduces a silhouette coefficient. The Silhouette Coefficient (S) is used for real-time validation. It is a core indicator for quantitatively evaluating the reasonableness of clustering results, simultaneously measuring intra-cluster compactness and inter-cluster separation. It is calculated after each clustering iteration. Value (range) ). Decision logic: If When (indicating sufficient separation between cluster boundaries and clear boundaries), dynamic clustering is used. When the condition "(indicating insufficient inter-cluster separation and blurred boundaries)" occurs, it means that the inter-cluster separation of the current data distribution is insufficient, and the dynamic clustering boundaries are unreliable. At this point, the system triggers a Fallback Rule: it stops using the current blurred dynamic clustering boundaries and automatically reverts to a 'preset static threshold based on large datasets' (e.g., directly using a preset threshold). (As a criterion for inefficient gaze). This mechanism ensures that even in mixed states where data distribution is not significant, the system can still maintain basic intervention functions based on conservative static criteria, preventing misclassification.

[0078] To avoid repeated jumps (oscillations) in system decisions near the boundaries of clusters A and B during dynamic clustering, a time window mechanism is introduced: only when consecutive... If all data points within a time window (e.g., 3 consecutive seconds) fall into cluster B (inefficient region) and the centroid distance exceeds a set threshold, then a state switch is confirmed and strategy C is activated.

[0079] The data points were divided into two significant clusters by a clustering algorithm (such as K-Means): Cluster A: Located in the upper right corner, representing "high-frequency sampling mode". Although users in this group may subjectively report difficulty, their high OCL1 value indicates that the objective input is effective, and they are judged to be in an "effective learning state".

[0080] Cluster B: Located in the lower left corner, it represents the "inefficient gaze pattern." This indicates that the learner is not effectively absorbing information and is judged as an "inefficient state."

[0081] The above thresholds can all be adjusted according to the actual situation.

[0082] Figure 5 illustrates a schematic diagram of the zone of proximal development (ZPD) motivational boundary detection model based on electrodermal signals. The curves in the figure show the nonlinear mapping relationship between task difficulty / physiological arousal (horizontal axis) and learning performance (vertical axis).

[0083] Learning performance refers to the normalized reading comprehension test score. In practice, the system inserts instant comprehension tests (such as fill-in-the-blank and multiple-choice questions) at reading task nodes (or paragraph intervals). The value on the vertical axis represents the accuracy rate or standard score of that test. (), is used to represent the learner's current level of understanding and mastery of the content.

[0084] Inflection point: When the difficulty / arousal exceeds a certain threshold (i.e., the motivation boundary), the curve reverses. The determination of the inflection point here is based on the logic of correlation inversion, i.e., the classic inverted U-shaped curve characteristic.

[0085] Linear growth zone (left side): An increase in electrical activity (EDA) is detected, while reading comprehension score (Performance) remains either increasing or stable. At this point, pressure is considered motivation, and the zone is considered effective.

[0086] Nonlinear inhibition region (right side): Electrodermal arousal (EDA) was monitored to continue to rise (even surge), but reading comprehension scores showed a significant decline.

[0087] Boundary Determination: The system detects this divergence where "physiological arousal continues to rise while cognitive scores decline in the opposite direction," marking it as the ZPD motivation boundary (i.e., the inflection point). This means that excessively high physiological arousal has begun to impair cognitive performance (such as the appearance of anxious thought block), and the circuit breaker protection must be triggered. Figure 6 shows the flowchart of the four-level adaptive feedback control strategy for multimodal feature mapping. The system executes the following decision logic according to priority: Level 1 Judgment (Motivation Circuit Breaker Protection): Based on whether the skin conductance detection exceeds the ZPD motivation boundary or based on the visual computing module detecting continuous high-intensity negative emotional characteristics (such as frustration or boredom). If so, execute Strategy D, forcibly switching to a low-arousal "comfort zone" task to restore motivation.

[0088] The second level of judgment (behavioral strategy guidance): Based on eye-tracking clustering detection, determine whether "inefficient long fixations" exist. If so, execute strategy C, using visual guidance lines on the screen to induce learners to adopt an efficient "high-frequency sampling" strategy.

[0089] Level 3 Judgment (Semantic / Syntactic Assistance): Based on anomalies in the parietal Beta or fNIRS ratio. If so, execute Strategy B, highlighting keywords or providing real-time annotations to lower the barrier to understanding.

[0090] Level 4 Decision (Memory Unloading): Based on parietal leaf Theta exceeding the limit. If so, execute strategy A, automatically breaking down long and complex sentences or reducing the speech playback speed to alleviate working memory load.

[0091] An embodiment of the present invention provides a multimodal intelligent learning guidance method driven by neurophysiological behavior, comprising the following steps: Collect multimodal data during the user's learning process, including multiple types of data such as electroencephalogram (EEG) data, cerebral blood oxygen metabolism data, eye movement data, electrodermal conductance data, and facial expression data; The multimodal data is decoded in parallel to obtain multiple features, and the decoded features are divided into cognitive state dimension features and emotional state dimension features. The cognitive state dimension features are used to represent the cognitive state of the user during the learning process, and the emotional state dimension features are used to represent the motivation level and emotional state of the user during the learning process. The task difficulty parameters and interaction form parameters in the learning task are dynamically adjusted based on the cognitive state dimension features and the emotional state dimension features. The system calls up teaching content that meets the requirements based on the adjusted task difficulty parameters, and presents the called teaching content using the adjusted interactive format parameters.

[0092] The working principle and technical effects of adaptive learning methods are the same as those of the adaptive learning systems described above, and will not be repeated here.

[0093] Those skilled in the art will readily understand that the above description is merely a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.

Claims

1. A neurophysiological behavior-driven multimodal intelligent learning guidance system, characterized in that, include: The multimodal data acquisition module is used to collect multimodal data during the user's learning process. The multimodal data includes multiple types of data such as electroencephalogram (EEG) data, cerebral blood oxygen metabolism data, eye movement data, skin conductance data, and facial expression data. The multi-dimensional feature decoding and three-dimensional state space feature construction module is used to perform parallel decoding on the multimodal data to obtain multiple features, and divide the decoded features into cognitive state dimension features and emotional state dimension features. The cognitive state dimension features are used to represent the cognitive state of the user during the learning process, and the emotional state dimension features are used to represent the motivation level and emotional state of the user during the learning process. The adaptive control module is used to dynamically adjust the task difficulty parameters and interaction form parameters in the learning task based on the cognitive state dimension features and the emotional state dimension features. The teaching content presentation module is used to call up teaching content that meets the requirements based on the adjusted task difficulty parameters, and to present the called teaching content according to the adjusted interaction format parameters.

2. The neurophysiological behavior-driven multimodal intelligent learning guidance system as described in claim 1, characterized in that, The cognitive state dimension features include: working memory load, which characterizes the brain's information storage pressure during the learning process; semantic integration load, which characterizes the brain's semantic integration pressure; syntactic metabolism ratio, which characterizes the brain's syntactic decoding pressure; cognitive locking coefficient, which characterizes whether the brain network is in an overloaded locked state; objective cognitive load comprehensive features, which characterize cognitive load pressure; and visual strategy state, which characterizes whether the brain is in a high-frequency sampling mode or an inefficient gaze mode.

3. The neurophysiological behavior-driven multimodal intelligent tutoring system as described in claim 2, characterized in that, Decoding to obtain working memory load includes the following steps: extracting the power spectral density of the 4-8 Hz band of the parietal lobe channel of EEG data, calculating the standard score of the power spectral density, and using the standard score of the power spectral density as the working memory load; Decoding to obtain semantic integration load includes the following steps: extracting the power spectral density of the 13-30Hz band of the parietal and midline channels in the EEG data, calculating the standard score of the power spectral density, and using the standard score of the power spectral density as the semantic integration load; Decoding to obtain the syntactic metabolic ratio includes the following steps: calculating the ratio of the concentration changes of oxyhemoglobin to deoxyhemoglobin in the left inferior frontal gyrus, and using this ratio as the syntactic metabolic ratio; Decoding to obtain the cognitive lock-in coefficient includes the following steps: calculating the Pearson correlation coefficient between frontal and parietal EEG data as the cognitive lock-in coefficient; Decoding to obtain the comprehensive objective cognitive load features includes the following steps: calculating the comprehensive objective cognitive load features based on the number of fixations, blinks, and pupil size in the eye movement data. The calculation formula is as follows: ,in In order to objectively understand the comprehensive characteristics of the load, For the number of fixations, The number of blinks, The size of the pupil; Decoding to obtain the visual strategy state includes the following steps: K-Means clustering based on eye-tracking data to identify it as a "high-frequency sampling pattern" or an "inefficient gaze pattern".

4. The neurophysiological behavior-driven multimodal intelligent tutoring system as described in claim 3, characterized in that, The step of identifying K-Means clustering based on eye-tracking data as either a "high-frequency sampling pattern" or an "inefficient gaze pattern" includes the following steps: Each time new eye movement data is collected, the eye movement data is standardized so that the horizontal axis represents the number of gazes and the vertical axis represents the comprehensive characteristics of objective cognitive load. The data points are dynamically clustered and divided into two categories: the first category is "high-frequency sampling pattern" data and the second category is "inefficient gaze pattern" data. The silhouette coefficient of the cluster is calculated after each clustering. Value, if S If the data points for n consecutive time windows all fall into the second category, it indicates that the visual strategy state has been switched to "inefficient gaze mode". The preset contour coefficient threshold is n, where n is the preset data point count threshold. ,and > '' indicates that the visual strategy state has been switched to "inefficient gaze mode". 'For default settings' Threshold.

5. The neurophysiological behavior-driven multimodal intelligent tutoring system as described in claim 1, characterized in that, The emotional state dimension features include: skin arousal level, which characterizes the degree of physiological activation, and negative emotion confidence level, which characterizes the probability of frustration or boredom.

6. The neurophysiological behavior-driven multimodal intelligent tutoring system as described in claim 1, characterized in that, Interaction parameters include visual guide line parameters and syntactic highlighting parameters, while task difficulty parameters include text decomposition parameters, annotation parameters, and motivational repair task parameters.

7. The neurophysiological behavior-driven multimodal intelligent tutoring system as described in claim 6, characterized in that, The step of dynamically adjusting the task difficulty parameters and interaction form parameters in the learning task based on the cognitive state dimension features and the emotional state dimension features includes the following steps: If the skin arousal level is continuously rising and reading performance declines beyond the user's zone of proximal development motivation boundary, or if the confidence level of negative emotions continuously exceeds the preset negative emotion confidence level threshold, then the motivation repair task parameters will be adjusted to repair the user's motivation. If the visual strategy state is detected to switch to "inefficient gaze mode", adjust the visual guide line parameters to generate a visual guide line. If the syntactic metabolism ratio is consistently lower than the preset syntactic metabolism ratio threshold, or the semantic integration load is consistently higher than the preset semantic integration load threshold, the syntactic highlighting parameters and annotation parameters will be adjusted to highlight and provide real-time annotations. If the working memory load exceeds the preset working memory load threshold or the cognitive lockout coefficient exceeds the preset cognitive lockout coefficient threshold, adjust the text decomposition parameters to reduce the difficulty of text comprehension.

8. The neurophysiological behavior-driven multimodal intelligent tutoring system as described in claim 7, characterized in that, The monitoring of continuously rising skin arousal levels and declining reading performance beyond the user's zone of proximal development motivation boundary includes the following steps: We continuously collect reading performance data for each user, using the reading performance at each sampling point as the vertical axis and the corresponding skin conductance arousal level as the horizontal axis. The skin conductance arousal level when the reading performance changes from rising to falling, which is the baseline skin conductance arousal threshold for each user, is used as the skin conductance arousal level when the skin conductance arousal level increases. Determine the motivation resilience coefficient for each user; Each user's skin charge arousal threshold is calculated based on each user's motivation resilience coefficient and each user's baseline skin charge arousal threshold. When a user's real-time skin charge arousal exceeds the skin charge arousal threshold, it is determined that the user has broken through the motivation boundary of the user's zone of proximal development.

9. The neurophysiological behavior-driven multimodal intelligent tutoring system as described in claim 7, characterized in that, The step of adjusting the visual guide line settings parameters to generate a visual guide line includes the following steps: The guide line's movement speed is set to ,in The speed at which the guide line moves. The current reading speed is captured in real time by the eye tracker. The pre-set traction coefficient is used; the color and / or transparency of the guide line dynamically change with the current eye movement state. When the fixation point is detected to be lagging behind the guide line, the line changes to a high-contrast color to enhance the cues. When the fixation point and the guide line are in good synchronization, it switches to a low-interference color.

10. A multimodal intelligent learning method driven by neurophysiological behavior, characterized in that, Including the following steps: Collect multimodal data during the user's learning process, including multiple types of data such as electroencephalogram (EEG) data, cerebral blood oxygen metabolism data, eye movement data, electrodermal conductance data, and facial expression data; The multimodal data is decoded in parallel to obtain multiple features, and the decoded features are divided into cognitive state dimension features and emotional state dimension features. The cognitive state dimension features are used to represent the cognitive state of the user during the learning process, and the emotional state dimension features are used to represent the motivation level and emotional state of the user during the learning process. The task difficulty parameters and interaction form parameters in the learning task are dynamically adjusted based on the cognitive state dimension features and the emotional state dimension features. The system calls up teaching content that meets the requirements based on the adjusted task difficulty parameters, and presents the called teaching content using the adjusted interactive format parameters.