A personalized learning process fusion multi-modal data assisted method and system

By collecting multimodal data to identify delay features and combining them with knowledge point association maps, contextualized guiding statements are generated, which solves the problem of not being able to accurately locate knowledge point extraction obstacles in existing technologies and improves the personalized problem-solving assistance effect of learning assistance systems.

CN121832780BActive Publication Date: 2026-06-26BEIJING AVIC FUTURE TECH GRP CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
BEIJING AVIC FUTURE TECH GRP CO LTD
Filing Date
2026-03-16
Publication Date
2026-06-26

AI Technical Summary

Technical Problem

Existing technologies cannot accurately pinpoint obstacles students encounter in retrieving knowledge points during problem-solving, resulting in learning assistance systems being unable to provide personalized and precise explanations of knowledge points, thus affecting problem-solving efficiency and knowledge point retrieval capabilities.

Method used

Multimodal behavioral data is collected in real time through eye tracking, keyboard interaction, and interface operation devices. Delay features are identified, and knowledge points that have been mastered but have retrieval difficulties are located by combining them with a pre-built knowledge point association graph. The knowledge points are then broken down into the smallest cognitive units by calling a standard knowledge base, generating contextualized guidance statements, and displaying them through augmented reality technology.

Benefits of technology

It enables precise intervention in students' knowledge retrieval obstacles during problem-solving, ensuring that guiding statements are accurately delivered, improving problem-solving efficiency and learning experience, and avoiding disruption to the problem-solving rhythm.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121832780B_ABST
    Figure CN121832780B_ABST
Patent Text Reader

Abstract

The present application relates to the technical field of personalized learning assistance, in particular, the present application relates to a kind of personalized learning process auxiliary method and system of fusion multi-modal data, the present application is through eye movement tracking, keyboard interaction record and interface operation monitoring equipment, real-time acquisition student's multi-modal behavior data when solving problem, according to pre-defined rule identification repeatedly looks and input pause feature, extract the difficulty of the delay feature of knowledge point;With the knowledge point label of current question and pre-constructed knowledge point association graph, the associated knowledge point that student has mastered but exists extraction obstacle is positioned;Call standard knowledge base and decompose the minimum cognitive unit of this knowledge point, fuse stem parameter and problem-solving context, reorganize as situational guiding sentence;In preset response time, based on eye movement trajectory prediction, guiding sentence is displayed in the stem area of student's gaze through augmented reality technology, synchronous trigger and short auditory prompt matching the type of obstacle, improve the pertinence and effectiveness of personalized learning assistance.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of personalized learning assistance technology, and more specifically, to a method and system for assisting the personalized learning process by integrating multimodal data. Background Technology

[0002] Personalized learning assistance technology is an important technology, specifically applied to the knowledge point retrieval assistance stage of students' problem-solving process. Its core is to improve learning efficiency by accurately identifying cognitive obstacles and providing targeted guidance, thus meeting students' core need for precise knowledge retrieval assistance in personalized problem-solving. During problem-solving, students may not be able to quickly recall knowledge points they have already mastered due to retrieval obstacles. These obstacles are presented through multimodal behavioral data such as eye movements, keyboard interactions, and interface operations. Because this multimodal data lacks systematic integration and targeted analysis, it is impossible to accurately locate the related knowledge points corresponding to the retrieval obstacles. Consequently, traditional learning assistance often provides generalized knowledge point explanations, which are difficult to tailor to the specific problem-solving scenario and achieve efficient guidance, affecting students' problem-solving pace and knowledge point retrieval ability. To solve this technical problem, we provide a personalized learning process assistance method and system that integrates multimodal data. Summary of the Invention

[0003] The purpose of this invention is to provide a method and system for assisting personalized learning processes by integrating multimodal data, so as to solve the problems mentioned in the background art.

[0004] To achieve the above objectives, one of the objectives of this invention is to provide a method for assisting a personalized learning process by fusing multimodal data, comprising the following steps:

[0005] S1. Collect students' multimodal behavior data in real time during the problem-solving process through eye-tracking devices, keyboard interaction recording devices, and interface operation monitoring devices. Based on predefined delay recognition rules, identify delay features that represent the difficulty in extracting knowledge points from the multimodal behavior data.

[0006] S2. Based on the delay features and the knowledge point tags corresponding to the current question, locate the related knowledge points that the student has mastered but has difficulty retrieving from the pre-constructed knowledge point association graph.

[0007] S3. Call the complete content of the related knowledge points in the standard knowledge base, break down the complete content into the smallest cognitive units that can be operated independently, and combine the specific parameters of the current question and the problem-solving context to reorganize the smallest cognitive units into contextualized guiding statements containing actual question stem data.

[0008] S4. Within a preset response time after recognizing the delayed feature, display the guiding statement in the question stem area currently being looked at by the user based on the real-time eye-tracking focus coordinates, and generate a short auditory cue that is bound to the keywords of the guiding statement.

[0009] The second objective of this invention is to provide a system for implementing a personalized learning process assistance method that integrates multimodal data as described in any one of the above-mentioned methods, comprising:

[0010] The data acquisition unit collects students' problem-solving behavior data in real time through eye-tracking, keyboard recording, and interface operation monitoring devices;

[0011] The delayed feature recognition unit identifies repeated viewing, input pauses, and cognitive stagnation behaviors based on predefined rules, and outputs delayed features that represent the difficulty in extracting knowledge points.

[0012] The related knowledge point location unit combines delay features with knowledge point association maps to filter basic knowledge points that have been mastered but have retrieval obstacles, and matches them with association paths for question comprehension or formula application.

[0013] The guidance statement generation unit calls the standard knowledge base to break down related knowledge points into the smallest cognitive units, combines them with question parameters to reorganize them into contextualized guidance statements, and binds them with keyword sound effects;

[0014] The multimodal guidance output unit dynamically renders guidance statements in the gaze area based on an eye-tracking prediction model, simultaneously triggers short auditory cues bound to keywords, and terminates display when the gaze moves away.

[0015] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0016] This invention collects multimodal behavioral data through eye tracking, keyboard interaction, and interface operation devices. Based on predefined rules, it accurately identifies delay features indicating difficulties in knowledge point retrieval, such as repeated viewing and input pauses. Combined with a pre-constructed knowledge point association map, it quickly locates related knowledge points that students have mastered but face retrieval obstacles, avoiding blind guidance. It calls upon a standard knowledge base to break down knowledge points into the smallest cognitive units, integrating question parameters and problem-solving context to reconstruct contextualized guidance statements. This overcomes the limitations of generic explanations. Utilizing eye-tracking prediction and augmented reality technology, the guidance statements are dynamically displayed in the student's gaze area, simultaneously triggering short auditory cues matching the type of obstacle. This ensures precise guidance without disrupting the problem-solving rhythm, achieving deep integration of multimodal data and precise intervention for retrieval obstacles. This allows learning assistance to be tailored to students' individual cognitive states, efficiently aiding knowledge point retrieval, improving problem-solving efficiency and learning experience, and helping students strengthen their knowledge application and retrieval abilities. Attached Figure Description

[0017] Figure 1 This is a flowchart illustrating the overall workflow of the present invention;

[0018] Figure 2 This is a schematic diagram of the overall structure of the present invention;

[0019] The meanings of the labels in the diagram are as follows:

[0020] 1. Data acquisition unit; 2. Delay feature recognition unit; 3. Related knowledge point location unit; 4. Guiding statement generation unit; 5. Multimodal guiding output unit. Detailed Implementation

[0021] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0022] Please see Figure 1 As shown, one of the objectives of this embodiment is to provide a method for assisting a personalized learning process by fusing multimodal data, including the following steps:

[0023] S1. Real-time collection of students' multimodal behavior data during problem-solving through eye-tracking devices, keyboard interaction recording devices, and interface operation monitoring devices; and identification of delayed features representing the difficulty in extracting knowledge points from the multimodal behavior data based on predefined delay recognition rules.

[0024] S2. Based on the delay characteristics and the knowledge point tags corresponding to the current question, locate the related knowledge points that the student has mastered but has difficulty retrieving from the pre-constructed knowledge point association map;

[0025] S3. Call the complete content of the related knowledge points in the standard knowledge base, break down the complete content into the smallest cognitive units that can be operated independently, and combine the specific parameters of the current question and the problem-solving context to reorganize the smallest cognitive units into contextualized guiding statements containing actual question data.

[0026] S4. Within a preset response time after recognizing the delayed feature, display the guiding statement in the question stem area currently being looked at by the user based on the real-time eye-tracking focus coordinates, and generate a short auditory cue that is bound to the keywords of the guiding statement.

[0027] The predefined delayed identification rules specifically include:

[0028] The system calculates the number of times the gaze point is repeatedly accessed in the same question stem block based on eye-tracking data. When the number of times the gaze point is repeatedly accessed is higher than the average number of times the current question is viewed, it is judged as repeated viewing behavior. The system counts the time interval between adjacent keystrokes based on keyboard interaction data. When the interval between consecutive mathematical symbol inputs is higher than the student's average input speed, it is judged as input pause behavior. The system monitors the time period without operation based on interface operation data. When the duration is higher than the threshold for coherent problem-solving operation, it is judged as cognitive stagnation behavior. Any of the above behaviors is output as a delayed feature representing the difficulty in extracting knowledge points.

[0029] The construction of the knowledge point association graph includes:

[0030] A tree-like hierarchical structure is constructed with the subject knowledge point system as the framework. The parent node represents a higher-order knowledge point, and the child node represents the basic knowledge point it depends on. Based on the cognitive psychology model, the extraction difficulty weight is marked for each knowledge point association path. The weight value is dynamically updated according to the average response delay when students call the knowledge point in historical data. A two-way error transmission link between knowledge points is established to record the statistical relationship of students' errors in solving problems with subsequent related knowledge points due to the failure to extract the preceding knowledge points.

[0031] The logic for determining which related knowledge points a student has mastered but has difficulty retrieving is as follows:

[0032] Extract the student's mastery score of the target knowledge point within a preset time period from the learning history database. If the score is higher than the mastery threshold, it is marked as mastered. Simultaneously analyze the real-time extraction difficulty weight of the knowledge point in the knowledge point association graph. When the weight is greater than the preset obstacle threshold and a delay feature is identified, it is determined that there is an extraction obstacle.

[0033] The process of locating related knowledge points specifically includes:

[0034] Based on the knowledge point tags of the current question, locate the root node in the knowledge point association graph, traverse the child nodes downwards along the graph, and filter out the basic knowledge points that simultaneously satisfy the state of being mastered and the state of having retrieval obstacles. Combine the identified delay feature type to match the association path. If the delay feature is repeated viewing behavior, prioritize locating knowledge points related to question understanding. If it is input pause behavior, locate knowledge points related to formula application.

[0035] Operations that can be broken down into the smallest independent cognitive units include:

[0036] The semantic structure of knowledge points in the standard knowledge base is analyzed, and each complete operation step is broken down into indivisible cognitive action units. The cognitive action units must meet the condition that they can be executed independently without additional explanation. Mathematical knowledge points are broken down according to the formula transformation steps, and textual knowledge points are broken down according to the logical reasoning nodes, and dependency relationship tags between units are generated.

[0037] Extract the specific parameter variables from the current problem, replace the abstract symbols in the obtained smallest cognitive unit with the actual problem stem data, determine the execution order of the smallest cognitive unit according to the problem-solving context, insert conjunctions to generate coherent operation instruction statements, delete descriptive content that is irrelevant to the current problem-solving stage, and retain core action verbs and key data.

[0038] An eye-tracking trajectory prediction model is established to predict the visual focus movement path for a preset duration based on historical eye-tracking data and the current gaze coordinates. Augmented reality overlay technology is used to dynamically render semi-transparent text boxes in the vicinity of the predicted path to display contextualized guiding statements. The display is immediately terminated when the eyeball leaves the question area.

[0039] The core action verbs in the generated contextualized guidance statements are extracted as keywords. Based on the determined knowledge point obstacle type, a pre-recorded sound effect fingerprint database is matched. Low-frequency prompts are played for conceptual understanding obstacles, and high-frequency prompts are played for operational execution obstacles. When the visual display is started, a prompt audio with a duration lower than the human reaction threshold is played synchronously.

[0040] Further explanation is needed: after collecting multimodal behavioral data of students during the problem-solving process through eye-tracking devices, keyboard interaction recording devices, and interface operation monitoring devices, in order to accurately capture cognitive signals representing difficulties in knowledge point retrieval from different dimensions, it is necessary to conduct targeted analysis of the three types of data based on predefined delay recognition rules. This allows for the inference of cognitive retrieval obstacles through behavioral performance. The specific implementation method is as follows:

[0041] The predefined delay identification rules specifically include the analysis logic for three types of data. Based on the laws of cognitive psychology and a large number of student problem-solving behavior samples, the predefined delay identification rules are a set of standardized judgment logics set for the three types of multimodal data. Each rule clearly defines the data processing method, judgment threshold and behavior definition to ensure that cognitive delay can be comprehensively identified from three dimensions: visual focus, input fluency and operation consistency. Delay features are abnormal behavioral signals exhibited by students because they cannot quickly extract the knowledge points required to solve problems. These signals indirectly reflect the difficulty of extraction at the cognitive level and are the core basis for subsequent location of related knowledge points. First, the number of times the gaze point is repeatedly accessed within the same question stem block is calculated based on eye-tracking data. The eye-tracking data consists of time-series data such as the student's gaze coordinates, gaze point position, and dwell time, recorded by the eye-tracking device at a sampling frequency of 100 Hz. The same question stem block is a local area of ​​the question automatically divided by the system according to semantic logic before solving the problem. Each block is clearly defined by page coordinates to ensure semantic consistency of the division. The number of times the gaze point is repeatedly accessed refers to the cumulative number of times the student's gaze leaves a question stem block and then returns to that block, reflecting the student's need for repeated confirmation of information in that block. The specific implementation process is as follows: First, all valid fixation points are screened from the eye-tracking data (excluding invalid data such as blinking and gaze drift). Fixation points are then categorized according to the boundary coordinates of the question's blocks, clarifying the block to which each fixation point belongs. Next, the fixation point sequence is traversed in chronological order, recording the trajectory of gaze switching between different blocks. When the gaze switches from block A to other blocks (such as blocks B or C) and then returns to block A, it is considered a repeated visit. The number of repeated visits for each question's block is accumulated. At the same time, the average number of times the current question is viewed is calculated (the average number of repeated visits for all question blocks). When the number of repeated visits for a certain question block is higher than this average, it is considered a repeated viewing behavior. This behavior usually occurs because students cannot extract relevant knowledge points to understand the question information and need to repeatedly review and confirm the core conditions. Therefore, it is an important manifestation of delayed features. Next, the keyboard interaction data was analyzed to determine the time interval between adjacent keystrokes. This data includes keystroke timestamps and key types (letters, numbers, mathematical symbols, etc.) collected by the keyboard recording device, with a focus on continuous input segments during the problem-solving process. The time interval between adjacent keystrokes refers to the time difference between two consecutive keystrokes, directly reflecting the fluency of input. The student's average input speed was calculated based on keyboard data from 30 previous problem-solving attempts, specifically the average interval between consecutive mathematical symbol inputs (mathematical symbol input is closely related to formula application and knowledge point retrieval, and is more susceptible to retrieval difficulties), thus avoiding misjudgments due to individual differences in input habits.The specific implementation process is as follows: Continuous input segments related to problem-solving are filtered from keyboard data (excluding ineffective inputs such as window switching and pauses in problem-solving). The input timestamps of continuous mathematical symbols (such as addition, subtraction, multiplication, division, parentheses, and equal signs) are extracted, and the time interval between adjacent symbols is calculated. The student's historical average input interval is used as a benchmark. When the current input interval of continuous mathematical symbols exceeds this benchmark, it is determined as an input pause. This pause often occurs when students encounter obstacles in deriving formulas or calculation steps, leading to an interruption in their thought process and subsequent input lag. Therefore, it is a typical characteristic of delay. Next, the interface operation data is monitored for periods of inactivity. Interface operation data consists of records of mouse clicks, scrolling, window switching, etc., collected by the interface monitoring device. Periods of inactivity are continuous time intervals during which no effective operations occur. The threshold for continuous problem-solving operations is a critical value calibrated based on a large number of continuous problem-solving behavior samples (verified and set at 10 seconds). This threshold excludes short pauses in normal thinking while capturing long periods of inactivity caused by cognitive lag. The specific implementation process is as follows: Extract the timestamps of all valid operations from the interface operation data, arrange them in chronological order, calculate the time difference between two adjacent operations, and obtain the duration of multiple inactive periods. When the duration of a certain inactive period exceeds the threshold for continuous problem-solving operations, it is judged as cognitive stagnation. This prolonged inactivity usually occurs because students fail to retrieve knowledge points, are unable to advance the problem-solving process, and fall into cognitive blockage, which is a direct behavioral manifestation of cognitive retrieval difficulties. Since repeated viewing, input pauses, and cognitive stagnation comprehensively cover the abnormal manifestations that may be caused by difficulties in knowledge point retrieval from the three dimensions of vision, input, and operation, and the three types of behaviors complement each other without omission, the occurrence of any of the above behaviors is used as a delayed feature output representing difficulties in knowledge point retrieval. This ensures accurate and comprehensive capture of cognitive retrieval obstacle signals, providing a reliable basis for subsequent problem root location based on knowledge point association graphs.

[0042] After identifying the delay features that indicate difficulty in extracting knowledge points, in order to accurately locate the core related knowledge points causing the delay, it is necessary to first pre-construct a structured knowledge point association graph. This graph needs to systematically sort out the dependencies between knowledge points, the differences in extraction difficulty, and the error propagation logic, providing a solid knowledge framework support for subsequently combining delay features to pinpoint extraction obstacles. The specific implementation method is as follows:

[0043] The construction of the knowledge point association graph revolves around three main dimensions: hierarchical relationship building, difficulty labeling extraction, and error propagation recording. Specifically, it involves first constructing a tree-like hierarchical structure based on the subject's knowledge point system. The subject's knowledge point system is a knowledge framework standardized by educational and teaching standards for the corresponding subject, covering a complete knowledge path from basic concepts to complex applications. For example, in mathematics, the knowledge points at various levels under branches such as numbers and algebra, and geometry. The tree-like hierarchical structure is a hierarchical data structure that mimics the branching growth of a tree, possessing clear hierarchical relationships and intuitively presenting the dependency logic between knowledge points. Higher-order knowledge points refer to complex knowledge points that require the integration of multiple basic knowledge points to understand or apply, such as the practical application of quadratic equations. Basic knowledge points are the prerequisites and foundational knowledge supporting higher-order knowledge points, such as the perfect square formula and factorization. These types of knowledge points are essential prerequisites for learning higher-order content. During the construction process, the subject knowledge system is first decomposed according to the logic of subject—knowledge module—higher-level knowledge point—basic knowledge point. The higher-level knowledge point is used as the parent node of the tree structure, and all the basic knowledge points it depends on are used as child nodes, forming a hierarchical relationship where the parent node governs the child nodes. For example, the comprehensive application of the Pythagorean theorem is used as the parent node, and its child nodes include basic knowledge points such as the definition of the Pythagorean theorem, the rules for determining the side length of a right triangle, and the rules for calculating the side length. For basic knowledge points with multiple dependencies (such as multiplication formulas that simultaneously support two higher-level knowledge points, factorization and algebraic simplification), a multi-parent node association mode is adopted to ensure that the basic knowledge point establishes an effective connection with all the higher-level knowledge points that depend on it. Finally, a tree-like hierarchical structure covering the entire subject knowledge network is formed, laying the foundation for subsequent location of association paths. Next, based on a cognitive psychology model, the extraction difficulty weight is labeled for each knowledge point association path. The cognitive psychology model is a quantitative model built with reference to memory retrieval theory and knowledge transfer rules. Its core is to combine factors such as the abstractness, association complexity, and application frequency of knowledge points to determine the cognitive difficulty for students to call up the knowledge points. The knowledge point association path refers to the one-way connection link from the parent node (higher-order knowledge point) to the child node (basic knowledge point). Each path corresponds to a set of higher-order knowledge point-basic knowledge point dependency relationships. The extraction difficulty weight is a coefficient that quantifies the ease with which students can quickly call up basic knowledge points under this path. The value ranges from 0 to 1. The larger the value, the higher the extraction difficulty and the more likely retrieval obstacles will occur.The specific annotation process is as follows: First, the cognitive characteristics of basic knowledge points in each associated path are analyzed using a cognitive psychology model. Abstract concepts (such as the definition of function monotonicity) are given higher initial weights than specific operational rules (such as the commutative law of addition), and multi-dimensional associated knowledge points (such as trigonometric function reduction formulas) are given higher initial weights than single associated knowledge points. Then, combined with a large amount of students' historical problem-solving data, the average response delay when the basic knowledge point is called is statistically analyzed (i.e., the time difference from the appearance of the question to the student's successful application of the knowledge point). The initial weights and the average response delay are normalized and integrated to obtain the final extraction difficulty weight. The weight value is not fixed but dynamically updated. Every 1,000 new calls to the knowledge point are accumulated, the average response delay is recalculated, and the extraction difficulty weight is adjusted accordingly to ensure that the weight can reflect the changes in the extraction difficulty of the knowledge point by the student group in real time, making the subsequent obstacle location more in line with the actual situation. Finally, a two-way error propagation link is established between knowledge points. This link records the correlation channels between knowledge points caused by extraction failures. It includes both the positive propagation from the failure to extract prerequisite basic knowledge points to subsequent errors in solving problems with higher-level knowledge points, and the reverse tracing from subsequent errors in solving problems with higher-level knowledge points to obstacles in extracting prerequisite basic knowledge points, thus achieving a two-way correlation of error causes. Prerequisite knowledge points refer to the basic knowledge points that are upstream in the knowledge dependency chain and provide support for subsequent knowledge points. For example, finding a common denominator for fractions is a prerequisite knowledge point for solving fractional equations. Statistical relationships refer to the frequency, probability, and other data of a certain error propagation path, which can quantify the correlation strength between the failure to extract prerequisite knowledge points and subsequent errors in solving problems with knowledge points. The specific process is as follows: First, a large amount of students' problem-solving error data is collected. Each data point includes the knowledge point tag corresponding to the incorrect question, the error type, and the root cause analysis of the error (determined through manual annotation or AI error analysis model). Then, cases of subsequent knowledge point problem-solving errors caused by the failure to extract the preceding knowledge points are selected from the data. For example, a student cannot extract the factorization knowledge point (preceding knowledge point), resulting in an error in solving a quadratic equation (subsequent knowledge point). For these cases, an error transmission link is established between the corresponding preceding and subsequent knowledge points, and the frequency of occurrence of the link is recorded. At the same time, a reverse tracing function is supported. When a problem-solving error is detected in a higher-order knowledge point, the preceding basic knowledge point that may be hindered by the link can be quickly located. As the error data continues to accumulate, the system will regularly update the statistical relationship of each link (such as frequency of occurrence and percentage), so that the link can accurately reflect the error transmission pattern between knowledge points, providing data support for subsequent location of extraction obstacles by combining delay features. Through a complete process of hierarchical structure construction, difficulty labeling extraction, and error propagation recording, the knowledge point association graph not only clearly presents the dependency logic between knowledge points, but also quantifies the differences in extraction difficulty and records the error propagation patterns. It can provide a comprehensive and reliable framework support for accurately locating related knowledge points that students have mastered but have extraction difficulties based on delay characteristics and question knowledge point tags.

[0044] After completing the construction of the knowledge point association graph and the identification of delayed features, the core task is to accurately distinguish between knowledge points that students have not mastered and knowledge points that they have mastered but have difficulty retrieving. The former requires supplementary basic learning, while the latter requires targeted guidance for retrieval. Therefore, it is necessary to identify the latter through clear judgment logic. The specific implementation method is as follows:

[0045] The logic for determining related knowledge points that students have mastered but face retrieval difficulties is based on a triple condition: mastery status verification, retrieval difficulty matching, and delayed feature triggering. Specifically, it first extracts the student's mastery score for the target knowledge point within a preset time period from the learning history database. The learning history database is a structured database storing the student's past learning data, including information such as practice completion status, incorrect answers, test scores, and review frequency for each knowledge point, providing data support for quantifying mastery. The target knowledge point is the basic knowledge point related to the current question in the knowledge point association graph, i.e., the related knowledge point to be determined. The preset time period is to ensure the mastery score is accurate. The time frame for the validity period is set (3 months based on teaching practice) to avoid the scoring failing to reflect the current level of mastery due to excessive time. The mastery score is a quantitative indicator calculated based on comprehensive learning history data, ranging from 0 to 100 points. A higher score indicates a more solid grasp. The calculation logic is as follows: practice accuracy accounts for 50%, test score accounts for 30%, and review frequency accounts for 20%, and the final score is obtained by weighted summation. The mastery threshold is the critical score (set at 70 points) that distinguishes between mastery and non-mastery. A score higher than this indicates that the student has the basic application ability of the knowledge point and does not need to relearn it; the only possible issue is at the retrieval level. The specific retrieval and judgment process is as follows: the system selects all learning data related to the student's knowledge point within the past 3 months from the learning history database based on the unique identifier of the target knowledge point, and calculates the mastery score according to the above weighting rules. If the score is higher than the mastery threshold of 70 points, the knowledge point is directly marked as mastered and proceeds to the subsequent retrieval obstacle judgment stage; if the score is lower than or equal to 70 points, it is judged as not mastered, eliminating the possibility of retrieval obstacles, and no further analysis is required. Next, the system synchronously analyzes the real-time extraction difficulty weight of the knowledge point in the knowledge point association graph. This real-time extraction difficulty weight is a core parameter of each association path in the graph, ranging from 0 to 1. It is dynamically updated based on the cognitive psychology model described earlier and the latest student data, directly reflecting the ease with which students can quickly extract the knowledge point. A higher value indicates a greater likelihood of stumbling or forgetting during extraction. Synchronous analysis means that while determining the student's mastery status, the system retrieves the corresponding weight data from the association graph using the knowledge point identifier, ensuring the efficiency of the judgment process without additional waiting. In practice, the system matches the target knowledge point's ID with nodes in the association graph to quickly obtain the corresponding real-time extraction difficulty weight. For example, the real-time extraction difficulty weight of the perfect square formula, a basic knowledge point, is 0.6, indicating that this knowledge point is of medium to high extraction difficulty for most students.Finally, when the weight is greater than the preset barrier threshold and a delay feature is identified, it is determined that there is a retrieval barrier. The preset barrier threshold is the critical value (calibrated to 0.5) that distinguishes between normal retrieval difficulty and excessively high retrieval difficulty. A weight greater than this value indicates that the knowledge point itself has a high retrieval difficulty, which is a potential cause of retrieval barrier. Retrieval barrier refers to the fact that students have the ability to apply knowledge points, but cannot quickly recall them when solving problems due to cognitive difficulties (such as temporary forgetting or poor association), which is manifested as delay features (repeated viewing, input pauses, cognitive stagnation, etc.). Delay features are the abnormal signals identified from multimodal behavioral data mentioned above, which directly reflect the cognitive blockage of students when solving problems. The specific judgment logic is as follows: Given that the target knowledge point has been marked as mastered, if its real-time extraction difficulty weight is higher than a preset obstacle threshold of 0.5, and the system has identified delay features during the current problem-solving process, then the triple conditions of mastery + high extraction difficulty + cognitive blockage are met, and the knowledge point is judged as a related knowledge point with extraction obstacles. If only the weight is higher than the threshold but there are no delay features, or there are delay features but the weight is lower than the threshold, it is not judged as an extraction obstacle. The former indicates that although the knowledge point is difficult to extract, the student successfully retrieved it this time; the latter indicates that the delay features may be caused by other factors (such as misunderstanding of the question stem or calculation errors), and are unrelated to the extraction of this knowledge point. Through this progressive judgment logic, interference from unmastered knowledge points is first eliminated, and then, combined with the extraction difficulty of the knowledge point itself and the student's real-time behavior, the core related knowledge points that are mastered but difficult to extract are accurately identified. This avoids misjudging unmastered knowledge points as extraction obstacles and also prevents the omission of real extraction problems, providing precise target guidance for the subsequent generation of targeted guidance statements.

[0046] After clarifying the logic for determining whether students have mastered the material but face retrieval difficulties, and identifying specific delay characteristics, in order to ensure that subsequent guiding statements accurately match cognitive pain points, it is necessary to specifically locate related knowledge points from the knowledge point association graph. The core is to focus on the current question requirements and, combined with the type of delay characteristics, pinpoint the root basic knowledge points. The specific implementation method is as follows: The process of locating related knowledge points follows the core logic of graph positioning—hierarchical traversal—conditional filtering—feature matching. Specifically, it includes first locating the root node in the knowledge point association graph based on the knowledge point label of the current question. The knowledge point label of the current question is the core high-order knowledge point identifier marked by the system according to the subject knowledge point system when the question is entered (such as the practical application of quadratic functions and the analysis of argumentative essays). It corresponds one-to-one with the parent node in the knowledge point association graph to ensure accurate positioning. The root node is the high-order knowledge point node in the graph that directly matches the label. It is the starting point for tracing down all dependent basic knowledge points. For example, if the current question label is "Comprehensive Application of the Pythagorean Theorem", then the parent node corresponding to "Comprehensive Application of the Pythagorean Theorem" in the graph is the root node. The specific location process is as follows: The system matches the unique identifier of the knowledge point tag in the question with the node ID in the knowledge point association graph to quickly locate the corresponding root node. At the same time, it loads all direct child nodes and associated path information of the root node to prepare for subsequent traversal. Then, it traverses the graph downwards along the child nodes. The child nodes are the basic knowledge point nodes that the root node (higher-order knowledge point) depends on. They are distributed hierarchically according to the knowledge dependency relationship. For example, the direct child nodes of the comprehensive application of the Pythagorean theorem include the definition of the Pythagorean theorem and the determination of right triangles, while the child nodes of the determination of right triangles include the determination of angle degree and the determination of side length relationship, forming a complete dependency chain. Downward traversal means starting from the root node and visiting each child node in the hierarchical order of direct child node to indirect child node to ensure that no basic knowledge points required by the higher-order knowledge point are missed. During the traversal, the knowledge point content and associated path information corresponding to each child node are recorded simultaneously to avoid repeated traversal or omission. The system then filters out basic knowledge points that simultaneously meet the criteria of being "mastered" and having "extraction obstacles." The filtering logic directly follows the judgment rules mentioned earlier. The system first retrieves the mastery scores of each sub-node (basic knowledge point) from the learning history database over the past three months. If the score is higher than the mastery threshold of 70, it is marked as "mastered." Next, the system extracts the real-time extraction difficulty weight of the sub-node from the knowledge point association graph. If the weight is greater than the preset obstacle threshold of 0.5, and a delay feature has been identified in the current problem-solving process, it is judged as having "extraction obstacles." Finally, the system filters out basic knowledge points that simultaneously meet the criteria of being "mastered" and having "extraction obstacles," excluding unmastered knowledge points (requiring supplementary learning) and mastered knowledge points without extraction obstacles (requiring no guidance), ensuring that the identified knowledge points are the core ones that truly require guidance for extraction.Finally, the identified delay feature types are combined to match the associated paths. The associated paths are hierarchical connection links from the root node to the selected basic knowledge points. Each path corresponds to the dependency relationship between higher-level knowledge points and basic knowledge points. The delay feature types are the repeated viewing behavior, input pause behavior, or cognitive stagnation behavior identified above. Different types correspond to different cognitive pain points and need to be matched with targeted associated knowledge points. The specific matching rules are as follows: If the delay feature is repeated viewing behavior, it means that the student repeatedly reviews the question stem section. The core pain point is the inability to extract relevant knowledge points to understand the question stem information. Therefore, priority is given to identifying knowledge points related to question stem comprehension. These knowledge points are the basic knowledge points that help interpret the question stem conditions and clarify the problem-solving objectives, such as keyword definition, conditional logical relationship analysis, and interpretation of data parameters. Their association path directly corresponds to the cognitive stage of question stem comprehension. If the delay feature is input pause behavior, it means that the student is stuck when inputting mathematical symbols and derivation steps. The core pain point is the obstacle to formula application or extraction of operation rules. Therefore, knowledge points related to formula application are identified. These knowledge points include formula transformation rules, operation steps, standard symbol application, etc. Their association path directly corresponds to the operational stage of problem-solving derivation. If the delay feature is cognitive stagnation behavior (excessive time without operation), then, combining the identification results of the first two types of knowledge points, priority is given to selecting the related knowledge points with the highest extraction difficulty weight to ensure that the guidance can quickly overcome cognitive stagnation. Through this series of targeted positioning operations, the basic knowledge points directly related to the current question are locked in by relying on the hierarchical dependency relationship of the knowledge point association graph. At the same time, the core of the extraction obstacle is accurately focused by delayed feature type matching, avoiding blind guidance. This provides a precise target for subsequent knowledge point decomposition and generation of contextualized guidance statements, ensuring that the guidance can directly address the students' cognitive pain points and efficiently assist in knowledge point extraction.

[0047] After identifying the relevant knowledge points that students have mastered but face retrieval difficulties, to ensure that subsequent guiding statements accurately adapt to the problem-solving scenario and directly address cognitive pain points, it is necessary to first break down the complete content of the relevant knowledge points into the smallest recombinable cognitive units. These units must be independently executable, fitting the current question parameters while restoring the core operational logic of the knowledge points. The specific implementation method is as follows: The operation of breaking down the knowledge points into the smallest independently operable cognitive units follows a core process of semantic parsing, classification and decomposition, and dependency annotation. Specifically, this includes first parsing the semantic structure of knowledge points in the standard knowledge base. The standard knowledge base is a structured database that stores the content of standardized knowledge points in various disciplines, containing authoritative information such as the definition, operation steps, and application scenarios of knowledge points, ensuring the accuracy of the decomposition basis. The semantic structure is the internal logical composition of the knowledge point content, including elements such as core concepts, operation actions, logical relationships, and sequence. For example, the known conditions of a mathematical formula—formula application—result derivation logic, and the concept definition—logical reasoning—conclusion chain in text analysis. The smallest cognitive unit is a basic cognitive module that cannot be further divided after decomposition. Each unit corresponds to a specific and single cognitive action and is the smallest unit that constitutes a complete knowledge point. During the parsing process, the system first calls the complete content of the related knowledge point in the standard knowledge base. It then uses a natural language processing model to semantically segment the text, identifying core actions, key data, logical connectors, and the order of operations. For example, the complete content of the knowledge point on solving fractional equations, after semantic parsing, can extract core actions such as removing denominators, removing parentheses, moving terms, combining like terms, and reducing the coefficient to 1, as well as the sequential logic between these actions. Subsequently, each complete operation step is broken down into indivisible cognitive action units. A cognitive action unit is the smallest operational instruction that can be executed independently without additional explanation; that is, students can complete the corresponding action simply by understanding the description of the unit, without needing to consult other information. Indivisible means that the unit cannot be broken down into smaller sub-actions that still have independent operational meaning. For example, removing denominators cannot be broken down into two units: finding the denominator and eliminating the denominator, because the two must be executed together to have complete operational meaning. However, removing denominators (multiplying both sides by the least common denominator) meets the requirement of independent execution.Differentiated decomposition logic is adopted for different types of knowledge points: Mathematical knowledge points are decomposed according to the formula transformation steps. Mathematical knowledge points are centered on formula application and calculation derivation. The decomposition must be close to the specific calculation process. For example, the knowledge point of transforming the vertex form of a quadratic function is decomposed into cognitive action units such as extracting the coefficient of the quadratic term, completing the square (adding and subtracting the square of half of the coefficient of the linear term), arranging it into a perfect square form, and simplifying the constant term. Each unit corresponds to a specific step in formula transformation to ensure the continuity of operation. Textual knowledge points are decomposed according to logical reasoning nodes. Textual knowledge points are centered on concept understanding and logical analysis. The decomposition must be closely linked to the reasoning chain. For example, the knowledge point of determining the thesis of an argumentative essay is decomposed into cognitive action units such as identifying the core topic of the question, selecting key sentences that reflect the author's point of view, eliminating evidence, background and other auxiliary information, integrating the viewpoint and forming the thesis. Each unit corresponds to a key node in logical reasoning to ensure the progression of thinking. After decomposition, dependency tags between units need to be generated. These tags are identifiers that mark the execution order of each cognitive action unit, clarifying the sequential, causal, or parallel relationships between units to ensure logical coherence when subsequently reorganizing guiding statements. Tag types include prerequisite dependencies (unit A must be executed before unit B can be executed), parallel dependencies (units A and B can be executed synchronously or without a strict order), and causal dependencies (execution of unit A is a prerequisite for the result of unit B). The specific generation process is as follows:

[0048] Based on the semantic structure and operational logic of the knowledge points, the relationships between each cognitive action unit and other units are analyzed. For example, in mathematical knowledge points, the prerequisite dependency label for the "transferring terms" unit is "removing parentheses," and in textual knowledge points, the prerequisite dependency label for integrating viewpoints to form arguments is "selecting key sentences that reflect the author's viewpoint." These dependencies are then linked to the corresponding cognitive action units in the form of labels, ultimately forming a structured decomposition result of cognitive action units + dependency label. This provides a flexible and orderly foundational module for subsequently reorganizing contextualized guiding statements based on question parameters. This targeted decomposition method retains the core operational logic of the knowledge points while breaking down complex content into simple, easy-to-understand, and independently executable minimum units. This facilitates the replacement of abstract symbols based on specific question parameters and allows for adjustment of the unit execution order according to the problem-solving context, ensuring that the guiding statements are both consistent with the essence of the knowledge points and adaptable to the current problem-solving scenario.

[0049] After breaking down related knowledge points into the smallest independent cognitive units, in order to ensure that the guiding statements are completely free from abstraction and fit the current problem-solving scenario, these units need to be reorganized and optimized in combination with the actual information of the question and the progress of problem-solving. The core is to replace abstract symbols with real data from the question stem, sort the units according to the problem-solving logic, and integrate the content with coherent expressions to ensure that the guidance can directly assist students in solving the problem. The specific implementation method is as follows: The reorganization process follows a progressive logic of parameter extraction, symbol replacement, sequence calibration, statement integration, and redundancy reduction. First, the specific parameter variables in the current question are extracted. Specific parameter variables refer to the actual data, limiting conditions, core objects, and other information that can be directly used to solve the problem, which are clearly given in the question stem. In mathematical questions, this includes specific values, known conditions, and variable names. In textual questions, this includes key concepts, logical relationships, and limited ranges. For example, in the mathematical question "If the two legs of a right triangle are 3 and 4, find the length of the hypotenuse", the two legs of the right triangle are 3 and 4, and the textual question "Analyze the role of the scene of the father seeing his son off in the movie 'The Back View' in expressing the theme", the theme expressed by the father seeing his son off is a specific parameter variable. During the extraction process, the system uses a natural language processing model to parse the current question stem text, identify and filter all parameters with practical significance, and classify and label them according to numerical, conditional, and object types. Simultaneously, it records the position and relationships of each parameter in the question stem, ensuring comprehensive extraction and clear classification, providing accurate material for subsequent replacement of abstract symbols. Next, the abstract symbols in the obtained smallest cognitive units are replaced with actual question stem data. Abstract symbols are general placeholders retained when breaking down the smallest cognitive units. In mathematical units, they are mostly generalized expressions of general variables and operators in formulas; in textual units, they are mostly general references to concepts. For example, in mathematical units, variables a and b are substituted into the formula ab; in textual units, X is used to analyze logical relationships based on the core concept X—both are abstract symbols. The actual question stem data refers to the specific parameter variables extracted earlier. The core principle of replacement is one-to-one correspondence and semantic fit, ensuring that the replaced unit still meets the requirement of being independently executable and requiring no additional explanation. In practice, the system first identifies the abstract symbols and their meanings in each smallest cognitive unit, and then matches semantically consistent actual data from the question parameters for replacement. For example, in math units, if the parameters in the question are the lengths of the right-angled legs 3 and 4, the system replaces them with the sum of squares of 3 and 4. In text units, the system finds the key plot points in the text that embody the core concept Y. If the core concept in the question is the father's love, the system replaces it with finding the key plot points in "The Back View" that embody the father's love. After replacement, the semantic integrity of the unit needs to be verified to ensure that students can directly understand the operational requirements.Then, the execution order of the smallest cognitive units is determined based on the problem-solving context. The problem-solving context refers to the student's current problem-solving progress, the completed steps, the specific steps where they are stuck, and the problem-solving logic itself (such as the calculation process of a math problem or the analysis approach of a text problem). For example, when solving the right triangle problem mentioned above, the student has completed the step of determining the triangle type but is stuck on the step of applying the formula. The problem-solving context is that the right triangle determination has been completed, and the hypotenuse needs to be found. The execution order refers to the sequential execution logic of each smallest cognitive unit, which must simultaneously meet two conditions: first, the dependency label of the unit itself (such as determining the right-angled side value before calculating the sum of squares); and second, the actual needs of the problem-solving context (such as prioritizing the execution of formula-related units when the student is stuck on formula application). The specific determination process is as follows: The system first retrieves the dependency tags of each unit to sort out the basic order of preceding units to subsequent units; then, combined with the student's current problem-solving behavior data (such as eye movement focusing on the formula area, keyboard input remaining on the calculation steps), it identifies the stuck link and adjusts the basic order, prioritizing units directly related to the stuck link and secondarily arranging auxiliary units; for example, when solving a right triangle, the adjusted order is to determine the values ​​of the right-angled sides 3 and 4, to perform the sum of squares on 3 and 4, to calculate the arithmetic square root of the sum of squares, and to obtain the length of the hypotenuse, ensuring that the order fits the problem-solving process and can directly guide students to overcome their obstacles. Next, connectors are inserted to generate coherent operation instruction statements. Connectors are words used to connect the smallest cognitive units and make the statements logically fluent. In mathematical operations, connectors such as first, then, then, and finally are often used to indicate order, while in text analysis, connectors such as first, then, combined, and then are often used to indicate logic. The selection of connectors needs to be appropriate for the relationship between units (sequential relationship, causal relationship, progressive relationship). In the specific integration process, the replaced cognitive units are arranged sequentially according to the adjusted execution order. Corresponding connecting words are inserted between the units to string together independent operation instructions into complete sentences. For example, after stringing together math units, the first step is to determine the values ​​of the right-angled sides as 3 and 4. Then, the sum of squares of 3 and 4 is calculated. Finally, the arithmetic square root of the sum of squares is calculated to obtain the length of the hypotenuse. After stringing together text units, the first step is to find the specific scene in "The Back View" where the father sees his son off to the car. Then, the father's behavioral characteristics are analyzed in combination with the details of the scene. Finally, the role of the scene in expressing the theme of "father's love" is deduced, ensuring that the sentences are coherent and natural, without any awkward splicing.Finally, descriptive content irrelevant to the current problem-solving stage is deleted, retaining only core action verbs and key data. Descriptive content irrelevant to the current problem-solving stage refers to statements explaining the background and principles of knowledge points within the smallest cognitive unit. This content is retained during the breakdown to ensure unit integrity but does not need to be presented when reorganizing the guiding statements. Examples include explanatory content such as the principle of the sum of squares (a² + b²) in a math unit or the core idea of ​​an article in a text unit. Core action verbs refer to verbs in the unit that represent specific operations, such as "determine," "calculate," "find," "analyze," and "derive," which are crucial for guiding students to perform the operations. Key data refers to the actual parameters in the replaced question stem, representing the specific objects of the operation. During the deletion process, the system screens the concatenated statements sentence by sentence, eliminating all explanatory and background statements, retaining only the core structure of core action verbs + key data. Simultaneously, the system verifies the conciseness and operability of the statements to ensure that the final generated contextualized guiding statements are free of redundant information and clearly convey the operational requirements, allowing students to immediately understand what to do and what to use. Through this series of reorganizations and optimizations, the smallest cognitive unit has been transformed from an abstract, general operational module into personalized guiding statements that fit the current question and the problem-solving progress. This retains the core operational logic of the knowledge points while completely breaking free from the constraints of abstract symbols. It can directly and accurately assist students in extracting related knowledge points and advancing the problem-solving process, avoiding the guidance from becoming a mere formality.

[0050] After reorganizing the contextualized guidance statements, to ensure the guidance accurately matches the user's current visual focus without interfering with normal problem-solving, it's necessary to lock the display position using eye-tracking prediction and achieve unobstructed presentation using augmented reality technology. A dynamic termination mechanism is also implemented to avoid redundant interference. The specific implementation method is as follows: After generating contextualized guidance statements containing actual question data, the core is to ensure the guidance accurately adapts to the user's eye movement rhythm. Therefore, an eye-tracking prediction model is first established. This model is based on time-series data patterns and is specifically designed to predict the user's eye movement path in the near future, ensuring the guidance statements are pre-positioned in the area the eye is about to reach. Historical eye-tracking data represents the user's past problem-solving experiences. The eye-tracking time-series data accumulated during the process includes fixation point sequences, movement speeds, dwell times, and region switching frequencies under different question types. It also integrates common eye-tracking patterns of students working on similar questions (such as prioritizing data areas in math questions and focusing on keyword areas in text questions), providing a dual reference for the model. The current fixation point coordinates are the two-dimensional pixel coordinates of the user's current gaze point, collected in real time by the eye-tracking device, serving as the starting point for prediction. The future preset duration is a reasonable prediction window (set to 0.5 seconds) calibrated through experiments. Too short a duration will lead to display lag, while too long a duration is prone to prediction deviations due to sudden changes in gaze. The visual focus movement path is a continuous coordinate sequence that the gaze may pass through within the next 0.5 seconds, output by the model, intuitively reflecting the direction and range of gaze movement. The model building process is as follows: First, eye-tracking data from the past three months of user learning history is extracted and categorized by question type (mathematics, Chinese, physics, etc.). Valid trajectory segments are then selected (excluding invalid data such as blinking and device interference). Next, the common eye-tracking trajectory data of a large number of students on similar questions is integrated to extract the gaze movement patterns at different problem-solving stages (question reading, step derivation, and answer input) as the model's prior knowledge. Subsequently, a temporal prediction algorithm is used to train the model, allowing it to learn the mapping relationship between the current gaze point, historical movement patterns, question type, and future trajectory. During training, parameters are continuously optimized using a validation set to ensure that the deviation between the predicted path and the actual gaze is controlled within a preset range. After the model is deployed, it receives the current gaze point coordinates transmitted by the eye-tracking device in real time. Combined with the current question type, it quickly outputs the visual focus movement path for the next 0.5 seconds, providing accurate basis for guiding sentence positioning.Subsequently, the augmented reality overlay technology is adopted to dynamically render semi-transparent text boxes in the adjacent area of the predicted path to display contextual guiding statements. The augmented reality overlay technology is a technology that seamlessly integrates virtual text information with the real question interface. Without changing the original interface layout, it only overlays and displays guiding content in the specified area to ensure visual coherence. The adjacent area refers to the range of 5-10 pixels on both sides of the predicted path, which not only ensures that the guiding statements can be quickly captured by the user but also does not block the content of the stem being currently gazed at, thus avoiding interfering with normal problem-solving. Dynamic rendering means that the position of the text box will be adjusted in real-time as the predicted path is updated. If the model detects a change in the direction of eye movement, it will immediately recalculate the adjacent area and update the position of the text box. The semi-transparent text box is used to further reduce the occlusion effect, with the transparency set at 60% (ensuring that the guiding statements are clearly readable while still allowing the content of the stem below to be seen through the text box). The width of the text box is adapted to the length of the guiding statement, and the height is automatically adjusted according to the font size. The font style is selected to be consistent with that of the stem to avoid visual abruptness. The specific display process is as follows: After obtaining the predicted path, the system quickly calculates the coordinate range of the adjacent area and determines the pixel coordinates of the upper left and lower right corners of the text box. Subsequently, the recombined contextual guiding statement is filled into the text box, and the text box is overlaid on the question interface through the augmented reality rendering engine, with the rendering delay controlled within 0.1 second to ensure synchronization with eye movement. During the display process, the model updates the predicted path every 0.1 second, and the position of the text box is dynamically adjusted accordingly, always remaining in the adjacent area of the eye movement trajectory, enabling the user to see the guiding content without having to deliberately shift their line of sight. Finally, a display termination mechanism is set: The display is immediately terminated when it is detected that the eyeball has left the stem area. The stem area is the core scope of the question defined by the system according to semantic logic (including the stem text, data parameters, question statements, etc.), and is clearly defined by the preset page coordinate boundaries. The detection logic for leaving the stem area is as follows: If the user's fixation point coordinates exceed the boundary range of the stem area and do not return within 0.3 seconds continuously, it is determined that the user has left. When terminating the display, the system quickly removes the overlaid semi-transparent text box and restores the original state of the question interface to avoid interference caused by the continuous presence of the guiding statement when the user is focusing on other areas. Through the complete process of model prediction and positioning + augmented reality overlay + dynamic termination, the guiding statements can not only precisely fit the rhythm of the user's eye movement and be presented in the most perceptible area but also minimize interference with the problem-solving process, ensuring the effectiveness and practicality of the guidance, allowing students to quickly obtain knowledge points and extraction guidance without interrupting their problem-solving thinking.

[0051] After visually displaying contextualized guiding statements using augmented reality overlay technology, to further enhance the perceptual effect of the guidance—allowing students to quickly grasp the guiding signals without interfering with their normal problem-solving process—it is necessary to simultaneously generate short auditory cues that match the core guiding principles. Through dual stimulation of visual and auditory senses, this precisely assists in the extraction of knowledge points. The specific implementation method is as follows: First, extract the core action verbs from the generated contextualized guiding statements as keywords. Core action verbs are the core words in the guiding statements that represent specific operational requirements, directly reflecting the cognitive or operational behaviors that students need to perform. For example, in guiding statements, "analyzing the quantitative relationships in the question stem and then substituting them into the formula to calculate the result," or "identifying and organizing the argument sentences in the text and sorting out the logical relationships," these are all core action verbs. Keywords here specifically refer to core action verbs, whose function is to establish the connection between the guiding statements and the auditory cues, ensuring that the auditory signals accurately respond to the core requirements of the visual guidance. The specific extraction process is as follows:

[0052] The system uses a natural language processing model to perform semantic analysis on the recombined contextualized guidance statements, filter out all verbs that represent actions, and then prioritize the verbs according to their core importance in the statement (whether they directly determine the direction of the operation). The verb ranked first is selected as the keyword. If the statement contains only one action verb, it is directly used as the keyword to ensure that the keyword can accurately summarize the core operation of the guidance statement. Next, based on the determined knowledge point obstacle type, a pre-recorded sound effect fingerprint database is used to match the obstacle. The knowledge point obstacle type is the obstacle attribute identified when locating related knowledge points earlier, divided into conceptual understanding obstacles and operational execution obstacles: Conceptual understanding obstacles refer to students' inability to extract basic concepts, definitions, logical relationships, etc., resulting in their inability to understand the question stem or the connotation of knowledge points. For example, they may not understand the definition of the monotonicity of a function and therefore cannot analyze the conditions in the question stem. Operational execution obstacles refer to students' inability to extract formula applications, standardized steps, and calculation rules, resulting in their inability to proceed with the problem-solving operation. For example, they may know the definition of the Pythagorean theorem but not know how to substitute the side lengths for calculation. The pre-recorded sound effect fingerprint database is a collection of short prompt sounds stored in the system, categorized by obstacle type. Each prompt sound has a unique sound effect fingerprint (i.e., audio feature identifier) ​​for quick matching and retrieval. The prompt sounds in the sound effect database have all been optimized to avoid being harsh or lengthy and to ensure that they do not interfere with problem-solving. Matching refers to matching the obstacle type with the preset classification tags in the sound effect fingerprint database to quickly retrieve the corresponding prompt sound. The specific matching process is as follows: The system first reads the knowledge point obstacle type tags determined in the previous text. If the tag is a concept understanding obstacle, then it matches the low-frequency prompt sound under the concept understanding category in the sound effect fingerprint library; if the tag is an operation execution obstacle, then it matches the high-frequency prompt sound under the operation execution category. The category tags in the sound effect fingerprint library correspond one-to-one with the obstacle type tags to ensure that the matching is without deviation. The call response time is controlled within 0.05 seconds to achieve synchronization with the visual display. For conceptual comprehension difficulties, low-frequency prompts are played. These are low-pitched, soothing sounds with frequencies between 100-300 Hz. Such sounds are not abrupt and can convey guidance without interrupting students' thinking, suitable for scenarios requiring calm and focused thought, such as a low, 0.1-second sound. For operational difficulties, high-frequency prompts are played. These are crisp, clear sounds with frequencies between 800-1200 Hz. These sounds are highly recognizable and can quickly remind students to focus on operational steps, suitable for scenarios requiring clear action, such as a crisp, 0.1-second sound. The duration of both types of prompts is set below the human reaction threshold (0.1 seconds, calibrated experimentally). The human reaction threshold is the shortest audio duration that humans can perceive without causing interference; exceeding this threshold can easily distract attention, while being below it may result in imperceptibility. The 0.1-second setting achieves the effect of perceiving the prompt without interrupting the train of thought.Finally, a prompt audio is played synchronously when the visual display starts. Synchronized playback means that the audio signal is output simultaneously the moment the augmented reality text box begins to render and display, ensuring that the visual guidance and auditory prompts are perfectly timed, allowing students to quickly associate the auditory signal with the visual guidance. The playback channel is by default through the built-in speaker of the student's problem-solving device, with the volume set to 30%-50% of the system volume (manually adjustable) to avoid interference from excessively high volume or inability to perceive from excessively low volume. If the student does not pay attention to the visual guidance during playback (e.g., their gaze is not fixed on the question stem area), the system will not repeat the prompt audio; it will only be triggered once synchronously during the first visual display, ensuring the necessity and appropriateness of the auditory prompts, strengthening the guidance effect without creating redundant interference. Through the process of keyword extraction—obstacle type matching—precise sound effect playback, auditory prompts and visual guidance complement each other, solving the problem that visual guidance may be overlooked, and adapting to different cognitive scenarios through differentiated sound effects, making personalized assistance more comprehensive and accurate. This helps students quickly capture guidance signals and extract related knowledge points without interrupting their problem-solving rhythm.

[0053] The second objective of this invention is to provide a system for implementing a personalized learning process assistance method that integrates multimodal data as described in any one of the above-mentioned methods, comprising:

[0054] Data acquisition unit 1 collects students' problem-solving behavior data in real time through eye-tracking, keyboard recording, and interface operation monitoring devices;

[0055] The delayed feature recognition unit 2 identifies repeated viewing, input pauses and cognitive stagnation behaviors based on predefined rules, and outputs delayed features that represent the difficulty in extracting knowledge points;

[0056] The third unit for locating related knowledge points combines delay features with knowledge point association maps to screen basic knowledge points that have been mastered but have retrieval obstacles, and matches them with association paths for question comprehension or formula application.

[0057] The guidance statement generation unit 4 calls the standard knowledge base to break down related knowledge points into the smallest cognitive units, combines them with question parameters to reorganize them into contextualized guidance statements, and binds them with keyword sound effects;

[0058] The multimodal guidance output unit 5 dynamically renders guidance statements in the gaze area based on the eye-tracking prediction model, synchronously triggers short auditory cues bound to keywords, and terminates the display when the gaze moves away.

[0059] This invention uses eye-tracking, keyboard interaction recording, and interface operation monitoring devices to collect multimodal behavioral data of students during problem-solving in real time. Based on predefined rules, it identifies delay features indicating difficulty in knowledge point retrieval, such as repeated viewing and input pauses. Combining the knowledge point tags of the current question with a pre-constructed knowledge point association map, it locates related knowledge points that the student has mastered but faces retrieval difficulties with. It then calls upon a standard knowledge base to break down these knowledge points into their smallest cognitive units, integrates the question stem parameters and problem-solving context, and reassembles them into contextualized guiding statements. Within a preset response time, based on eye-tracking trajectory prediction, it displays the guiding statements in the question stem area that the student is looking at using augmented reality technology, simultaneously triggering short auditory cues matching the type of obstacle, thus improving the targeting and effectiveness of personalized learning assistance.

[0060] The foregoing has shown and described the basic principles, main features, and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited to the above embodiments. The embodiments and descriptions in the specification are merely preferred examples and are not intended to limit the invention. Various changes and modifications can be made to the invention without departing from its spirit and scope, and all such changes and modifications fall within the scope of the present invention as claimed. The scope of protection of the present invention is defined by the appended claims and their equivalents.

Claims

1. A personalized learning process assistance method that integrates multimodal data, characterized in that: Includes the following steps: S1. Collect students' multimodal behavior data in real time during the problem-solving process through eye-tracking devices, keyboard interaction recording devices, and interface operation monitoring devices. Based on predefined delay recognition rules, identify delay features that represent the difficulty in extracting knowledge points from the multimodal behavior data. S2. Based on the delay features and the knowledge point tags corresponding to the current question, locate the related knowledge points that the student has mastered but has difficulty retrieving from the pre-constructed knowledge point association graph. S3. Call the complete content of the related knowledge points in the standard knowledge base, break down the complete content into the smallest cognitive units that can be operated independently, and combine the specific parameters of the current question and the problem-solving context to reorganize the smallest cognitive units into contextualized guiding statements containing actual question stem data. The operations of breaking down the cognitive unit into independent operations include: The semantic structure of knowledge points in the standard knowledge base is analyzed, and each complete operation step is broken down into indivisible cognitive action units. The cognitive action units must meet the condition that they can be executed independently without additional explanation. Mathematical knowledge points are broken down according to formula transformation steps, and textual knowledge points are broken down according to logical reasoning nodes, and dependency relationship tags between units are generated. Extract the specific parameter variables from the current question, replace the abstract symbols in the obtained minimum cognitive unit with the actual question data, determine the execution order of the minimum cognitive unit according to the problem-solving context, insert conjunctions to generate coherent operation instruction statements, delete descriptive content that is irrelevant to the current problem-solving stage, and retain core action verbs and key data; S4. Within a preset response time after recognizing the delayed feature, display the guiding statement in the question stem area currently being looked at by the user based on the real-time eye-tracking focus coordinates, and generate a short auditory cue that is bound to the keywords of the guiding statement.

2. The personalized learning process assistance method that integrates multimodal data according to claim 1, characterized in that: The predefined delayed identification rules specifically include: The system calculates the number of times the gaze point is repeatedly accessed in the same question stem block based on eye-tracking data. When the number of times the gaze point is repeatedly accessed is higher than the average number of times the current question is viewed, it is judged as repeated viewing behavior. The system counts the time interval between adjacent keystrokes based on keyboard interaction data. When the interval between consecutive mathematical symbol inputs is higher than the student's average input speed, it is judged as input pause behavior. The system monitors the time period without operation based on interface operation data. When the duration is higher than the threshold for coherent problem-solving operation, it is judged as cognitive stagnation behavior. Any of the above behaviors is output as a delayed feature representing the difficulty in extracting knowledge points.

3. The personalized learning process assistance method that integrates multimodal data according to claim 2, characterized in that: The construction of the knowledge point association graph specifically includes: A tree-like hierarchical structure is constructed with the subject knowledge point system as the framework. The parent node represents a higher-order knowledge point, and the child node represents the basic knowledge point it depends on. Based on the cognitive psychology model, the extraction difficulty weight is marked for each knowledge point association path. The weight value is dynamically updated according to the average response delay when students call the knowledge point in historical data. A two-way error transmission link between knowledge points is established to record the statistical relationship of students' errors in solving problems with subsequent related knowledge points due to the failure to extract the preceding knowledge points.

4. The personalized learning process assistance method that integrates multimodal data according to claim 3, characterized in that: The logic for determining the related knowledge points that students have mastered but have difficulty retrieving is as follows: Extract the student's mastery score of the target knowledge point within a preset time period from the learning history database. If the score is higher than the mastery threshold, it is marked as mastered. Simultaneously analyze the real-time extraction difficulty weight of the knowledge point in the knowledge point association graph. When the weight is greater than the preset obstacle threshold and a delay feature is identified, it is determined that there is an extraction obstacle.

5. The personalized learning process assistance method that integrates multimodal data according to claim 4, characterized in that: The process of locating and associating knowledge points specifically includes: Based on the knowledge point tags of the current question, locate the root node in the knowledge point association graph, traverse the child nodes downwards along the graph, and filter out the basic knowledge points that simultaneously satisfy the state of being mastered and the state of having extraction obstacles. Combine the identified delay feature type to match the association path. If the delay feature is repeated viewing behavior, prioritize locating knowledge points related to question understanding. If it is input pause behavior, locate knowledge points related to formula application.

6. The personalized learning process assistance method that integrates multimodal data according to claim 1, characterized in that: An eye-tracking trajectory prediction model is established to predict the visual focus movement path for a preset duration based on historical eye-tracking data and the current gaze coordinates. Augmented reality overlay technology is used to dynamically render semi-transparent text boxes in the vicinity of the predicted path to display contextualized guiding statements. The display is immediately terminated when the eyeball leaves the question area.

7. The personalized learning process assistance method that integrates multimodal data according to claim 1, characterized in that: The core action verbs in the generated contextualized guidance statements are extracted as keywords. Based on the determined knowledge point obstacle type, a pre-recorded sound effect fingerprint database is matched. Low-frequency prompts are played for conceptual understanding obstacles, and high-frequency prompts are played for operational execution obstacles. When the visual display is started, a prompt audio with a duration lower than the human reaction threshold is played synchronously.

8. A system for implementing a personalized learning process assistance method comprising any one of claims 1-7, characterized in that, include: The data acquisition unit (1) is used to collect students' multimodal behavior data in real time through eye-tracking device, keyboard interaction recording device and interface operation monitoring device, and to identify delayed features that are difficult to extract knowledge points from the multimodal behavior data based on predefined delay recognition rules; The delayed feature identification unit (2) is used to construct and dynamically update a knowledge point association graph containing a tree-like hierarchical structure, extraction difficulty weights and bidirectional error transmission links. Based on the delayed features and the current question knowledge point labels, combined with historical mastery data, it locates the associated knowledge points that students have mastered but have extraction obstacles. The related knowledge point positioning unit (3) is used to call the standard knowledge base, decompose the content of the related knowledge points into the smallest cognitive units that can be operated independently, and combine the specific parameters of the question with the problem-solving context to reorganize the smallest cognitive units into contextualized guiding statements. The decomposition into the smallest cognitive units includes parsing the semantic structure of the knowledge points, splitting the operation steps into indivisible cognitive action units that can be executed independently and do not require additional explanation conditions, decomposing mathematical and textual knowledge points according to the formula transformation steps and logical reasoning nodes respectively, and generating inter-unit dependency relationship tags. The reorganization includes extracting the specific parameters of the question to replace abstract symbols, determining the execution order according to the problem-solving context, and inserting connecting words to generate coherent operation instruction statements. The guidance statement generation unit (4) is used to dynamically display the contextualized guidance statement in the question stem area currently being looked at by the user within a preset response time after recognizing the delayed features, based on the real-time eye-tracking focus coordinates. The multimodal guidance output unit (5) is used to extract the core action verbs in the contextualized guidance statement as keywords, match the sound effect library according to the knowledge point obstacle type, and generate and synchronously play short auditory prompts bound to the visual display.

Citation Information

Patent Citations

  • Generative and multi-modal sensing integrated agent learning system

    CN120277628A

  • Remote education data processing system

    CN120471275A