An AI multi-modal fusion-based learning behavior intervention method for a companion robot

By using multimodal fusion technology, the problem of aligning multimodal behavioral data across time scales in the learning process of learning companion robots has been solved. This has enabled the stable storage of learning behavior data and the orderly generation of cognitive segments, ensuring the accuracy of intervention trigger timing and the traceability of intervention results.

CN122433908APending Publication Date: 2026-07-21GUANGDONG HUANYU PRECISION TECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
GUANGDONG HUANYU PRECISION TECHNOLOGY CO LTD
Filing Date
2026-04-28
Publication Date
2026-07-21

AI Technical Summary

Technical Problem

In the learning scenarios of learning companion robots, existing technologies face difficulties in aligning multimodal behavioral data across time scales. Overlapping cognitive state cues lead to distortion in segment judgment and deviation in intervention trigger timing, resulting in bias in cognitive input judgment, lag in state switching recognition, poor continuity of intervention path, and difficulty in tracing results.

Method used

By collecting multimodal learning behavior data, preprocessing and storing it to build a behavior database, using multimodal event time-series difference data to analyze cross-modal time delay mismatch, performing segment resegmentation and boundary correction, reorganizing cognitive segments based on synchronous segments and behavioral cue data, and integrating multimodal behavior ratios and time delay mismatch data to perform cognitive input separation analysis, achieving state determination and intervention trigger determination, and executing graded intervention and closed-loop verification.

Benefits of technology

It achieves continuous association and stable storage of multimodal learning behavior data, orderly generation of cognitive segment packages and stable establishment of indexes for the same segment, hierarchical control of intervention trigger timing and closed-loop verification of intervention results, and solves the problems of mixed dimensions of multi-source data, fragmented learning behavior cues and single intervention path.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122433908A_ABST
    Figure CN122433908A_ABST
Patent Text Reader

Abstract

The application discloses a kind of based on AI multimodal fusion's learning behavior intervention method of companion robot, it is related to intelligent decision-making technical field.The learning behavior intervention method of companion robot based on AI multimodal fusion, including S1, acquisition companion robot learning behavior intervention data and pre-processing;S2, by multimodal event time series difference data is carried out cross-modal time delay mismatch analysis;S3, based on synchronization fragment, correction fragment and behavior clue data are carried out cognitive fragment reorganization, constructs cognitive fragment package and establishes same fragment index;S4, fusion multimodal behavior ratio and time delay mismatch data are carried out cognitive input separation degree analysis;S5, around cognitive separation value, behavior accumulation and time delay mismatch data are carried out intervention trigger determination.Solve the problem that multiple modal behavior data is difficult to align across time scale in companion scene, cognitive state clue overlap causes fragment determination distortion and intervention trigger opportunity deviation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of intelligent decision-making technology, specifically to a method for intervening in the learning behavior of a learning companion robot based on AI multimodal fusion. Background Technology

[0002] Learning companion robots are designed for scenarios involving learning companionship, problem explanation, answer guidance, and learning status recognition. Existing technologies typically combine facial videos, eye movements, head posture, writing actions, voice clips, and problem presentation records to perceive learning behavior. They extract behavioral data such as the duration of gaze on the problem, the number of gaze shifts, the timing of head posture changes, the duration of continuous writing, the moment of starting a pen, the moment of ending a semantic segment, the moment of starting the first answer, the keyword repetition matching results, the duration of pauses, the number of consecutive incorrect answers, and the duration of pauses without answering. After unifying and organizing the multi-source behavioral data, it is written into a behavior record database for subsequent learning status analysis, cognitive engagement judgment, and intervention trigger determination.

[0003] For example, invention patent CN114860893B discloses an intelligent decision-making method and apparatus based on multimodal data fusion and reinforcement learning. The method includes: acquiring an intelligent decision-making task comprising language instructions and visual information; encoding the language instructions and visual information to obtain language encoding vectors and visual encoding vectors, thus obtaining multimodal data; based on a multimodal fusion method, obtaining multimodal fused data according to the multimodal data; inputting the multimodal data into a distance-optimized language understanding model; when determining whether the environmental state corresponds to the language instruction, providing immediate language rewards to the reinforcement learning agent; inputting the multimodal fused data into a reinforcement learning algorithm; and based on the reinforcement learning algorithm and the immediate language rewards, outputting actions and completing the intelligent decision. This enables the agent to understand natural language instructions and accelerates learning by providing language rewards, thereby quickly completing the task.

[0004] For example, the invention patent with announcement number CN112307982B discloses a human behavior recognition method based on interleaved enhanced attention networks. This method addresses the problems of existing technologies neglecting local information and being easily interfered with by a large amount of redundant background information and irrelevant information in videos, resulting in insufficient behavior recognition capabilities. The implementation steps are as follows: (1) generating a training set; (2) obtaining low-level feature maps and high-level feature maps; (3) constructing a hierarchical complementary attention module; (4) constructing a local enhanced attention module; (5) building a classification network; (6) constructing an interleaved enhanced attention network; (7) constructing the loss function of the interleaved enhanced attention network; (8) training the interleaved enhanced attention network; and (9) recognizing behaviors in video images. The construction of the interleaved enhanced attention network and its loss function can improve the accuracy of behavior recognition.

[0005] However, changes in facial representations, gaze retraction, head alignment, pen-writing actions, semantic comprehension results, and initial response behavior during the learning process often occur at different time scales, with the former usually appearing first and the latter usually appearing later. Fixed time window alignment methods easily lead to incorrect binding of physiological behavioral changes in the previous time period with semantic and response results in the subsequent time period, resulting in cross-modal temporal mismatch. Simultaneously, distraction, tension, stagnation, and insufficient engagement exhibit strong similarities in behavioral manifestations. Changes in question gaze, writing continuity, keyword repetition, brow retraction, pauses, and consecutive incorrect answers also tend to overlap. In scenarios with noise interference, modal absence, and lighting changes, state boundaries become even more difficult to distinguish. As a result, existing technologies are prone to increased cognitive engagement judgment bias, delayed state switching recognition, shifted intervention trigger points, decreased continuity of intervention paths, and difficulties in result tracing.

[0006] Therefore, there is an urgent need for a learning behavior intervention method for learning companion robots based on AI multimodal fusion. Summary of the Invention

[0007] Technical problems to be solved

[0008] To address the shortcomings of existing technologies, this invention provides a learning behavior intervention method for learning companion robots based on AI multimodal fusion, which solves the problems of difficulty in aligning multimodal behavioral data across time scales, distortion of segment judgment caused by overlapping cognitive state cues, and deviation of intervention triggering timing in learning companion scenarios.

[0009] Technical solution

[0010] To achieve the above objectives, the present invention provides the following technical solution: a learning behavior intervention method for a learning companion robot based on AI multimodal fusion, comprising: S1, collecting learning behavior intervention data of the learning companion robot, preprocessing the learning behavior intervention data, storing it, and constructing a learning companion behavior database; S2, performing cross-modal time delay mismatch analysis using multimodal event time sequence difference data, and performing segment resegmentation, segment boundary correction, and writing of segments to be reviewed based on the time delay analysis results; S3, reorganizing cognitive segments based on synchronous segments, corrected segments, and behavioral cue data, constructing cognitive segment packages, and establishing a segment index; S4, performing cognitive input separation analysis by fusing multimodal behavior ratios and time delay mismatch data, and performing state determination, segment classification, and state marking operations based on the separation analysis results; S5, performing intervention trigger determination based on cognitive separation values, behavior accumulation, and time delay mismatch data, and performing graded intervention, closed-loop verification, and continuous tracking operations based on the intervention determination results.

[0011] Furthermore, the specific steps for collecting learning behavior intervention data of the learning companion robot are as follows: Collect learning behavior intervention data of the learning companion robot by acquiring the learner's facial video stream, gaze movement video stream, and head posture video stream through the forward-facing camera; acquiring the writing action video stream through the desktop view camera; acquiring the speech stream through the sound pickup unit; and acquiring the question display time, question stem broadcast start time, question stem broadcast end time, prompt trigger time, and answer submission time through the question presentation recording link; obtaining the total number of target keywords for the current question from the question presentation recording link or question bank records; counting the number of times the current question group has been answered from the question submission records and answer result records; and counting the number of frame groups in the current segment from the video slice results of the current learning segment.

[0012] Further, the specific steps for preprocessing and storing the learning behavior intervention data of the learning companion robot and constructing the learning companion behavior database are as follows: For the facial video stream, perform facial key point localization, gaze trajectory tracking, and head posture change extraction; and extract the number of eyebrow contractions based on the displacement of the eyebrow key point, the change in dense optical flow between the eyebrows, and the response of the eyebrow action unit to obtain the duration of gaze on the question, the peak time of gaze fallback, the peak time of head return, the number of non-task gaze deviations, and the number of eyebrow contractions; For the desktop perspective video stream, perform pen tip recognition, trajectory continuity extraction, and pen starting point detection to obtain the continuous writing duration and the pen starting time; For the speech stream, perform speech activity detection, question stem alignment, keyword matching, and pause detection, and combine the results of lip key point movement changes, speech activity detection results, and pause duration thresholds to obtain the semantic segment end time and semantic segment duration. The study analyzed the duration of each ratio, including the initial response start time, the number of matching words for keyword paraphrasing, and the duration of mouth hesitation pauses. It also analyzed the alignment of question submission records and response result records with the execution order and segment association to obtain the number of consecutive incorrect answers and the longest pause without a response. Based on the semantic segment duration, the total number of target keywords for the current question, the number of frames in the current segment, and the number of times the current question group has been answered, the study obtained the following ratios: question gaze dwell time ratio, continuous writing time ratio, mouth hesitation pause time ratio, longest pause without a response ratio, keyword paraphrasing matching ratio, non-task gaze deviation frequency ratio, glabellar contraction frequency ratio, and consecutive incorrect answers ratio. After performing range normalization, outlier removal, and missing value imputation on all ratio parameters, the data was stored by learner ID, subject ID, question ID, and segment ID to construct a learning companion behavior database. A cognitive segment record table was then constructed within this database.

[0013] Furthermore, the specific steps for cross-modal time delay mismatch analysis using multimodal event temporal difference data are as follows: Obtain the multimodal event temporal difference data for the j-th learning segment. This data includes the semantic segment end time, the peak time of gaze fallback in the question text, the peak time of head alignment, the pen start time, the first answer start time, and the duration of the semantic segment; divide the absolute value of the difference between the peak time of gaze fallback in the question text and the end time of the semantic segment by the duration of the semantic segment to obtain the gaze-semantic time difference ratio; divide the absolute value of the difference between the peak time of head alignment and the end time of the semantic segment by the duration of the semantic segment to obtain the posture-semantic time difference ratio; divide the absolute value of the difference between the peak time of head alignment and the end time of the semantic segment by the duration of the semantic segment to obtain the posture-semantic time difference ratio; and divide the pen start time... The absolute value of the difference between the end time of the semantic segment and the end time of the semantic segment is divided by the duration of the semantic segment to obtain the writing semantic time difference ratio; the absolute value of the difference between the start time of the first response and the end time of the semantic segment is divided by the duration of the semantic segment to obtain the response semantic time difference ratio; the maximum value of the four time difference ratios is obtained by finding the maximum value, and the minimum value of the four time difference ratios is obtained by finding the minimum value; the four time difference ratios are added by one and then multiplied, and the fourth root operation is performed on the product result to obtain the overall time difference aggregate value; the difference between the maximum time difference ratio and the minimum time difference ratio is added by one to obtain the time difference discrete value; the overall time difference aggregate value and the time difference discrete value are multiplied to obtain the time delay mismatch value of the j-th learning segment.

[0014] Furthermore, the specific steps for performing segment resegmentation, segment boundary correction, and writing of segments to be reviewed based on the latency analysis results are as follows: By comparing the latency mismatch value with the mismatch threshold in real time, when the latency mismatch value is less than the mismatch threshold, the current learning segment is marked as a synchronous segment and directly sent to the cognitive segment reorganization process; when the latency mismatch value is greater than or equal to the mismatch threshold, the current learning segment is marked as a mismatched segment, and the earliest time among the peak time of gaze fallback, peak time of head return, and pen start time is used as the left boundary, and the latest time among the end time of semantic segment and the start time of first answer is used as the right boundary, the current learning segment is resegmented; the previous segment is generated. The system retrieves key question stem voice commands. The key question stem is a sentence in the question stem text that contains the target keyword of the current question. It collects continuous short segments from when the user looks at the question again to when they start writing again. At the same time, it calls up adjacent learning segments of the same question with a time delay mismatch value less than the mismatch threshold as time delay references and corrects the boundary of the current learning segment. Based on the corrected segment boundary, it re-extracts behavioral parameters. If the recalculated time delay mismatch value is still greater than or equal to the mismatch threshold, the current learning segment is marked as a segment to be temporarily deferred and written into the manual review sequence. It also generates a review prompt data package containing the segment start and end time, time delay mismatch value, ratio of consecutive incorrect answers, and ratio of the longest pause without answering.

[0015] Furthermore, the specific steps for reorganizing cognitive segments based on synchronous segments, revised segments, and behavioral cue data, constructing cognitive segment packages, and establishing segment indexes are as follows: Receive synchronous and revised segments; establish segment anchor sequence based on the semantic segment end time, the first answer start time, the peak time of gaze fallback on the question, and the pen start time; include the ratio of question gaze dwell time, the ratio of continuous writing time, and the ratio of keyword paraphrasing matching in the input candidate segment set; include the ratio of non-task gaze deviation frequency in the distraction candidate segment set; include the ratio of glabellar contraction frequency and the ratio of mouth hesitation pause duration in the tension candidate segment set; and include the ratio of consecutive incorrect answers and the ratio of the longest no-answer pause duration in the stagnation candidate segment set; and select time delay mismatches around adjacent learning segments of the same question. Learning segments with values ​​less than the mismatch threshold are designated as master anchor segments. Centered on the master anchor segment, when the segment interval is less than the splicing limit and the anchor point offset is less than the offset limit, forward splicing is performed on the front side of the master anchor segment, and backward splicing is performed on the back side of the master anchor segment. The middle p% of the total duration of the spliced ​​cognitive segment package is taken as the segment center segment. Overlapping input cues within the same anchor point interval are written into the segment center segment. After deducting the segment center segment from the total duration of the cognitive segment package, the remaining part is divided equally before and after as segment edge segments. Distraction cues, tension cues, and stagnation cues are written into the segment edge segments first to form a cognitive segment package. A segment index is established for the recombined cognitive segment package, and the segments are written into the cognitive segment record table according to the learner number, question number, segment number, and segment sequence number.

[0016] Furthermore, the specific steps for analyzing cognitive input separation by integrating multimodal behavior ratios and time delay mismatch data are as follows: Obtain the following ratios for the j-th cognitive segment package: question fixation duration ratio, continuous writing duration ratio, keyword paraphrasing matching ratio, non-task gaze deviation frequency ratio, glabellar contraction frequency ratio, mouth hesitation pause duration ratio, and time delay mismatch value; Add one to the question fixation duration ratio, continuous writing duration ratio, and keyword paraphrasing matching ratio respectively, and then perform cubic root aggregation to obtain the input support term; Add one to the non-task gaze deviation frequency ratio and glabellar contraction frequency ratio respectively, and then perform square root aggregation to obtain the interference diffusion term; Sum the mouth hesitation pause duration ratio and time delay mismatch value, add two, and perform natural logarithmic operation to obtain the hysteresis amplification value; Divide the input support term by the interference diffusion term and the hysteresis amplification value, and then sum them to obtain the cognitive separation value.

[0017] Further, the specific steps for performing state determination, fragment classification, and state marking operations based on the separation degree analysis results are as follows: By comparing the cognitive separation value with the separation threshold in real time, when the cognitive separation value is greater than or equal to the separation threshold, the current cognitive fragment package is marked as a clear input fragment, maintaining the boundary of the cognitive fragment package and the fragment order unchanged; when the cognitive separation value is less than the separation threshold, the current cognitive fragment package is marked as a state-entangled fragment, and the following values ​​are extracted from the n consecutive cognitive fragments before and after the current question: the ratio of the frequency of eyebrow contraction, the ratio of the duration of mouth hesitation pause, the ratio of the frequency of non-task gaze deviation, the ratio of keyword paraphrasing matching, the ratio of consecutive incorrect answers, and the ratio of the longest pause without answering, respectively, and the corresponding median values ​​are calculated; if the ratio of the frequency of eyebrow contraction is less than the threshold... If the frequency ratio of glabellar contraction is greater than the median value, and the frequency ratio of mouth hesitation pauses is greater than the median value, and the frequency ratio of non-task gaze deviation is less than the median value, then it is marked as a tense segment; if the frequency ratio of non-task gaze deviation is greater than the median value, and the frequency ratio of keyword paraphrasing is less than the median value, then it is marked as a distracted segment; if the frequency ratio of consecutive incorrect answers is greater than the median value, and the frequency ratio of the longest pause without answering is greater than the median value, then it is marked as a stagnant segment; the remaining cognitive segments are marked as mixed segments, and the segment labels and cognitive dissociation values ​​are sent together to the intervention triggering process.

[0018] Furthermore, the specific steps for determining the intervention trigger based on cognitive separation value, behavioral accumulation, and time delay mismatch data are as follows: Obtain the cognitive separation value, consecutive incorrect answer ratio, longest no-answer pause ratio, and time delay mismatch value of the j-th cognitive segment package; Subtract the cognitive separation value of the current cognitive segment package from the cognitive separation value of the previous cognitive segment package; When the difference is greater than zero, use the difference as the state drop value; when the difference is less than or equal to zero, set the state drop value to zero; Add one to the cognitive separation value and perform a reciprocal operation to obtain the current clarity compensation term; Summate the state drop value, the current clarity compensation term, and one to obtain the state aggregation term; Add one to the consecutive incorrect answer ratio to obtain the error accumulation term; Add two to the longest no-answer pause ratio and perform a natural logarithmic operation to obtain the stagnation amplification; Multiply the state aggregation term, the error accumulation term, and the stagnation amplification to obtain the numerator aggregation value; Add one to the time delay mismatch value to obtain the denominator adjustment value; Divide the numerator aggregation value by the denominator adjustment value to obtain the intervention ready value.

[0019] Furthermore, the specific steps for implementing tiered intervention, closed-loop verification, and continuous tracking based on the intervention judgment results are as follows: By comparing the intervention readiness value and the intervention threshold, which includes T1 and T2, when the intervention readiness value < T1, the current cognitive segment boundary, question stem playback order, and single prompt length remain unchanged; when T1 ≤ intervention readiness value < T2, a first-level intervention is implemented: for tense segments, the previous key question stem is replayed at a slower pace and the single prompt text length is shortened; for distraction segments, the control display module highlights and enlarges the key lines on the question and simultaneously enlarges the keyword font; for stagnant segments, a one-step prompt text switch is implemented and irrelevant prompt content is hidden; for mixed segments, the key conditional sentence is replayed and a one-sentence prompt is superimposed; when the intervention readiness value < T1 < T2, the intervention is implemented at a higher level. When the threshold is ≥T2, a secondary intervention is implemented. The current learning content is broken down into two consecutive steps. The key sentence and the first step prompt are displayed first. The second step prompt is displayed after the signal of the start time of writing or the start time of the first answer is re-acquired. The keyword repetition matching ratio, the non-task gaze deviation frequency ratio, the ratio of consecutive incorrect answers, and the ratio of the longest pause without answering are continuously tracked in the next two cognitive segment packages. When the keyword repetition matching ratio in the next two cognitive segment packages is less than the corresponding value before the intervention, and the non-task gaze deviation frequency ratio, the ratio of consecutive incorrect answers, and the ratio of the longest pause without answering are all greater than the corresponding value before the intervention, the next cognitive segment package is marked as a continuous intervention segment, and the secondary intervention path is continued.

[0020] Beneficial effects

[0021] The present invention has the following beneficial effects:

[0022] (1) This invention collects facial video streams, eye movement video streams, head posture video streams, writing action video streams, speech streams, and question presentation records. It also performs unified processing on the question gaze dwell time, continuous writing time, mouth hesitation pause time, number of keyword repetition matching words, number of non-task eye gaze shifts, number of eyebrow contractions, number of consecutive incorrect answers, and longest pause without answering. It also removes outliers and fills in missing values, so that multi-source learning behavior data can form comparable inputs at the same segment scale. This achieves the effect of continuous association and stable storage of multi-modal learning behavior data, effectively solving the problems of mixed dimensions of multi-source data, obvious missing interference, and unstable early input foundation in the prior art.

[0023] (2) This invention establishes a sequence of segment anchor points around the end time of semantic segments, the start time of the first answer, the peak time of the question gaze fall, and the start time of writing. It combines synchronous segments, corrected segments, input candidate segment sets, distracted candidate segment sets, tense candidate segment sets, and stagnant candidate segment sets to perform forward splicing, backward splicing, and rearrangement of central and edge segments. This enables the continuous reconstruction of behavioral cues between adjacent learning segments, thereby achieving the effect of orderly generation of cognitive segment packages and stable establishment of the same segment index. It effectively solves the problems of fragmented learning behavioral cues, disconnection between segments, and difficulty in coherently expressing the state evolution process in the prior art.

[0024] (3) In this invention, the cognitive separation value, state drop value, ratio of consecutive wrong answers, ratio of longest pause without answering, and time delay mismatch value are used to form the intervention readiness value. The hierarchical intervention path is switched according to T1 and T2, so that the first-level intervention and the second-level intervention form corresponding changes in the triggering time, prompt length, question replay and step breakdown. This achieves the effect of hierarchical control of intervention triggering timing and effectively solves the problems of premature intervention, late intervention, single intervention level and insufficient matching degree of intervention action in the prior art.

[0025] (4) This invention continuously tracks the keyword repetition matching ratio, non-task gaze deviation frequency ratio, consecutive incorrect answer ratio and longest pause without answering ratio after graded intervention, and performs continuous intervention segment marking and path follow-up based on subsequent segment change results, so that behavioral changes before and after intervention can be written back to the same judgment link, thereby achieving the effect of closed-loop verification and continuous tracking of intervention results, effectively solving the problems of difficulty in verifying intervention results, insufficient correlation between segments before and after intervention and insufficient traceability of intervention process in the prior art.

[0026] Of course, any product implementing this invention does not necessarily need to achieve all of the advantages described above at the same time. Attached Figure Description

[0027] Figure 1 This is a flowchart of a learning behavior intervention method for a learning companion robot based on AI multimodal fusion according to the present invention;

[0028] Figure 2 This is a time-series distribution curve of the multimodal cognitive state of the learning companion robot of this invention;

[0029] Figure 3 This is a schematic diagram of the three-dimensional surface visualization of the multimodal cognitive separation situation of the present invention;

[0030] Figure 4 This is a flowchart illustrating the cognitive state grading intervention process of the learning companion robot of the present invention. Detailed Implementation

[0031] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0032] Please see Figures 1-4 This invention provides a technical solution: a learning behavior intervention method for a learning companion robot based on AI multimodal fusion, comprising: S1, collecting learning behavior intervention data of the learning companion robot, preprocessing the learning behavior intervention data, storing it, and constructing a learning companion behavior database; S2, performing cross-modal time delay mismatch analysis through multimodal event time sequence difference data, and performing segment resegmentation, segment boundary correction, and writing of segments to be reviewed based on the time delay analysis results; S3, reorganizing cognitive segments based on synchronous segments, corrected segments, and behavioral cue data, constructing cognitive segment packages, and establishing a segment index; S4, performing cognitive input separation analysis by fusing multimodal behavior ratios and time delay mismatch data, and performing state determination, segment classification, and state marking operations based on the separation analysis results; S5, performing intervention trigger determination based on cognitive separation value, behavior accumulation, and time delay mismatch data, and performing graded intervention, closed-loop verification, and continuous tracking operations based on the intervention determination results.

[0033] Specifically, the steps for collecting learning behavior intervention data from the learning companion robot are as follows: Acquire the learner's facial video stream, gaze movement video stream, and head posture video stream through a forward-facing camera; acquire the writing action video stream through a desktop-view camera; acquire the speech stream through a sound pickup unit; and acquire the question display time, question stem broadcast start time, question stem broadcast end time, prompt trigger time, and answer submission time through the question presentation recording link; obtain the total number of target keywords for the current question from the question presentation recording link or question bank records. The total number of target keywords for the current question is obtained after pre-labeling each question in the question bank with keywords, representing the question... The number of keywords in the core semantic information, where synonyms are counted as the same keyword after unified mapping; the number of times the current question group has been answered is counted from the question submission record and answer result record, where the current question group is a set of consecutive practice questions of the same type to which the current question belongs, and the number of times the current question group has been answered is the total number of times the learner has submitted answers to each question in the current question group before entering the current learning segment; the number of frame groups in the current segment is counted from the video slice results of the current learning segment, where the number of frame groups in the current segment is the total number of image frames obtained after sampling the video stream corresponding to the current learning segment at equal intervals according to a preset fixed frame rate.

[0034] This implementation plan acquires a complete multimodal data stream of learners' facial and eye movements, handwriting dynamics, speech interaction timing, and question presentation and response records. It also clarifies the specific definitions and statistical methods for three benchmark parameters: the total number of target keywords for the current question, the number of responses to the current question group, and the number of frames in the current segment. This provides a unified time normalization benchmark and counting denominator for subsequent multimodal behavior ratio calculations, ensuring that the extraction process of ratios such as question gaze dwell time, continuous writing time, keyword paraphrasing and matching, non-task gaze deviation frequency, glabellar contraction frequency, consecutive incorrect answers, mouth hesitation pause duration, and longest pause without response has clear quantitative references and comparability. This guarantees the data consistency and assessment effectiveness of each ratio parameter in the cognitive segment record table of the learning companion behavior database across segments, questions, and learners.

[0035] Specifically, the steps for preprocessing and storing the learning behavior intervention data of the learning companion robot and constructing the learning companion behavior database are as follows: The steps involve performing facial key point localization, gaze trajectory tracking, and head pose change extraction on the facial video stream, and extracting the number of eyebrow contractions based on the eyebrow key point displacement, the change in dense optical flow between the eyebrows, and the response of the eyebrow action unit. The eyebrow key point displacement is the change in Euclidean distance between key point coordinates in the eyebrow region between adjacent frames; the change in dense optical flow between the eyebrows is the average amplitude of the optical flow vector of all pixels in the eyebrow region; and the response of the eyebrow action unit is based on the facial motion coding system... The detection response value of the AU4 action unit is defined as follows: when any two of the following exceed their respective preset thresholds—the displacement of the eyebrow keypoint, the change in the dense optical flow between the eyebrows, and the response value of the eyebrow action unit—and the duration of the contraction reaches the minimum duration threshold, it is considered an eyebrow contraction. This yields the following data: question gaze dwell time, peak time of question gaze fallback, peak time of head return to center, number of non-task gaze deviations, and number of eyebrow contractions. Pen tip recognition, trajectory continuity extraction, and pen start point detection are performed on the desktop view video stream to obtain the continuous writing duration and pen start time. Speech activity detection, question stem alignment, keyword matching, and pause detection are performed on the speech stream, and combined with the results of lip keypoint movement changes, speech activity detection results, and pause duration thresholds, the following data is obtained. The semantic segment end time, semantic segment duration, first response start time, number of keyword paraphrased matching words, and mouth hesitation pause duration are obtained. The pause duration threshold is an adaptive threshold pre-defined based on learner age and grade or an empirical threshold obtained through statistics from the same grade group. The execution order of question submission records and answer result records is aligned and the segments are associated to obtain the number of consecutive incorrect answers and the longest pause without answering. Based on the semantic segment duration, the question fixation duration, continuous writing duration, mouth hesitation pause duration, and longest pause without answering duration are divided by the semantic segment duration to obtain the question fixation duration ratio, continuous writing duration ratio, mouth hesitation pause duration ratio, and longest pause without answering duration, respectively. The following ratios are used: 1. **Keyword paraphrasing ratio:** The number of keyword paraphrasing matches is divided by the total number of target keywords in the current question. This ratio is zero when the total number of target keywords in the current question is zero. 2. **Non-task gaze shift frequency ratio and glabellar contraction frequency ratio:** The number of non-task gaze shifts and glabellar contraction counts are divided by the number of frame groups in the current segment. This ratio is zero when the number of frame groups in the current segment is zero. 3. **Continuous incorrect answer ratio:** The number of consecutive incorrect answers is divided by the number of consecutive incorrect answers in the current question group. This ratio is zero when the number of consecutive incorrect answers in the current question group is zero.All ratio parameters were normalized using range normalization, with the normalization range set between zero and one. Outlier removal and missing value imputation were performed. Outlier removal used the three-sigma criterion to remove outliers exceeding three standard deviations above or below the mean. Missing value imputation used the mean of corresponding ratios of adjacent segments of the same question type for the same learner. After processing, the data was stored by learner ID, subject ID, question ID, and segment ID, and a learning support behavior database was constructed. A cognitive segment record table was then built within this database.

[0036] In this implementation plan, the calculation rules and judgment criteria for various feature parameters are clarified through the refined extraction and standardization of multimodal learning behavior data. This effectively avoids calculation anomalies caused by the denominator being zero during ratio calculation. At the same time, the normalization of data, outlier removal, and missing value filling are completed, ensuring the standardization, consistency, and reliability of all behavioral feature parameters. Finally, a structured learning behavior database is formed and a cognitive segment record table is constructed, providing complete and standardized data support for subsequent cognitive segment reconstruction, cognitive state quantitative analysis, and intervention strategy determination.

[0037] Specifically, the steps for cross-modal time delay mismatch analysis using multimodal event temporal difference data are as follows: Obtain the multimodal event temporal difference data for the j-th learning segment. This data includes the semantic segment end time, the peak time of gaze fallback on the question, the peak time of head alignment, the pen start time, the first response start time, and the duration of the semantic segment. Divide the absolute value of the difference between the peak time of gaze fallback on the question and the end time of the semantic segment by the duration of the semantic segment to obtain the gaze-semantic time difference ratio. Divide the absolute value of the difference between the peak time of head alignment and the end time of the semantic segment by the duration of the semantic segment to obtain the posture-semantic time difference ratio. Divide the absolute value of the difference between the pen start time and the end time of the semantic segment by the duration of the semantic segment to obtain the writing-semantic time difference ratio. Divide the absolute value of the difference between the first response start time and the end time of the semantic segment by the duration of the semantic segment to obtain the response-semantic time difference ratio. Maximize the value of the four time difference ratios. To obtain the minimum time difference ratio, the minimum value is obtained by finding the minimum of the four time difference ratios. Then, the four time difference ratios are multiplied by one, and the fourth root operation is performed on the product to obtain the overall time difference aggregate value. The addition of one to each ratio prevents excessive decay of the product when the ratio is less than one. The fourth root operation gives the overall time difference aggregate value the same dimensional scale as each time difference ratio, and provides a smoother response in the low ratio range compared to square root or logarithmic transformations, suppressing the excessive influence of extreme fluctuations in individual time difference ratios on the overall aggregate value. The difference between the maximum and minimum time difference ratios is added by one to obtain the time difference discrete value. Adding one ensures that the discrete term remains non-zero even when the four time difference ratios are completely equal. The overall time difference aggregate value is multiplied by the time difference discrete value to obtain the time delay mismatch value of the j-th learning segment. This time delay mismatch value maintains scale stability as the semantic segment duration changes and simultaneously reflects the average offset and dispersion of cross-modal events.

[0038] The specific formula for calculating the time delay mismatch value is as follows:

[0039] ;

[0040] In the formula, Indicates the first The time delay mismatch value of each learning segment is used to characterize the overall offset and dispersion of visual, posture, writing and response events relative to semantic events within the same learning segment; Indicates the first The fixation semantic time difference ratio of each learning segment reflects the degree of deviation of the fixation behavior on the question relative to the end time of the semantic segment; Indicates the first The ratio of the pose-semantic time difference of each learning segment reflects the degree of deviation of the head pose correction behavior relative to the end time of the semantic segment; Indicates the first The ratio of the writing semantic time difference of each learning segment reflects the degree of deviation of the writing start behavior relative to the end time of the semantic segment; Indicates the first The ratio of the semantic time difference of the response to each learning segment reflects the degree of deviation of the first response behavior relative to the end time of the semantic segment; Indicates the first The maximum value among the four time difference ratios of a learning segment reflects the type of behavior with the most obvious deviation within the current learning segment. Indicates the first The minimum of the four time difference ratios for a learning segment reflects the type of behavior with the lightest offset within the current learning segment.

[0041] In this implementation scheme, the time delay mismatch value of the j-th learning segment is obtained by calculating the gaze semantic time difference ratio, the posture semantic time difference ratio, the writing semantic time difference ratio, and the response semantic time difference ratio, and multiplying the overall time difference aggregate value of the four time difference ratios with the time difference discrete value. This achieves a unified quantification of the degree and dispersion of cross-modal temporal offset between audiovisual behavior and semantic events. Moreover, the generated time delay mismatch value has scale stability for the duration change of semantic segments, and can provide a consistent comparison benchmark for subsequent segment synchronization determination and boundary correction.

[0042] Specifically, the steps for performing segment resegmentation, segment boundary correction, and writing of segments to be reviewed based on the latency analysis results are as follows: The latency mismatch value and mismatch threshold are compared in real time. The mismatch threshold is a pre-defined judgment boundary based on the statistical distribution of latency mismatch values ​​in synchronous segments among learners of the same grade. When the latency mismatch value is less than the mismatch threshold, the current learning segment is marked as a synchronous segment and directly sent to the cognitive segment reorganization process. When the latency mismatch value is greater than or equal to the mismatch threshold, the current learning segment is marked as a mismatched segment. The earliest time among the peak time of gaze fallback, peak time of head alignment, and pen start time is used as the left boundary, and the latest time among the end time of semantic segment and the start time of first answer is used as the right boundary to resegment the current learning segment. A voice retrieval command for the previous key question stem is generated, controlling the voice playback unit of the learning companion robot to play back the key question stem once. The key question stem is... The dry text contains sentences containing the target keywords of the current question, and collects short, continuous segments from when the learner re-focuses on the question to when they start writing again. Simultaneously, it calls upon adjacent learning segments of the same question with a latency mismatch value less than the mismatch threshold as latency references. The average of the gaze semantic latency ratio, posture semantic latency ratio, writing semantic latency ratio, and answer semantic latency ratio of this reference segment is used as the target offset. The boundary of the current learning segment is then corrected by shrinking or expanding in the direction of the target offset. Based on the corrected segment boundary, behavioral parameters are re-extracted and the latency mismatch value is recalculated. If the recalculated latency mismatch value is still greater than or equal to the mismatch threshold, the current learning segment is marked as a temporarily suspended segment, written into the manual review sequence, and a review prompt data package containing the segment's start and end times, latency mismatch value, consecutive incorrect answer ratio, and the ratio of the longest pause without answering is generated. This data package is then manually reviewed by the teacher after class, and the teacher can annotate the behavioral status label of the segment.

[0043] In this implementation scheme, the learning segments are synchronized and mismatched by comparing the time delay mismatch value with the mismatch threshold in real time. Under mismatch conditions, the segments are re-segmented with the extreme moment of the multimodal event as the boundary, and the key question stem is replayed to collect retry short segments. At the same time, the boundary of the current segment is directionally corrected by using the time difference statistics of adjacent synchronized segments. For segments that still exceed the threshold after correction, a review prompt data packet is generated and transferred to the manual review sequence. This realizes the automatic identification, local repair and marking and diversion of cross-modal time mismatches, providing a time consistency basis for subsequent cognitive segment reorganization.

[0044] Specifically, the steps for reorganizing cognitive segments based on synchronous segments, modified segments, and behavioral cue data, constructing cognitive segment packages, and establishing segment indexes are as follows: Receive synchronous and modified segments from the learning companion behavior database; based on preprocessed multimodal behavioral temporal features, establish segment anchor sequence using the semantic segment end time, first answer start time, question gaze drop peak time, and pen start time to achieve temporal alignment and benchmark unification of multi-source behavioral data; classify and aggregate according to the cognitive state attributes corresponding to the behavioral features, including the ratio of question gaze dwell time, the ratio of continuous writing time, etc. Keyword paraphrasing matching ratios are included in the input candidate segment set; non-task gaze deviation frequency ratios are included in the distraction candidate segment set; glabellar contraction frequency ratios and mouth hesitation pause duration ratios are included in the tension candidate segment set; and consecutive incorrect answer ratios and longest no-answer pause duration ratios are included in the stagnation candidate segment set. Reorganization and screening are conducted on adjacent learning segments under the same question. Learning segments with a time delay mismatch value less than the mismatch threshold are selected as main anchor segments. Using the main anchor segment as the core benchmark, under the condition that the segment interval between adjacent segments is less than the splicing limit and the corresponding anchor point offset is less than the offset limit, the main anchor segments are reorganized. The adjacent segments before the main anchor segment are spliced ​​forward, and the adjacent segments after the main anchor segment are spliced ​​backward. The middle p% of the total duration of the spliced ​​cognitive segment package is taken as the segment center segment. The input cues that overlap within the same anchor point interval are written into the segment center segment. Here, p is the center segment ratio parameter, which is a real number between 20 and 80. It is used to determine the concentrated storage interval of the core input cues in the cognitive segment package. The typical value of p is 50, which means that the middle 50% interval of the total duration of the cognitive segment package is taken as the segment center segment, and the input cues that overlap within the same anchor point interval are written into the segment center segment. The remaining portion of the total cognitive segment package after deducting the central segment is divided equally into segment edge segments. Distraction cues, tension cues, and stagnation cues are prioritized and written into the segment edge segments. Feature partitioning storage is used to distinguish and control core input features from interference features, forming a complete and standardized cognitive segment package. A dedicated segment index is established for the cognitive segment package that has been reorganized and feature partitioned to ensure the traceability of data association within the segment. The relevant data of the cognitive segment package is completely written into the cognitive segment record table in a unified format of learner number, question number, segment number, and segment sequence number.

[0045] In this implementation scheme, by receiving synchronous and corrected segments and establishing a segment anchor sequence, the temporal alignment and benchmark unification of multimodal learning behavior segments are achieved. At the same time, candidate segment sets are classified according to cognitive state attributes for various feature ratios. Cognitive segments are recombined by selecting main anchor segments and splicing before and after execution. The division of central and peripheral segments enables partitioned storage of input cues and interference cues such as distraction, tension, and stagnation. Furthermore, a segment index is established and cognitive segment packages are written into the cognitive segment record table in a standardized manner. This effectively ensures the structure, standardization, and traceability of cognitive segments, clarifies the storage logic of different cognitive features, and provides standardized and reusable segment data support for subsequent quantitative assessment of cognitive states, determination of state drops, and precise intervention triggering.

[0046] Specifically, the steps for analyzing cognitive input separation by integrating multimodal behavior ratios and time delay mismatch data are as follows: Obtain the following ratios for the j-th cognitive segment: question fixation duration ratio, continuous writing duration ratio, keyword paraphrasing matching ratio, non-task gaze deviation frequency ratio, glabellar contraction frequency ratio, mouth hesitation / pause duration ratio, and time delay mismatch value; Add one to each of the question fixation duration ratio, continuous writing duration ratio, and keyword paraphrasing matching ratio, and then perform cubic root aggregation to obtain the input support term. Each addition is used as a smoothing constant to avoid excessive attenuation of the input support term due to zero-value multiplication. The cubic root operation normalizes the product of the three input dimensions to the same dimensional scale as a single ratio, making the input support term dimensionless and reflecting the synergistic enhancement effect of fixation, writing, and paraphrasing; Add one to each of the non-task gaze deviation frequency ratio and glabellar contraction frequency ratio, and then perform square root aggregation to obtain the interference diffusion term. Adding one to each term as a smoothing constant, the square root operation brings the interference diffusion term and the input support term into a comparable numerical range. This interference diffusion term is dimensionless and characterizes the degree of joint interference of distracted gaze deviation and tense facial expression on cognitive input. The sum of the ratio of mouth hesitation pause duration and the delay mismatch value, plus two, and then performing a natural logarithmic operation, yields the hysteresis amplification. The addition of two serves as a smoothing constant to ensure that the natural logarithmic input domain is greater than zero. The natural logarithmic transformation makes the hysteresis amplification exhibit a compressive response to the combined effect of mouth pause and delay mismatch, thereby suppressing excessive expansion of the denominator in the high hysteresis range. This hysteresis amplification, after being dimensionless, reflects the damping effect of expression hysteresis and cross-modal mismatch on the separation of input signals. Dividing the input support term by the sum of the interference diffusion term and the hysteresis amplification yields the cognitive separation value, which is a dimensionless index whose magnitude directly reflects the clarity of separation of the cognitive input signal relative to distraction interference and behavioral hysteresis.

[0047] The specific formula for calculating the cognitive dissociation value is as follows:

[0048] ;

[0049] In the formula, The cognitive separation value represents the degree of separation between the cognitive engagement signal and the distraction signal, tension signal, and time delay mismatch signal. The ratio of gaze duration to the question text reflects the learner's sustained visual engagement with the current question text. It represents the ratio of continuous writing time, reflecting the degree to which learners maintain explicit response actions around the current question; The keyword paraphrase matching ratio reflects the learner's level of following and reproducing the core information of the current question; This represents the ratio of non-task gaze deviation frequency, reflecting how frequently learners detach themselves from the question within the current learning segment; The ratio of the frequency of glabellar contractions reflects the frequency with which learners exhibit tense, strained, or hesitant facial expressions during the current learning segment. The ratio of the duration of hesitation and pause at the mouth reflects the degree of delay, pause and uncertainty in the learner's speech expression. It represents the time delay mismatch value, reflecting the overall offset and dispersion of visual, posture, writing and response events relative to semantic events.

[0050] Table 1 shows the cognitive state evaluation table of the learning segment of the learning companion robot based on multimodal fusion in this embodiment. At the first monitoring time, the ratio of question gaze dwell time was 0.72, the ratio of continuous writing time was 0.68, the ratio of keyword paraphrasing and matching was 0.85, the ratio of non-task gaze deviation frequency was 0.12, the ratio of brow contraction frequency was 0.08, the ratio of mouth hesitation and pause time was 0.15, and the time delay mismatch was 0.28, resulting in a calculated cognitive separation value of 1.86. At the second monitoring time, the ratio of question gaze dwell time was 0.45, the ratio of continuous writing time was 0.41, the ratio of keyword paraphrasing and matching was 0.52, the ratio of non-task gaze deviation frequency was 0.18, the ratio of brow contraction frequency was 0.36, the ratio of mouth hesitation and pause time was 0.42, and the time delay mismatch was 0.56, resulting in a calculated cognitive separation value of 0.71. At the third monitoring time point, the ratio of question fixation duration was 0.38, the ratio of continuous writing duration was 0.33, the ratio of keyword paraphrasing and matching was 0.40, the ratio of non-task gaze deviation frequency was 0.45, the ratio of glabellar contraction frequency was 0.15, the ratio of mouth hesitation and pause duration was 0.31, and the time delay mismatch was 0.62, resulting in a calculated cognitive dissociation value of 0.53. At the fourth monitoring time point, the ratio of question fixation duration was 0.29, the ratio of continuous writing duration was 0.25, the ratio of keyword paraphrasing and matching was 0.31, the ratio of non-task gaze deviation frequency was 0.22, the ratio of glabellar contraction frequency was 0.29, the ratio of mouth hesitation and pause duration was 0.58, and the time delay mismatch was 0.75, resulting in a calculated cognitive dissociation value of 0.39.

[0051] Table 1. Cognitive State Assessment Table of Learning Companion Robot at Monitoring Time Based on Multimodal Fusion

[0052] Monitoring time 1 0.72 0.68 0.85 0.12 0.08 0.15 0.28 1.86 2 0.45 0.41 0.52 0.18 0.36 0.42 0.56 0.71 3 0.38 0.33 0.40 0.45 0.15 0.31 0.62 0.53 4 0.29 0.25 0.31 0.22 0.29 0.58 0.75 0.39

[0053] like Figure 2 As shown, this is a time-series distribution curve of the multimodal cognitive state of the learning companion robot provided in this application embodiment. The horizontal axis represents learning time, and the vertical axis represents cognitive separation value. (See Table 1 for more details.) Figure 4 As can be seen, the cognitive separation value at the first monitoring time was 1.86, corresponding to a highly focused state with excellent input indicators and minimal interference and time delay mismatch; the cognitive separation value at the second monitoring time dropped to 0.71, corresponding to a mildly distracted and tense state with declining input indicators and increased interference and hesitation; the cognitive separation value at the third monitoring time was 0.53, corresponding to a significantly distracted state with low input indicators and significant non-task gaze deviation; and the cognitive separation value at the second monitoring time was as low as 0.39, corresponding to a cognitive stagnation state with weak input indicators and prominent hesitation and time delay mismatch.

[0054] like Figure 3The figure shows a three-dimensional surface visualization of the multimodal cognitive separation state provided in this application embodiment. The horizontal axis represents the ratio of question fixation duration to the duration of attention, the vertical axis represents the time delay mismatch value, and the vertical axis represents the cognitive separation value. As can be seen from the figure, in the region where the ratio of question fixation duration to the duration of attention is high and the time delay mismatch value is low, the surface rises significantly, and the cognitive separation value reaches its peak, indicating that the learner's multimodal behavior and semantic events are highly synchronized, and the cognitive engagement signal is clearly identifiable. As the time delay mismatch value increases, the surface rapidly decreases along this direction. Even if the ratio of question fixation duration to the duration of attention remains high, the cognitive separation value also decreases significantly, reflecting that the disconnect between behavior and semantics makes it difficult to separate the engagement state. In the region where the ratio of question fixation duration to the duration of attention is low and the time delay mismatch value is high, the surface drops to its lowest point, corresponding to non-engagement states such as distraction, tension, or stagnation. The overall trend of this surface intuitively reflects the joint response characteristics of the cognitive separation value to the intensity of engagement and the degree of time delay mismatch, verifying the effective ability of this application to distinguish between cognitive engagement and distraction states under multimodal asynchronous conditions.

[0055] Table 1 provides a direct and quantitative comparison of the numerical differences in the following parameters at four monitoring times: the ratio of gaze duration on the question, the ratio of continuous writing duration, the ratio of keyword paraphrasing and matching, the ratio of non-task gaze deviation frequency, the ratio of glabellar contraction frequency, the ratio of mouth hesitation and pause duration, the time delay mismatch value, and the cognitive separation value. This accurately highlights a strong positive correlation between the cognitive separation value and input-related characteristic parameters, and a strong negative correlation with interference-related and time delay mismatch parameters. This fully verifies that the cognitive input separation assessment module of this application can effectively quantify the cognitive input state and interference separation effect of the learning companion robot's multimodal learning behavior. It provides intuitive data support and scientific basis for subsequent intervention triggering, graded intervention execution, and closed-loop verification, effectively ensuring the accuracy and timeliness of the learning companion robot's learning behavior intervention.

[0056] In this implementation scheme, a dimensionless input support term is obtained by performing cubic root aggregation on the ratio of topic gaze dwell time, continuous writing time, and keyword paraphrasing matching of the j-th cognitive segment package. A dimensionless interference diffusion term is obtained by performing square root aggregation on the ratio of non-task gaze deviation frequency and the ratio of glabellar contraction frequency. A dimensionless hysteresis amplification is obtained by summing the ratio of mouth hesitation pause time and the time delay mismatch value and performing natural logarithmic operation. The input support term is then divided by the sum of the interference diffusion term and the hysteresis amplification to obtain a dimensionless cognitive separation value. Thus, under a unified dimensional framework, the degree of separation clarity of the cognitive input signal relative to distraction interference and behavioral hysteresis is quantitatively characterized.

[0057] Specifically, the steps for performing state determination, segment classification, and state marking operations based on the separation degree analysis results are as follows: By comparing the cognitive separation value with the separation threshold in real time, where the separation threshold is a pre-defined judgment boundary based on the statistical distribution of cognitive separation values ​​of learners of the same grade on clearly inputted segments, when the cognitive separation value is greater than or equal to the separation threshold, the current cognitive segment package is marked as a clearly inputted segment, maintaining the boundaries of the cognitive segment package and the segment order unchanged; when the cognitive separation value is less than the separation threshold, the current cognitive segment package is marked as a state-entangled segment, and the following are extracted from n consecutive cognitive segments before and after the current question: the ratio of the frequency of eyebrow contraction, the ratio of the duration of mouth hesitation pauses, the ratio of the frequency of non-task gaze deviation, the keyword repetition matching ratio, the ratio of consecutive incorrect answers, and the ratio of the longest pause without answering. The n consecutive cognitive segments are the nearest cognitive segment packages belonging to the current question, taken forward and backward on the time axis, centered on the current cognitive segment package. The sliding window formed by the problem has a preset positive integer value for n, which does not exceed the total number of cognitive segment packages generated for the current question. The median value of each ratio parameter in the n consecutive cognitive segments is calculated. If the ratio of the frequency of eyebrow contraction is greater than the median value of the frequency of eyebrow contraction, and the ratio of the duration of mouth hesitation is greater than the median value of the duration of mouth hesitation, and the ratio of the frequency of non-task gaze deviation is less than the median value of the frequency of non-task gaze deviation, then it is marked as a tense segment. If the ratio of the frequency of non-task gaze deviation is greater than the median value of the frequency of non-task gaze deviation, and the ratio of the keyword paraphrasing matching is less than the median value of the keyword paraphrasing matching, then it is marked as a distracted segment. If the ratio of consecutive incorrect answers is greater than the median value of the ratio of consecutive incorrect answers, and the ratio of the longest pause without answering is greater than the median value of the ratio of the longest pause without answering, then it is marked as a stagnant segment. The remaining cognitive segment packages are marked as mixed segments, and the segment labels and cognitive separation values ​​are sent to the intervention triggering process together.

[0058] This implementation plan clearly distinguishes between clearly focused segments and entangled segments by comparing cognitive separation values ​​and separation thresholds in real time. It explicitly defines the statistical window as the eight most recent cognitive segments for the current learner on the same topic, unifying the median calculation standard and resolving the issue of non-reproducible results caused by an unclear median statistical window. Simultaneously, by extracting multiple feature ratios and calculating corresponding medians, it accurately classifies and labels entangled segments, clearly distinguishing four segment types: tension, distraction, stagnation, and mixed. This standardizes the segment classification criteria, ensuring consistency and repeatability in the statistics and classification of various feature parameters. This effectively achieves accurate identification and classification of cognitive segments, providing a clear and traceable basis for subsequent intervention triggering processes. It further improves the quantitative judgment system for cognitive states, ensuring the targetedness and accuracy of subsequent intervention strategies.

[0059] Specifically, the steps for intervention trigger determination based on cognitive separation value, behavioral accumulation, and time delay mismatch data are as follows: Obtain the cognitive separation value, consecutive incorrect answer ratio, longest pause without answering ratio, and time delay mismatch value for the j-th cognitive segment package; subtract the cognitive separation value of the current cognitive segment package from the cognitive separation value of the previous cognitive segment package; when the difference is greater than zero, use the difference as the state drop value to characterize the decrease in cognitive input state from the previous cognitive segment package to the current cognitive segment package; when the difference is less than or equal to zero, the state drop value is zero, indicating no decline in cognitive input; increment the cognitive separation value by one and then perform a reciprocal. The calculation yields the current clarity compensation term, where adding one serves as a smoothing constant to avoid abnormal reciprocal calculations when the cognitive dissociation value is zero. Furthermore, the current clarity compensation term has a larger value when the cognitive dissociation value is low and a smaller value when the cognitive dissociation value is high, thus exerting a greater intervention tendency on low-input states. The state drop value, the current clarity compensation term, and one are summed to obtain the state aggregation term, where adding one serves as a baseline constant, ensuring that the state aggregation term maintains a non-zero base greater than one when there is no drop and the cognitive dissociation value is high. One is added to the ratio of consecutive incorrect answers to obtain the error accumulation term, which also serves as a smoothing constant to ensure that the error accumulation term is not less than one. The ratio of the longest no-response pause duration is incremented by two, and then a natural logarithmic operation is performed to obtain the stagnation amplification. The increment of two serves as a smoothing constant to ensure that the natural logarithmic input domain is greater than one, resulting in a positive output value. The natural logarithmic transformation compresses the high stagnation interval, preventing excessive expansion of the numerator aggregation value when the ratio of the longest no-response pause duration is extremely large, thereby improving the numerical stability of the intervention readiness value. The state aggregation term, error accumulation term, and stagnation amplification are multiplied to obtain the numerator aggregation value. This multiplicative aggregation creates a combined amplifying effect on the four intervention triggers: state drop, insufficient current clarity, continuous error accumulation, and prolonged stagnation. In other words, the prominence of any one of these triggers can enhance the effect. The intervention readiness value is significantly increased. The delay mismatch value is incremented by one to obtain the denominator adjustment value, where incrementing avoids division by zero, and the denominator adjustment value increases with the increase of the delay mismatch value. This suppresses the intervention readiness value under high cross-modal mismatch conditions, preventing false intervention triggering due to poor signal quality. The numerator aggregation value is divided by the denominator adjustment value to obtain the intervention readiness value, which is a dimensionless index. Its value comprehensively reflects the urgency of the current cognitive segment package entering the intervention process. Furthermore, the denominator adjustment mechanism actively suppresses high-delay mismatch segments, ensuring the rationality and stability of intervention triggering under behavioral semantic synchronization conditions.

[0060] The specific formula for calculating the intervention readiness value is as follows:

[0061] ;

[0062] In the formula, This represents the intervention readiness value, used to characterize the urgency of the current learning segment entering the intervention process; This represents the state drop term, used to characterize the degree of decline in engagement state between the current learning segment and the previous learning segment; This represents the cognitive separation value, reflecting the degree of separation between the cognitive engagement signal and the distraction signal, tension signal, and time delay mismatch signal. The percentage of consecutive incorrect answers reflects the degree of continuous difficulty a learner encounters in the current question set. The ratio of the longest pause without a response reflects the degree of stagnation of learners in the current learning segment; It represents the time delay mismatch value, reflecting the degree of offset and dispersion between multimodal events and semantic events within the current learning segment.

[0063] In this implementation plan, the intervention readiness value is generated by calculating the state drop value, current clarity compensation item, error accumulation item, stagnation amplification amount and denominator adjustment value, and then generating the value through multiplication aggregation and division operations. Under the dimensionless framework, this value comprehensively reflects the combined effect of cognitive input decline, current input insufficiency, continuous error accumulation, stagnation extension and time delay mismatch inhibition on the urgency of intervention, thereby providing a unified quantitative criterion for triggering subsequent graded interventions.

[0064] Specifically, the steps for implementing tiered intervention, closed-loop verification, and continuous follow-up based on the intervention determination results are as follows: By comparing the intervention readiness value and the intervention threshold, such as... Figure 4This is a flowchart illustrating the hierarchical intervention process for the learning companion robot's cognitive state in this embodiment. The intervention thresholds include a first intervention threshold T1 and a second intervention threshold T2. These thresholds are pre-defined hierarchical boundaries based on the distribution of intervention readiness values ​​among learners of the same grade and feedback data on intervention effects. When the intervention readiness value < T1, the boundaries of the current cognitive segment, the order of question presentation, and the length of each prompt remain unchanged. When T1 ≤ intervention readiness value < T2, a first-level intervention is executed: for tense segments, the previous key question is replayed at a slower pace and the length of each prompt is shortened; for distraction segments, the display module highlights key lines on the question and simultaneously enlarges the keyword font; for stagnant segments, a one-step prompt text switch is executed and irrelevant prompt content is hidden; for mixed segments, the key conditional sentence is replayed and a one-sentence prompt is superimposed. When the intervention readiness value ≥ T2, a second-level intervention is executed, changing the current learning content... The process is divided into two consecutive steps. First, the key sentence and the first-step prompt are displayed. After re-acquiring the signal of the start time of writing or the start time of the first answer, the second-step prompt is displayed. The keyword repetition matching ratio, non-task gaze deviation frequency ratio, consecutive incorrect answer ratio, and longest pause without answering ratio in the next two cognitive segment packages are continuously tracked. The corresponding values ​​before the intervention are the values ​​of the corresponding parameters in the most recent cognitive segment package before the intervention is triggered. When the keyword repetition matching ratio in the next two cognitive segment packages is less than their respective corresponding values ​​before the intervention, and the non-task gaze deviation frequency ratio, consecutive incorrect answer ratio, and longest pause without answering ratio are all greater than their respective corresponding values ​​before the intervention, it indicates that the learner's input indicators have not shown an improvement trend after the secondary intervention, and the cognitive difficulty continues. In this case, the next cognitive segment package is marked as a continuous intervention segment, and the secondary intervention path is continued.

[0065] In this implementation plan, the intervention readiness value is compared with the first and second intervention thresholds in a hierarchical manner. In the low readiness interval, the boundaries of the current cognitive segment package and the prompting strategy remain unchanged. In the medium readiness interval, first-level intervention operations such as slowing down the playback of the question stem, highlighting the key lines, switching the prompts one step at a time, or overlaying the key conditional sentences are executed according to the tension, distraction, stagnation, or mixed state. In the high readiness interval, a second-level intervention of step-by-step presentation of learning content is triggered, and prompts are advanced after the writing or answering signal is re-acquired. At the same time, the keyword repetition matching ratio, the non-task gaze deviation frequency ratio, the ratio of consecutive wrong answers, and the ratio of the longest pause without answering are continuously tracked for the next two cognitive segment packages. When the direction of change of the above four ratios indicates that the input indicators have not improved after the intervention, the next cognitive segment package is marked as a continuous intervention segment and the second-level intervention path is continued, thereby realizing the hierarchical intervention triggering and closed-loop tracking of intervention effects based on cognitive separation assessment.

[0066] It should be noted that, in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article, or apparatus.

[0067] The preferred embodiments of the present invention disclosed above are merely illustrative of the invention. These preferred embodiments do not exhaustively describe all details, nor do they limit the invention to the specific implementations described. Clearly, many modifications and variations can be made based on the content of this specification. This specification selects and specifically describes these embodiments to better explain the principles and practical applications of the invention, thereby enabling those skilled in the art to better understand and utilize the invention. The invention is limited only by the claims and their full scope and equivalents.

Claims

1. A method for intervening in the learning behavior of a learning companion robot based on AI multimodal fusion, characterized in that, Includes the following steps: S1. Collect learning behavior intervention data of the learning companion robot, preprocess the learning behavior intervention data of the learning companion robot, store it, and build a learning companion behavior database; S2 performs cross-modal time delay mismatch analysis using multimodal event time sequence difference data, and performs segment resegmentation, segment boundary correction, and segment writing operations based on the time delay analysis results; S3, based on synchronous fragments, corrected fragments and behavioral cue data, reorganizes cognitive fragments, constructs cognitive fragment packages and establishes a fragment index; S4 integrates multimodal behavior ratios and time delay mismatch data to perform cognitive input separation analysis, and performs state determination, segment classification and state labeling operations based on the separation analysis results; S5 uses cognitive separation value, behavioral accumulation and time delay mismatch data to determine intervention triggers, and performs graded intervention, closed-loop verification and continuous tracking operations based on the intervention determination results.

2. The method for intervening in the learning behavior of a learning companion robot based on AI multimodal fusion according to claim 1, characterized in that: The specific steps for collecting intervention data on the learning behavior of the learning companion robot are as follows: Data on learning behavior intervention of the learning companion robot is collected: facial video stream, gaze movement video stream, and head posture video stream are acquired through a forward-facing camera; writing action video stream is acquired through a desktop-view camera; speech stream is acquired through a sound pickup unit; and the question display time, question stem broadcast start time, question stem broadcast end time, prompt trigger time, and answer submission time are acquired through the question presentation record link. The total number of target keywords for the current question is obtained from the question presentation record link or question bank records. The number of times the current question group has been answered is counted from the question submission records and answer result records. The number of frame groups in the current segment is counted from the video slice results of the current learning segment.

3. The method for intervening in the learning behavior of a learning companion robot based on AI multimodal fusion according to claim 1, characterized in that: The specific steps for preprocessing and storing the learning behavior intervention data of the learning companion robot and constructing the learning companion behavior database are as follows: The system performs facial landmark localization, gaze trajectory tracking, and head pose change extraction on the facial video stream. Based on the displacement of the eyebrow key point, the change of the dense optical flow between the eyebrows, and the response of the eyebrow action unit, the number of eyebrow contractions is extracted to obtain the duration of gaze on the task, the peak time of gaze fallback, the peak time of head return, the number of non-task gaze deviations, and the number of eyebrow contractions. Pen tip recognition, trajectory continuity extraction, and pen starting point detection are performed on the desktop-view video stream to obtain the continuous writing duration and pen start time. Speech activity detection, question stem alignment, keyword matching, and pause detection are performed on the speech stream. Combined with the results of lip key point movement changes, speech activity detection results, and pause duration thresholds, the semantic segment end time, semantic segment duration, first answer start time, number of keyword paraphrased matching words, and mouth hesitation pause duration are obtained. Sequence alignment and segment association are performed on the question submission record and answer result record to obtain the number of consecutive incorrect answers and the longest pause without answering. Based on the semantic segment duration, the total number of target keywords in the current question, the number of frames in the current segment, and the number of times the current question group has been answered, the following ratios are obtained: question gaze dwell time ratio, continuous writing duration ratio, mouth hesitation pause duration ratio, longest pause without answering duration ratio, keyword paraphrasing matching ratio, non-task gaze deviation frequency ratio, glabellar contraction frequency ratio, and consecutive incorrect answer ratio. After performing range normalization, outlier removal, and missing value completion on all ratio parameters, the learning companion behavior database is stored and constructed according to learner ID, subject ID, question ID, and segment ID. A cognitive segment record table is then constructed within the learning companion behavior database.

4. The method for intervening in the learning behavior of a learning companion robot based on AI multimodal fusion according to claim 1, characterized in that: The specific steps for cross-modal time delay mismatch analysis using multimodal event time series difference data are as follows: Obtain the multimodal event temporal difference data for the j-th learning segment. The multimodal event temporal difference data includes the semantic segment end time, the peak time of gaze fallback in the question, the peak time of head return to center, the pen start time, the first answer start time, and the duration of the semantic segment. The absolute value of the difference between the peak time of gaze fallback and the end time of the semantic segment is divided by the duration of the semantic segment to obtain the gaze semantic time difference ratio; the absolute value of the difference between the peak time of head alignment and the end time of the semantic segment is divided by the duration of the semantic segment to obtain the posture semantic time difference ratio; the absolute value of the difference between the start time of pen placement and the end time of the semantic segment is divided by the duration of the semantic segment to obtain the writing semantic time difference ratio; the absolute value of the difference between the start time of the first response and the end time of the semantic segment is divided by the duration of the semantic segment to obtain the response semantic time difference ratio; the maximum value of the four time difference ratios is obtained by finding the maximum value, and the minimum value is obtained by finding the minimum value; the four time difference ratios are added by one and then multiplied, and the fourth root operation is performed on the product to obtain the overall time difference aggregate value; the difference between the maximum and minimum time difference ratios is added by one to obtain the time difference discrete value; the overall time difference aggregate value and the time difference discrete value are multiplied to obtain the time delay mismatch value of the j-th learning segment.

5. The method for intervening in the learning behavior of a learning companion robot based on AI multimodal fusion according to claim 1, characterized in that: The specific steps for performing segment resegmentation, segment boundary correction, and writing of segments to be reviewed based on the latency analysis results are as follows: By comparing the time delay mismatch value with the mismatch threshold in real time, when the time delay mismatch value is less than the mismatch threshold, the current learning segment is marked as a synchronous segment and directly sent into the cognitive segment reorganization process; When the time delay mismatch value is greater than or equal to the mismatch threshold, the current learning segment is marked as a mismatch segment. The earliest of the following times is used as the left boundary: the peak time of gaze fallback, the peak time of head return, and the time of pen start. The latest of the following times is used as the right boundary: the end time of semantic segment and the start time of first answer. The current learning segment is then re-segmented. A voice command for retrieving the previous key question stem is generated. The key question stem is a sentence in the question stem text that contains the target keyword of the current question. Continuous short segments between re-gazing at the question and re-writing are collected. At the same time, learning segments adjacent to the same question with a time delay mismatch value less than the mismatch threshold are called as time delay references to correct the boundary of the current learning segment. Behavioral parameters are re-extracted based on the corrected segment boundary. If the recalculated time delay mismatch value is still greater than or equal to the mismatch threshold, the current learning segment is marked as a segment for deferred judgment, written into the manual review sequence, and a review prompt data package containing the segment start and end time, time delay mismatch value, ratio of consecutive incorrect answers, and ratio of the longest pause without answering is generated.

6. The method for intervening in the learning behavior of a learning companion robot based on AI multimodal fusion according to claim 1, characterized in that: The specific steps for reorganizing cognitive fragments based on synchronized fragments, corrected fragments, and behavioral cue data, constructing cognitive fragment packages, and establishing a fragment index are as follows: The system receives synchronized and corrected segments, and establishes a segment anchor sequence based on the semantic segment end time, the first answer start time, the peak time of gaze fallback in the question text, and the pen start time. It then categorizes the ratio of gaze dwell time in the question text, the ratio of continuous writing time, and the ratio of keyword paraphrasing matching into the input candidate segment set; the ratio of non-task gaze deviation frequency into the distraction candidate segment set; the ratio of glabellar contraction frequency and the ratio of mouth hesitation pause duration into the tension candidate segment set; and the ratio of consecutive incorrect answers and the longest pause without answering into the stagnation candidate segment set. Finally, it selects learning segments with a latency mismatch value less than the mismatch threshold from adjacent learning segments of the same question as the main anchor segment, using the main anchor segment as the center. When the interval between segments is less than the splicing limit and the anchor point offset is less than the offset limit, forward splicing is performed on the front side of the main anchor segment, and backward splicing is performed on the back side of the main anchor segment. The middle p% of the total duration of the spliced ​​cognitive segment package is taken as the segment center segment. The input cues that overlap within the same anchor point interval are written into the segment center segment. After deducting the segment center segment from the total duration of the cognitive segment package, the remaining part is divided equally before and after as the segment edge segment. Distraction cues, tension cues, and stagnation cues are written into the segment edge segment first to form a cognitive segment package. A segment index is established for the recombined cognitive segment package, and the segment is written into the cognitive segment record table according to the learner number, question number, segment number, and segment sequence number.

7. The method for intervening in the learning behavior of a learning companion robot based on AI multimodal fusion according to claim 1, characterized in that: The specific steps for analyzing cognitive input separation by fusing multimodal behavior ratios and time delay mismatch data are as follows: Obtain the following ratios for the j-th cognitive segment: gaze duration ratio, continuous writing duration ratio, keyword paraphrase matching ratio, non-task gaze deviation frequency ratio, glabellar contraction frequency ratio, mouth hesitation / pause duration ratio, and time delay mismatch value. Add one to each of the following ratios and perform cubic root aggregation to obtain the input support term. Add one to each of the non-task gaze deviation frequency ratio and glabellar contraction frequency ratio, and perform square root aggregation to obtain the interference diffusion term. Sum the mouth hesitation / pause duration ratio and time delay mismatch value, add two, and perform natural logarithmic calculation to obtain the hysteresis amplification value. Divide the input support term by the interference diffusion term and the hysteresis amplification value, and sum them to obtain the cognitive dissociation value.

8. The method for intervening in the learning behavior of a learning companion robot based on AI multimodal fusion according to claim 1, characterized in that: The specific steps for performing state determination, fragment classification, and state labeling operations based on the separation degree analysis results are as follows: By comparing the cognitive separation value with the separation threshold in real time, when the cognitive separation value is greater than or equal to the separation threshold, the current cognitive segment package is marked as a clear segment, and the boundaries of the cognitive segment package and the segment order remain unchanged. When the cognitive dissociation value is less than the dissociation threshold, the current cognitive segment is marked as a state-entangled segment. From the n consecutive cognitive segments before and after the current question, the following ratios are extracted: the frequency ratio of eyebrow contractions, the frequency ratio of mouth hesitation / pause duration, the frequency ratio of non-task gaze deviation, the keyword paraphrasing matching ratio, the ratio of consecutive incorrect answers, and the ratio of the longest pause without answering. The corresponding medians are calculated for each. If the frequency ratio of eyebrow contractions is greater than the median of the frequency ratio of eyebrow contractions, and the frequency ratio of mouth hesitation / pause duration is greater than the median of the frequency ratio of mouth hesitation / pause duration, and the frequency ratio of non-task gaze deviation... If the ratio of non-task gaze deviation frequency is less than the median value, it is marked as a tense segment; if the ratio of non-task gaze deviation frequency is greater than the median value, and the keyword paraphrase matching ratio is less than the median value, it is marked as a distracted segment; if the ratio of consecutive incorrect answers is greater than the median value, and the ratio of the longest pause without answering is greater than the median value, it is marked as a stagnant segment; the remaining cognitive segments are marked as mixed segments, and the segment labels and cognitive dissociation values ​​are sent together to the intervention triggering process.

9. The method for intervening in the learning behavior of a learning companion robot based on AI multimodal fusion according to claim 1, characterized in that: The specific steps for determining intervention triggers based on cognitive separation values, behavioral accumulation, and time delay mismatch data are as follows: Obtain the cognitive separation value, consecutive incorrect answer ratio, longest pause duration without answering ratio, and time delay mismatch value of the j-th cognitive segment package; subtract the cognitive separation value of the current cognitive segment package from the cognitive separation value of the previous cognitive segment package; when the difference is greater than zero, use the difference as the state drop value; when the difference is less than or equal to zero, set the state drop value to zero. Add one to the cognitive separation value and then perform a reciprocal operation to obtain the current clarity compensation term; sum the state drop value, the current clarity compensation term, and one to obtain the state aggregation term; Add one to the ratio of consecutive incorrect answers to obtain the error accumulation term; add two to the ratio of the longest pause without answering and then perform a natural logarithmic operation to obtain the stagnation amplification; multiply the state aggregation term, the error accumulation term, and the stagnation amplification to obtain the numerator aggregation value; add one to the time delay mismatch value to obtain the denominator adjustment value; divide the numerator aggregation value by the denominator adjustment value to obtain the intervention readiness value.

10. The method for intervening in the learning behavior of a learning companion robot based on AI multimodal fusion according to claim 1, characterized in that: The specific steps for implementing tiered intervention, closed-loop verification, and continuous follow-up based on the intervention determination results are as follows: By comparing the intervention readiness value and the intervention threshold, which includes T1 and T2, when the intervention readiness value < T1, the current cognitive segment package boundary, the question stem broadcast order and the single prompt length remain unchanged; When T1 ≤ Intervention Readiness Value < T2, Level 1 Intervention is executed. For tense segments, the previous key question stem is replayed at a slower pace and the length of the single prompt text is shortened. For distracted segments, the control display module highlights the key lines of the question and simultaneously enlarges the font of the keywords. For stagnant segments, a one-step prompt text switch is executed and irrelevant prompt content is hidden. For mixed segments, the key conditional sentence is replayed and a one-sentence prompt is superimposed. When the intervention readiness value is ≥ T2, a secondary intervention is implemented. The current learning content is divided into two consecutive steps. The key sentence and the first step prompt are displayed first. After the signal of the writing start time or the first answer start time is reacquired, the second step prompt is displayed. The keyword repetition matching ratio, non-task gaze deviation frequency ratio, consecutive wrong answer ratio, and longest no-answer pause duration ratio are continuously tracked in the next two cognitive segment packages. When the keyword repetition matching ratio in the next two cognitive segment packages is less than the corresponding value before the intervention, and the non-task gaze deviation frequency ratio, consecutive wrong answer ratio, and longest no-answer pause duration ratio are all greater than the corresponding value before the intervention, the next cognitive segment package is marked as a continuous intervention segment, and the secondary intervention path is continued.

Citation Information

Patent Citations

  • Human Behavior Recognition Method Based on Interleaved Reinforced Attention Network

    CN112307982B

  • Intelligent decision-making methods and devices based on multimodal data fusion and reinforcement learning

    CN114860893B