Multi-terminal online tutoring collaborative management system and method based on learning behavior analysis

CN122887656APending Publication Date: 2026-10-09HAIDAO (SHENZHEN) EDUCATION TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611392323.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-09-09
Publication Date
2026-10-09

AI Technical Summary

Technical Problem

[0003]然而,现有的在线辅导通常只实现基础的屏幕共享或单向操作录制,缺乏对用户跨终端学习行为的深度感知与理解能力,使得在跨终端交互场景下,因注意力碎片化导致的隐蔽性操作障碍无法被及时感知与定位的问题,即当学习方在操作过程中遇到困难,例如在编程调试、复杂公式推导或软件操作环节出现反复无效的尝试、频繁的焦点切换、长时间停滞后又快速回退等行为时,现有的系统既无法自动识别这种“操作困扰”状态,也无法从混乱的交互模式中准确定位造成困扰的具体内容对象和操作序列

Benefits of technology

本申请提供的基于学习行为分析的多终端在线辅导协同管理系统及方法,可在注意力碎片化导致的隐蔽性操作障碍影响下实现多终端在线辅导协同;首先,将多终端持续数据流监测与焦点对象属性提取作为整个协同管理流程的触发起点,通过在交互层实时采集原始交互事件流并进行连续性检测,能够在会话进行过程中动态判断各终端是否处于用户持续操作状态,避免对闲置或断连终端的无效分析,同时提取的焦点对象属性特征使得用户当前操作的界面元素从无意义的屏幕坐标点被提升为具有业务语义与空间位置双重属性的可计算实体;其次,以时间戳序列为索引、以焦点对象属性特征为内容构建的注意力焦点流,可将分布在多个终端上的离散交互事件串联为表达注意力跨设备跳转的全局时序图,在此基础上通过滑动时间窗口下的终端切换频率与焦点对象语义跳变率双维度联合分析,能够从物理设备切换的异常节奏和操作内容语义的碎片化程度两个层面综合判定用户是否陷入操作困扰,并通过偏离指数序列的阈值分割截取出非稳态交互特征流,将用户在多终端协同中因界面认知混乱或操作路径不明确而表现出的异常行为模式从正常交互流中剥离出来。然后,通过从非稳态交互特征流回溯原始交互事件并提取未产生有效结果输入的重复性操作序列,实现了对操作困扰行为从宏观切换模式到微观操作内容的逐层聚焦,操作载荷相似度计算与辅导会话状态机有效迁移校验构成的双重验证机制,确保提取出的重复操作是真正未推动任务进展的无效交互而非正常的多次订正尝试,再结合问题发生窗口的界面上下文快照与焦点对象匹配,将用户反复尝试但失败的界面元素及其辅导语义信息与具体操作证据封装为协同请求,使得系统能够以结构化数据包的形式精确描述用户当前遇到的具体问题与行为表现。最后,通过会话控制信令解析与辅导方交互行为活跃度检测实现协同对象的智能识别,避免了在多辅导方并发在线场景下协同请求的广播式分发造成的干扰,基于交互行为活跃度的筛选机制能够将协同请求精准投递至当前正在积极参与辅导的辅导方终端,同时将非稳态交互特征流与协同请求一同推送展示,使得辅导方在接收到协同通知时能够同步了解用户陷入困扰的严重程度与持续时间,进而根据困扰程度合理调整介入时机与辅导策略,形成从自动感知到精准投递再到信息充分展示的完整协同闭环;综上所述,本申请提供的技术方案可在注意力碎片化导致的隐蔽性操作障碍影响下实现多终端在线辅导协同。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122887656A_ABST
    Figure CN122887656A_ABST
Patent Text Reader

Abstract

The application provides a multi-terminal online tutoring collaborative management system and method based on learning behavior analysis, relates to the technical field of intelligent education management, and the method comprises the following steps: extracting the timestamp sequence of each terminal interaction behavior and the attribute characteristics of the corresponding focus object; determining an attention focus stream for describing the jumping of user attention among multiple terminals, and extracting a non-steady-state interaction feature stream from the attention focus stream; extracting a repetitive operation sequence from the non-steady-state interaction feature stream, and determining a collaborative request carrying specific problem context and behavior evidence; identifying the tutoring party terminal being collaborated in the current tutoring session according to the collaborative request, and pushing the collaborative request and the corresponding non-steady-state interaction feature stream summary to the identified tutoring party terminal for display. The technical scheme provided by the application can realize multi-terminal online tutoring collaboration under the influence of hidden operation obstacles caused by attention fragmentation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of intelligent education management technology, and more specifically, to a multi-terminal online tutoring collaborative management system and method based on learning behavior analysis. Background Technology

[0002] With the popularization of distance education and blended learning, learners and tutors are often in a multi-device parallel operating environment. For example, learners may simultaneously watch instructional videos on a tablet, operate a programming environment or exercise system on a laptop, while the tutor's terminal may simultaneously display courseware, a whiteboard, and a mirror image of the learner's screen. In this multi-device collaborative learning scenario, the user's attention frequently jumps between multiple devices and multiple interface areas, forming a complex flow of interactive behaviors.

[0003] However, existing online tutoring typically only achieves basic screen sharing or one-way operation recording, lacking the ability to deeply perceive and understand users' cross-device learning behavior. This makes it difficult to promptly detect and locate hidden operational obstacles caused by fragmented attention in cross-device interaction scenarios. Specifically, when learners encounter difficulties during operation, such as repeated ineffective attempts, frequent focus switching, or rapid rewinding after prolonged pauses in programming debugging, complex formula derivation, or software operation, existing systems cannot automatically identify this "operational distress" state, nor can they accurately locate the specific content object and operation sequence causing the distress from the chaotic interaction pattern. This leads to two significant technical problems: First, tutors struggle to capture the truly difficult behavioral signals requiring intervention in real time from the vast amount of information on the learner's screen, often relying on the learner's proactive questions or their own experience to judge, resulting in delayed and unpredictable responses. Second, even if tutors perceive potential difficulties, they lack the ability to reconstruct the complete context of the learner's attention shifts across multiple devices, making it difficult to obtain a snapshot of the problem with precise behavioral evidence that can be directly used for diagnosis. This has resulted in inaccurate allocation of tutoring resources and low efficiency in collaborative intervention.

[0004] Therefore, there is an urgent need for a technical solution that can automatically capture the attention focus shifts between multiple terminals, identify the non-steady-state interaction patterns of users in operational difficulties, and automatically generate collaborative requests carrying problem context and behavioral evidence. Finally, it can proactively push a refined summary of the difficulties to the tutor, so as to achieve multi-terminal online tutoring collaboration under the influence of hidden operational obstacles caused by attention fragmentation, thereby improving the perception sensitivity and collaborative response efficiency of remote tutoring. Summary of the Invention

[0005] This application provides a multi-terminal online tutoring collaborative management system and method based on learning behavior analysis, which can realize multi-terminal online tutoring collaboration under the influence of hidden operational obstacles caused by attention fragmentation.

[0006] Firstly, this application provides a multi-terminal online tutoring collaborative management method based on learning behavior analysis, comprising the following steps: Extract the timestamp sequence of each terminal's interactive behavior and the attribute features of the corresponding focus object in the current tutoring session; Using the timestamp sequence of each terminal as an index, an attention focus flow is constructed based on the attribute features of the corresponding focus object to show the user's attention jumping between multiple terminals. The attention focus flow is then analyzed based on the interaction pattern of sliding time window to obtain the non-steady-state interaction feature flow when the user is in an operational confusion state. Extract repetitive operation sequences that do not produce valid input results from the non-steady-state interaction feature stream, and generate collaborative requests carrying specific problem context and behavioral evidence based on the repetitive operation sequences and the content objects in the interface where they occur; The collaboration request identifies the tutoring terminal currently collaborating in the tutoring session, and the collaboration request and the corresponding non-steady-state interaction feature stream summary are pushed to the identified tutoring terminal for display.

[0007] In some embodiments, identifying the tutoring terminal that is currently collaborating in the tutoring session based on the collaboration request specifically includes: Obtain the session control signaling record of the current tutoring session, parse the role tags of each terminal that has joined the current tutoring session from the session control signaling record, and form a terminal role mapping relationship; The system queries the terminal role mapping relationship for terminals with the role tag "tutor" to form a candidate tutor terminal set. The system then performs real-time detection on the online status of each terminal in the candidate tutor terminal set and filters out online tutor terminals that are currently in a session connection state. The activity level of the online tutoring terminal's interactive behavior within a preset time window is detected, and online tutoring terminals whose interactive behavior activity exceeds a preset assistance threshold are identified as tutoring terminals that are currently collaborating.

[0008] In some embodiments, extracting the timestamp sequence of each terminal interaction behavior and the attribute features of the corresponding focus object in the current tutoring session specifically includes: The system collects the original interactive event streams of all terminals with established session connections in the tutoring session in real time at the interaction layer, and performs continuity detection on the original interactive event streams to obtain the data stream status identifier of each terminal. When the data stream status identifier is characterized as a continuous stream, the interactive behavior record carrying the operation timing is parsed from the original interactive event stream of the corresponding terminal, and the interactive behavior record is organized into a timestamp sequence according to the order of occurrence. For each interaction in the timestamp sequence, determine the active interface element in the tutorial interface when the current interaction occurs, mark the interface element as the focus object, and extract the node type identifier, semantic tag and spatial coordinates of the focus object in the document object model to form the attribute features of the corresponding focus object.

[0009] In some embodiments, performing continuity detection on the original interactive event stream to obtain the data stream status identifier of each terminal specifically includes: Using a preset time window as the granularity of data collection, the original interactive event stream is segmented and the event arrival rate of interactive events within each time window is calculated. The event arrival rate is compared with a preset silence threshold. When the event arrival rate is lower than the silence threshold, the time window is marked as a silent window; otherwise, it is marked as an active window. The alternating distribution characteristics of the silent window and the active window within the sliding observation period are determined, and the continuity of the data stream is determined based on the alternating distribution characteristics. If there are no continuous silent windows within the sliding observation period and the proportion of active windows exceeds a preset ratio, the data stream status of the corresponding terminal is determined as a continuous stream.

[0010] In some embodiments, constructing an attention focus flow that allows user attention to jump between multiple terminals based on the attribute features of the corresponding focus object, using the timestamp sequence of each terminal as an index, specifically includes: Align the timestamp sequence of each terminal with the attribute features of the corresponding focus object according to the timestamp to form a unified timeline record carrying terminal identifier and focus attribute information; The unified timeline records are merged from multiple sources, and the focus objects from different terminals are concatenated into a global focus transfer sequence according to the order in which the interaction occurs. Extract the terminal identifiers of adjacent focus objects at each focus switch in the global focus transfer sequence to form directed transfer pairs that express attention jumping across devices between terminals, and organize all directed transfer pairs into an attention focus stream in chronological order.

[0011] In some embodiments, performing interaction pattern analysis based on time window sliding on the attention focus flow to obtain the non-steady-state interaction feature flow when the user is in an operational distress state specifically includes: A sliding time window of a preset step size is applied to the attention focus stream, and the directed transition pair sub-sequences within each sliding window are extracted sequentially. The terminal switching frequency and the semantic jump rate of the focus object between adjacent directed transition pairs in the directed transition pair sub-sequences are calculated. The terminal switching frequency and the semantic jump rate of the focus object are input into a pre-constructed steady-state deviation metric function, and the deviation index, which characterizes the degree of deviation of the attention shift mode from the steady state within the current sliding window, is output. The deviation indices of each sliding window on the attention focus flow are arranged in time sequence to obtain a deviation index sequence. Continuous intervals in the deviation index sequence where the deviation index exceeds a preset perturbation threshold are extracted. The directed transition pair sub-sequences corresponding to these continuous intervals are marked as non-steady-state interaction feature flows in which the user is in an operational distress state.

[0012] In some embodiments, extracting the repetitive operation sequence that does not produce valid input from the non-steady-state interaction feature stream specifically includes: Based on the timestamp interval and terminal identifier carried by each directed transition pair in the non-steady-state interaction feature stream, the original interaction event stream of the corresponding terminal within the timestamp interval is extracted back to form a set of interaction events to be analyzed. From the set of interactive events to be analyzed, identify consecutive clusters of interactive events with the same type of interactive events and the same focus object, and calculate the similarity of the operation loads of each interactive event in the consecutive clusters of interactive events to obtain an operation load similarity sequence. The continuous interactive events and their operational payloads in the operational payload similarity sequence that have a similarity exceeding a preset repetition threshold are organized into candidate repetitive operational sequences according to their occurrence time. Identify whether a valid state transition occurs in the coaching session state machine after each operation in the candidate repetitive operation sequence, and determine the candidate repetitive operation sequence that does not trigger a valid state transition as a repetitive operation sequence that does not produce a valid result input.

[0013] Secondly, this application provides a multi-terminal online tutoring collaborative management system based on learning behavior analysis, used to execute a multi-terminal online tutoring collaborative management method based on learning behavior analysis. The system includes: The acquisition module is used to extract the timestamp sequence of each terminal's interactive behavior and the attribute features of the corresponding focus object in the current tutoring session; The processing module is used to construct an attention focus flow of user attention jumping between multiple terminals based on the attribute features of the corresponding focus object, using the timestamp sequence of each terminal as an index, and to perform interaction pattern analysis on the attention focus flow based on time window sliding to obtain the non-steady-state interaction feature flow when the user is in an operational confusion state. The processing module is also used to extract repetitive operation sequences that have not produced valid input results from the non-steady-state interaction feature stream, and generate collaborative requests carrying specific problem context and behavioral evidence based on the repetitive operation sequences and the content objects in the interface where they occur. The execution module is used to identify the tutoring terminal that is currently collaborating in the tutoring session based on the collaboration request, and push the collaboration request and the corresponding non-steady-state interaction feature stream summary to the identified tutoring terminal for display.

[0014] Thirdly, this application provides a computer device, the computer device including a memory and a processor, the memory storing code, and the processor being configured to acquire the code and execute the above-described multi-terminal online tutoring collaborative management method based on learning behavior analysis.

[0015] Fourthly, this application provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the above-described multi-terminal online tutoring collaborative management method based on learning behavior analysis.

[0016] The technical solutions provided by the embodiments disclosed in this application have the following beneficial effects: The multi-terminal online tutoring collaborative management system and method based on learning behavior analysis provided in this application can achieve multi-terminal online tutoring collaboration under the influence of hidden operational obstacles caused by attention fragmentation. First, the continuous monitoring of multi-terminal data streams and the extraction of focus object attributes are used as the triggering point for the entire collaborative management process. By collecting the original interaction event stream in real time at the interaction layer and performing continuous detection, it is possible to dynamically determine whether each terminal is in a continuous user operation state during the session, avoiding invalid analysis of idle or disconnected terminals. At the same time, the extracted focus object attribute features elevate the interface elements currently operated by the user from meaningless screen coordinates to measurable elements with both business semantics and spatial location attributes. First, it calculates entities; second, it constructs an attention focus flow based on timestamp sequences as indexes and focus object attribute features as content. This can string together discrete interaction events distributed across multiple terminals into a global time sequence graph that expresses attention jumps across devices. Based on this, through joint analysis of terminal switching frequency and focus object semantic jump rate under a sliding time window, it can comprehensively determine whether users are trapped in operation confusion from two aspects: the abnormal rhythm of physical device switching and the degree of fragmentation of operation content semantics. Furthermore, by segmenting the non-steady-state interaction feature flow through threshold segmentation of deviation from the exponential sequence, it can separate the abnormal behavior patterns exhibited by users in multi-terminal collaboration due to interface cognition confusion or unclear operation paths from the normal interaction flow. Then, by tracing back the original interaction events from the non-steady-state interaction feature stream and extracting repetitive operation sequences that did not produce effective input results, the system achieves a layer-by-layer focus on operational distress behaviors from macro-level switching modes to micro-level operation content. The dual verification mechanism, consisting of operation load similarity calculation and effective transition verification of the tutoring session state machine, ensures that the extracted repetitive operations are truly invalid interactions that did not advance the task, rather than normal multiple correction attempts. Furthermore, by combining the interface context snapshot of the problem occurrence window with the focus object matching, the interface elements that the user repeatedly tried but failed to access, along with their tutoring semantic information and specific operation evidence, are encapsulated into a collaborative request. This allows the system to accurately describe the specific problems and behavioral performance currently encountered by the user in the form of structured data packets. Finally, intelligent identification of collaborative objects is achieved through session control signaling parsing and detection of the interactive behavior activity of tutors. This avoids interference caused by the broadcast distribution of collaborative requests in scenarios where multiple tutors are online concurrently. The filtering mechanism based on interactive behavior activity can accurately deliver collaborative requests to the terminals of tutors who are currently actively participating in tutoring. At the same time, the non-steady-state interaction feature stream is pushed and displayed together with the collaborative request, so that when the tutor receives the collaborative notification, it can simultaneously understand the severity and duration of the user's distress. Then, it can reasonably adjust the intervention timing and tutoring strategy according to the degree of distress, forming a complete collaborative closed loop from automatic perception to accurate delivery to full display of information. In summary, the technical solution provided in this application can realize multi-terminal online tutoring collaboration under the influence of hidden operational obstacles caused by attention fragmentation. Attached Figure Description

[0017] Figure 1 This is an exemplary flowchart of a multi-terminal online tutoring collaborative management method based on learning behavior analysis, according to some embodiments of this application; Figure 2 This is a schematic diagram illustrating the principle of the interaction mode analysis shown in some embodiments of this application; Figure 3 This is a schematic diagram of the structure of a multi-terminal online tutoring collaborative management system based on learning behavior analysis, according to some embodiments of this application; Figure 4 This is a schematic diagram of the structure of a computer device that implements a multi-terminal online tutoring collaborative management method based on learning behavior analysis, according to some embodiments of this application. Detailed Implementation

[0018] To better understand the technical solution of this application, the technical solution of this application will be described in detail below with reference to the accompanying drawings and specific embodiments.

[0019] refer to Figure 1 The figure is an exemplary flowchart of a multi-terminal online tutoring collaborative management method based on learning behavior analysis, according to some embodiments of this application. The figure mainly includes the following steps: In step S101, the timestamp sequence of each terminal interaction behavior and the attribute features of the corresponding focus object are extracted in the current tutoring session. The focus object is the content element that is interacted with in the tutoring interface.

[0020] In some embodiments, the extraction of the timestamp sequence of each terminal interaction behavior and the attribute features of the corresponding focus object in the current tutoring session is achieved through the following steps: The system collects the original interactive event streams of all terminals with established session connections in the tutoring session in real time at the interaction layer, and performs continuity detection on the original interactive event streams to obtain the data stream status identifier of each terminal. When the data stream status identifier is characterized as a continuous stream, the interactive behavior record carrying the operation timing is parsed from the original interactive event stream of the corresponding terminal, and the interactive behavior record is organized into a timestamp sequence according to the order of occurrence. For each interaction in the timestamp sequence, determine the active interface element in the tutorial interface when the current interaction occurs, mark the interface element as the focus object, and extract the node type identifier, semantic tag and spatial coordinates of the focus object in the document object model to form the attribute features of the corresponding focus object.

[0021] In specific implementation, firstly, by injecting event listeners into the interaction layer of each terminal's operating system, the system collects raw interactive event streams generated by all terminals with established session connections in real time based on the input event interface provided by the device system. These raw interactive event streams refer to the temporal set of underlying input events triggered by the user on the terminal, such as mouse clicks, keyboard keystrokes, and touch gestures. The system then performs continuity detection on these raw interactive event streams to obtain the data stream status identifier for each terminal. Next, after determining that the data stream status identifier is a continuous stream, the system uses an event parser to extract the timestamp and operation type fields from the raw interactive event streams of the corresponding terminal, based on the metadata structure of the input events. Each event is then encapsulated as an interactive behavior record carrying the operation time sequence. These interactive behavior records are structured data units containing the event occurrence time, operation type, and reference to the original event object. Finally, all interactive behavior records are sorted in ascending order of timestamp field values ​​to form a timestamp sequence. Finally, for each interaction record in the timestamp sequence, the DOM node query interface in the browser developer protocol is called to obtain the interface elements in the current interaction state that are in the focused state or mouse hover state at the time of the current interaction. The interface elements in the active state refer to the interactive interface elements that have the focus cursor, are highlighted, or are directly pointed to by the mouse pointer. The interface element is marked as the focus object. By traversing the node path of the focus object in the document object model tree, the node type identifier is extracted to distinguish the element types such as buttons, input boxes, and text paragraphs. Semantic tags are extracted to obtain the ARIA role attributes or text content identifiers bound to the element. Spatial coordinates are extracted to obtain the boundary rectangle position of the element in the interface coordinate system. The node type identifier, semantic tags, and spatial coordinates together constitute the attribute features of the focus object.

[0022] It should be noted that, in this application, the attribute features of the focus object refer to a set of structured attributes used to uniquely describe the interactive semantics and spatial location of the focus object. The focus object refers to the interactive content element that the user directly interacts with on the tutoring interface during the tutoring session through mouse hovering, clicking to focus, touch pressing, or switching using the keyboard Tab key. This content element corresponds to a specific DOM node in the Document Object Model tree, including specific forms such as buttons, text input boxes, drop-down selectors, graphical question options, code editing areas, or draggable teaching aid components. The focus object is determined in real time through the focus state identifier provided by the browser rendering engine or the current focus node reference returned by the system-level accessibility interface. Determining the attribute features of the focus object can provide structured metadata for quantifying user interaction intentions in the construction of cross-terminal attention focus streams. That is, the three types of attribute features provided by the attribute features of the focus object together constitute a complete characterization of the user's current attention focus, enabling discrete interactive events generated on different terminals to be aligned and compared under a unified semantic-spatial dimension, providing a computable feature basis for the construction of directed transfer pairs in the attention focus stream and the subsequent identification of non-steady-state interaction patterns.

[0023] In some embodiments, the continuity detection of the original interactive event stream to obtain the data stream status identifier of each terminal is achieved through the following steps: Using a preset time window as the granularity of data collection, the original interactive event stream is segmented and the event arrival rate of interactive events within each time window is calculated. The event arrival rate is compared with a preset silence threshold. When the event arrival rate is lower than the silence threshold, the time window is marked as a silent window; otherwise, it is marked as an active window. The alternating distribution characteristics of the silent window and the active window within the sliding observation period are determined, and the continuity of the data stream is determined based on the alternating distribution characteristics. If there are no continuous silent windows within the sliding observation period and the proportion of active windows exceeds a preset ratio, the data stream status of the corresponding terminal is determined as a continuous stream.

[0024] In specific implementation, a fixed-duration window-based arrival rate statistical method is used to detect the continuity of the original interactive event stream. First, the system clock is divided into continuous time windows of 2 seconds as the collection granularity. The collection granularity refers to the smallest time unit for segmenting and statistically analyzing the event stream. The number of interactive events arriving within each 2-second time window is accumulated by an event counter. The accumulated result is divided by the window duration to obtain the event arrival rate of that time window. The event arrival rate refers to the average rate at which interactive events are generated within a unit time window. For example, if 16 mouse movement and keyboard click events are collected within a 2-second window, the event arrival rate is 8 times / second. Then, the event arrival rate of each time window is compared with a preset silence threshold. The silence threshold is set to 0.5 times / second based on the empirical distribution of operation pause intervals in human-computer interaction research. When the event arrival rate of a time window is less than 0.5 times / second, the time window is marked as a silent window. The silent window refers to the time segment where user interaction behavior is below the detectable level. Conversely, when the event arrival rate is greater than or equal to 0.5 times / second, the time window is marked as an active window. The active window refers to the time segment where there is obvious user interaction behavior. Finally, after completing the window marking, a 30-second sliding observation period is introduced. This sliding observation period moves backward on the time axis in 1-second steps. Within each observation period, the occurrence frequency and continuous distribution of silent and active windows are counted to obtain alternating distribution characteristics. The alternating distribution characteristics refer to the arrangement pattern and proportion of silent and active windows within the observation period. Specifically, it is counted whether there are more than 3 consecutive silent windows (i.e., covering more than 6 seconds) within the 30-second observation period, and the proportion of active windows to the total number of windows is calculated. If there are no more than 3 consecutive silent windows within the observation period and the proportion of active windows exceeds the preset 70%, it is determined that the terminal has maintained a continuous operating rhythm rather than being in an intermittent idle state within the current observation period, and the data stream status of the corresponding terminal is identified as a continuous stream.

[0025] It should be noted that, in this application, the data stream status identifier refers to a status flag value used to indicate whether the terminal is currently in a continuous interactive state, an intermittent interactive state, or a silent state.

[0026] In step S102, using the timestamp sequence of each terminal as an index, an attention focus flow is constructed based on the attribute features of the corresponding focus object to show the user's attention jumping between multiple terminals. The attention focus flow is then analyzed based on the interaction pattern of sliding time window to obtain the non-steady-state interaction feature flow when the user is in an operational confusion state.

[0027] In some embodiments, the following steps are used to construct the attention focus flow that allows user attention to jump between multiple terminals, based on the attribute features of the corresponding focus object and indexed by the timestamp sequence of each terminal: Align the timestamp sequence of each terminal with the attribute features of the corresponding focus object according to the timestamp to form a unified timeline record carrying terminal identifier and focus attribute information; The unified timeline records are merged from multiple sources, and the focus objects from different terminals are concatenated into a global focus transfer sequence according to the order in which the interaction occurs. Extract the terminal identifiers of adjacent focus objects at each focus switch in the global focus transfer sequence to form directed transfer pairs that express attention jumping across devices between terminals, and organize all directed transfer pairs into an attention focus stream in chronological order.

[0028] In specific implementation, firstly, the timestamp sequence of interactive behaviors from various terminals is associated with the attribute features of the corresponding focus object. Using an equi-join method based on timestamp hashing, the timestamp of each interactive behavior record is used as the join key and left-out join is performed with the attribute feature tuple of the focus object with the same timestamp. This expands each interactive behavior record into a wide table record containing terminal identifier, event occurrence time, node type identifier, semantic label, and spatial coordinates. The wide table record sets of each terminal together constitute a unified timeline record carrying terminal identifier and focus attribute information. The unified timeline record refers to the complete event log after aligning the discrete interactive events and focus attributes of multiple terminals according to the time dimension. Then, a multi-source time-series merging operation is performed on the unified timeline records. A multi-way merge sort algorithm is used to maintain an event pointer for each terminal, pointing to its own unified timeline record. Each event pointer initially points to its earliest event record. Each time, the timestamp values ​​of the records pointed to by the current event pointers of each terminal are compared. The record with the smallest timestamp is selected and appended to the end of the global focus transfer sequence, and the corresponding event pointer is moved one position to the right. This process is repeated until the event pointers of all terminals have been traversed. The global focus transfer sequence refers to the total order sequence formed by strictly sorting the focus objects of all terminals according to the time of interaction from earliest to latest. For example, if a user clicks on the graphic area of ​​a geometry problem on a tablet and then clicks on line 15 of the code editor on a PC 1.2 seconds later, the global focus transfer sequence will sequentially include the geometric focus object on the tablet and the code line focus object on the PC. Finally, after obtaining the global focus transfer sequence, iterate sequentially from the head of the global focus transfer sequence. Each pair of adjacent focus objects constitutes a transfer step. Extract the terminal identifiers of the current focus object and the next focus object. If the terminal identifier of the current focus object is different from the terminal identifier of the next focus object, construct a directed transfer pair with the terminal identifier of the current focus object as the source and the terminal identifier of the next focus object as the destination. The directed transfer pair refers to the basic data structure representing a cross-device switching event where the user's attention jumps from one terminal to another. The directed transfer pair also contains the semantic labels of the source focus object and the semantic labels of the destination focus object, as well as the time interval of the switching. Store all the extracted directed transfer pairs in a list structure in chronological order of the switching occurrence. This list is the attention focus stream.

[0029] It should be noted that, in this application, attention focus flow refers to the temporal expression of user attention switching behavior across multiple devices. Existing technologies typically collect operation events from each device independently, failing to correlate the temporal continuity between interactive behaviors on different devices. This solution, however, aligns discrete interactive events from multiple devices with the semantic and spatial attributes of the focus object through a unified timeline recording. Then, a global focus transfer sequence is constructed through multi-source temporal merging, connecting clicks, drags, and inputs scattered across tablets, PCs, and mobile phones into a continuous migration trajectory of user attention in cross-device collaborative scenarios. Based on this, the extracted directed transfer pairs directly quantify the direction, frequency, and semantic association of attention jumps in the form of directed edges from the source terminal to the destination terminal, achieving the continuity and intent of user operations when switching between multiple devices during tutoring. Figure 1 Consistent fine-grained representation. This abstract transformation from device operation logs to attention transition graphs allows the collaborative behavior patterns that were originally implicitly scattered across multiple terminal events to be explicitly extracted into a computable data structure, breaking through the technical limitations of existing technologies that can only analyze behavior within a single terminal and cannot capture the dynamic migration of attention across terminals.

[0030] In some embodiments, the interaction pattern analysis based on time window sliding is performed on the attention focus flow to obtain the non-steady-state interaction feature flow when the user is in an operational distress state. This is achieved through the following steps: A sliding time window of a preset step size is applied to the attention focus stream, and the directed transition pair sub-sequences within each sliding window are extracted sequentially. The terminal switching frequency and the semantic jump rate of the focus object between adjacent directed transition pairs in the directed transition pair sub-sequences are calculated. The terminal switching frequency and the semantic jump rate of the focus object are input into a pre-constructed steady-state deviation metric function, and the deviation index, which characterizes the degree of deviation of the attention shift mode from the steady state within the current sliding window, is output. The deviation indices of each sliding window on the attention focus flow are arranged in time sequence to obtain a deviation index sequence. Continuous intervals in the deviation index sequence where the deviation index exceeds a preset perturbation threshold are extracted. The directed transition pair sub-sequences corresponding to these continuous intervals are marked as non-steady-state interaction feature flows in which the user is in an operational distress state.

[0031] In specific implementation, firstly, a sliding time window of length 15 directed transition pairs is defined on the constructed attention focus stream. The window moves sequentially from the starting point to the ending point of the attention focus stream, sliding one directed transition pair at a time. The sliding time window is a moving data frame that extracts a local subsequence of fixed length from a continuous time series for analysis. After each shift, all directed transition pairs contained within the current window are extracted to form a directed transition pair subsequence. The directed transition pair subsequence is a continuous segment of directed transition pairs extracted from the attention focus stream by the sliding window. For each directed transition pair subsequence, the number of times the source terminal identifier and the destination terminal identifier change is counted and divided by the total number of transition pairs within the window to obtain the terminal switching frequency. The terminal switching frequency is... This refers to the average probability that user attention switches between terminals within a single directed transfer pair. For example, if there are 6 back-and-forth jumps between the tablet and PC within a window with 15 transfers, the frequency is 0.4. Simultaneously, the semantic tags of the destination focus object in adjacent directed transfer pairs are compared with the semantic tags of the source focus object in the next directed transfer pair. The proportion of times the semantic tags change relative to the total number of semantic tag comparisons is calculated as the focus object semantic jump rate. The focus object semantic jump rate refers to the proportion of changes in the semantic category of the tutoring content associated with the attention switch. Specifically, a string fuzzy matching method based on edit distance is used to determine whether two semantic tags belong to the same semantic category. When the edit distance exceeds 30% of the tag character length, it is determined as a semantic jump. Then, after obtaining the terminal switching frequency and the semantic jump rate of the focus object, they are input into a steady-state deviation measurement function constructed based on a logistic regression binary classification model. The training process of this steady-state deviation measurement function is to collect attention focus flow samples under normal collaborative tutoring scenarios and attention focus flow samples under known operational distress scenarios, extract the terminal switching frequency and semantic jump rate of the focus object of each sliding window in the samples as feature vectors, and use whether it is an operational distress state as a label for supervised training to obtain the weight coefficients and bias terms of logistic regression. During deployment, the terminal switching frequency and semantic jump rate of the focus object of the current window are multiplied by the corresponding weight coefficients and the bias term is added. Then, the deviation index is mapped to a value between 0 and 1 through the Sigmoid function. The deviation index refers to the degree of deviation of the attention transfer mode from the normal collaborative mode in the current sliding window. The closer the deviation index is to 1, the closer it is to the operational distress state.Finally, the deviation indices calculated for all sliding windows along the attention focus flow are stored sequentially in a time series array according to the temporal position of the windows, forming a deviation index sequence. The deviation index sequence refers to the vector of deviation indices generated after the sliding windows move along the attention focus flow, arranged in time. Then, the deviation index sequence is thresholded with a preset perturbation threshold of 0.7 as the boundary. The original directed transfer pairs of multiple adjacent sliding windows in the attention focus flow corresponding to the intervals where the deviation index continuously exceeds 0.7 are spliced ​​together to obtain the non-steady-state interaction feature flow in which the user is in an operational distress state. The operational distress state refers to the cognitive and interactive blockage state exhibited by the user when interacting with content elements in the tutoring interface during multi-terminal collaborative tutoring, where the user's operation behavior fails to achieve the expected goal.

[0032] refer to Figure 2 This diagram illustrates the principle of interaction pattern analysis based on some embodiments of this application. The horizontal timeline with arrows at the bottom of the diagram sequentially arranges user interaction behavior nodes such as click, swipe, input, pause, and return. Time windows W1, W2, and W3, sliding sequentially to the right along the timeline and overlapping each other, segment attention focus flow data containing multi-terminal timestamp sequences and focus object attributes from different time periods, and analyze the interaction patterns one by one. By relying on the overlapping sliding of windows, continuous interaction details are fully preserved, avoiding the problem of fragmented and lost timeline behaviors. The results of all window analysis are summarized to generate the diagram on the right. The fluctuating non-steady-state interaction feature flow within the box represents a normal interaction state where users learn smoothly and their attention is stable, with gentle curves. Large oscillations and peaks in the curve represent disordered operation states where users frequently switch devices, repeat invalid operations, and backtrack repeatedly, indicating that users are struggling with learning operations. The overall process uses time-series segmented sliding analysis to determine the interaction mode, accurately distinguishing between steady-state learning behavior and abnormal interaction behavior under obstruction. Finally, a standardized non-steady-state interaction feature flow is output, laying a quantitative data foundation for subsequent extraction of invalid and repetitive operation sequences and generation of tutoring and collaboration requests carrying problem context and behavioral evidence.

[0033] It should be noted that, in this application, the non-steady-state interaction feature flow refers to the set of directed shift pairs in which the user's attention exhibits abnormal switching patterns across multiple terminals during the period of operational distress. Existing technologies typically only monitor whether the user has repeatedly clicked or timed out on a single terminal, and cannot distinguish whether these abnormal operations are due to interface cognitive confusion during multi-device collaboration or difficulty in understanding the single task itself. However, this solution applies a sliding time window to the attention focus flow and simultaneously calculates two cross-device behavioral indicators: terminal switching frequency and semantic jump rate of the focus object. It incorporates the user's physical switching rhythm across multiple terminals and the semantic coherence of the operation content into a unified steady-state deviation measurement function for fusion evaluation. This dual-dimensional joint analysis enables users to stably identify and quantify their frequent device switching behavior due to not being able to find the correct operation entry point across multiple terminals, as well as their blind attempts between different semantic functional areas due to unclear operation goals, as deviation indices. By using time-series threshold segmentation, non-steady-state interaction feature streams in the state of operational distress can be accurately extracted from the continuous behavioral stream. This provides precise behavioral positioning anchors for the subsequent generation of collaborative requests carrying specific problem context and behavioral evidence, overcoming the problems of missed detection and misjudgment of user operational distress in existing technologies in multi-terminal collaborative scenarios.

[0034] In step S103, repetitive operation sequences that do not produce valid input results are extracted from the non-steady-state interaction feature stream, and a collaborative request carrying specific problem context and behavioral evidence is generated based on the repetitive operation sequence and the content object in the interface where it occurs.

[0035] In some embodiments, extracting repetitive operation sequences that do not produce valid inputs from the non-steady-state interaction feature stream is achieved through the following steps: Based on the timestamp interval and terminal identifier carried by each directed transition pair in the non-steady-state interaction feature stream, the original interaction event stream of the corresponding terminal within the timestamp interval is extracted back to form a set of interaction events to be analyzed. From the set of interactive events to be analyzed, identify consecutive clusters of interactive events with the same type of interactive events and the same focus object, and calculate the similarity of the operation loads of each interactive event in the consecutive clusters of interactive events to obtain an operation load similarity sequence. The continuous interactive events and their operational payloads in the operational payload similarity sequence that have a similarity exceeding a preset repetition threshold are organized into candidate repetitive operational sequences according to their occurrence time. Identify whether a valid state transition occurs in the coaching session state machine after each operation in the candidate repetitive operation sequence, and determine the candidate repetitive operation sequence that does not trigger a valid state transition as a repetitive operation sequence that does not produce a valid result input.

[0036] In specific implementation, firstly, the directed transition pairs contained in the non-steady-state interaction feature stream are parsed. Each directed transition pair records the start and end timestamps of the cross-terminal handover, as well as the source terminal identifier and the destination terminal identifier in its data structure. The original interaction event streams of the terminals corresponding to these two terminal identifiers within the interval from the start timestamp to the end timestamp are retrieved from the event log storage. Since the directed transition pairs may have overlapping time intervals, all retrieved event records are deduplicated by event ID to form a set of interaction events to be analyzed. The set of interaction events to be analyzed refers to the union of all original interaction events that occurred on the terminals involved in the non-steady-state period. Then, the set of interactive events to be analyzed is clustered into continuous homogeneous events. Specifically, a double-key sorting and grouping method based on event type and operation object identifier is adopted. First, the interactive events in the set are grouped by the event type field, and then a second grouping is performed within each group by the focus object identifier. After sorting the event records in each group in ascending order by timestamp, the continuity is detected by the adjacent event time interval threshold method. If the time interval between two adjacent events is less than the preset 1.5-second operation continuity threshold, they are classified into the same continuous interactive event cluster. The continuous interactive event cluster refers to the clustering unit of interactive events with the same event type, the same operation object, and close temporal proximity. For example, if a user clicks the same "Submit Answer" button 5 times consecutively with each click interval between 0.3 and 0.8 seconds, these 5 click events constitute a continuous interactive event cluster. After obtaining each cluster of consecutive interactive events, the similarity of the operation payload fields carried by all interactive events within each cluster is calculated pairwise. The operation payload field refers to the specific operation parameters passed when the interactive event occurs, such as the key character sequence of keyboard keystrokes, the screen coordinate offset of mouse clicks, or the sliding trajectory vector of touch gestures. The similarity calculation adopts a sequence distance metric method based on dynamic time warping. After converting the two operation payloads into feature vectors, the DTW distance is calculated. Then, the DTW distance is mapped to a similarity value in the interval of 0 to 1 through a Gaussian kernel function to obtain the operation payload similarity. The operation payload similarity sequence is obtained by arranging the operation payload similarity of all adjacent event pairs within the cluster in chronological order. The operation payload similarity sequence refers to the temporal set of the similarities of each adjacent operation payload in the consecutive interactive event cluster.Finally, threshold segmentation is performed on the operation load similarity sequence. Continuous interaction events with operation load similarity exceeding a preset repetition threshold of 0.85, along with their operation loads, are arranged in ascending order according to their original timestamps, forming candidate repetitive operation sequences. These candidate repetitive operation sequences refer to interaction event sequences with highly similar and consecutively occurring operations suspected of being duplicated or invalid. The occurrence time of each operation event in the candidate repetitive operation sequence is used as a query point to check whether a preset effective state transition has occurred before and after that time in the tutoring session state machine. The tutoring session state machine is a finite state automaton that maintains the progress of the tutoring task. Effective state transitions include changes in the question-answering state from... The system employs predefined state transition rules, such as transitioning from unanswered to answered, from compilation error to compilation success in code editing, and from unviewed to viewed in knowledge point browsing. If no valid state transition occurs in any of the operation events in a candidate repetitive operation sequence, then all operations in that sequence are deemed not to have made a substantial contribution to the progress of the tutoring task, and are thus identified as a repetitive operation sequence that has not produced a valid result input. In addition, candidate repetitive operation sequences in which the number of operations that have not triggered a valid state transition exceeds a preset invalid threshold can also be identified as repetitive operation sequences that have not produced a valid result input. This will not be elaborated further here.

[0037] It should be noted that, in this application, repetitive operation sequences that do not produce valid input results refer to segments of user behavior that repeatedly perform the same or highly similar operations but fail to advance the tutoring session. The identification of repetitive operation sequences that do not produce valid input results involves upgrading the existing coarse detection method, which relies solely on abnormal operation frequency or intervals, to a precise extraction method based on dual verification of operation content similarity and task progress effectiveness. This solution introduces operation load similarity calculation based on dynamic time warping within continuous interaction event clusters, quantifying the degree of repetition of each operation at the operation parameter level. Then, candidate repetitive operation sequences are compared and verified against the effective state transition rules of the tutoring session state machine. Whether the state machine undergoes a preset task progress jump is used as an objective criterion for whether the operation produces valid input results. This precisely isolates invalid interaction segments with highly repetitive content that consistently fail to advance the tutoring task, providing reliable behavioral anchors for generating collaborative requests with precise problem contexts. This allows the tutor to directly locate the specific operation steps that the user repeatedly tried but failed to achieve.

[0038] In some embodiments, generating a collaborative request carrying specific problem context and behavioral evidence based on the repetitive operation sequence and the content object in the interface where it occurs is achieved through the following steps: The problem occurrence window is defined by the start and end timestamps of the repetitive operation sequence. Each content object in the rendering state in the tutoring interface within the problem occurrence window and its node attributes in the document object model are captured to form an interface context snapshot. The attribute features of the focus object targeted by each operation in the repetitive operation sequence are matched with the content objects in the interface context snapshot to extract the tutoring semantic information carried by the target content object that has an interactive mapping relationship with the repetitive operation sequence. Using the tutoring semantic information as the problem context, and the operation type, operation load, and repetition frequency of the repetitive operation sequence as behavioral evidence, a collaborative request carrying the problem context and the behavioral evidence is generated.

[0039] In specific implementation, firstly, the earliest timestamp of the operation event in the repetitive operation sequence is extracted as the start timestamp, and the latest timestamp is extracted as the end timestamp. This start and end timestamps form a closed interval as the problem occurrence window, which refers to the time range boundary where the user performs repeated invalid operations. By calling the DOMSnapshot interface in the browser developer tools protocol, passing in the start and end timestamps of the problem occurrence window, the system requests the return of all visible DOM nodes and their computed style attributes in the tutoring interface rendering tree within this time period. The returned DOM node set is serialized into a structured record containing node type, node attributes, text content, and the coordinates of the boundary rectangle relative to the viewport, forming an interface context snapshot. This interface context snapshot is a complete copy of the tutoring interface's rendering state within the problem occurrence time period. For example, during the period when the user repeatedly clicks a disabled "Next" button, the interface context snapshot records that the button's disabled property is true, the button text is "Next," and its parent container is a question answering area. The system retrieves complete interface structure information, including nodes, question text nodes within the same area, and selected option nodes. Then, it extracts the attribute features of the focus object associated with each operation event in the repetitive operation sequence. Using the node type identifier and spatial coordinates of the focus object as matching conditions, node association matching is performed in the interface context snapshot. The matching process employs a dual matching strategy based on DOM tree path and spatial location. First, a set of candidate nodes of the same type is filtered in the interface context snapshot based on the node type identifier of the focus object. Then, the intersection-union ratio (IUR) between the boundary rectangle coordinates of each candidate node and the spatial coordinates of the focus object is calculated in this candidate node set. The candidate node with the largest IUR exceeding the matching threshold of 0.8 is selected as the matching node. If the IUR based on spatial coordinates fails to uniquely determine the matching node, the system further compares the DOM tree path of the focus object recorded before the non-steady-state interaction feature flow occurs with the DOM tree path of the candidate node, using a path similarity algorithm based on tree edit distance to determine the most suitable match. Optimal matching is used to accurately extract target content objects that have an interactive mapping relationship with the repetitive operation sequence from the interface context snapshot. These target content objects refer to interface elements that the user repeatedly attempts to interact with but does not receive a response during repetitive operations, along with their associated tutoring content nodes. After locking the target content object, its text content, ARIA semantic attributes, and the text content of its adjacent sibling nodes are extracted to constitute tutoring semantic information. This tutoring semantic information refers to textual information with tutoring business meaning, such as the specific tutoring question number, knowledge point name, and interface area function description involved in the user's current operation context. For example, in a scenario where a user repeatedly attempts to input code in a locked code editing area, the tutoring semantic information includes the question number corresponding to the code editing area, the current locked status prompt text of the code editing area, and the description of the programming task to which the editing area belongs. Finally, a JSON-formatted collaborative request message body is constructed. The extracted tutoring semantic information is filled into the question context field of the message body, and the operation type enumeration value of the repetitive operation sequence, the specific operation payload array for each operation, and the statistical value of the number of times the operation is repeated are filled into the behavior evidence field.

[0040] It should be noted that, in this application, a collaborative request refers to a standardized help data package that carries a description of the specific problem the user is currently encountering and evidence of the operational behavior supporting that judgment. Determining a collaborative request can transform the lagging help mechanism in traditional online tutoring systems, which relies on users to actively initiate or tutors to passively observe and discover problems, into an active collaborative triggering mechanism in which the system automatically captures operational difficulties and generates precise problem context and quantified behavioral evidence. This allows the tutor's response to focus on problem-solving itself rather than problem localization, significantly shortening the response delay from when the user is in an operational predicament to when the tutor intervenes effectively.

[0041] In step S104, the tutoring terminal that is currently collaborating in the tutoring session is identified according to the collaboration request, and the collaboration request and the corresponding non-steady-state interaction feature stream summary are pushed to the identified tutoring terminal for display.

[0042] In some embodiments, identifying the tutoring terminal that is currently collaborating in the tutoring session based on the collaboration request is achieved through the following steps: Obtain the session control signaling record of the current tutoring session, parse the role tags of each terminal that has joined the current tutoring session from the session control signaling record, and form a terminal role mapping relationship; The system queries the terminal role mapping relationship for terminals with the role tag "tutor" to form a candidate tutor terminal set. The system then performs real-time detection on the online status of each terminal in the candidate tutor terminal set and filters out online tutor terminals that are currently in a session connection state. The activity level of the online tutoring terminal's interactive behavior within a preset time window is detected, and online tutoring terminals whose interactive behavior activity exceeds a preset assistance threshold are identified as tutoring terminals that are currently collaborating.

[0043] In specific implementation, firstly, the signaling query interface provided by the session management service of the current tutoring session initiates an HTTP GET request with the session identifier as the request parameter to obtain all session control signaling records from the establishment of the session to the current time. The session control signaling records refer to the control message logs transmitted during the lifecycle of the tutoring session using the WebRTC signaling protocol or the WebSocket session management protocol, including the time sequence records of control events such as terminal joining the session, leaving the session, role assignment, media stream establishment and release. Each signaling message in the session control signaling records is traversed, and the terminal identifier field and role tag field carried in the message body are extracted to establish a mapping entry from the terminal identifier to the role tag. All mapping entries are summarized to form the terminal role mapping relationship. The terminal role mapping relationship refers to the dictionary structure of the correspondence between the unique identifier of the terminal device and the role it assumes in the tutoring session. The value of the role tag includes two categories: tutor and tutored. For example, if the terminal identifier in a certain signaling message record is "T-003" and the role tag is "tutor", then a key-value pair "T-003" pointing to "tutor" is added in the mapping relationship. Then, all entries in the terminal role mapping relationship are traversed, and terminal identifiers whose role tags are equal to those of the tutor are filtered out. These terminal identifiers are added to a list structure to form a candidate tutor terminal set. The candidate tutor terminal set refers to the set of identifiers of all terminals whose identity is set as tutor in the current tutoring session. For each terminal identifier in this set, a terminal online status query request is sent to the heartbeat detection interface of the session management service to obtain the timestamp of the last heartbeat message sent by the terminal. The time difference between the current system time and the heartbeat timestamp is calculated. If the time difference is less than a preset 15-second heartbeat timeout threshold, the terminal is determined to be in a session connection state. Terminal identifiers in the session connection state are filtered out from the candidate tutor terminal set to form an online tutor terminal list. The online tutor terminal refers to a terminal whose role is tutor and which is currently maintaining a valid network connection with the tutoring session server. Finally, the interactive activity of the online tutor terminals within a preset time window is detected. Online tutor terminals whose interactive activity exceeds a preset assistance threshold of 0.5 are identified as tutor terminals that are currently collaborating. The tutor terminals that are currently collaborating refer to actual tutoring execution terminals that are not only online at the moment but are also participating in tutoring interaction.

[0044] In some embodiments, the activity level of the online tutoring terminal's interactive behavior within a preset time window is determined using the following steps: Collect the tutoring party interaction event stream generated by the online tutoring party terminal within a preset time window, and obtain the event occurrence frequency of the interaction events in the tutoring party interaction event stream; Extract the attribute features of the focus object associated with each interaction event from the tutor's interaction event stream, identify the target interaction event with tutoring guidance intent based on the semantic tags in the attribute features, and calculate the proportion of the target interaction event in the tutor's interaction event stream. The frequency of the events and the proportion of the intentional events are weighted and fused to obtain the interactive activity level of the online tutoring terminal.

[0045] In practice, firstly, all interactive events generated by the terminal within the last 30 seconds are captured within a preset time window. These events include mouse movement, mouse clicks, keyboard input, screen touch, voice input, and whiteboard annotation. The captured interactive events are then arranged in ascending order by timestamp to form a tutoring interactive event stream. This tutoring interactive event stream refers to the temporal set of all interactive events generated by the tutoring terminal within the specified time window. The tutoring interactive event stream is traversed, and the total number of interactive events in the event stream is accumulated using an event counter. The accumulated result is used as the event frequency. The event frequency refers to the absolute number of interactive events generated on the tutoring terminal within the preset time window. For example, if the tutor performs 12 mouse clicks, 8 keyboard inputs, and 5 whiteboard annotations within 30 seconds, the event frequency is 25. Then, for each interaction event in the tutoring interaction event flow, the node reference of the currently activated interface element in the Document Object Model is returned by calling the focus acquisition interface of the terminal operating system at the time of the event. The node type identifier and semantic tags are read from this node reference to form the attribute features of the focus object. These attribute features refer to structured data describing the type and business meaning of the target interface element currently being operated by the tutor. The semantic tags from the attribute features of each interaction event are input into an intent classifier built based on the bag-of-words model and TF-IDF features. This intent classifier is pre-trained using labeled corpora containing semantic tag samples corresponding to various tutoring guidance operations, such as "question explanation area" and "answer annotation box". Tags such as "code highlighting line", "knowledge point label", "student answer area", and "drawing tool" are labeled as positive examples with tutoring guidance intent, while tags such as "volume adjustment", "settings menu", "conversation list", and "minimize button" are labeled as negative examples without tutoring guidance intent. After training, the classifier outputs a binary classification judgment result for the input semantic labels, marking the interactive events judged to have tutoring guidance intent as target interactive events. The target interactive event refers to the operation event performed by the tutor on the tutoring content itself that has teaching guidance significance. The number of target interactive events is counted and divided by the frequency of event occurrence to obtain the intention event ratio, which refers to the proportion of interactive events with tutoring guidance intent in all interactive events. Finally, the frequency of events is processed dimensionlessly using the Min-Max normalization method. The upper and lower bounds of the normalization are set to the preset upper limit of 50 times and the lower limit of 0 times for the active frequency of the tutor. The normalized frequency of events is multiplied by a weighting coefficient of 0.4, and the proportion of intentional events is multiplied by a weighting coefficient of 0.6. The two weighted results are added together to obtain the interactive behavior activity of the online tutor terminal. The interactive behavior activity refers to the quantitative score value of the current participation in tutoring of the terminal after comprehensively considering the number of tutor operations and the orientation of the operation tutoring. The value range is a real number between 0 and 1.

[0046] In some embodiments, the following steps are used to push the collaborative request and the corresponding non-steady-state interaction feature stream summary to the identified tutor terminal for display: The non-steady-state interaction feature stream is compressed to extract the peak value of terminal switching frequency, the mean value of semantic jump rate of focus object, and the duration of non-steady state in the non-steady-state interaction feature stream, forming a summary of the non-steady-state interaction feature stream. The problem context and behavioral evidence carried in the collaboration request are combined with the non-steady-state interaction feature stream summary to form a message body, generating a collaboration notification message for the tutoring terminal. Through the session maintenance channel of the current tutoring session, the collaborative notification message is pushed to the identified tutoring terminal, triggering the tutoring terminal to render and display the problem context, behavioral evidence, and non-steady-state interaction feature stream summary in the tutoring interface.

[0047] In specific implementation, firstly, all directed transition pairs contained in the non-steady-state interaction feature stream are traversed, aggregated, and statistically analyzed. The terminal switching frequency within each window is recalculated using a sliding time window as the unit. The maximum value of the terminal switching frequency across all windows is taken as the peak value of the terminal switching frequency. The peak value of the terminal switching frequency refers to the extreme value of the switching frequency at the moment when the user's switching behavior between terminals is most intense during the non-steady-state period. Simultaneously, the semantic jump events of the focus objects involved in all directed transition pairs in the non-steady-state interaction feature stream are counted. The average semantic jump rate of the focus objects is obtained by dividing the number of times the semantic label changes by the total number of directed transition pairs. The mean semantic jump rate refers to the overall average level of semantic coherence of user operation content during the entire non-steady-state period. The non-steady-state duration is obtained by subtracting the timestamp of the first directed transition pair from the timestamp of the last directed transition pair in the non-steady-state interaction feature stream. The non-steady-state duration refers to the time span during which the user is in an operational confusion state. The peak terminal switching frequency, the mean semantic jump rate of the focus object, and the duration are encapsulated into a lightweight non-steady-state interaction feature stream summary object. The non-steady-state interaction feature stream summary refers to the set of key feature indicators formed after statistical compression of the non-steady-state interaction feature stream. Then, the problem context field and behavior evidence field are extracted from the generated collaborative request message body and combined with the non-steady-state interaction feature stream summary object according to the preset collaborative notification message template. This template is defined using the Protocol Buffers serialization format and includes the message type identifier and timestamp in the message header, the tutored terminal identifier, the problem context structure, the behavior evidence array, and the feature summary structure in the message body. Each part of the data is written into the corresponding binary data segment according to the field number to complete the message body assembly and generate a collaborative notification message for the tutored terminal. The collaborative notification message refers to a push data message that uniformly encapsulates the user's problem description, operation behavior evidence, and distress level summary. Finally, messages are delivered through the session maintenance channel negotiated and determined during the establishment of the current tutoring session. If the session is based on the WebRTC architecture, the DataChannel is used to send messages in a reliable and ordered mode. If it is based on the WebSocket long connection architecture, the message frames of the long connection are pushed. The serialized collaborative notification message byte stream is delivered from the server to the client SDK of the identified tutoring terminal. After receiving the message, the client SDK deserializes and restores the content of each field, and calls the pre-built collaborative notification display component in the tutoring interface. This display component renders an interactive information card in the sidebar or floating window area of ​​the tutoring interface. The card renders the question number and interface area description text in the question context, the operation type icon and repetition number logo in the behavioral evidence, and three statistical values ​​in the non-steady-state interaction feature stream summary from top to bottom, thus completing the structured display of the question context, behavioral evidence, and non-steady-state interaction feature stream summary.

[0048] Furthermore, in another aspect of this application, in some embodiments, this application provides a multi-terminal online tutoring collaborative management system based on learning behavior analysis, referencing... Figure 3 The figure is a schematic diagram of the structure of a multi-terminal online tutoring collaborative management system based on learning behavior analysis, according to some embodiments of this application. The multi-terminal online tutoring collaborative management system based on learning behavior analysis includes: an acquisition module 201, a processing module 202, and an execution module 203, which are described below: The acquisition module 201 in this application is mainly used to extract the timestamp sequence of each terminal's interactive behavior and the attribute features of the corresponding focus object in the current tutoring session. Processing module 202 in this application is mainly used to construct an attention focus flow of user attention jumping between multiple terminals based on the attribute features of the corresponding focus object, using the timestamp sequence of each terminal as an index, and to perform interaction pattern analysis on the attention focus flow based on time window sliding, so as to obtain the non-steady-state interaction feature flow of the user in an operational confusion state. The processing module 202 is further configured to extract repetitive operation sequences that have not produced valid input results from the non-steady-state interaction feature stream, and generate a collaborative request carrying specific problem context and behavioral evidence based on the repetitive operation sequence and the content object in the interface where it occurs. The execution module 203 in this application is mainly used to identify the tutoring terminal that is currently collaborating in the tutoring session according to the collaboration request, and push the collaboration request and the corresponding non-steady-state interaction feature stream summary to the identified tutoring terminal for display.

[0049] In addition, this application also provides a computer device, the computer device including a memory and a processor, the memory storing code, and the processor being configured to acquire the code and execute the above-described multi-terminal online tutoring collaborative management method based on learning behavior analysis.

[0050] In some embodiments, reference Figure 4 The figure is a schematic diagram of the structure of a computer device implementing a multi-terminal online tutoring collaborative management method based on learning behavior analysis, according to some embodiments of this application. The multi-terminal online tutoring collaborative management method based on learning behavior analysis in the above embodiments can... Figure 4 The computer device shown is used to implement this, and the computer device includes at least one processor 301, a communication bus 302, a memory 303, and at least one communication interface 304.

[0051] The processor 301 can be a general-purpose central processing unit (CPU), an application-specific integrated circuit (ASIC), or one or more devices used to control the execution of the multi-terminal online tutoring collaborative management method based on learning behavior analysis in this application.

[0052] The communication bus 302 can be used to transmit information between the aforementioned components.

[0053] The memory 303 may be a read-only memory (ROM) or other type of static storage device capable of storing static information and instructions, random access memory (RAM) or other type of dynamic storage device capable of storing information and instructions, or electrically erasable programmable read-only memory (EEPROM), compact disc read-only memory (CD-ROM) or other optical disc storage, optical disc storage (including compressed optical discs, laser discs, optical discs, digital universal optical discs, Blu-ray discs, etc.), magnetic disks or other magnetic storage devices, or any other medium capable of carrying or storing desired program code in the form of instructions or data structures and accessible by a computer, but not limited thereto. The memory 303 may exist independently and be connected to the processor 301 via the communication bus 302. The memory 303 may also be integrated with the processor 301.

[0054] The memory 303 stores program code for executing the scheme of this application, and its execution is controlled by the processor 301. The processor 301 executes the program code stored in the memory 303. The program code may include one or more software modules. In the above embodiments, the determination of the multi-terminal online tutoring collaborative management method based on learning behavior analysis can be achieved by the processor 301 and one or more software modules in the program code in the memory 303.

[0055] Communication interface 304 uses any transceiver-like device for communicating with other devices or communication networks, such as Ethernet, radio access network (RAN), wireless local area networks (WLAN), etc.

[0056] In a specific implementation, as one example, a computer device may include multiple processors, each of which may be a single-core (single-CPU) processor or a multi-core (multi-CPU) processor. Here, a processor may refer to one or more devices, circuits, and / or processing cores used to process data (e.g., computer program instructions).

[0057] The aforementioned computer device can be a general-purpose computer device or a special-purpose computer device. In specific implementations, the computer device can be a desktop computer, a portable computer, a network server, a handheld digital assistant (PDA), a mobile phone, a tablet computer, a wireless terminal device, a communication device, or an embedded device. This application does not limit the type of computer device.

[0058] In addition, this application also provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the above-described multi-terminal online tutoring collaborative management method based on learning behavior analysis.

[0059] Although preferred embodiments of this application have been described, those skilled in the art, once they have learned the basic inventive concept, can make other changes and modifications to these embodiments.

[0060] Obviously, those skilled in the art can make various modifications and variations to this application without departing from the spirit and scope of this application.

Claims

1. A multi-terminal online tutoring collaborative management method based on learning behavior analysis, characterized in that, Includes the following steps: Extract the timestamp sequence of each terminal's interactive behavior and the attribute features of the corresponding focus object in the current tutoring session; Using the timestamp sequence of each terminal as an index, an attention focus flow is constructed based on the attribute features of the corresponding focus object to show the user's attention jumping between multiple terminals. The attention focus flow is then analyzed based on the interaction pattern of sliding time window to obtain the non-steady-state interaction feature flow when the user is in an operational confusion state. Extract repetitive operation sequences that do not produce valid input results from the non-steady-state interaction feature stream, and generate collaborative requests carrying specific problem context and behavioral evidence based on the repetitive operation sequences and the content objects in the interface where they occur; The collaboration request identifies the tutoring terminal currently collaborating in the tutoring session, and the collaboration request and the corresponding non-steady-state interaction feature stream summary are pushed to the identified tutoring terminal for display.

2. The method as described in claim 1, characterized in that, The specific tutoring terminal that is currently collaborating in the tutoring session, as identified by the collaboration request, includes: Obtain the session control signaling record of the current tutoring session, parse the role tags of each terminal that has joined the current tutoring session from the session control signaling record, and form a terminal role mapping relationship; The system queries the terminal role mapping relationship for terminals with the role tag "tutor" to form a candidate tutor terminal set. The system then performs real-time detection on the online status of each terminal in the candidate tutor terminal set and filters out online tutor terminals that are currently in a session connection state. The activity level of the online tutoring terminal's interactive behavior within a preset time window is detected, and online tutoring terminals whose interactive behavior activity exceeds a preset assistance threshold are identified as tutoring terminals that are currently collaborating.

3. The method as described in claim 1, characterized in that, Extracting the timestamp sequence of each terminal's interactive behavior and the attribute features of the corresponding focus object in the current tutoring session specifically includes: The system collects the original interactive event streams of all terminals with established session connections in the tutoring session in real time at the interaction layer, and performs continuity detection on the original interactive event streams to obtain the data stream status identifier of each terminal. When the data stream status identifier is characterized as a continuous stream, the interactive behavior record carrying the operation timing is parsed from the original interactive event stream of the corresponding terminal, and the interactive behavior record is organized into a timestamp sequence according to the order of occurrence. For each interaction in the timestamp sequence, determine the active interface element in the tutorial interface when the current interaction occurs, mark the interface element as the focus object, and extract the node type identifier, semantic tag and spatial coordinates of the focus object in the document object model to form the attribute features of the corresponding focus object.

4. The method as described in claim 3, characterized in that, Continuity detection of the original interactive event stream to obtain the data stream status identifier of each terminal specifically includes: Using a preset time window as the granularity of data collection, the original interactive event stream is segmented and the event arrival rate of interactive events within each time window is calculated. The event arrival rate is compared with a preset silence threshold. When the event arrival rate is lower than the silence threshold, the time window is marked as a silent window; otherwise, it is marked as an active window. The alternating distribution characteristics of the silent window and the active window within the sliding observation period are determined, and the continuity of the data stream is determined based on the alternating distribution characteristics. If there are no continuous silent windows within the sliding observation period and the proportion of active windows exceeds a preset ratio, the data stream status of the corresponding terminal is determined as a continuous stream.

5. The method as described in claim 1, characterized in that, Using the timestamp sequence of each terminal as an index, the attention focus flow that allows user attention to jump between multiple terminals is constructed based on the attribute features of the corresponding focus object. Specifically, this includes: Align the timestamp sequence of each terminal with the attribute features of the corresponding focus object according to the timestamp to form a unified timeline record carrying terminal identifier and focus attribute information; The unified timeline records are merged from multiple sources, and the focus objects from different terminals are concatenated into a global focus transfer sequence according to the order in which the interaction occurs. Extract the terminal identifiers of adjacent focus objects at each focus switch in the global focus transfer sequence to form directed transfer pairs that express attention jumping across devices between terminals, and organize all directed transfer pairs into an attention focus stream in chronological order.

6. The method as described in claim 1, characterized in that, The interaction pattern analysis based on time window sliding is performed on the attention focus flow to obtain the non-steady-state interaction feature flow when the user is in an operational distress state, specifically including: A sliding time window of a preset step size is applied to the attention focus stream, and the directed transition pair sub-sequences within each sliding window are extracted sequentially. The terminal switching frequency and the semantic jump rate of the focus object between adjacent directed transition pairs in the directed transition pair sub-sequences are calculated. The terminal switching frequency and the semantic jump rate of the focus object are input into a pre-constructed steady-state deviation metric function, and the deviation index, which characterizes the degree of deviation of the attention shift mode from the steady state within the current sliding window, is output. The deviation indices of each sliding window on the attention focus flow are arranged in time sequence to obtain a deviation index sequence. Continuous intervals in the deviation index sequence where the deviation index exceeds a preset perturbation threshold are extracted. The directed transition pair sub-sequences corresponding to these continuous intervals are marked as non-steady-state interaction feature flows in which the user is in an operational distress state.

7. The method as described in claim 1, characterized in that, Extracting repetitive operation sequences that do not produce valid inputs from the non-steady-state interaction feature stream specifically includes: Based on the timestamp interval and terminal identifier carried by each directed transition pair in the non-steady-state interaction feature stream, the original interaction event stream of the corresponding terminal within the timestamp interval is extracted back to form a set of interaction events to be analyzed. From the set of interactive events to be analyzed, identify consecutive clusters of interactive events with the same type of interactive events and the same focus object, and calculate the similarity of the operation loads of each interactive event in the consecutive clusters of interactive events to obtain an operation load similarity sequence. The continuous interactive events and their operational payloads in the operational payload similarity sequence that have a similarity exceeding a preset repetition threshold are organized into candidate repetitive operational sequences according to their occurrence time. Identify whether a valid state transition occurs in the coaching session state machine after each operation in the candidate repetitive operation sequence, and determine the candidate repetitive operation sequence that does not trigger a valid state transition as a repetitive operation sequence that does not produce a valid result input.

8. A multi-terminal online tutoring collaborative management system based on learning behavior analysis, used to execute the multi-terminal online tutoring collaborative management method based on learning behavior analysis as described in any one of claims 1 to 7, characterized in that, The system includes: The acquisition module is used to extract the timestamp sequence of each terminal's interactive behavior and the attribute features of the corresponding focus object in the current tutoring session; The processing module is used to construct an attention focus flow of user attention jumping between multiple terminals based on the attribute features of the corresponding focus object, using the timestamp sequence of each terminal as an index, and to perform interaction pattern analysis on the attention focus flow based on time window sliding to obtain the non-steady-state interaction feature flow when the user is in an operational confusion state. The processing module is also used to extract repetitive operation sequences that have not produced valid input results from the non-steady-state interaction feature stream, and generate collaborative requests carrying specific problem context and behavioral evidence based on the repetitive operation sequences and the content objects in the interface where they occur. The execution module is used to identify the tutoring terminal that is currently collaborating in the tutoring session based on the collaboration request, and push the collaboration request and the corresponding non-steady-state interaction feature stream summary to the identified tutoring terminal for display.

9. A computer device, characterized in that, The computer device includes a memory and a processor, the memory storing code, and the processor being configured to retrieve the code and execute the multi-terminal online tutoring collaborative management method based on learning behavior analysis as described in any one of claims 1 to 7.

10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the multi-terminal online tutoring collaborative management method based on learning behavior analysis as described in any one of claims 1 to 7.