Digital network interactive teaching method and system
By analyzing students' head movements and pupil changes, and combining this with the teaching progress to generate relevant interactive questions, the problem of misjudgment of inattention in remote teaching was solved. This enabled precise intervention and continuity of the learning process, thereby improving the effectiveness of remote teaching.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- LIANYUNGANG NORMAL COLLEGE
- Filing Date
- 2026-01-28
- Publication Date
- 2026-05-01
AI Technical Summary
Existing remote teaching systems cannot accurately assess the quality of students' visual attention focus, and are prone to misjudging students looking down at notes as distraction, leading to unnecessary interventions. Furthermore, existing intervention methods are disconnected from the teaching content and fail to naturally guide students' attention back.
By acquiring video streams of students' heads and analyzing information on head movements and pupil changes, attentional characteristics are formed. In conjunction with the teaching progress, a recognition benchmark is dynamically established, interactive questions related to the teaching content are generated, and guidance is provided through voice and visual elements.
It enables precise capture and adaptive intervention of students' attention, avoids unnecessary teaching interference, improves the accuracy of inattention identification and the relevance of teaching content, and enhances the continuity of learning and the effect of knowledge consolidation.
Smart Images

Figure CN121963326A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of digital network teaching technology, and in particular to a digital network interactive teaching method and system. Background Technology
[0002] In real-world remote learning scenarios, students may briefly look down or deviate from the screen while thinking about questions or taking notes, which is part of the effective learning process; however, they may also have their eyes unfocused and their heads in a non-engaged posture for extended periods due to fatigue or external distractions.
[0003] Existing technologies lack continuous and detailed assessment of the quality of visual attention focus, relying solely on coarse-grained posture information, which easily leads to misjudgments. For example, a student's act of looking down at notes may be misjudged as distraction, triggering intervention that disrupts the continuity of learning and deviates from the original intention of teaching assistance. Furthermore, even if distraction is identified, existing intervention methods are often disconnected from the current teaching content, resulting in rigid interventions that fail to naturally and smoothly guide students' attention back to the specific teaching context. Summary of the Invention
[0004] This invention provides a digital network interactive teaching method and system to solve the problems mentioned in the background art.
[0005] To achieve the above objectives, the present invention provides a digital network interactive teaching method, comprising:
[0006] S1. Obtain the video stream containing the student's head collected by the teaching client, and extract the head movement trajectory and pupil change information;
[0007] S2. Correlate head movement trajectories with pupil changes to form attentional characteristics that reflect cognitive states;
[0008] S3. Based on the historical trends of attention characteristics and the current teaching progress, dynamically establish the identification criteria for inattention;
[0009] S4. When the recognition benchmark indicates inattention, generate an interactive question directly related to the teaching content being taught.
[0010] S5. Drives the teaching client to play the audio of the interactive questions and detects students' confirmatory tactile actions based on the video stream;
[0011] S6. When a confirmatory haptic action is detected, control the teaching client to display the full content of the interactive question on the demonstration interface.
[0012] Preferably, the step of correlating and analyzing head movement trajectories with pupil change information to form attentional features reflecting cognitive states includes:
[0013] The head movement trajectory and pupil change information are synchronized in time to obtain a synchronized movement sequence and a synchronized pupil sequence with time consistency.
[0014] Based on the extraction of the coordinated change relationship between the synchronous motion sequence and the synchronous pupil sequence, a description of the coordinated change characterizing the degree of head and eye movement coordination is obtained;
[0015] Based on the pre-defined attention mapping relationship, the description of collaborative changes is transformed into the corresponding attention state identifier, thereby obtaining attention features that reflect the cognitive state.
[0016] Preferably, the step of dynamically establishing a benchmark for identifying inattention based on the historical trends of attention characteristics and the current teaching progress includes:
[0017] Determine the current teaching stage type based on the teaching event markers that are synchronized with the current teaching progress;
[0018] Based on the mapping relationship between the benchmark rigor and the different teaching stage types, determine the current required benchmark rigor;
[0019] Based on the baseline rigor, the attention characteristics generated during the current teaching progress period are fitted with a baseline to directly obtain the identification baseline for inattention.
[0020] Preferably, the step of fitting a baseline to the attentional characteristics generated during the current teaching progress period based on the baseline strictness to directly obtain the identification baseline for inattention dissipation includes:
[0021] The distribution characteristics of attention generated during the current teaching progress period are analyzed to obtain a characteristic distribution description representing the degree of attention concentration at the current stage;
[0022] Based on the baseline strictness, determine the sensitivity parameters for interpreting the feature distribution description;
[0023] Based on feature distribution description and sensitivity parameters, a dynamic discrimination condition for instantaneous inattention is constructed as a recognition benchmark.
[0024] Preferably, generating an interactive question directly related to the currently taught teaching content includes:
[0025] Semantic structure analysis is performed on the currently taught teaching content to extract at least one core concept entity and its relationships, resulting in a knowledge graph fragment of the current teaching content;
[0026] Select a target concept entity from the knowledge graph fragment that has not appeared in the recent interaction history to determine the focus of the interaction assessment;
[0027] Based on the interactive assessment focus and its relationship with knowledge graph fragments, question stems are constructed to form interactive questions directly related to the teaching content.
[0028] Preferably, after constructing the question stem, the method further includes:
[0029] Based on the deviation patterns of attention characteristics from the identification benchmark, and in conjunction with the current teaching stage, determine the intervention tendency category of inattention;
[0030] Based on the intervention tendency category and the focus of the interactive assessment, select appropriate guiding expression elements;
[0031] By combining guiding elements with the constructed question stem, interactive questions with guiding functions are formed.
[0032] Preferably, the method of driving the teaching client to play the audio of the interactive question includes:
[0033] During the initial period when the audio of the question begins to play, features of the students' real-time facial orientation and gaze direction are extracted based on the video stream to obtain the initial attention feedback features during the audio playback phase.
[0034] Based on the initial attention feedback characteristics, the subsequent playback rhythm and repetition strategy of the question stem audio are dynamically adjusted to form adaptive audio playback instructions.
[0035] Based on the adaptive voice playback instructions, control the teaching client to play the audio of the question.
[0036] Preferably, the method of detecting students' confirmatory somatosensory actions based on video stream includes:
[0037] Throughout the entire playback of the question's audio and within a preset time period after its end, the student's head posture sequence in the video stream is continuously matched and analyzed with a preset response intention action library to obtain a preliminary judgment of response intention.
[0038] By combining the initial attention feedback features extracted during the speech playback phase, the confidence level of the preliminary response intention judgment is verified to obtain a multimodal confirmation result.
[0039] When the multimodal confirmation result represents the intention to confirm, a dynamic attention focus mapping is established in the video stream based on the teaching content area related to the interactive question, and a confirmatory somatosensory action judgment is obtained to trigger the next operation.
[0040] Preferably, the control teaching client displays the complete content of the interactive questions on the demonstration interface, including:
[0041] Based on the semantic structure of the interactive questions and the possible types of attention distraction, determine the targeted content display logic;
[0042] Based on the content display logic, the complete content of the interactive question is visually decomposed and spatially laid out to obtain a dynamic content display solution.
[0043] According to the dynamic content display scheme, the visual elements are serialized and rendered on the demonstration interface to complete the display of the full content of the interactive questions.
[0044] To address the above problems, the present invention also provides a digital network interactive teaching system, the system comprising:
[0045] The data acquisition module is used to acquire video streams containing students' heads collected by the teaching client and extract head movement trajectories and pupil change information;
[0046] The attention feature formation module is used to correlate and analyze head movement trajectories with pupil change information to form attention features that reflect cognitive state.
[0047] The baseline establishment module is used to dynamically establish a baseline for identifying inattention based on the historical trends of attention characteristics and the current teaching progress.
[0048] The interactive question generation module is used to generate an interactive question that is directly related to the teaching content being taught, based on the current teaching content, when the recognition benchmark determines that the student is inattentive.
[0049] The confirmatory somatosensory action confirmation module is used to drive the teaching client to play the audio of the interactive questions and detect the student's confirmatory somatosensory actions based on the video stream;
[0050] The interactive question display module is used to control the teaching client to display the full content of the interactive question on the demonstration interface when a confirmatory haptic action is detected.
[0051] Compared with the prior art, the present invention has the following beneficial effects:
[0052] 1. By correlating and analyzing head movement trajectories with pupil changes to form attentional characteristics, the system accurately captures the coordinated changes in head and eye movements, effectively distinguishing between effective behaviors and genuine inattention during the learning process, ensuring the accuracy of attention assessment. Furthermore, by dynamically establishing identification benchmarks in conjunction with the teaching progress, the system adapts attentional inattention assessment to the attentional needs of different teaching stages, ensuring precise intervention timing, avoiding unnecessary teaching interference, and effectively maintaining the continuity of the learning process.
[0053] 2. Based on the semantic analysis of the current teaching content, relevant interactive questions are generated and guided expression elements are incorporated. This not only achieves a deep integration of attention intervention and teaching content, but also guides students to quickly connect with the knowledge context. By dynamically adjusting the audio playback strategy of the questions and combining multimodal data to verify confirmatory haptic actions, the effectiveness of interactive question delivery and the reliability of response judgment are improved, further enhancing the attention recall effect. At the same time, it helps students consolidate the current teaching focus and achieves synergistic effect between attention intervention and knowledge learning. Attached Figure Description
[0054] Figure 1 A flowchart of a digital network interactive teaching method provided by the present invention;
[0055] Figure 2 This invention provides a modular structure diagram of a digital network interactive teaching system.
[0056] The realization of the objective, functional features and advantages of the present invention will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation
[0057] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of the present invention.
[0058] Example 1, referring to Figure 1 The diagram shown is a flowchart illustrating a digital network interactive teaching method according to an embodiment of the present invention. In this embodiment, the digital network interactive teaching method includes:
[0059] S1. Obtain the video stream containing the student's head collected by the teaching client, and extract the head movement trajectory and pupil change information;
[0060] S2. Correlate head movement trajectories with pupil changes to form attentional characteristics that reflect cognitive states;
[0061] S3. Based on the historical trends of attention characteristics and the current teaching progress, dynamically establish the identification criteria for inattention;
[0062] S4. When the recognition benchmark indicates inattention, generate an interactive question directly related to the teaching content being taught.
[0063] S5. Drives the teaching client to play the audio of the interactive questions and detects students' confirmatory tactile actions based on the video stream;
[0064] S6. When a confirmatory haptic action is detected, control the teaching client to display the full content of the interactive question on the demonstration interface.
[0065] In this embodiment, a teaching tablet with a high-definition camera is used as the client. After the course is started, the tablet camera automatically acquires the video stream of the student's front and calls the camera acquisition permission through system permissions to ensure that the video covers the entire head area of the student.
[0066] Specifically, regarding head motion trajectory extraction, 68 facial feature points, such as the corners of the eyes and the jaw angle, can be detected using the Dlib algorithm. The coordinate changes of each point between frames are recorded in real time to generate three-dimensional motion trajectory data of the head's up-down, left-right, and pitch movements.
[0067] Dlib is an open-source, cross-platform machine learning and computer vision library that encapsulates a large number of mature and efficient computer vision algorithms, specifically designed for scenarios such as face detection, feature point extraction, object tracking, and image recognition.
[0068] Specifically, regarding the extraction of pupil change information, the eye region in the video stream can be segmented into an area of interest (ROI), the pupil center can be located using a grayscale thresholding method, the pupil diameter and center offset can be calculated for each frame, and dynamic information on pupil contraction, dilation, and gaze deviation can be extracted simultaneously. This information is then stored in alignment with head movement trajectory data by timestamp for subsequent analysis.
[0069] In the implementation of this invention, the head movement trajectory and pupil change information are correlated and analyzed to form attentional features reflecting cognitive state, including:
[0070] The head movement trajectory and pupil change information are synchronized in time to obtain a synchronized movement sequence and a synchronized pupil sequence with time consistency.
[0071] Based on the extraction of the coordinated change relationship between the synchronous motion sequence and the synchronous pupil sequence, a description of the coordinated change characterizing the degree of head and eye movement coordination is obtained;
[0072] Based on the pre-defined attention mapping relationship, the description of collaborative changes is transformed into the corresponding attention state identifier, thereby obtaining attention features that reflect the cognitive state.
[0073] In practice, taking online live-streaming classes of advanced mathematics as the implementation scenario, the teaching client collects video streams of students' heads, and divides the screen into dedicated areas such as formula derivation area and example calculation area. Using the video stream acquisition frame rate as a unified time reference, the spatial position and movement direction of the head relative to the dedicated area in each frame are first extracted to form the original information of the head movement trajectory.
[0074] Then, the pupil diameter change and gaze direction within the same frame are extracted to form the original pupil change information. Then, the timestamps of the two types of information are matched one by one according to the time node of the acquisition frame, so that each head movement trajectory data point corresponds to the pupil change data point at the same time node. Finally, they are integrated to form a synchronous motion sequence and a synchronous pupil sequence with corresponding data of the advanced mathematics teaching area.
[0075] Furthermore, we first analyze the correlation and matching degree between the head movement trend towards the advanced mathematics-specific region in the synchronous motion sequence and the change trend of the pupil gaze corresponding region in the synchronous pupil sequence. When the head is stably facing the formula derivation area and the pupil gaze region is in the core position, it is determined to be positive coordination.
[0076] If the head deviates from the designated area or the pupils gaze at unrelated areas such as the desktop or the edge of the screen, it is determined to be non-coordinated. The coordination determination results of each time point are then integrated to form a description of the coordination changes of head and eye movement during different teaching periods of the university advanced mathematics live class, including the type of head and eye movement coordination, the duration of positive coordination, and the specific nodes that trigger non-coordinated movements.
[0077] Furthermore, the attention mapping relationship can be established by collecting head movement and eye movement coordination data of first-year engineering students in the scenario of online live-streaming classes of advanced mathematics in universities. Attention is divided into three categories: concentrated, slightly distracted, and heavily distracted according to the proportion of positive coordination duration and the frequency of non-coordination triggers. After multiple rounds of actual testing and calibration in the same advanced mathematics scenario and removing non-attention-related data such as students flipping through electronic textbooks and briefly drinking water, the attention mapping relationship is established.
[0078] Finally, the collaborative type, duration, and trigger frequency in the collaborative change description are matched one by one with the mapping relationship, thereby transforming them into attention state identifiers that include the current advanced mathematics teaching period, dedicated area collaborative data, and attention state level. Finally, these attention state identifiers are integrated to obtain attention features that reflect the student's cognitive state.
[0079] In summary, this solution eliminates the need for single-frame interpolation by using frame-block level synchronous alignment, thus avoiding the collaborative analysis errors caused by noise in single-frame data. At the same time, the overall feature fit analysis can capture the continuous collaborative patterns of head and eye movements, rather than fragmented single-frame data statistics, making the description of collaborative changes more consistent with the actual cognitive state changes of students, thereby achieving a dual improvement in the anti-interference and continuity of attention feature recognition.
[0080] On the other hand, by establishing a personalized baseline based on the student's initial state of focus, the identification bias caused by the innate differences in the student's eye movement and head movement habits in the standardized mapping relationship is completely eliminated, making the determination of attention characteristics more suitable for the student's personal characteristics and achieving personalized and accurate identification.
[0081] In embodiments of the present invention, a benchmark for identifying inattention is dynamically established based on the historical trends of attention characteristics and the current teaching progress, including:
[0082] Determine the current teaching stage type based on the teaching event markers that are synchronized with the current teaching progress;
[0083] Based on the mapping relationship between the benchmark rigor and the different teaching stage types, determine the current required benchmark rigor;
[0084] Based on the baseline rigor, the attention characteristics generated during the current teaching progress period are fitted with a baseline to directly obtain the identification baseline for inattention.
[0085] In embodiments of the present invention, baseline fitting is performed on the attention features generated during the current teaching progress period to directly obtain the identification benchmark for inattention dispersion, including:
[0086] The distribution characteristics of attention generated during the current teaching progress period are analyzed to obtain a characteristic distribution description representing the degree of attention concentration at the current stage;
[0087] Based on the baseline strictness, determine the sensitivity parameters for interpreting the feature distribution description;
[0088] Based on feature distribution description and sensitivity parameters, a dynamic discrimination condition for instantaneous inattention is constructed as a recognition benchmark.
[0089] To quantify the dynamic discrimination conditions, a dynamic threshold function can be introduced. Assume that N attention feature values are collected during the current teaching period, forming a set. .
[0090] The formula for the dynamic discrimination threshold is as follows:
[0091] ;
[0092] In the formula, This represents the set of attentional features generated within the current teaching progress period. Represents the set of attention features The arithmetic mean, Represents the set of attention features standard deviation Represents the sensitivity coefficient. Represents the sensitivity coefficient function. This indicates a dynamic threshold.
[0093] When the instantaneous attentional characteristic is less than the dynamic threshold, it is instantaneously judged as inattentiveness.
[0094] In practice, taking online live-streaming courses in higher mathematics as the implementation scenario, the teaching client can synchronously record teaching event markers that are linked to the current teaching progress. These markers correspond to specific aspects of the higher mathematics course, such as formula derivation, example calculation, theorem review, and exercise Q&A.
[0095] By comparing the event markers corresponding to the current progress, the current teaching stage type is determined. For example, if the current progress is explaining the derivation of the integral formula of multivariable functions, the teaching stage type is determined to be the formula derivation stage. The feature distribution description should include the collaborative data and attention state ratio information of the formula derivation area corresponding to this stage.
[0096] Furthermore, attentional characteristic data from different teaching stages of live-streamed advanced mathematics courses for first-year engineering students can be collected, covering all aspects such as formula derivation and problem-solving. A baseline rigor level can be set according to the difficulty of knowledge at each stage and the students' attentional needs. The formula derivation and theorem review stages have high attentional requirements, so the baseline rigor is set to high, while the problem-solving stage is set to medium. After N rounds of real-world testing and calibration in the same scenario, and after removing abnormal data caused by equipment lag, a mapping relationship between different teaching stage types and baseline rigor can be established. Then, based on the currently determined formula derivation stage type, the required baseline rigor can be matched from the mapping relationship.
[0097] Furthermore, based on high-level benchmark rigor, all attention features generated during the current multivariate function integral formula derivation period can be extracted. The frequency and duration of concentrated, slightly distracted, and heavily distracted states among these features can be statistically analyzed. The clustering range and dispersion of feature values can be analyzed to obtain a feature distribution description representing the degree of attention concentration at this stage. This description needs to be associated with head and eye movement coordination data in the formula derivation area to clarify the feature value distribution intervals corresponding to different coordination states.
[0098] Furthermore, based on the high-level benchmark rigor, a sensitivity parameter for interpreting the feature distribution description can be determined. When the benchmark rigor is high, the sensitivity parameter is adjusted to a level that can capture slight attentional deviations, ensuring accurate perception of students' attentional changes during the formula derivation stage. This parameter is obtained by matching the correspondence between the benchmark rigor and the sensitivity parameter, and has been verified through multiple rounds of testing to be suitable for the attentional judgment requirements during the advanced mathematics formula derivation stage.
[0099] Furthermore, all attention feature values within the current formula derivation period can be collected to form a set. The arithmetic mean of all feature values in the set can be calculated as the mean. Then, the deviation of each feature value from the mean can be calculated to obtain the standard deviation. The sensitivity coefficient is set according to the high-level benchmark strictness. The higher the benchmark strictness, the larger the sensitivity coefficient value. The sensitivity coefficient function increases as the sensitivity coefficient increases.
[0100] The dynamic threshold is the product of the mean, the sensitivity coefficient function, and the standard deviation. It defines the boundary for judging inattention during the formula derivation stage. The larger the sensitivity coefficient, the lower the dynamic threshold, and the easier it is to judge inattention. Based on these parameters, dynamic discrimination conditions are constructed as the recognition benchmark for inattention.
[0101] Furthermore, the instantaneous attention characteristics of students during the formula derivation stage can be extracted in real time. These characteristics are associated with the students' collaborative state data facing the formula derivation area. When the instantaneous attention characteristics are less than the dynamic threshold, the student is judged to be inattentive. If the instantaneous attention characteristics are greater than or equal to the dynamic threshold, the student is judged to be focused. At the same time, valid behaviors that are not inattentive, such as students briefly flipping through electronic handouts or looking down to take notes, are excluded, thus clarifying the contextual judgment boundary.
[0102] In summary, this solution deeply integrates historical trends in attention characteristics with teaching progress and dynamically adjusts the rigor of the benchmark in conjunction with the type of teaching stage. This avoids the problem of mismatch between traditional fixed benchmarks and teaching scenarios and knowledge difficulty, allowing the recognition benchmark to accurately adapt to the attention requirements of different teaching stages.
[0103] Furthermore, by analyzing feature distribution and adapting sensitivity parameters, dynamic discrimination conditions are constructed, abandoning the single threshold judgment mode. This not only captures the instantaneous changes in attention distraction but also avoids misjudgments caused by fluctuations in normal individual behavior, thus improving the accuracy and flexibility of attention distraction recognition.
[0104] Overall, dynamic benchmarks can adapt to changes in students' attention levels in real time as the teaching progresses, quickly respond to differences in attention needs at different teaching stages, provide a scientific basis for accurately generating interactive questions and timely intervention in attention deficit, and help teaching interventions be implemented efficiently.
[0105] In embodiments of the present invention, an interactive question directly related to the currently taught teaching content is generated, including:
[0106] Semantic structure analysis is performed on the currently taught teaching content to extract at least one core concept entity and its relationships, resulting in a knowledge graph fragment of the current teaching content;
[0107] Select a target concept entity from the knowledge graph fragment that has not appeared in the recent interaction history to determine the focus of the interaction assessment;
[0108] Based on the interactive assessment focus and its relationship with knowledge graph fragments, question stems are constructed to form interactive questions directly related to the teaching content.
[0109] In embodiments of the present invention, after constructing the question stem, the method further includes:
[0110] Based on the deviation patterns of attention characteristics from the identification benchmark, and in conjunction with the current teaching stage, determine the intervention tendency category of inattention;
[0111] Based on the intervention tendency category and the focus of the interactive assessment, select appropriate guiding expression elements;
[0112] By combining guiding elements with the constructed question stem, interactive questions with guiding functions are formed.
[0113] The candidate elements are quantitatively evaluated using the following guiding element matching degree function, and the one with the highest matching degree is selected as the optimal option. The formula for calculating the guiding element matching degree is as follows:
[0114] ;
[0115] In the formula, This indicates the score for the matching degree of the guiding element. Indicates the category of intervention tendency. The default category label for the guide element. Represents the category correlation function, This indicates the focus of the interactive assessment. This represents a vector of related knowledge points that guide the element. Represents the semantic similarity function. This represents the category matching weight coefficient. This represents the semantic relevance weight coefficient.
[0116] In practice, taking online live-streaming classes of advanced mathematics as the implementation scenario, when it is determined that students' attention is scattered based on the recognition benchmark, semantic structure analysis can be performed on the content currently being taught.
[0117] For example, when teaching the application of partial derivatives in multivariable differential calculus, we can break down the semantic logic of the content and extract the core conceptual entities, such as the geometric meaning of partial derivatives, the chain rule of partial derivatives of composite functions, and the solution of partial derivatives of implicit functions.
[0118] At the same time, the relationships between concepts are sorted out. For example, solving the partial derivative of a composite function requires relying on the definition of partial derivative, and solving the partial derivative of an implicit function can be extended by combining the chain rule. This leads to a knowledge graph fragment of the current teaching content that includes the relevant concepts and relationships of the application of partial derivatives in calculus, as well as the teaching content of the derivation of related formulas.
[0119] Furthermore, the core conceptual entities involved in the most recent three rounds of interactive questions can be retrieved. For example, if the previous interactions have examined the chain rule of partial derivatives of composite functions and the geometric meaning of partial derivatives, the solution of implicit function partial derivatives that has not appeared in the knowledge graph fragments can be selected as the target conceptual entity. This concept can be identified as the focus of the interactive examination, ensuring that the focus is in line with the current teaching progress and avoiding repeated examination.
[0120] Furthermore, based on the interactive examination focus of solving implicit partial derivatives, and combined with its relationship with the chain rule in the knowledge graph fragment, we can construct question stems, such as "How to solve the first-order partial derivatives of an implicit function determined by a bivariate equation using the chain rule of partial derivatives of composite functions?", forming an interactive question directly related to the current teaching content on the application of partial derivatives.
[0121] Furthermore, mild inattention can be identified based on the deviation pattern of attention features from the recognition benchmark, such as feature values slightly below the dynamic threshold and short duration.
[0122] Then, considering the current teaching stage as the concept application and explanation stage, the intervention category for distracted attention is determined to be the guidance and reinforcement category. Non-distracted effective behaviors such as students looking down to record formulas or briefly flipping through electronic handouts are excluded, and the boundaries of the scenario-based judgment are clarified.
[0123] Furthermore, based on the interactive examination focus of guiding reinforcement intervention tendency categories and solving implicit function partial derivatives, suitable guiding expression elements can be selected from the preset guiding element library. For example, review-style expression elements that relate to the content already taught in class, such as "combining the chain rule formula we just derived," can ensure that the elements both meet the intervention needs and fit the examination focus.
[0124] Finally, the candidate guiding expression elements can be quantitatively evaluated through weighted calculation. The category matching degree weight coefficient and semantic relevance degree weight coefficient can be preset based on the advanced mathematics teaching scenario.
[0125] Among them, the category matching degree weight coefficient is higher than the semantic relevance degree weight coefficient, because the adaptability of intervention tendency takes precedence over semantic relevance. The category relevance function judges the degree of fit between the preset category label of the guiding element and the intervention tendency of the guidance reinforcement category, and the semantic similarity function judges the semantic closeness between the knowledge points associated with the guiding element and the focus of the implicit function partial derivative solution.
[0126] Specifically, the matching score is a weighted result of the two, and its significance is to select the most suitable guiding element. The larger the weight coefficient, the more significant the impact of the corresponding item on the score.
[0127] Finally, based on the weighted results, the element with the highest matching score is selected and combined with the constructed question stem, such as "Combining the chain rule formula we just derived, how to solve the first-order partial derivative of the implicit function determined by the bivariate equation using the chain rule of partial derivatives of the composite function?" This forms an interactive question with guiding function. The weight coefficients are calibrated through multiple rounds of testing and verification in the same scenario of advanced mathematics to ensure that they are suitable for teaching intervention needs.
[0128] In summary, this solution deeply integrates the results of attention distraction assessment with real-time teaching content, abandons the traditional general interactive format that is detached from the teaching context, extracts the core concepts and relationships of the current teaching content through semantic parsing, accurately identifies knowledge points that have not been repeatedly tested as the focus of interaction, and constructs questions accordingly.
[0129] Overall, this approach ensures that interactive questions are closely related to the current content being taught, avoiding irrelevant questions from distracting students or disrupting the teaching pace. It also effectively draws students' attention back to the current knowledge points, creating a synergy between attention intervention and knowledge consolidation.
[0130] At the same time, by matching guiding statements with the degree of attention deficit, the suitability and guiding effect of the questions are further improved. This allows interactive intervention to not only quickly bring students' attention back, but also to strengthen their understanding and memory of the current teaching focus through the questions. This removes attentional obstacles for the smooth progress of subsequent teaching content and significantly improves the pertinence and effectiveness of digital network interactive teaching.
[0131] In embodiments of the present invention, driving the teaching client to play the audio of the interactive question includes:
[0132] During the initial period when the audio of the question begins to play, features of the students' real-time facial orientation and gaze direction are extracted based on the video stream to obtain the initial attention feedback features during the audio playback phase.
[0133] Based on the initial attention feedback characteristics, the subsequent playback rhythm and repetition strategy of the question stem audio are dynamically adjusted to form adaptive audio playback instructions.
[0134] Based on the adaptive voice playback instructions, control the teaching client to play the audio of the question.
[0135] In embodiments of the present invention, detecting students' confirmatory somatosensory actions based on video streams includes:
[0136] Throughout the entire playback of the question's audio and within a preset time period after its end, the student's head posture sequence in the video stream is continuously matched and analyzed with a preset response intention action library to obtain a preliminary judgment of response intention.
[0137] By combining the initial attention feedback features extracted during the speech playback phase, the confidence level of the preliminary response intention judgment is verified to obtain a multimodal confirmation result.
[0138] When the multimodal confirmation result represents the intention to confirm, a dynamic attention focus mapping is established in the video stream based on the teaching content area related to the interactive question, and a confirmatory somatosensory action judgment is obtained to trigger the next operation.
[0139] In practice, taking online live-streaming classes of advanced mathematics as the implementation scenario, the current interactive problem is related to solving the partial derivatives of implicit functions. The teaching client can be driven to play the audio of the interactive problem. During the initial period of playback, the real-time facial orientation and gaze direction features of students are extracted based on the video stream. For example, it can detect whether the face is facing the formula derivation area on the screen and whether the gaze is focused on the core formula position in that area.
[0140] If the face is facing the screen and the gaze is steadily fixed on the formula derivation area, the initial attention feedback feature is characterized as good feedback. If the face is off the screen or the gaze is scattered on the desktop, the feedback is characterized as weak. Thus, the initial attention feedback feature of the voice playback stage is obtained. This feature is associated with the current teaching progress of partial derivatives and the visual data of the formula derivation area.
[0141] Furthermore, based on the extracted initial attention feedback features, the subsequent playback rhythm and rereading strategy of the question stem audio can be dynamically adjusted. If the initial feedback features indicate weak feedback, the playback rhythm can be slowed down and the keyword dwell time can be extended.
[0142] At the same time, the content related to the focus of advanced mathematics examinations, such as "chain rule", "implicit function" and "first-order partial derivative", is repeated. If the feedback is good, the normal rhythm is maintained, thus forming an adaptive voice playback instruction. The instruction includes rhythm parameters, repeated keywords and corresponding playback duration, which fits the current teaching scenario of partial derivative application.
[0143] Furthermore, based on the adaptive voice playback command, the teaching client can be controlled to play the question's audio. For example, the combined audio of the question can be played in its entirety according to the adjusted rhythm, and the display status of the formula derivation area on the screen can be synchronized to ensure that the audio playback is synchronized with the display of the formulas related to the implicit function partial derivatives in the area, accurately conveying the core information of the interactive question.
[0144] Furthermore, the system can continuously match and analyze the student head posture sequence in the video stream with a preset response intention action library throughout the entire playback of the question's audio and for a preset time period after its end.
[0145] It should be noted that this response intention action database can be constructed by collecting the response habits of students in online calculus classes. It includes confirmation intention actions such as nodding and looking up to steadily gaze at the formula derivation area, as well as negative intention actions such as shaking the head and looking down to avoid the screen. After multiple rounds of testing and calibration in the same calculus scenario to eliminate abnormal actions, it is finalized. For example, if two consecutive nodding actions are detected, the initial response intention is judged to be that there is a confirmation intention.
[0146] Furthermore, the confidence level of the initial response intention judgment can be verified by combining the initial attention feedback features extracted during the audio playback phase. If the initial judgment is a confirmation intention and the initial feedback features indicate good attention, the confidence level is increased. If the initial judgment is a confirmation intention but the initial feedback features indicate that the gaze is continuously deviating from the formula derivation area, the confidence level is decreased. This excludes non-effective response behaviors such as students unconsciously nodding or looking down to record formulas, clarifies the contextual judgment boundary, and thus obtains multimodal confirmation results.
[0147] Finally, when the multimodal confirmation result represents the intention to confirm, a dynamic attention focus mapping can be established in the video stream based on the teaching content area of the formula derivation area related to the problem of solving the partial derivatives of implicit functions. The formula derivation area is used as the core mapping area to accurately associate the teaching content corresponding to the interactive question, thereby obtaining the confirmatory haptic action judgment used to trigger the next operation, ensuring that the judgment result is deeply bound to the current advanced mathematics teaching content.
[0148] In summary, this solution constructs a multimodal interactive closed loop of "voice playback - real-time feedback - action confirmation," abandoning the fragmented interaction mode of traditional one-way voice broadcasting or single action detection. By simultaneously extracting initial attention feedback features such as students' facial orientation and gaze direction during the initial period of the question's voice playback, the solution dynamically adjusts the voice playback rhythm and repetition strategy. At the same time, it combines continuous matching analysis of students' head posture sequences throughout and after the voice playback, and incorporates the initial attention feedback features for confidence verification, forming a multi-dimensional confirmation result.
[0149] Overall, this solution dynamically adapts the audio playback format based on real-time attention feedback, enabling the audio to reach students with scattered attention more accurately, avoiding students missing key information due to fixed playback patterns, and effectively improving the effectiveness of audio delivery.
[0150] On the other hand, by using multimodal data fusion to verify confirmatory somatosensory actions, we can avoid the misjudgments that may occur with single-action detection, such as students’ unconscious actions being misjudged as confirmation intentions. We can also accurately capture students’ true response intentions and ensure the reliability of action judgment. Ultimately, we can achieve efficient connection between interactive question delivery and student response confirmation. This not only ensures the continuity of interactive intervention, but also provides a precise trigger for subsequent complete question presentations, further enhancing the timeliness of attention recall and the closed-loop effect of interactive teaching.
[0151] In embodiments of the present invention, controlling the teaching client to display the complete content of interactive questions on the demonstration interface includes:
[0152] Based on the semantic structure of the interactive questions and the possible types of attention distraction, determine the targeted content display logic;
[0153] Based on the content display logic, the complete content of the interactive question is visually decomposed and spatially laid out to obtain a dynamic content display solution.
[0154] According to the dynamic content display scheme, the visual elements are serialized and rendered on the demonstration interface to complete the display of the full content of the interactive questions.
[0155] In practice, taking online live-streaming classes in higher mathematics as the implementation scenario, when a confirmatory haptic action is detected, the teaching client can be controlled to display the complete content of interactive problems related to solving implicit function partial derivatives on the demonstration interface. Combining the semantic structure of the interactive problem with the possible types of mild attention dissipation, a targeted content display logic can be determined. For example, the related chain rule formula can be presented first to lay the foundation for knowledge points, and then the complete question can be displayed to help students quickly connect with the content they have learned and adapt to the need for recalling mild attention dissipation.
[0156] Furthermore, based on the established content display logic, the complete content of the interactive question can be visually decomposed and spatially laid out, resulting in visual elements such as chain rule formula graphics, question stem text, and core keyword annotations. The keyword annotations target content such as "implicit function" and "first-order partial derivative." The layout plan places the formula above the formula derivation area on the screen, the question stem text below, and the keywords in a prominent style, resulting in a dynamic content display scheme that is relevant to the advanced mathematics teaching area and adapted to attention recall.
[0157] Finally, following the dynamic content display scheme, the visual elements can be rendered and presented sequentially on the demonstration interface. First, the chain rule formula is rendered and paused briefly, then the question text is rendered step by step, and core keywords are highlighted in sync. This ensures that the order of element presentation matches the logical connection of knowledge points, completes the display of the full content of the interactive question, and the display effect accurately corresponds to the teaching content in the formula derivation area, helping students focus on the core issues.
[0158] In summary, this solution first determines the presentation logic based on the type of attention deficit and the semantics of the question, which makes the presentation of the question more in line with the students' attention recall needs and avoids comprehension obstacles caused by the disconnect between the presentation logic and the students' cognitive state.
[0159] Furthermore, by breaking down and scientifically arranging visual elements, the core test points and related teaching content are highlighted, enhancing visual guidance and helping students whose attention has just been refocused to quickly refocus on key information.
[0160] Overall, this solution achieves efficient delivery of the complete content of interactive questions, ensuring that students can clearly understand the core of the questions and consolidating the attention recall effect through display formats adapted to teaching scenarios. This allows interactive intervention to form a complete closed loop of "voice wake-up - action confirmation - visual focus", further enhancing the coherence and effectiveness of digital network interactive teaching.
[0161] Example 2, as Figure 2 The diagram shown is a modular structure diagram of a digital network interactive teaching system provided by the present invention, which includes:
[0162] The data acquisition module 101 is used to acquire the video stream containing the student's head collected by the teaching client, and extract the head movement trajectory and pupil change information;
[0163] Attention feature formation module 102 is used to correlate and analyze head movement trajectory with pupil change information to form attention features that reflect cognitive state;
[0164] The baseline establishment module 103 is used to dynamically establish a baseline for identifying inattention based on the historical trends of attention characteristics and the current teaching progress.
[0165] The interactive question generation module 104 is used to generate an interactive question that is directly related to the teaching content, based on the teaching content being taught, when the inattention is determined based on the recognition benchmark.
[0166] The confirmatory somatosensory action confirmation module 105 is used to drive the teaching client to play the audio of the interactive question and detect the student's confirmatory somatosensory actions based on the video stream.
[0167] The interactive question display module 106 is used to control the teaching client to display the full content of the interactive question on the demonstration interface when a confirmatory haptic action is detected.
[0168] In the several embodiments provided by this invention, it should be understood that the disclosed methods and systems can be implemented in other ways. For example, the system embodiments described above are merely illustrative; for instance, the division of modules is only a logical functional division, and other division methods may be used in actual implementation.
[0169] The modules described as separate components may or may not be physically separate. The components shown as modules may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs.
[0170] Furthermore, the functional modules in the various embodiments of the present invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or in the form of hardware plus software functional modules.
[0171] It will be apparent to those skilled in the art that the present invention is not limited to the details of the exemplary embodiments described above, and that the present invention can be implemented in other specific forms without departing from the spirit or essential characteristics of the present invention.
[0172] The embodiments of this application can acquire and process relevant data based on artificial intelligence technology. Artificial intelligence is the theory, method, technology, and application system that uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use that knowledge to obtain optimal results.
[0173] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the technical solutions of the present invention.
Claims
1. A digital network interactive teaching method, characterized in that, The method includes: S1. Obtain the video stream containing the student's head collected by the teaching client, and extract the head movement trajectory and pupil change information; S2. Correlate head movement trajectories with pupil changes to form attentional characteristics that reflect cognitive states; S3. Based on the historical trends of attention characteristics and the current teaching progress, dynamically establish the identification criteria for inattention; S4. When the recognition benchmark indicates inattention, generate an interactive question directly related to the teaching content being taught. S5. Drives the teaching client to play the audio of the interactive questions and detects students' confirmatory tactile actions based on the video stream; S6. When a confirmatory haptic action is detected, control the teaching client to display the full content of the interactive question on the demonstration interface.
2. The digital network interactive teaching method as described in claim 1, characterized in that, The process of correlating head movement trajectories with pupil changes to form attentional features reflecting cognitive states includes: The head movement trajectory and pupil change information are synchronized in time to obtain a synchronized movement sequence and a synchronized pupil sequence with time consistency. Based on the extraction of the coordinated change relationship between the synchronous motion sequence and the synchronous pupil sequence, a description of the coordinated change characterizing the degree of head and eye movement coordination is obtained; Based on the pre-defined attention mapping relationship, the description of collaborative changes is transformed into corresponding attention state identifiers, thereby obtaining attention features that reflect cognitive states.
3. The digital network interactive teaching method as described in claim 1, characterized in that, The method of dynamically establishing a benchmark for identifying inattention based on the historical trends of attention characteristics and the current teaching progress includes: Determine the current teaching stage type based on the teaching event markers that are synchronized with the current teaching progress; Based on the mapping relationship between the benchmark rigor and the different teaching stage types, determine the current required benchmark rigor. Based on the baseline rigor, the attention characteristics generated during the current teaching progress period are fitted with a baseline to directly obtain the identification baseline for inattention.
4. The digital network interactive teaching method as described in claim 3, characterized in that, The process of fitting a baseline to the attention characteristics generated during the current teaching progress period based on the baseline rigor directly yields the identification baseline for inattention, including: The distribution characteristics of attention generated during the current teaching progress period are analyzed to obtain a characteristic distribution description representing the degree of attention concentration at the current stage; Based on the baseline strictness, determine the sensitivity parameters for interpreting the feature distribution description; Based on feature distribution description and sensitivity parameters, a dynamic discrimination condition for instantaneous inattention is constructed as a recognition benchmark.
5. The digital network interactive teaching method as described in claim 1, characterized in that, The process of generating an interactive question directly related to the currently taught teaching content includes: Semantic structure analysis is performed on the currently taught teaching content to extract at least one core concept entity and its relationships, resulting in a knowledge graph fragment of the current teaching content; Select a target concept entity from the knowledge graph fragment that has not appeared in the recent interaction history to determine the focus of the interaction assessment; Based on the interactive assessment focus and its relationship with knowledge graph fragments, question stems are constructed to form interactive questions directly related to the teaching content.
6. The digital network interactive teaching method as described in claim 5, characterized in that, After constructing the question stem, the process also includes: Based on the deviation patterns of attention characteristics from the identification benchmark, and in conjunction with the current teaching stage, determine the intervention tendency category of inattention; Based on the intervention tendency category and the focus of the interactive assessment, select appropriate guiding expression elements; By combining guiding elements with the constructed question stem, interactive questions with guiding functions are formed.
7. The digital network interactive teaching method as described in claim 1, characterized in that, The driver for the teaching client to play the audio of the interactive questions includes: During the initial period when the audio of the question begins to play, features of the students' real-time facial orientation and gaze direction are extracted based on the video stream to obtain the initial attention feedback features during the audio playback phase. Based on the initial attention feedback characteristics, the subsequent playback rhythm and repetition strategy of the question stem audio are dynamically adjusted to form adaptive audio playback instructions. Based on the adaptive voice playback instructions, control the teaching client to play the audio of the question.
8. The digital network interactive teaching method as described in claim 7, characterized in that, The method of detecting students' confirmatory somatosensory movements based on video streams includes: Throughout the entire playback of the question's audio and within a preset time period after its end, the student's head posture sequence in the video stream is continuously matched and analyzed with a preset response intention action library to obtain a preliminary judgment of response intention. By combining the initial attention feedback features extracted during the speech playback phase, the confidence level of the preliminary response intention judgment is verified to obtain a multimodal confirmation result. When the multimodal confirmation result represents the intention to confirm, a dynamic attention focus mapping is established in the video stream based on the teaching content area related to the interactive question, and a confirmatory somatosensory action judgment is obtained to trigger the next operation.
9. The digital network interactive teaching method as described in claim 1, characterized in that, The control and teaching client displays the complete content of interactive questions on the demonstration interface, including: Based on the semantic structure of the interactive questions and the possible types of attention distraction, determine the targeted content display logic; Based on the content display logic, the complete content of the interactive question is visually decomposed and spatially laid out to obtain a dynamic content display solution. According to the dynamic content display scheme, the visual elements are serialized and rendered on the demonstration interface to complete the display of the full content of the interactive questions.
10. A digital network interactive teaching system for implementing the digital network interactive teaching method according to any one of claims 1-9, characterized in that, The system includes: The data acquisition module is used to acquire video streams containing students' heads collected by the teaching client and extract head movement trajectories and pupil change information; The attention feature formation module is used to correlate and analyze head movement trajectories with pupil change information to form attention features that reflect cognitive state. The baseline establishment module is used to dynamically establish a baseline for identifying inattention based on the historical trends of attention characteristics and the current teaching progress. The interactive question generation module is used to generate an interactive question that is directly related to the teaching content being taught, based on the current teaching content, when the recognition benchmark determines that the student is inattentive. The confirmatory somatosensory action confirmation module is used to drive the teaching client to play the audio of the interactive questions and detect the student's confirmatory somatosensory actions based on the video stream; The interactive question display module is used to control the teaching client to display the full content of the interactive question on the demonstration interface when a confirmatory haptic action is detected.