A concentration state detection method and device, a storage medium and an electronic device

CN122802706APending Publication Date: 2026-09-22浙江海亮科技有限公司
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610995478.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-07-06
Publication Date
2026-09-22

AI Technical Summary

Technical Problem

[0005]有鉴于此,本申请提供了一种专注状态检测方法、装置、存储介质及电子设备,主要目的在于改善目前现有技术在录播课授课的过程中,讲解焦点会在授课屏幕的不同位置发生迁移,教学场景会根据不同授课流程发生变化;使用这种方式无法适配动态变化的讲解焦点和教学场景,会造成专注状态误判、漏判等情况,导致专注状态检测的准确度低的技术问题

Benefits of technology

[0016]Using the above technical solution, this application provides a method, device, storage medium, and electronic device for detecting focus state, comprising: acquiring multi-dimensional course information of recorded courses and multi-dimensional behavioral information of students learning recorded courses; extracting teaching behavior characteristics of teachers explaining recorded courses and semantic features of courseware from the multi-dimensional course information of recorded courses that students need to learn; dividing the recorded courses into scenarios based on the extracted teaching behavior characteristics and courseware semantic features to obtain the target course scenario of the recorded courses; and dividing the recorded course screen into multiple target focus areas based on the teacher's voice heat, handwriting heat, and gaze heat in the recorded courses within a target time period, with the teacher's gaze center as the base. The system accurately analyzes the changes in teachers' focus across multiple target focus areas to construct a focus map corresponding to the recorded lessons. Within the target course scenario, it extracts students' attention areas from multi-dimensional behavioral information of students learning the recorded lessons. Based on the overlap between the attention areas and the focus map, as well as the overlap of focus time, it analyzes the degree of focus deviation of students relative to the focus map, thus obtaining the target focus information for students learning the recorded lessons. Based on this target focus information, it detects students' actual learning status and generates corresponding focus learning prompts when the actual learning status is unfocused, reminding students to focus on the recorded lessons. Compared with existing technologies, this application divides the screen focus area based on teachers' multi-dimensional heat and constructs a focus map for recorded lessons, determining standardized focus reference data suitable for the dynamic shift of teachers' focus during explanations. By combining spatial overlap information of focus areas and temporal overlap information to analyze the degree of students' focus shift and obtain target focus information, it achieves the quantification of the matching degree between students' actual focus and standard focus from both spatial and temporal dimensions, improving the accuracy of determining the matching degree. By detecting students' learning status based on target focus information and generating targeted prompts, it effectively reduces the probability of misjudgment and omission of students' focus status in dynamic teaching scenarios and under dynamic explanation focus, improving the accuracy of students' focus status detection in recorded lessons.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122802706A_ABST
    Figure CN122802706A_ABST
Patent Text Reader

Abstract

This application discloses a method, device, storage medium, and electronic device for detecting focus status, relating to the field of educational technology. The method includes: acquiring multi-dimensional course information of recorded lessons and multi-dimensional behavioral information of students learning the recorded lessons; extracting teaching behavior features and semantic features of the courseware to obtain the target course scene of the recorded lessons; dividing the recorded lesson screen into multiple target focus areas based on the teacher's voice heat, handwriting heat, and gaze heat to construct a focus map; obtaining the student's target focus information based on the overlap information of the focus areas and focus time between the learning gaze areas and the focus map; and detecting the student's actual learning status based on the student's target focus information in the recorded lessons. This application improves the accuracy of detecting student focus status in recorded lessons by constructing a focus map of the recorded lessons and matching the student's focus information with the focus map.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of educational technology, and in particular to a method, apparatus, storage medium, and electronic device for detecting attentional state. Background Technology

[0002] Pre-recorded lessons are lectures that teachers record in advance, allowing students to watch and learn independently via their devices. This pre-recorded lesson format is widely used across various educational stages, and student focus during these lessons is a crucial indicator of the quality of online learning.

[0003] Currently, by collecting students' facial features or eye movements and combining them with students' performance within a fixed teaching screen area, the learning behavior of students is analyzed, and their concentration during the viewing of recorded lessons is detected.

[0004] However, during the delivery of recorded lessons, the focus of the explanation will shift to different positions on the teaching screen, and the teaching scenario will change according to different teaching processes. This method cannot adapt to the dynamically changing focus of explanation and teaching scenario, which will cause misjudgment or omission of focus status, resulting in low accuracy of focus status detection. Summary of the Invention

[0005] In view of this, this application provides a method, device, storage medium, and electronic device for detecting focus state. The main purpose is to improve the technical problem that in the current technology, during the teaching of recorded courses, the focus of the explanation will shift to different positions on the teaching screen, and the teaching scene will change according to different teaching processes. Using this method cannot adapt to the dynamically changing focus of explanation and teaching scene, which will cause misjudgment or omission of focus state, resulting in low accuracy of focus state detection.

[0006] Firstly, this application provides a method for detecting a state of focus, including: Obtain multi-dimensional course information of recorded courses and multi-dimensional behavioral information of students learning the recorded courses; From the multi-dimensional course information of the recorded courses that students need to learn, the teaching behavior characteristics of the teacher explaining the recorded courses and the semantic features of the courseware of the recorded courses are extracted. Based on the extracted teaching behavior characteristics and the semantic features of the courseware, the recorded courses are divided into scenarios to obtain the target course scenarios of the recorded courses. In the target course scenario, based on the teacher's voice popularity, handwriting popularity, and gaze popularity in the recorded course within the target time, the recorded course screen is divided into multiple target focus areas. Using the teacher's gaze center as a reference, the focus change information of the teacher in the multiple target focus areas is analyzed to construct the focus map corresponding to the recorded course. In the target course scenario, the student's learning gaze area is extracted from the multi-dimensional behavioral information of the student learning the recorded course. Based on the overlap information of the learning gaze area with the focus area and the focus time overlap information of the focus map, the student's focus level is analyzed relative to the focus map, and the target focus information of the student learning the recorded course is obtained. Based on the student's target focus information when learning the recorded course, the system detects the student's actual learning status and generates a focus learning prompt message for the student when the actual learning status is unfocused, in order to remind the student to focus on learning the recorded course.

[0007] Optionally, in the target course scenario, based on the teacher's voice popularity, handwriting popularity, and gaze popularity in the recorded course within the target time period, the recorded course screen is divided into multiple target focus areas. Using the teacher's gaze center as a reference, the teacher's focus changes in the multiple target focus areas are analyzed to construct a focus map corresponding to the recorded course, including: When the target course scenario is a teaching scenario, the Gaussian heat distribution of the audio explanation content in the recorded courseware is analyzed, with the center coordinate of the courseware corresponding to the teacher's audio explanation content in the recorded course as the center, to obtain the audio heat of the teacher's explanation of the recorded course. Using the center coordinates of the handwriting in the recorded lesson as the center, the Gaussian heat distribution of the handwriting in the recorded lesson courseware is analyzed to obtain the handwriting heat during the process of the teacher writing the handwriting. Using the coordinates of the teacher's gaze center corresponding to the gaze area in the recorded lesson as the center, the Gaussian heat distribution of the gaze area in the recorded lesson courseware is analyzed to obtain the teacher's gaze heat when gazing at the recorded lesson screen; Threshold binarization processing is performed on the voice heat, handwriting heat, and gaze heat respectively to remove low heat noise regions and select effective voice heat regions, handwriting heat regions, and gaze heat regions. Determine the first overlapping intersection region where any two regions of the voice heat region, the handwriting heat region, and the gaze heat region overlap, and the second overlapping intersection region where any three regions overlap, and merge the gaze heat region, the first overlapping intersection region, and the second overlapping intersection region to obtain the initial teaching focus region corresponding to the teacher's teaching focus in the recorded lesson; The initial teaching focus area is expanded to compensate for the missing area caused by the teacher's teaching focus shift in the recorded lesson, resulting in an expanded teaching focus area. The expanded teaching focus area is then filtered for connected components to obtain a connected region containing the coordinates of the teacher's gaze center. This connected region is then determined as the primary focus area among the multiple target focus areas. Based on the teacher's focus shift and / or courseware page switching process in the recorded course, the student's shift between the main focus areas corresponding to the target time and historical time is analyzed to obtain the transition area among the multiple target focus areas. The screen area outside the main focus area and the transition area within the screen range of the recorded course is determined as the irrelevant area among the multiple target focus areas. Using the teacher's gaze center as a reference, the teacher's attention changes in multiple target attention areas are analyzed to construct an attention map corresponding to the recorded lesson.

[0008] Optionally, the step of analyzing the teacher's attentional variation information in multiple target attention areas based on the teacher's gaze center to construct an attentional map corresponding to the recorded lesson includes: Using the coordinates of the teacher's gaze center corresponding to the teacher's gaze center as the reference center, the recorded lesson screen is divided into regions based on the target radius and the coordinates of the teacher's gaze center, resulting in multiple gaze blocks. The heat values ​​of the multiple gaze blocks covered within the target radius are binned and accumulated. The changes in the heat values ​​of the multiple gaze blocks at different distances relative to the coordinates of the teacher's gaze center are analyzed to obtain heat distribution data based on the teacher's gaze center. The heat distribution data is used to characterize the teacher's focus at different screen positions. Obtain the target course scene at the target time, analyze the duration and scene sequence of the target course scene, and obtain the time sequence segmentation information of the target course scene; The coordinate range of the target focus area on the screen of the recorded course, the heat distribution data of the target focus area, and the time segmentation information of the target course scene are respectively temporally correlated and range fused. The focus change data of the teacher's focus in multiple target focus areas over time are analyzed to construct the focus map corresponding to the recorded course.

[0009] Optionally, in the target course scenario, the student's learning gaze region is extracted from the multi-dimensional behavioral information of the student learning the recorded course, and based on the overlap information of the learning gaze region with the focus region and the focus time overlap information of the focus map, the student's learning focus level and the focus shift level based on the focus map are analyzed to obtain the student's target focus information for learning the recorded course, including: When the target course scenario is a teaching scenario, the student's learning gaze area is extracted from the multi-dimensional behavioral information of the student learning the recorded course. The student's learning gaze area is then matched with the main focus area, the transition area, and the irrelevant area in the focus map. The degree of overlap between the student's actual gaze area and the gaze area corresponding to the teaching focus of the recorded course is analyzed to obtain the first region overlap information of the student's learning gaze area overlapping with the main focus area and the second region overlap information of the student's learning gaze area overlapping with the transition area. The student's gaze time sequence data is matched with the time sequence change data in the attention map to analyze the progress difference between the student's actual learning progress and the standard progress of the recorded course, thereby obtaining the attention time overlap information between the student's actual learning time and the standard learning time of the recorded course. Based on the first region overlap information, the second region overlap information, and the focus time overlap information, the spatial and temporal deviations of the student's actual learning focus level relative to the standard focus level in the focus map are analyzed to obtain the student's focus shift level when learning the recorded course. The overlap information of the first region, the overlap information of the second region, and the overlap information of the focus time are fused based on the degree of focus offset to analyze the actual focus of the student learning the recorded course and obtain the target focus information of the student learning the recorded course.

[0010] Optionally, when the target course scenario is a teaching scenario, before extracting the student's learning gaze region from the multi-dimensional behavioral information of the student learning the recorded course, and performing region matching between the student's learning gaze region and the main focus area, the transition area, and the irrelevant area in the focus map, and analyzing the degree of overlap between the student's actual gaze region and the gaze region corresponding to the teaching focus of the recorded course, and obtaining the first region overlap information of the student's learning gaze region overlapping with the main focus area, and the second region overlap information of the student's learning gaze region overlapping with the transition area, the method further includes: Extract the student's gaze time sequence within the target time period, and the teacher's teaching focus time sequence from the attention map; Based on the teacher's teaching focus time sequence, the student's gaze time sequence is time-shifted, and the standard gaze delay range between the student and the teacher is analyzed to obtain the time offset between the student's gaze time sequence and the teacher's teaching focus time sequence. The student's gaze timing sequence is time-shifted based on the time offset to align the student's gaze timing sequence with the teacher's teaching focus timing sequence, so that the student's gaze progress and the teacher's teaching progress remain synchronized within the standard gaze delay range.

[0011] Optionally, the step of detecting the student's actual learning status based on the target focus information of the student learning the recorded course, and generating a focus learning prompt message for the student when the actual learning status is an unfocused state, to remind the student to focus on learning the recorded course, includes: When the target course scenario is a teaching scenario, based on the target focus information of the student learning the recorded course, the actual learning state of the student learning the recorded course is detected. When the target focus information meets the focus state threshold condition and the first region overlap information meets the main focus area threshold condition, the actual learning state is determined as a focus state. If the target focus information does not meet the focus state threshold condition and the second region overlap information meets the transition zone threshold condition, the actual learning state is determined as a transition state. When the target focus information does not meet the focus state threshold condition, the first region overlap information does not meet the main focus area threshold condition, and the second region overlap information does not meet the transition area threshold condition, the confidence level of the target focus information and the duration of the target focus information are analyzed. If the confidence level of the target focus information does not meet the confidence level condition and the duration of the target focus information exceeds the duration of inattentiveness, the actual learning state is determined to be an inattentive state, and a focus learning prompt message corresponding to the student is generated to remind the student to focus on learning the recorded course.

[0012] Optionally, the method further includes: When the target course scenario is a student self-study scenario, the student's voice features are extracted from the multi-dimensional behavioral information of the student learning the recorded course, and the student's voice features are analyzed to determine whether the student makes any speech unrelated to learning, thereby detecting the student's actual learning status in the recorded course. When the target course scenario is a practice scenario, the handwriting features and gaze center coordinates of the student are extracted from the multi-dimensional behavioral information of the student learning the recorded course. The handwriting area and gaze area of ​​the student on the recorded course screen are analyzed to determine whether they are in the same screen area, thereby detecting the student's actual learning status in the recorded course.

[0013] Optionally, before obtaining the multi-dimensional course information of the recorded course and the multi-dimensional behavioral information of students learning the recorded course, the method further includes: Acquire camera image frames and device configuration parameters of the target device, wherein the target device includes the recording device for the teacher to record the recorded lesson and the learning device for the student to learn the recorded lesson; Face detection and facial landmark localization are performed on the camera image frames to extract the eye area images and eye landmark sets of the teacher and the student; Based on the extracted eye region image and the set of eye key points, the eye region features and head postures corresponding to the teacher and the student are analyzed to obtain the gaze regions corresponding to the gaze lines of the teacher and the student when looking at the target device. Based on the device configuration parameters of the target device, the gaze positions of the teacher and the student on the screen of the target device are analyzed to obtain the coordinates of the gaze point on the screen of the target device, and the coordinates of the gaze point are mapped to screen pixel coordinates through coordinate transformation. Based on the screen gaze point coordinates and gaze error data, a two-dimensional Gaussian distribution model of the gaze region of the teacher and the student is constructed, and an isodense elliptical region that meets the gaze threshold condition is selected from the two-dimensional Gaussian distribution model, and the isodense elliptical region is determined as the initial gaze region. Abnormal jump rejection is performed on the initial gaze regions of the teacher and the student to remove gaze regions with abnormal jump migration. The removed initial gaze regions are then subjected to temporal smoothing to obtain the teacher gaze region corresponding to the teacher and the student gaze region corresponding to the student.

[0014] Secondly, this application provides a focus state detection device, comprising: The acquisition module is configured to acquire multi-dimensional course information of the recorded courses and multi-dimensional behavioral information of students learning the recorded courses; The extraction module is configured to extract the teaching behavior characteristics of the teacher explaining the recorded course and the semantic features of the courseware from the multi-dimensional course information of the recorded course that students need to learn; and to divide the recorded course into scenarios based on the extracted teaching behavior characteristics and the semantic features of the courseware to obtain the target course scenario of the recorded course. The construction module is configured to, in the target course scenario, divide the recorded course screen into multiple target focus areas based on the teacher's voice heat, handwriting heat and gaze heat in the recorded course within the target time period, and, with the teacher's gaze center as the reference, analyze the teacher's focus change information in the multiple target focus areas to construct the focus map corresponding to the recorded course. The analysis module is configured to extract the student's learning gaze area from the multi-dimensional behavioral information of the student learning the recorded course in the target course scenario, and analyze the student's learning focus degree relative to the focus map based on the overlap information of the learning gaze area and the focus area overlap information of the focus map, so as to obtain the target focus information of the student learning the recorded course. The detection module is configured to detect the student's actual learning status in the recorded course based on the student's target focus information, and generate a focus learning prompt message for the student when the actual learning status is unfocused, so as to remind the student to focus on learning the recorded course.

[0015] Thirdly, this application provides an electronic device, including a storage medium, a processor, and a computer program stored on the storage medium and executable on the processor, wherein the processor executes the computer program to implement the focus state detection method described in the first aspect.

[0016] Using the above technical solution, this application provides a method, device, storage medium, and electronic device for detecting focus state, comprising: acquiring multi-dimensional course information of recorded courses and multi-dimensional behavioral information of students learning recorded courses; extracting teaching behavior characteristics of teachers explaining recorded courses and semantic features of courseware from the multi-dimensional course information of recorded courses that students need to learn; dividing the recorded courses into scenarios based on the extracted teaching behavior characteristics and courseware semantic features to obtain the target course scenario of the recorded courses; and dividing the recorded course screen into multiple target focus areas based on the teacher's voice heat, handwriting heat, and gaze heat in the recorded courses within a target time period, with the teacher's gaze center as the base. The system accurately analyzes the changes in teachers' focus across multiple target focus areas to construct a focus map corresponding to the recorded lessons. Within the target course scenario, it extracts students' attention areas from multi-dimensional behavioral information of students learning the recorded lessons. Based on the overlap between the attention areas and the focus map, as well as the overlap of focus time, it analyzes the degree of focus deviation of students relative to the focus map, thus obtaining the target focus information for students learning the recorded lessons. Based on this target focus information, it detects students' actual learning status and generates corresponding focus learning prompts when the actual learning status is unfocused, reminding students to focus on the recorded lessons. Compared with existing technologies, this application divides the screen focus area based on teachers' multi-dimensional heat and constructs a focus map for recorded lessons, determining standardized focus reference data suitable for the dynamic shift of teachers' focus during explanations. By combining spatial overlap information of focus areas and temporal overlap information to analyze the degree of students' focus shift and obtain target focus information, it achieves the quantification of the matching degree between students' actual focus and standard focus from both spatial and temporal dimensions, improving the accuracy of determining the matching degree. By detecting students' learning status based on target focus information and generating targeted prompts, it effectively reduces the probability of misjudgment and omission of students' focus status in dynamic teaching scenarios and under dynamic explanation focus, improving the accuracy of students' focus status detection in recorded lessons. Attached Figure Description

[0017] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.

[0018] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0019] Figure 1A flowchart illustrating a focus state detection method provided in an embodiment of this application is shown; Figure 2 A flowchart illustrating a focus state detection method provided in an embodiment of this application is shown; Figure 3 This paper shows a schematic diagram of the structure of a focus state detection device provided in an embodiment of this application; Figure 4 A schematic diagram of the structure of an electronic device provided in an embodiment of this application is shown. Detailed Implementation

[0020] The embodiments of this application will now be described in more detail with reference to the accompanying drawings. It should be noted that, unless otherwise specified, the embodiments and features described herein can be combined with each other.

[0021] To address the issue that existing technologies cannot adapt to dynamically changing focus points and teaching scenarios during recorded lectures, leading to misjudgments and missed detections of focus states, this embodiment provides a focus state detection method, such as... Figure 1 As shown, the method includes: Step 101: Obtain multi-dimensional course information of the recorded courses and multi-dimensional behavioral information of students learning the recorded courses.

[0022] In this embodiment of the application, the recorded course can be a teaching course recorded by the teacher, and students can learn the recorded course through terminals such as tablets and computers.

[0023] In the embodiments of this application, multi-dimensional course information can be the original multi-source data set of recorded courses recorded and uploaded by the teacher. Multi-dimensional course information can include recorded course audio data, recorded course video data, courseware image sequences, teacher's handwritten whiteboard writing and page-turning action data, teacher's eye gaze at Region of Interest (ROI) data, etc.

[0024] For example, the audio data of the recorded course can include the teacher's real-time explanation and ambient background sound; the video data of the recorded course can include video of the entire teaching process and screen recording video; the courseware image sequence can include courseware page images such as examples, question stems, charts, options, and explanations; the teacher's handwritten notes and page-turning action data can include handwritten coordinate sequences, timestamps, and page switching events; and the teacher's eye gaze ROI data can include the coordinates, center, area, and corresponding timestamp of the gaze area.

[0025] In the embodiments of this application, multi-dimensional behavioral information can be behavioral information collected in real time during the student's viewing of the recorded course. Multi-dimensional behavioral information can include student facial state data, student gaze ROI sequence data, student voice feature data, and student handwriting feature data.

[0026] For example, student facial state data can include information such as face and head posture; student gaze ROI sequence data can be mapped to tablet screen coordinates and include a timestamp; student voice feature data can include voice intensity and irrelevant voice recognition results; and student handwriting feature data can include writing coordinates and writing intensity temporal information.

[0027] In this embodiment of the application, multi-dimensional course raw data of recorded courses uploaded by teachers and multi-dimensional behavioral data of students watching recorded courses can be obtained. A global timeline can be established to complete timestamp alignment and raw data collection, providing a data foundation for subsequent feature extraction, scene segmentation and attention detection.

[0028] Step 102: Extract the teaching behavior characteristics of teachers explaining the recorded lessons and the semantic features of the recorded lessons from the multi-dimensional course information of the recorded lessons that students need to learn. Based on the extracted teaching behavior characteristics and semantic features of the recorded lessons, classify the recorded lessons into scenarios to obtain the target course scenarios of the recorded lessons.

[0029] In this embodiment of the application, the teaching behavior features can be normalized behavioral quantification indicators obtained by preprocessing multi-dimensional course information. The teaching behavior features can include speech intensity features, handwriting intensity features, and gaze focus features.

[0030] For example, the speech intensity feature can be obtained by normalizing the effective speech segment after audio noise reduction. The value range of the speech intensity feature can be [0,1]. The speech intensity feature can be used to characterize the teacher's speaking activity level.

[0031] For example, the handwriting intensity feature can be obtained by normalizing after smoothing the handwriting time sequence aggregation. The value range of the handwriting intensity feature can be [0,1]. The handwriting intensity feature can be used to characterize the teacher's writing activity.

[0032] For example, the gaze focus feature can be obtained from the time-series stability and area statistics of the teacher's gaze ROI. The value range of the gaze focus feature can be [0,1]. The gaze focus feature can be used to characterize the teacher's gaze concentration.

[0033] In this embodiment of the application, the semantic features of the courseware can be normalized confidence features extracted through automatic speech recognition (ASR) semantic recognition and courseware optical character recognition (OCR) layout analysis. The semantic features of the courseware can include teaching semantic confidence, practice semantic confidence, and self-study reading semantic confidence. The value range of the semantic features of the courseware can be [0,1].

[0034] For the embodiments of this application, the target course scenario can be divided into teaching scenario, self-study scenario (i.e., student self-study scenario in the embodiments of this application), practice scenario, etc.

[0035] In this embodiment, multi-dimensional course information can undergo preprocessing processes such as format unification, audio noise reduction, video stabilization, courseware regularization, handwriting filtering, and gaze ROI smoothing alignment; normalized speech intensity features, handwriting intensity features, gaze focus features, and courseware semantic cues can be extracted; a fixed time window can be set, and the scores for teaching scenario, practice scenario, and self-study scenario can be calculated respectively using a weighted formula, with the maximum score taken as the current window scenario label.

[0036] For example, the formula for calculating scores in a teaching scenario can be as shown in Formula 1, where, The weights can represent the speech intensity features. The weights can represent the gaze focus feature. The weights can represent the handwriting intensity features. can represent confidence weights, v can represent normalized speech intensity features, g can represent normalized gaze focus features, and p can represent normalized handwriting intensity features. This can represent the confidence level of the teaching semantics; the formula for calculating the practice scenario score can be shown in Formula 2, where, The weights can represent the speech intensity features. The weights can represent the gaze focus feature. The weights can represent the handwriting intensity features. can represent confidence weights, v can represent normalized speech intensity features, g can represent normalized gaze focus features, and p can represent normalized handwriting intensity features. This can represent the semantic confidence of practice; the formula for calculating the self-study scenario score is shown in Formula 3, where v can represent the normalized speech intensity feature and p can represent the normalized handwriting intensity feature. The weights can represent the gaze focus feature. It can indicate a low-intensity, stable gaze tendency (used in self-study scenarios). It can represent the confidence weight. This can represent the semantic confidence of self-study reading; by combining temporal smoothing and the shortest duration constraint to correct the scene boundary, the scene segmentation result SceneSeg by time segment is output. SceneSeg can be shown in Formula 4, where k can represent the scene number. This can represent the start time of the k-th scene. This can represent the end time of the k-th scene. It can include teaching scenarios, self-study scenarios, and practice scenarios.

[0037] (Formula 1) (Formula 2) (Formula 3) (Formula 4) Step 103: In the target course scenario, based on the teacher's voice heat, handwriting heat and gaze heat in the recorded course within the target time, the recorded course screen is divided into multiple target focus areas. Using the teacher's gaze center as the benchmark, the focus change information of the teacher in multiple target focus areas is analyzed to construct the focus map corresponding to the recorded course.

[0038] In the embodiments of this application, the voice heat, handwriting heat, and gaze heat can all be constructed using a two-dimensional Gaussian kernel negative exponential decay model to construct the screen space heat distribution, which is used to characterize the importance of teaching at different positions on the screen.

[0039] In this embodiment of the application, the target focus area may include a main focus area, a transition area, and an irrelevant area. The areas do not overlap and can completely cover the recorded course screen. The main focus area can be the area that the teacher focuses on during the lecture. The transition area can be the area generated by the teacher's focus shift or the switching of courseware pages. The irrelevant area can be the non-lecture area on the recorded course screen other than the main focus area and the transition area.

[0040] In the embodiments of this application, the teacher's gaze center can be the gaze center coordinates obtained by smoothing the time sequence of the teacher's gaze ROI.

[0041] In this embodiment of the application, the focus map can be a spatiotemporal reference map constructed by integrating the coordinate range of the target focus area, heat distribution data, and scene temporal segmentation information, and includes two-dimensional spectral features.

[0042] Step 104: In the target course scenario, extract the student's learning gaze area from the multi-dimensional behavioral information of the student learning the recorded course, and based on the overlap information of the learning gaze area and the attention area and attention time of the attention map, analyze the degree of attention shift of the student's learning attention level relative to the attention map, and obtain the target attention information of the student learning the recorded course.

[0043] In this embodiment, the learning gaze area can be the screen gaze area (ROI) of a student during the learning of a recorded lesson, and the shape of the learning gaze area can be an isodense ellipse.

[0044] In the embodiments of this application, the focus area overlap information can be calculated using the intersection-union ratio (IUGR) quantization method. The focus area overlap information can characterize the degree of spatial overlap and matching between the student's gaze area and the main focus area and transition area.

[0045] In this embodiment of the application, the focus time overlap information can be the overlap information between the student's learning time and the standard teaching time obtained after aligning the student's gaze sequence with the teacher's teaching focus sequence, which is used to characterize the matching degree between the student's actual learning progress and the standard teaching progress.

[0046] In the embodiments of this application, the degree of focus deviation can be obtained by quantifying the combined spatial deviation and temporal rhythm deviation. The degree of focus deviation can be used to characterize the extent to which a student's focus deviates from the standard focus range.

[0047] In this embodiment of the application, the target focus information can be a comprehensive student focus score obtained by integrating spatial overlap, temporal overlap, and focus shift.

[0048] Step 105: Based on the student's target focus information for learning the recorded course, detect the student's actual learning status for learning the recorded course, and generate corresponding focus learning prompts for the student when the actual learning status is not focused, so as to remind the student to focus on learning the recorded course.

[0049] In the embodiments of this application, the actual learning state may include a focused state, a transitional state, and a non-focused state. For example, a focused state can be a normal learning state in which the student's gaze is highly aligned with the main focus area and the learning pace is synchronized with the standard teaching pace; a transitional state can be a pending state in which the student is not aligned with the main focus area but falls into the transitional tolerance area and there is a reasonable delay; a non-focused state can be a distracted state in which the student is out of the main focus area and the transitional area and deviates from the teaching focus for a long time.

[0050] In the embodiments of this application, the focus learning prompt information can be a pop-up prompt triggered by a tablet device, a voice reminder, or other information, to remind users of unfocused behavior.

[0051] Compared with existing technologies, this embodiment divides the screen focus area based on teachers' multi-dimensional heat and constructs a focus map for recorded lessons to determine standardized focus reference data that adapts to the dynamic shift of teachers' focus. By combining spatial overlap information of focus areas and temporal overlap information to analyze the degree of students' focus shift and obtain target focus information, it achieves the quantification of the matching degree between students' actual focus and standard focus from both spatial and temporal dimensions, improving the accuracy of determining the matching degree. By detecting students' learning status based on target focus information and generating targeted prompts, it effectively reduces the probability of misjudgment and omission of students' focus status in dynamic teaching scenarios and under dynamic explanation focus, thereby improving the accuracy of detecting students' focus status in recorded lessons.

[0052] As an optional approach, when performing the task of "dividing the recorded course screen into multiple target focus areas based on the teacher's voice heat, handwriting heat, and gaze heat within a target time frame, and analyzing the teacher's focus changes across these areas using the teacher's gaze center as a baseline to construct a focus map corresponding to the recorded course," the following methods can be used, but are not limited to: Figure 2 As shown, the method includes: Step 201: When the target course scenario is a teaching scenario, take the center coordinate of the courseware corresponding to the teacher's voice explanation content in the recorded course as the center, analyze the Gaussian heat distribution of the voice explanation content in the recorded courseware, and obtain the voice heat of the teacher's recorded course explanation.

[0053] In this embodiment, the center coordinates of the courseware can be obtained by recognizing keywords through speech ASR and locating the center of areas such as the question stem and analysis on the courseware OCR layout. The speech popularity can be calculated using a two-dimensional Gaussian negative exponential model. The formula for calculating speech popularity is shown in Formula 5, where (x, y) can represent any pixel coordinate on the recorded course screen. This can represent the center coordinates of the courseware content corresponding to the teacher's voice explanation at time t. For example, when the teacher says, "Let's look at the first question," the rectangular area of ​​the first question's stem on the screen can be located using ASR speech recognition and courseware OCR layout analysis. These are the coordinates of the center point of this rectangle.

[0054] (Formula 5) Step 202: Using the center coordinates of the handwriting in the recorded lesson as the center, analyze the Gaussian heat distribution of the handwriting in the recorded lesson courseware to obtain the handwriting heat during the teacher's handwriting process.

[0055] In this embodiment of the application, the coordinates of the handwriting center can be taken as the current pen tip coordinates or the short-term handwriting cluster center, and have undergone time-series smoothing processing; the handwriting heat calculation formula can be as shown in Formula 6, where (x,y) can represent any pixel coordinates on the recorded course screen. This can represent the coordinates of the center of the smoothed handwritten stroke at time t. It can be the coordinates of the teacher's pen tip in the current frame, or it can be the cluster center of all handwritten strokes in the last second, adapting to continuous writing scenarios.

[0056] (Formula 6) Step 203: Using the coordinates of the teacher's gaze center corresponding to the gaze area in the recorded lesson as the center, analyze the Gaussian heat distribution of the gaze area in the recorded lesson courseware to obtain the gaze heat of the teacher when gazing at the recorded lesson screen.

[0057] In this embodiment, the coordinates of the teacher's gaze center are the stable trajectory coordinates after low-pass filtering and smoothing; the formula for calculating gaze intensity can be as shown in Formula 7, where, It can represent the coordinates of the teacher's gaze center after time-series smoothing at time t.

[0058] (Formula 7) Step 204: Perform threshold binarization processing on the speech heat, handwriting heat and gaze heat respectively, remove low heat noise areas, and select effective speech heat areas, handwriting heat areas and gaze heat areas.

[0059] For the embodiments of this application, the formula for filtering effective speech heat regions through threshold binarization processing can be shown in Formula 8, where... It can represent the area of ​​voice intensity. It can represent the voice popularity value at any coordinate on the screen. This can represent the speech popularity threshold; the formula for filtering valid handwriting popularity regions is shown in Formula Nine, where, It can represent the area of ​​handwriting heat. It can represent the handwriting heat value at any coordinate on the screen. This can represent the handwriting heat threshold; the formula for filtering out effective gaze heat regions can be shown in Formula 10, where, It can indicate the area of ​​focus and heat. It can represent the gaze intensity value at any coordinate on the screen. It can represent the gaze heat threshold.

[0060] (Formula 8) (Formula Nine) (Formula 10) Step 205: Determine the first overlapping intersection area of ​​any two areas in the speech heat area, handwriting heat area, and gaze heat area, and the second overlapping intersection area of ​​any three areas. Then merge the gaze heat area, the first overlapping intersection area, and the second overlapping intersection area to obtain the initial teaching focus area corresponding to the teacher's teaching focus in the recorded lesson.

[0061] In this embodiment of the application, the calculation formula for the second overlapping intersection region of the three regions—voice heat region, handwriting heat region, and gaze heat region—can be as shown in Formula 11, wherein, This can represent the second overlapping intersection region. It can indicate the area of ​​focus and heat. It can represent the area of ​​voice intensity. It can represent the area of ​​handwriting heat.

[0062] (Formula Eleven) In this embodiment, a conditional judgment function can be used to count multimodal overlapping pixels. At least the bimodal coverage area can include a first overlapping area, and the calculation formula for at least the bimodal coverage area can be as shown in Formula Twelve, wherein... This can represent a conditional judgment function, where the value within the square brackets is 1 if the condition is true and 0 if it is false; for example, the first overlapping region between the voice heat region and the handwriting heat region can be... and The values ​​are 1, The value is 0; it can be determined that the screen area a student should focus on at a certain moment when learning a recorded lesson can be determined by the fact that at least two of the corresponding screen areas of voice behavior, handwriting behavior, and gaze behavior are the same area.

[0063] (Formula 12) Step 206: Perform region expansion processing on the initial teaching focus area to compensate for the region loss caused by the teacher's teaching focus shift in the recorded lesson, and obtain an expanded teaching focus area. Then, perform connected component filtering on the expanded teaching focus area to obtain a connected region containing the coordinates of the teacher's gaze center. The connected region is determined as the main focus area among multiple target focus areas.

[0064] For the embodiments of this application, region dilation can employ, but is not limited to, morphological dilation algorithms. The calculation formula for the principal focus region can be as shown in Formula Thirteen, wherein... It can represent the main area of ​​focus. This can be represented as the main region, which is the area where the gaze intensity region intersects with the second overlapping region. This can represent the first overlapping intersection region. It can represent the radius of morphological expansion scale. When filtering connected regions, priority is given to retaining connected regions that contain the teacher's gaze center, while discarding other scattered regions.

[0065] (Formula Thirteen) Step 207: Based on the areas where the teacher shifts focus during the recorded lesson and / or switches courseware pages, analyze the student's shift between the main focus areas corresponding to the target time and historical time, obtain the transition areas among multiple target focus areas, and determine the screen area outside the main focus area and transition area within the recorded lesson screen as the irrelevant area among multiple target focus areas.

[0066] In this embodiment, the transition region is generated by expanding the current and historical focus regions. The calculation formula for the transition region is shown in Formula Fourteen, where... It can represent a transition zone. It can represent the region of primary focus at time t. It can represent history The main focus area at all times It can represent the radius of the expansion scale.

[0067] (Formula Fourteen) In this embodiment, the irrelevant area is the area on the recorded course screen other than the main focus area and the transition area. The formula for calculating the irrelevant area is shown in Formula 15, where, It can represent a transition zone. It can represent the main area of ​​focus. It can represent a transition zone. It can represent the screen area of ​​a recorded lesson.

[0068] (Formula Fifteen) Step 208: Using the teacher's gaze center as a benchmark, analyze the changes in the teacher's focus in multiple target focus areas and construct a focus map corresponding to the recorded lesson.

[0069] In this embodiment of the application, based on the teacher's gaze center, the planar heat can be converted into a radius heat spectrum. The coordinate range of the target focus area on the recorded course screen, the heat distribution data of the target focus area, and the temporal segmentation information of the target course scene are temporally correlated and range-fused to generate a focus map.

[0070] Optionally, when performing the task of "analyzing the teacher's attention changes in multiple target attention areas based on the teacher's gaze center and constructing a attention map corresponding to the recorded lesson," the following methods can be used, but are not limited to: using the teacher's gaze center coordinates as the reference center, dividing the recorded lesson screen into multiple attention blocks based on the target radius and extending outwards from the teacher's gaze center coordinates; accumulating the heat values ​​of the multiple attention blocks covered within the target radius in buckets, and analyzing the heat values ​​of the multiple attention blocks at different distances relative to the teacher's gaze center coordinates. The data is analyzed to obtain heat distribution data based on the teacher's gaze center. This heat distribution data is used to characterize the teacher's focus at different screen positions. The target course scene at the target time is obtained, and the duration and sequence of the target course scene are analyzed to obtain the temporal segmentation information of the target course scene. The coordinate range of the target focus area on the recorded course screen, the heat distribution data of the target focus area, and the temporal segmentation information of the target course scene are temporally correlated and range-fused respectively. The focus change data of the teacher's focus in multiple target focus areas over time is analyzed to construct the focus map corresponding to the recorded course.

[0071] For the embodiments of this application, the formula for calculating the target radius can be as shown in Formula Sixteen, where, It can represent the target radius at time t. It can represent the coordinates of the teacher's gaze center after time-series smoothing at time t.

[0072] (Formula Sixteen) In this embodiment of the application, the heat values ​​of multiple gaze blocks covered within the target radius are accumulated by binning, and the calculation formula can be shown in Formula 17, where, It can represent the fusion total heat value (fusion heat of voice heat, handwriting heat and gaze heat) at the pixel coordinate (x,y) of the recorded course screen at time t. It can represent the target radius at time t. This can represent the k-th radius binning interval. This can represent the k-th radius binning interval. It can represent the screen area of ​​a recorded lesson. This can represent a conditional function; the value inside the square brackets is 1 if the condition is true and 0 if the condition is false. It can be used to implement bucket accumulation, that is, during the calculation process, only the heat value of screen pixels that fall exactly within the k-th radius bucket interval from the center point is accumulated, and all screen pixels that do not belong to this interval are automatically filtered out.

[0073] (Formula 17) In this embodiment of the application, the distribution data obtained from bucket accumulation can be normalized to obtain heat distribution data. The calculation formula for the heat distribution data is shown in Formula 18, where... It can represent the heat distribution data before normalization. It can represent the total cumulative fusion heat value corresponding to the j-th radius bucket at time t, where j can represent all radius buckets. This can represent the k-th radius binning interval. It can represent the j-th radius bucket interval.

[0074] (Formula 18) In this embodiment, the target course scene at the target time is obtained, and the duration and scene order of the target course scene are analyzed to obtain the temporal segmentation information of the target course scene. This can be done by determining the target course scene of the recorded course at each moment, sorting the duration and scene order of the target course scene, and obtaining the teaching moment of the recorded course and the teaching scene, self-study scene, or practice scene corresponding to the teaching moment. The coordinate range of the target focus area on the screen of the recorded course, the heat distribution data of the target focus area, and the temporal segmentation information of the target course scene are temporally correlated and range fused respectively. The focus change data of the teacher in multiple target focus areas changes over time is analyzed. The focus change data can be a focus change image, where the horizontal axis of the image can be the class time, the vertical axis can be the focus, and the background of the image can be multiple target focus areas divided based on the temporal segmentation information.

[0075] As an optional approach, when performing the task of "extracting students' learning gaze regions from multi-dimensional behavioral information of students learning recorded lessons in a target course scenario, and analyzing students' learning focus level and focus shift based on the overlap information of focus regions and focus time on the focus map to obtain students' target focus information for learning recorded lessons," the following methods can be used, but are not limited to: When the target course scenario is a teaching scenario, extracting students' learning gaze regions from multi-dimensional behavioral information of students learning recorded lessons, and performing region matching between students' learning gaze regions and the main focus area, transition area, and irrelevant area in the focus map, analyzing the degree of overlap between the actual gaze region of students learning recorded lessons and the gaze region corresponding to the teaching focus of the recorded lessons, to obtain students' learning gaze regions. The system analyzes the overlap information of the first region (overlapping with the main focus area) and the second region (overlapping with the transition area); it matches the students' gaze time sequence data with the time sequence change data in the focus map to analyze the progress difference between the students' actual learning progress and the standard progress of the recorded course, thus obtaining the focus time overlap information between the students' actual learning time and the standard learning time of the recorded course; based on the first region overlap information, the second region overlap information, and the focus time overlap information, it analyzes the spatial and temporal deviations of the students' actual learning focus level relative to the standard focus level in the focus map, thus obtaining the students' focus deviation degree in learning the recorded course; based on the focus deviation degree, it fuses the first region overlap information, the second region overlap information, and the focus time overlap information to analyze the students' actual focus level in learning the recorded course, thus obtaining the students' target focus information in learning the recorded course.

[0076] For the embodiments of this application, the formula for calculating the overlap information of the first region where the student's learning gaze area overlaps with the main focus area can be as shown in Formula Nineteen, where... This can represent the smoothed student gaze ROI. This can represent the primary focus area; the formula for calculating the overlap information of the second region where the student's learning focus area overlaps with the transition area can be shown in Formula 20, where, This can represent the smoothed student gaze ROI. It can represent a transition region.

[0077] (Formula 19) (Formula 20) In this embodiment of the application, the formula for calculating the gaze center of the student's gaze at the POI can be shown in Formula 21, wherein, This can represent the smoothed student gaze ROI. It can represent the student's focus of attention during learning.

[0078] (Formula 21) In this embodiment, the extracted student learning gaze region (ROI) undergoes radius mapping and spectral matching. Using the time-smoothed teacher gaze center coordinates in the attention map as the reference center, the deviation radius of the student learning gaze region's center coordinates relative to this reference center is calculated. The formula for calculating the deviation radius is shown in Formula 22. It can represent the student's focus of attention during learning. It can represent the teacher's gaze center after smoothing.

[0079] (Formula 22) In this embodiment of the application, the deviation radius is mapped to the radius bucket interval, and the normalized heat distribution spectrum corresponding to the current recording time is matched. The spectrum hit probability corresponding to the current student's gaze position is obtained by a lookup table matching method. The calculation formula for the spectrum hit probability is shown in Formula 23, where, It can represent the radius-divided interval. This can represent the deviation radius. Based on the overlap information of the first region, the overlap information of the second region, and the spectral hit probability, the spatial deviation of the student's actual learning focus level relative to the standard focus level in the focus map is determined.

[0080] (Formula 23) In this embodiment, the effective attention span of students and the effective teaching span of teachers are compared frame by frame and time window by time window. Periods in which the actual effective learning time of students completely overlap with the standard teaching time of teachers are selected, thus determining the synchronization difference between the actual learning pace of students and the standard teaching pace of the recorded course. Attention span overlap information is generated by combining the percentage of overlap and the time sequence deviation.

[0081] As an optional approach, before performing the following steps: "In the case of a teaching scenario, extract students' learning gaze regions from multi-dimensional behavioral information of students learning recorded lessons, and match these regions with the main focus area, transition area, and irrelevant area in the focus map, analyze the degree of overlap between the actual gaze region of students learning recorded lessons and the gaze region corresponding to the teaching focus of the recorded lessons, and obtain the first region overlap information of the student's learning gaze region overlapping with the main focus area and the second region overlap information of the student's learning gaze region overlapping with the transition area," the following method can be used, but is not limited to: extracting the student's gaze time sequence within the target time and the teacher's teaching focus time sequence in the focus map; performing time translation on the student's gaze time sequence based on the teacher's teaching focus time sequence, analyzing the standard gaze delay range between the student and the teacher, and obtaining the time offset between the student's gaze time sequence and the teacher's teaching focus time sequence; performing time translation on the student's gaze time sequence based on the time offset, and aligning the student's gaze time sequence with the teacher's teaching focus time sequence to keep the student's gaze progress synchronized with the teacher's teaching progress within the standard gaze delay range.

[0082] In this embodiment, the formula for calculating the time offset (i.e., the optimal time offset) between the student's gaze timing sequence and the teacher's teaching focus timing sequence can be as shown in Formula 24, where W can represent a fixed time window. This can represent the weight of the overlapping information in the first region. The weights can represent the probability of hitting the spectrum. This can represent the overlap information of the first region. It can represent the probability of hitting the spectrum. It can represent a time offset. It can represent the maximum allowed latency of the system.

[0083] (Formula 24) In this embodiment of the application, the student's gaze timing data within the time analysis window and the teacher's standard gaze timing data in the attention map are extracted. Based on the optimal time offset, the overall time sequence translation correction of the student's gaze timing sequence is performed to eliminate reasonable delay deviations in the student's listening process and to align the student's learning timing with the teacher's standard teaching timing.

[0084] In this embodiment of the application, the standard gaze delay range can be the reaction lag delay range of students receiving teaching knowledge points, following courseware with their eyes and the teacher's focus shift when watching recorded lessons; for example, the standard gaze delay range in this embodiment of the application can include the delay time range caused by students' normal slight dragging behavior while following the lesson, and the delay time range can be a certain time period or a specific time.

[0085] In this embodiment of the application, after aligning the student learning timeline with the teacher's standard teaching timeline, it is necessary to calculate the mean window value of the main focus area, the mean window value of the transition area, and the mean spectral hit probability. The formula for calculating the mean window value of the main focus area can be shown in Formula 25, where W can represent a fixed time window. It can represent the time offset (i.e., the optimal time offset). This can represent the overlap information of the first region after time alignment; the formula for calculating the window mean of the transition region can be shown in Formula 26, where W can represent a fixed time window. It can represent the time offset (i.e., the optimal time offset). This can represent the overlap information of the second region after temporal alignment; the formula for calculating the mean spectral hit probability is shown in Formula 27, where W can represent a fixed time window. It can represent the time offset (i.e., the optimal time offset). It can represent the overlap information of the second region after time alignment.

[0086] (Formula 25) (Formula 26) (Formula 27) In this embodiment, the overlap information of the first region, the overlap information of the second region, and the overlap information of attention time are fused based on the degree of attention shift. This can be achieved by fusing the window mean of the main attention region, the window mean of the transition region, and the mean of the spectrum hit probability to obtain the target attention information of the student learning the recorded course. The calculation formula for the target attention information can be as shown in Formula 28, where... The weights can represent the window mean of the main focus area. It can represent the window mean of the main focus area. The weights can represent the window mean of the transition zone. The weights can represent the mean of the spectrum hit probability. It can represent the window mean of the transition zone. It can represent the mean of the spectrum hit probability. The weight of the window mean in the main focus area can be greater than the weight of the window mean in the transition area and the weight of the mean of the spectrum hit probability.

[0087] (Formula 28) As an optional approach, when executing the function of "detecting the student's actual learning status based on the student's target focus information in the recorded course, and generating corresponding focus prompts to remind the student to focus on the recorded course when the actual learning status is not focused," the following method can also be used, but is not limited to: When the target course scenario is a teaching scenario, detecting the student's actual learning status based on the student's target focus information in the recorded course; and determining the actual learning status as focused when the target focus information meets the focus status threshold condition and the first region overlap information meets the main focus area threshold condition; in the case of the target... If the focused information does not meet the focused state threshold condition and the overlapping information in the second region meets the transition zone threshold condition, the actual learning state is determined to be a transition state. If the target focused information does not meet the focused state threshold condition, the overlapping information in the first region does not meet the main focused zone threshold condition, and the overlapping information in the second region does not meet the transition zone threshold condition, the confidence level and duration of the target focused information are analyzed. If the confidence level of the target focused information does not meet the confidence level condition and the duration of the target focused information exceeds the unfocused duration, the actual learning state is determined to be an unfocused state, and corresponding focused learning prompts are generated for the student to remind them to focus on the recorded lessons.

[0088] In the embodiments of this application, when the target course scenario is a teaching scenario, the actual learning state of the student learning the recorded course is detected based on the student's target focus information. If the target focus information is greater than or equal to the focus state threshold and the window mean of the main focus area is greater than or equal to the main focus area threshold, the actual learning state is determined as a focus state.

[0089] In the embodiments of this application, when the target focus information is less than the focus state threshold and the window mean of the transition zone is greater than or equal to the transition zone threshold, the actual learning state is determined as the transition state.

[0090] In this embodiment of the application, the focus confidence score can be mapped to the [0,1] interval by the Sigmoid function, and the focus confidence score can be used to characterize the reliability of the non-focus judgment.

[0091] In this embodiment, when the target focus information is less than the focus state threshold, the window mean of the main focus area is less than the main focus area threshold, and the window mean of the transition area is less than the transition area threshold, the confidence level of the target focus information and the duration of the target focus information are analyzed.

[0092] In the embodiments of this application, when the confidence level of the target focus information is less than the confidence threshold and the duration of the target focus information exceeds the duration of inattentiveness, the actual learning state is determined to be an inattentive state, and corresponding focus learning prompt information is generated for the student to remind the student to focus on learning the recorded course.

[0093] As an optional approach, but not limited to the following methods, the method includes: when the target course scenario is a student self-study scenario, extracting the student's voice features from the multi-dimensional behavioral information of the student learning the recorded course, analyzing whether the student makes learning-irrelevant speech, and detecting the student's actual learning status in the recorded course; when the target course scenario is a practice scenario, extracting the student's handwriting features and the student's gaze center coordinates from the multi-dimensional behavioral information of the student learning the recorded course, analyzing whether the student's handwriting area and gaze area on the recorded course screen are in the same screen area, and detecting the student's actual learning status in the recorded course.

[0094] In this embodiment, analyzing whether a student emits learning-irrelevant speech can be done by setting a speech threshold. When the intensity of a student's speech feature is greater than the speech threshold, it can be determined that learning-irrelevant speech is present, and the student's attention state can be identified as unfocused. When the intensity of a student's speech feature is less than or equal to the speech threshold, the student's attention state can be identified as focused.

[0095] In this embodiment of the application, analyzing whether the handwriting area and the gaze area of ​​a student on the recorded course screen are in the same screen area can be done by setting a distance threshold. When the Euclidean distance between the student's handwriting center and gaze center is less than the distance threshold and the handwriting intensity meets the standard, the student's focus state can be determined as focused; otherwise, it can be determined as unfocused.

[0096] As an optional approach, before performing the step of "obtaining multi-dimensional course information of recorded lessons and multi-dimensional behavioral information of students learning recorded lessons," the following methods can also be used, but are not limited to: acquiring camera image frames and device configuration parameters of the target device, including the recording device used by the teacher to record the lesson and the learning device used by the student to learn the lesson; performing face detection and facial key point localization on the camera image frames, extracting eye area images and sets of eye key points for the teacher and student; based on the extracted eye area images and sets of eye key points, analyzing the eye area features and head postures corresponding to the teacher and student respectively, to obtain the gaze areas corresponding to the gaze lines of the teacher and student looking at the target device; and analyzing the teaching... The gaze positions of teachers and students on the target device screen are determined, and the coordinates of the gaze point on the target device screen are obtained. The screen gaze point coordinates are then mapped to screen pixel coordinates through coordinate transformation. Based on the screen gaze point coordinates and gaze error data, a two-dimensional Gaussian distribution model of the gaze region of teachers and students is constructed. Iso-density elliptical regions that meet the gaze threshold conditions are selected from the two-dimensional Gaussian distribution model and determined as the initial gaze regions. Abnormal jumps are removed from the initial gaze regions of teachers and students to eliminate gaze regions with abnormal jump migrations. The removed initial gaze regions are then subjected to temporal smoothing to obtain the teacher's gaze region and the student's learning gaze region.

[0097] In this embodiment, the device configuration parameters may include camera intrinsic parameters, screen planar parameters, and camera-screen relative position calibration parameters.

[0098] In this embodiment, based on the extracted eye region image and eye key point set, the eye region features and head posture corresponding to the teacher and student are analyzed to obtain the gaze regions corresponding to the gaze lines of the teacher and student on the target device. The formula for calculating the gaze line can be shown in Formula 29, where o(t) can represent the gaze line emission point and g(t) can represent the gaze vector. It can represent scalar parameters greater than 0.

[0099] (Formula 29) For the embodiments of this application, the calculation formula for the intersection of the gaze line and the recorded course screen (i.e., the coordinates of the gaze point on the target device screen in the embodiments of this application) can be as shown in Formula 30 and Formula 31. In Formula 30, o(t) can represent the gaze line emission point, and g(t) can represent the gaze vector. It can represent the plane equation of the recorded course screen. It can represent extremely small positive values ​​to prevent the denominator from being zero. It can represent a scalar parameter of gaze. This can represent the screen's normal vector; in Equation 31, It can represent the coordinates of the screen gaze point, o(t) can represent the gaze exit point, and g(t) can represent the gaze vector. It can represent a scalar parameter of gaze.

[0100] (Formula 30) (Formula 31) In this embodiment, the screen gaze point coordinates are mapped to screen pixel coordinates through coordinate transformation. The formula for calculating the screen pixel coordinates is shown in Formula 32, where... It can represent a transformation matrix. It can represent the projection function. It can represent the coordinates of the screen gaze point.

[0101] (Formula 32) For the embodiments of this application, the two-dimensional Gaussian distribution model can be expressed as: The formula for calculating the initial gaze area is shown in Formula 33, where u can represent the coordinates of any point on the recorded lesson screen. It can represent the screen area of ​​a recorded lesson. It can be expressed as an equal density threshold (i.e., the gaze threshold in the embodiments of this application). It can represent screen pixel coordinates.

[0102] (Formula 33) In this embodiment of the application, the calculation formula for temporal smoothing can be as shown in Formulas 34 and 35. Formula 34 applies a low-pass filter to the gaze centers of the teacher ROI and student ROI to suppress jitter, wherein... It can represent the smoothing coefficient, with a value range of (0,1]. This can represent the original ROI center coordinates at time t. This can represent the smoothed ROI center coordinates at time t-1. This can represent the smoothed center coordinates of the ROI at time t; Formula 35 smooths the scale (uncertainty / radius / area) of the ROI, where, It can represent the smoothing coefficient. This can represent the original ROI scale parameter at time t. It can represent the scale parameter of the smoothed ROI at time t-1. It can represent the smoothed ROI scale parameter at time t.

[0103] (Formula Thirty-Four) (Formula 35) Compared with existing technologies, this embodiment quantifies the importance of teaching at different screen locations by employing a two-dimensional Gaussian decay model to calculate the teacher's voice heat, handwriting heat, and gaze heat. It also improves the reliability of determining the primary focus area by performing threshold binarization, region fusion, and dilation processing on multimodal heat data. Furthermore, it enhances the completeness and rationality of screen focus area division in dynamic teaching scenarios by combining teacher focus migration and courseware page switching characteristics to define transition and irrelevant areas. Finally, it constructs teacher focus heat spectra in both spatial and temporal dimensions by performing radius-based binning heat accumulation and normalization based on the teacher's gaze center. It improves the accuracy of quantifying student spatial focus deviation features by mapping student gaze ROIs to deviation radii and matching them with heat spectra to obtain spectrum hit probabilities. Finally, it effectively eliminates focus detection errors within the standard delay range by solving for the optimal time offset to align teacher and student gaze sequences. Finally, it improves the adaptability and accuracy of student focus detection in different course scenarios by differentiating focus detection conditions for teaching, self-study, and practice scenarios.

[0104] Furthermore, as Figure 1 and Figure 2 The specific implementation of the method shown in this embodiment provides a focus state detection device, such as... Figure 3 As shown, the device includes: an acquisition module 31, an extraction module 32, a construction module 33, an analysis module 34, and a generation module 35.

[0105] Module 31 is configured to acquire multi-dimensional course information of recorded courses and multi-dimensional behavioral information of students learning recorded courses. The extraction module 32 is configured to extract the teaching behavior characteristics of the teacher explaining the recorded course and the semantic features of the courseware from the multi-dimensional course information of the recorded course that students need to learn. Based on the extracted teaching behavior characteristics and courseware semantic features, the recorded course is divided into scenarios to obtain the target course scenario of the recorded course. Module 33 is configured to divide the recorded course screen into multiple target focus areas based on the teacher's voice heat, handwriting heat and gaze heat in the recorded course within the target time period in the target course scenario, and analyze the teacher's focus change information in multiple target focus areas based on the teacher's gaze center, and construct the focus map corresponding to the recorded course. Analysis module 34 is configured to extract the student's learning gaze area from the multi-dimensional behavioral information of the student learning the recorded course in the target course scenario, and analyze the student's learning focus degree relative to the focus map based on the overlap information of the focus area and focus time of the focus area and focus map, so as to obtain the student's target focus information for learning the recorded course. The detection module 35 is configured to detect the student's actual learning status based on the student's target focus information when learning the recorded course, and generate corresponding focus learning prompts for the student when the actual learning status is not focused, so as to remind the student to focus on learning the recorded course.

[0106] In some examples of this embodiment, the construction module 33 is specifically configured to, when the target course scenario is a teaching scenario, analyze the Gaussian heat distribution of the audio explanation content in the recorded courseware, centered on the center coordinates of the courseware corresponding to the teacher's audio explanation content, to obtain the audio heat of the teacher's audio explanation in the recorded course; analyze the Gaussian heat distribution of the handwritten handwriting content in the recorded courseware, centered on the center coordinates of the handwriting content corresponding to the teacher's handwritten handwriting content, to obtain the handwriting heat during the teacher's handwriting process; analyze the Gaussian heat distribution of the gaze area in the recorded courseware, centered on the teacher's gaze center coordinates corresponding to the gaze area in the recorded course, to obtain the gaze heat of the teacher gazing at the recorded course screen; perform threshold binarization processing on the audio heat, handwriting heat, and gaze heat respectively, remove low-heat noise areas, and filter out effective audio heat areas, handwriting heat areas, and gaze heat areas; determine the first overlapping intersection area and three overlapping intersection areas of any two areas in the audio heat area, handwriting heat area, and gaze heat area. The second overlapping intersection region is identified, and the gaze intensity region, the first overlapping intersection region, and the second overlapping intersection region are merged to obtain the initial teaching focus region corresponding to the teacher's teaching focus in the recorded lesson. This initial teaching focus region is then expanded to compensate for the region loss caused by the teacher's focus shift during the recorded lesson, resulting in an extended teaching focus region. Connectivity filtering is then performed on this extended teaching focus region to obtain connected regions containing the teacher's gaze center coordinates. These connected regions are then identified as the primary focus regions among the multiple target focus regions. Based on the teacher's focus shift and / or courseware page switching processes during the recorded lesson, the migration of students between the primary focus regions corresponding to the target time and historical time is analyzed to obtain transition regions among the multiple target focus regions. The screen area within the recorded lesson screen, excluding the primary focus region and transition region, is identified as the irrelevant region among the multiple target focus regions. Using the teacher's gaze center as a reference, the teacher's focus changes across multiple target focus regions are analyzed to construct a focus map corresponding to the recorded lesson.

[0107] In some examples of this embodiment, the construction module 33 is further configured to use the coordinates of the teacher's gaze center as the reference center, and divide the recorded course screen into multiple gaze blocks based on the target radius and the coordinates of the teacher's gaze center. The heat values ​​of the multiple gaze blocks covered within the target radius are then bucketed and accumulated. The changes in the heat values ​​of the multiple gaze blocks at different distances relative to the coordinates of the teacher's gaze center are analyzed to obtain heat distribution data based on the teacher's gaze center. This heat distribution data is used to characterize the teacher's focus at different screen positions. The target course scene at the target time is obtained, and the duration and sequence of the target course scene are analyzed to obtain the temporal segmentation information of the target course scene. The coordinate range of the target focus area on the recorded course screen, the heat distribution data of the target focus area, and the temporal segmentation information of the target course scene are temporally correlated and range-fused respectively. The focus change data of the teacher's focus in multiple target focus areas over time is analyzed to construct a focus map corresponding to the recorded course.

[0108] In some examples of this embodiment, the analysis module 34 is specifically configured to, when the target course scenario is a teaching scenario, extract the student's learning gaze area from the multi-dimensional behavioral information of the student learning the recorded course, and perform region matching between the student's learning gaze area and the main focus area, transition area, and irrelevant area in the focus map, respectively, analyze the degree of overlap between the student's actual gaze area and the gaze area corresponding to the teaching focus of the recorded course, obtain the first region overlap information of the student's learning gaze area overlapping with the main focus area, and the second region overlap information of the student's learning gaze area overlapping with the transition area; and perform time-series matching between the student's gaze time series data and the time series change data in the focus map. Sequence matching is used to analyze the progress difference between students' actual learning progress and the standard progress of the recorded course, obtaining the attention time overlap information between the students' actual learning time and the standard learning time of the recorded course. Based on the overlap information of the first region, the second region, and the attention time overlap information, the spatial and temporal deviations of students' actual learning attention levels relative to the standard attention levels in the attention map are analyzed to obtain the students' attention deviation degree when learning the recorded course. Based on the attention deviation degree, the overlap information of the first region, the second region, and the attention time overlap information are fused to analyze the students' actual attention level when learning the recorded course, obtaining the students' target attention information when learning the recorded course.

[0109] In some examples of this embodiment, the analysis module 34 is further configured to extract the student's gaze time sequence within the target time and the teacher's teaching focus time sequence in the attention map; based on the teacher's teaching focus time sequence, perform time translation on the student's gaze time sequence, analyze the standard gaze delay range between the student and the teacher, and obtain the time offset between the student's gaze time sequence and the teacher's teaching focus time sequence; perform time translation on the student's gaze time sequence based on the time offset, and time align the student's gaze time sequence with the teacher's teaching focus time sequence so that the student's gaze progress and the teacher's teaching progress remain synchronized within the standard gaze delay range.

[0110] In some examples of this embodiment, the generation module 35 is specifically configured to, when the target course scenario is a teaching scenario, detect the student's actual learning state based on the student's target focus information for the recorded course; if the target focus information meets the focus state threshold condition and the first region overlap information meets the main focus area threshold condition, determine the actual learning state as a focus state; if the target focus information does not meet the focus state threshold condition but the second region overlap information meets the transition area threshold condition, determine the actual learning state as a transition state; if the target focus information does not meet the focus state threshold condition, the first region overlap information does not meet the main focus area threshold condition, and the second region overlap information does not meet the transition area threshold condition, analyze the confidence level and duration of the target focus information; if the confidence level of the target focus information does not meet the confidence level condition and the duration of the target focus information exceeds the inattentive duration, determine the actual learning state as an inattentive state, and generate corresponding focus learning prompt information for the student to remind the student to focus on the recorded course.

[0111] In some examples of this embodiment, the generation module 35 is further configured to: extract students' speech features from the multi-dimensional behavioral information of students learning the recorded course when the target course scenario is a student self-study scenario; analyze whether students make speech unrelated to learning; and detect the actual learning status of students learning the recorded course. When the target course scenario is a practice scenario, extract students' handwriting features and the coordinates of students' gaze center from the multi-dimensional behavioral information of students learning the recorded course; analyze whether the handwriting area and gaze area of ​​students on the recorded course screen are in the same screen area; and detect the actual learning status of students learning the recorded course.

[0112] In some examples of this embodiment, the acquisition module 31 is specifically configured to acquire camera image frames and device configuration parameters of the target device, including a recording device for teachers to record lessons and a learning device for students to learn lessons; perform face detection and facial key point localization on the camera image frames, extract eye area images and eye key point sets for teachers and students; based on the extracted eye area images and eye key point sets, analyze the eye area features and head postures corresponding to teachers and students respectively, and obtain the gaze areas corresponding to the gaze lines of teachers and students looking at the target device; according to the device configuration parameters of the target device, analyze the gaze positions of teachers and students looking at the screen of the target device, and obtain the gaze positions of teachers and students looking at the screen of the target device. The gaze coordinates on the target device's screen are determined, and these coordinates are mapped to screen pixel coordinates through coordinate transformation. Based on the screen gaze coordinates and gaze error data, a two-dimensional Gaussian distribution model of the teacher's and student's gaze regions is constructed. Iso-density elliptical regions that meet the gaze threshold conditions are selected from the two-dimensional Gaussian distribution model and defined as the initial gaze regions. Abnormal jumps are eliminated from the initial gaze regions of both teachers and students to remove gaze regions with abnormal jump migrations. The eliminated initial gaze regions are then subjected to temporal smoothing to obtain the teacher's gaze region and the student's learning gaze region.

[0113] Based on the above, Figure 1 and Figure 2 Accordingly, this embodiment also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the above-described method. Figure 1 and Figure 2 The method shown.

[0114] Based on this understanding, the technical solution of this application can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (such as CD-ROM, USB flash drive, mobile hard drive, etc.) and includes several instructions to cause a computer device (such as personal computer, server, or network device, etc.) to execute the methods of various implementation scenarios of this application.

[0115] like Figure 4 The diagram shown is a hardware structure schematic of an electronic device according to the present invention, comprising: At least one processor 401; and, Memory 402 is communicatively connected to at least one processor 401; wherein, The memory 402 stores instructions that can be executed by at least one processor, which enables the at least one processor to perform the attention state detection method as described above.

[0116] Figure 4Take a processor 401 as an example.

[0117] The electronic device may also include an input device 403 and a display device 404.

[0118] The processor 401, memory 402, input device 403, and display device 404 can be connected via a bus or other means. Figure 4 Taking the example of a connection between China and Israel via a bus.

[0119] Memory 402, as a non-volatile computer-readable storage medium, can be used to store non-volatile software programs, non-volatile computer-executable programs, and modules, such as the program instructions / modules corresponding to the focus state detection method in the embodiments of this application, for example, Figure 1 and Figure 2 The method flow is shown. The processor 401 executes various functional applications and data processing by running non-volatile software programs, instructions, and modules stored in the memory 402, thereby implementing the focus state detection method in the above embodiments.

[0120] Memory 402 may include a program storage area and a data storage area, wherein the program storage area may store the operating system and applications required for at least one function; the data storage area may store data created based on the use of the attention state detection method, etc. Furthermore, memory 402 may include high-speed random access memory and may also include non-volatile memory, such as at least one disk storage device, flash memory device, or other non-volatile solid-state storage device. In some embodiments, memory 402 may optionally include memory remotely located relative to processor 401, and these remote memories may be connected via a network to the apparatus performing the attention state detection method. Examples of such networks include, but are not limited to, the Internet, intranets, local area networks, mobile communication networks, and combinations thereof.

[0121] The input device 403 can receive user clicks and generate signal inputs related to user settings and function control of the attention state detection method. The display device 404 may include a display screen or other display device.

[0122] One or more modules are stored in memory 402, and when run by one or more processors 401, the focus state detection method in any of the above method embodiments is executed.

[0123] Optionally, the aforementioned physical devices may also include a user interface, a network interface, a camera, radio frequency (RF) circuitry, sensors, audio circuitry, a Wi-Fi module, etc. The user interface may include a display screen, input units such as a keyboard, etc., and optional user interfaces may also include USB interfaces, card reader interfaces, etc. The network interface may optionally include standard wired interfaces, wireless interfaces (such as Wi-Fi interfaces), etc.

[0124] Those skilled in the art will understand that the physical device structure provided in this embodiment does not constitute a limitation on the physical device, and may include more or fewer components, or combine certain components, or have different component arrangements.

[0125] The storage medium may also include an operating system and a network communication module. The operating system is a program that manages the hardware and software resources of the aforementioned physical device, supporting the operation of information processing programs and other software and / or programs. The network communication module is used to enable communication between the various components within the storage medium, as well as communication with other hardware and software in the information processing physical device.

[0126] Through the above description of the embodiments, those skilled in the art can clearly understand that this application can be implemented by means of software plus necessary general-purpose hardware platforms, or it can be implemented by hardware. By applying the solution of this embodiment, compared with the existing technology, this embodiment completes the screen focus area division based on the teacher's multi-dimensional heat and constructs a focus map of the recorded course to determine standardized focus reference data that is suitable for the dynamic migration of the teacher's explanation focus; by combining the spatial overlap information of focus area and the temporal overlap information of focus area to analyze the degree of student focus shift and obtain target focus information, it realizes the quantification of the matching degree between the student's actual focus and the standard focus from both spatial and temporal dimensions, thereby improving the accuracy of determining the matching degree; by detecting the student's learning status based on the target focus information and generating targeted prompt information, it effectively reduces the probability of misjudgment and omission of the student's focus status in dynamic teaching scenarios and dynamic explanation focus, thereby improving the accuracy of student focus status detection in recorded courses.

[0127] It should be noted that, in this document, relational terms such as "first" and "second" are used merely to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes the element.

[0128] The above are merely specific embodiments of this application, enabling those skilled in the art to understand or implement this application. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of this application. Therefore, this application is not to be limited to these embodiments, but is to be accorded the widest scope consistent with the principles and novel features claimed herein.

Claims

1. A method for detecting a state of focus, characterized in that, include: Obtain multi-dimensional course information of recorded courses and multi-dimensional behavioral information of students learning the recorded courses; From the multi-dimensional course information of the recorded courses that students need to learn, the teaching behavior characteristics of the teacher explaining the recorded courses and the semantic features of the courseware of the recorded courses are extracted. Based on the extracted teaching behavior characteristics and the semantic features of the courseware, the recorded courses are divided into scenarios to obtain the target course scenarios of the recorded courses. In the target course scenario, based on the teacher's voice popularity, handwriting popularity, and gaze popularity in the recorded course within the target time, the recorded course screen is divided into multiple target focus areas. Using the teacher's gaze center as a reference, the focus change information of the teacher in the multiple target focus areas is analyzed to construct the focus map corresponding to the recorded course. In the target course scenario, the student's learning gaze area is extracted from the multi-dimensional behavioral information of the student learning the recorded course. Based on the overlap information of the learning gaze area with the focus area and the focus time overlap information of the focus map, the student's focus level is analyzed relative to the focus map, and the target focus information of the student learning the recorded course is obtained. Based on the student's target focus information when learning the recorded course, the system detects the student's actual learning status and generates a focus learning prompt message for the student when the actual learning status is unfocused, in order to remind the student to focus on learning the recorded course.

2. The method according to claim 1, characterized in that, In the target course scenario, based on the teacher's voice popularity, handwriting popularity, and gaze popularity during the recorded course within the target time period, the recorded course screen is divided into multiple target focus areas. Using the teacher's gaze center as a reference, the focus change information of the teacher in the multiple target focus areas is analyzed to construct a focus map corresponding to the recorded course, including: When the target course scenario is a teaching scenario, the Gaussian heat distribution of the audio explanation content in the recorded courseware is analyzed, with the center coordinate of the courseware corresponding to the teacher's audio explanation content in the recorded course as the center, to obtain the audio heat of the teacher's explanation of the recorded course. Using the center coordinates of the handwriting in the recorded lesson as the center, the Gaussian heat distribution of the handwriting in the recorded lesson courseware is analyzed to obtain the handwriting heat during the process of the teacher writing the handwriting. Using the coordinates of the teacher's gaze center corresponding to the gaze area in the recorded lesson as the center, the Gaussian heat distribution of the gaze area in the recorded lesson courseware is analyzed to obtain the teacher's gaze heat when gazing at the recorded lesson screen; Threshold binarization processing is performed on the voice heat, handwriting heat, and gaze heat respectively to remove low heat noise regions and select effective voice heat regions, handwriting heat regions, and gaze heat regions. Determine the first overlapping intersection region where any two regions of the voice heat region, the handwriting heat region, and the gaze heat region overlap, and the second overlapping intersection region where any three regions overlap, and merge the gaze heat region, the first overlapping intersection region, and the second overlapping intersection region to obtain the initial teaching focus region corresponding to the teacher's teaching focus in the recorded lesson; The initial teaching focus area is expanded to compensate for the missing area caused by the teacher's teaching focus shift in the recorded lesson, resulting in an expanded teaching focus area. The expanded teaching focus area is then filtered for connected components to obtain a connected region containing the coordinates of the teacher's gaze center. This connected region is then determined as the primary focus area among the multiple target focus areas. Based on the teacher's focus shift and / or courseware page switching process in the recorded course, the student's shift between the main focus areas corresponding to the target time and historical time is analyzed to obtain the transition area among the multiple target focus areas. The screen area outside the main focus area and the transition area within the screen range of the recorded course is determined as the irrelevant area among the multiple target focus areas. Using the teacher's gaze center as a reference, the teacher's attention changes in multiple target attention areas are analyzed to construct an attention map corresponding to the recorded lesson.

3. The method according to claim 2, characterized in that, The process involves analyzing the teacher's attentional changes across multiple target attention areas, using the teacher's gaze center as a reference, to construct an attentional map corresponding to the recorded lesson. This includes: Using the coordinates of the teacher's gaze center corresponding to the teacher's gaze center as the reference center, the recorded lesson screen is divided into regions based on the target radius and the coordinates of the teacher's gaze center, resulting in multiple gaze blocks. The heat values ​​of the multiple gaze blocks covered within the target radius are binned and accumulated. The changes in the heat values ​​of the multiple gaze blocks at different distances relative to the coordinates of the teacher's gaze center are analyzed to obtain heat distribution data based on the teacher's gaze center. The heat distribution data is used to characterize the teacher's focus at different screen positions. Obtain the target course scene at the target time, analyze the duration and scene sequence of the target course scene, and obtain the time sequence segmentation information of the target course scene; The coordinate range of the target focus area on the screen of the recorded course, the heat distribution data of the target focus area, and the time segmentation information of the target course scene are respectively temporally correlated and range fused. The focus change data of the teacher's focus in multiple target focus areas over time are analyzed to construct the focus map corresponding to the recorded course.

4. The method according to claim 1, characterized in that, In the target course scenario, the student's learning gaze region is extracted from the multi-dimensional behavioral information of the student learning the recorded course. Based on the overlap information of the learning gaze region with the focus region and the focus time overlap information of the focus map, the student's learning focus level and focus shift based on the focus map are analyzed to obtain the student's target focus information for learning the recorded course, including: When the target course scenario is a teaching scenario, the student's learning gaze area is extracted from the multi-dimensional behavioral information of the student learning the recorded course. The student's learning gaze area is then matched with the main focus area, the transition area, and the irrelevant area in the focus map. The degree of overlap between the student's actual gaze area and the gaze area corresponding to the teaching focus of the recorded course is analyzed to obtain the first region overlap information of the student's learning gaze area overlapping with the main focus area and the second region overlap information of the student's learning gaze area overlapping with the transition area. The student's gaze time sequence data is matched with the time sequence change data in the attention map to analyze the progress difference between the student's actual learning progress and the standard progress of the recorded course, thereby obtaining the attention time overlap information between the student's actual learning time and the standard learning time of the recorded course. Based on the first region overlap information, the second region overlap information, and the focus time overlap information, the spatial and temporal deviations of the student's actual learning focus level relative to the standard focus level in the focus map are analyzed to obtain the student's focus shift level when learning the recorded course. The overlap information of the first region, the overlap information of the second region, and the overlap information of the focus time are fused based on the degree of focus offset to analyze the actual focus of the student learning the recorded course and obtain the target focus information of the student learning the recorded course.

5. The method according to claim 4, characterized in that, In the case where the target course scenario is a teaching scenario, before extracting the student's learning gaze region from the multi-dimensional behavioral information of the student learning the recorded course, and performing region matching between the student's learning gaze region and the main focus area, the transition area, and the irrelevant area in the focus map, and analyzing the degree of overlap between the student's actual gaze region and the gaze region corresponding to the teaching focus of the recorded course, and obtaining the first region overlap information of the student's learning gaze region overlapping with the main focus area, and the second region overlap information of the student's learning gaze region overlapping with the transition area, the method further includes: Extract the student's gaze time sequence within the target time period, and the teacher's teaching focus time sequence from the attention map; Based on the teacher's teaching focus time sequence, the student's gaze time sequence is time-shifted, and the standard gaze delay range between the student and the teacher is analyzed to obtain the time offset between the student's gaze time sequence and the teacher's teaching focus time sequence. The student's gaze timing sequence is time-shifted based on the time offset to align the student's gaze timing sequence with the teacher's teaching focus timing sequence, so that the student's gaze progress and the teacher's teaching progress remain synchronized within the standard gaze delay range.

6. The method according to claim 1, characterized in that, The step of detecting the student's actual learning status based on the target focus information of the student learning the recorded course, and generating a focus learning prompt message for the student when the actual learning status is not focused, to remind the student to focus on learning the recorded course, includes: When the target course scenario is a teaching scenario, based on the target focus information of the student learning the recorded course, the actual learning state of the student learning the recorded course is detected. When the target focus information meets the focus state threshold condition and the first region overlap information meets the main focus area threshold condition, the actual learning state is determined as a focus state. If the target focus information does not meet the focus state threshold condition and the second region overlap information meets the transition zone threshold condition, the actual learning state is determined as a transition state. When the target focus information does not meet the focus state threshold condition, the first region overlap information does not meet the main focus area threshold condition, and the second region overlap information does not meet the transition area threshold condition, the confidence level of the target focus information and the duration of the target focus information are analyzed. If the confidence level of the target focus information does not meet the confidence level condition and the duration of the target focus information exceeds the duration of inattentiveness, the actual learning state is determined to be an inattentive state, and a focus learning prompt message corresponding to the student is generated to remind the student to focus on learning the recorded course.

7. The method according to claim 6, characterized in that, The method further includes: When the target course scenario is a student self-study scenario, the student's voice features are extracted from the multi-dimensional behavioral information of the student learning the recorded course, and the student's voice features are analyzed to determine whether the student makes any speech unrelated to learning, thereby detecting the student's actual learning status in the recorded course. When the target course scenario is a practice scenario, the handwriting features and gaze center coordinates of the student are extracted from the multi-dimensional behavioral information of the student learning the recorded course. The handwriting area and gaze area of ​​the student on the recorded course screen are analyzed to determine whether they are in the same screen area, thereby detecting the student's actual learning status in the recorded course.

8. The method according to claim 1, characterized in that, Before acquiring multi-dimensional course information of the recorded lessons and multi-dimensional behavioral information of students learning the recorded lessons, the method further includes: Acquire camera image frames and device configuration parameters of the target device, wherein the target device includes the recording device for the teacher to record the recorded lesson and the learning device for the student to learn the recorded lesson; Face detection and facial landmark localization are performed on the camera image frames to extract the eye area images and eye landmark sets of the teacher and the student; Based on the extracted eye region image and the set of eye key points, the eye region features and head postures corresponding to the teacher and the student are analyzed to obtain the gaze regions corresponding to the gaze lines of the teacher and the student when looking at the target device. Based on the device configuration parameters of the target device, the gaze positions of the teacher and the student on the screen of the target device are analyzed to obtain the coordinates of the gaze point on the screen of the target device, and the coordinates of the gaze point are mapped to screen pixel coordinates through coordinate transformation. Based on the screen gaze point coordinates and gaze error data, a two-dimensional Gaussian distribution model of the gaze region of the teacher and the student is constructed, and an isodense elliptical region that meets the gaze threshold condition is selected from the two-dimensional Gaussian distribution model, and the isodense elliptical region is determined as the initial gaze region. Abnormal jump rejection is performed on the initial gaze regions of the teacher and the student to remove gaze regions with abnormal jump migration. The removed initial gaze regions are then subjected to temporal smoothing to obtain the teacher gaze region corresponding to the teacher and the student gaze region corresponding to the student.

9. A focus state detection device, characterized in that, include: The acquisition module is configured to acquire multi-dimensional course information of the recorded courses and multi-dimensional behavioral information of students learning the recorded courses; The extraction module is configured to extract the teaching behavior characteristics of the teacher explaining the recorded course and the semantic features of the courseware from the multi-dimensional course information of the recorded course that students need to learn; and to divide the recorded course into scenarios based on the extracted teaching behavior characteristics and the semantic features of the courseware to obtain the target course scenario of the recorded course. The construction module is configured to, in the target course scenario, divide the recorded course screen into multiple target focus areas based on the teacher's voice heat, handwriting heat and gaze heat in the recorded course within the target time period, and, with the teacher's gaze center as the reference, analyze the teacher's focus change information in the multiple target focus areas to construct the focus map corresponding to the recorded course. The analysis module is configured to extract the student's learning gaze area from the multi-dimensional behavioral information of the student learning the recorded course in the target course scenario, and analyze the student's learning focus degree relative to the focus map based on the overlap information of the learning gaze area and the focus area overlap information of the focus map, so as to obtain the target focus information of the student learning the recorded course. The detection module is configured to detect the student's actual learning status in the recorded course based on the student's target focus information, and generate a focus learning prompt message for the student when the actual learning status is unfocused, so as to remind the student to focus on learning the recorded course.

10. An electronic device, comprising a storage medium, a processor, and a computer program stored on the storage medium and executable on the processor, characterized in that, When the processor executes the computer program, it implements the method of any one of claims 1 to 8.