Multimedia interactive teaching machine control system
Through the multimedia interactive teaching machine control system, students' concentration is evaluated and learning reports are generated using the overlap rate of the line of sight focus area calculation, which solves the problems of insufficient line of sight focus capture and lagging interaction feedback in existing equipment, and realizes accurate teaching feedback and data-driven teaching optimization.
Patent Information
- Application Number
- CN202510773267.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-11
- Publication Date
- 2025-07-11
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Existing multimedia interactive teaching equipment cannot accurately capture the focus area of teachers and students' sights, cannot evaluate students' concentration in real time, interactive feedback is not timely, teaching data analysis is insufficient, and it is difficult to generate systematic course learning reports.
The data acquisition module obtains facial videos of teachers and students, uses the data analysis module to perform time alignment and focal area extraction, calculates the overlap rate of focal area of focal area to evaluate the student's concentration, and uses the interactive feedback module to perform sound reminders or highlight rendering to generate a course learning report.
Accurate assessment and timely feedback on students' concentration have been achieved, the intelligence and data-driven characteristics of teaching interaction have been improved, and the quality and efficiency of teaching have been improved.
Smart Images

Figure CN120295486A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of educational technology, and specifically, relates to a control system for a multimedia interactive teaching machine. Background Art
[0002] With the development of the economy, the educational field has been continuously promoting informatization construction, and multimedia interactive teaching equipment has emerged as the times require; as an important carrier, the functions of teaching machines have been gradually expanded, from the initial simple presentation of teaching content to the current integration of devices such as cameras, realizing real-time interaction between instructors and students.
[0003] Although the prior art has realized the sharing of teaching resources such as teaching videos and blackboard writings among teaching machines, there are still many deficiencies; firstly, most of the prior art can only perform simple presentation and sharing of teaching content, lacking accurate capture and analysis of the focus areas of the instructors' and students' lines of sight, and unable to evaluate the students' concentration in real time; secondly, the interactive feedback mechanism of the prior art is relatively lagging and one-sided, and cannot effectively remind in time when students' attention is distracted or abnormal behaviors occur, making it difficult to ensure the teaching effect; furthermore, the prior art fails to fully explore and utilize the data in the teaching process, and the data is often not deeply analyzed after being stored, and cannot generate a systematic course learning report to assist in teaching optimization; based on the above problems, the present invention provides a control system for a multimedia interactive teaching machine. Summary of the Invention
[0004] Aiming at the deficiencies of the prior art, the present invention provides a control system for a multimedia interactive teaching machine, which solves the problems of poor monitoring of students' concentration, untimely interactive feedback, and insufficient teaching data analysis in the prior art.
[0005] The purpose of the present invention can be achieved by the following technical solutions: A control system for a multimedia interactive teaching machine, the system includes the following: A data acquisition module, which interacts with the teaching machine, acquires the instructor's face video and the student's face video of the instructor and the student, and transmits them to the data storage module for storage; A data analysis module, which extracts the instructor's face video and the student's face video, performs time alignment, and determines the instructor's line-of-sight focus area associated with the instructor and the student's line-of-sight focus area associated with the student; Extract the start timestamp and end timestamp of the instructor's line-of-sight focus area, and obtain the student's line-of-sight focus area within the start timestamp and end timestamp; If the acquisition is successful, calculate the coincidence rate between the instructor's line-of-sight focus area and the student's line-of-sight focus area and evaluate the student's concentration; If the acquisition fails, it is regarded that the student has an abnormal behavior and is recorded; An interactive feedback module, if the student has an abnormal behavior, reminds the student through sound; If the coincidence rate is lower than the lowest coincidence rate threshold, the focus area of the instructor's line of sight is highlighted and rendered to remind the students. After the course ends, a course learning report of the students is generated and fed back to the instructor.
[0006] As a further solution of the present invention, the system further includes a data storage module, which stores any one of the calculation steps and analysis steps described in the solution, and stores the calculation results of any one of the calculation steps and the analysis results of the analysis steps in this solution.
[0007] As a further solution of the present invention, the data acquisition module interacts with the teaching machine in real time; The teaching machine includes a student machine and an instructor machine; The configurations of the student machine and the instructor machine are the same; Both the instructor machine and the student machine have numbers; The configuration includes a camera, a display, a host, and a communication protocol; The camera is integrated on the display and the position is fixed; The instructor machine shares the teaching video and the instructor's blackboard writing to the student machine.
[0008] As a further solution of the present invention, the specific way for the data acquisition module to obtain the instructor's face video and the student's face video of the instructor and the student is as follows: Determine the instructor in the course and mark it as ; Determine the total number of students in the course and record it as ; Sort the students according to the number sequence of the instructor machine to obtain the student sequence ; Obtain the instructor's instructor face video in the course ; Obtain the student face videos of all students in the student sequence and sort them in the order of the student sequence to obtain the student face video sequence .
[0009] As a further solution of the present invention, the specific way for the data analysis module to determine the instructor's associated instructor line-of-sight focus area and the student's associated student line-of-sight focus area is as follows: Determine any one student in the student sequence and extract the associated student face video , where i is a counting index and the value range is from 1 to ; Arrange the instructor face video in the order of video frames Split into a sequence of video frames of the instructor's face , where is the total number of video frames in Then extract the teaching video that is time-aligned with the video of the instructor's face ; ; Taking the first pixel point in the lower left corner of the screen as the origin, the lower boundary as the horizontal axis, and the left boundary as the vertical axis, construct the associated two-dimensional coordinate system, denoted as , which is also the two-dimensional coordinate system associated with the instructor's machine display screen; Extract the video frame , obtain the line-of-sight focus in and map it to , obtain the mapped point and extract the two-dimensional coordinates of the mapped point as the line-of-sight focus coordinates of the line-of-sight focus in in , marked as ; ; Repeat the above steps to obtain the set of line-of-sight focus coordinates of all video frames in the sequence of video frames of the instructor's face in : ; ; Perform clustering on to determine all the instructor's line-of-sight focus areas associated in and summarize them in the order of the timeline to obtain the set of instructor's line-of-sight focus areas associated with , where represents the total number of instructor's line-of-sight focus areas associated in the course; Similarly, obtain the student and the student's line-of-sight focus areas and the set of student's line-of-sight focus areas associated with other students in the course.
[0010] As a further solution of the present invention, the data analysis module performs clustering on to determine all the instructor's line-of-sight focus areas associated in in the following specific manner: Extract the line-of-sight focus coordinates in the set of line-of-sight focus coordinates , and use them as a reference, then extract backward and calculate , Euclidean distance between ; Obtain the Euclidean distance threshold preset by the operator . If , then and are clustered as the same group, and continue to obtain the line-of-sight focus coordinates backward for clustering judgment with the benchmark; If , then extract for clustering judgment. If f consecutive line-of-sight focus coordinates are extracted and the Euclidean distance from is greater than , then discard , use as the benchmark, continue to extract the line-of-sight focus coordinates backward for clustering judgment, where f is a value preset by the operator; If the total number of line-of-sight focus coordinates in the same group, including the benchmark, exceeds n, it is regarded as successful clustering of the same group. Use a fixed circle to cover and fit all the line-of-sight focus coordinates with successful clustering into a circular area, and regard it as the instructor's line-of-sight focus area associated with the line-of-sight focus coordinates with successful clustering of the same group. Extract the timestamp of the first line-of-sight focus coordinate in the same group as the start timestamp of the instructor's line-of-sight focus area, and then extract the timestamp of the last line-of-sight focus coordinate in the same group as the end timestamp of the instructor's line-of-sight focus area. Among them, the radius of the fixed circle is greater than the Euclidean distance threshold , and the values of n are preset by the operator; Determine the first instructor's line-of-sight focus area, denoted as , and continue to obtain the line-of-sight focus coordinates backward for clustering judgment until a line-of-sight focus coordinate that fails to cluster with the benchmark is obtained as a new benchmark, and repeat the above steps until the last line-of-sight focus coordinate in the set of line-of-sight focus coordinates is judged for clustering, obtaining All the instructor's line-of-sight focus areas associated in , and the subsequent instructor's line-of-sight focus areas are recorded in numerical order.
[0011] As a further solution of the present invention, the specific method for the data analysis module to calculate the coincidence rate between the instructor's line-of-sight focus area and the student's line-of-sight focus area and evaluate the student's concentration is as follows: S71. Extract the set of instructor's line-of-sight focus areas associated ; S72. Extract the set of student's line-of-sight focus areas associated with the student, where , among which represents The total number of the focus areas of the trainee's line of sight associated in the course; S73. Extraction The start timestamp and the end timestamp of all the focus areas of the instructor's line of sight in Then from extract the start timestamp and the end timestamp of all the focus areas of the trainee's line of sight; S74. Construct the time series trace line , whose time span is from the start time to the end time of the course; Mark each focus area of the instructor's line of sight in on in the form of a histogram according to the associated start timestamp and end timestamp; S75. Then mark the start timestamp and the end timestamp associated with each focus area of the trainee's line of sight in on in the form of a histogram; S76. Extract the timestamp interval during which the histogram mapped by any one focus area of the instructor's line of sight coincides with the histogram mapped by any one focus area of the trainee's line of sight from , denoted as ; S77. Then extract the focus areas of the instructor's line of sight that contain , and the focus areas of the trainee's line of sight that contain , and calculate the coincidence rate of the focus areas of the instructor's line of sight and the focus areas of the trainee's line of sight ; S78. Extract the coincidence rate rating standard preset by the operator: , and compare with the coincidence rate rating standard to obtain the coincidence rate rating associated in the timestamp interval ; Among them, , , the specific values of which are all determined by the operator in combination with the actual situation, and is the minimum coincidence rate threshold for judging that the trainee is listening to the class normally; Repeat the above steps to determine all the coincidence rate ratings in the course, calculate the concentration score based on the coincidence rate rating, and evaluate the concentration of the trainee .
[0012] As a further solution of the present invention, the data analysis module calculates the concentration score based on the coincidence rate rating and evaluates the concentration of the trainee The specific ways of the concentration are as follows: In the coincidence rate rating standard, the score for excellent is 20; the score for good is 15; the score for medium is 10; the score for poor is 5; Obtain the students Combine all the coincidence rate ratings associated with the students in the course with the scores corresponding to different coincidence rate ratings in the coincidence rate rating standard, and count The associated concentration scores; The concentration score is proportional to the concentration of the students.
[0013] As a further solution of the present invention, in the interactive feedback module, the specific ways of reminding the students include the following steps: In S76, if there is The histogram mapped by any one of the instructor's line of sight focus areas associated with The histogram of any one of the students' line of sight focus areas associated with has no coincidence, then it is determined that within the start timestamp and end timestamp associated with the instructor's line of sight focus area An abnormal behavior occurs, and through The student machine used A sound reminder is sent; In step S77, if the coincidence rate Is lower than the lowest coincidence rate threshold , then the instructor's line of sight focus area is highlighted to remind the students.
[0014] As a further solution of the present invention, in the interactive feedback module, the specific way of generating the course learning report of the students and feeding it back to the instructor is: Summarize the students Concentration scores, the number of abnormal behaviors in the course, and the students The start timestamp and end timestamp associated with each abnormal behavior generated by the students The course learning report associated with the students in the course; Repeat the above steps to generate The course learning reports associated with each of the students in, and transmit them to the instructor machine used by the instructor for the instructor to view.
[0015] The beneficial effects of the present invention: (1) The present invention constructs a closed-loop teaching interaction mechanism; uses a camera to collect the line-of-sight focus data of the instructor and students in real time, and accurately evaluates the students' concentration through time alignment and coincidence rate calculation, solving the pain point of difficult quantification of the learning state in traditional teaching; the overall system deeply integrates artificial intelligence technology with the teaching scenario, realizing a full-link closed-loop from behavior perception to teaching optimization, enhancing the intelligence and data-driven characteristics of teaching interaction while improving the students' concentration. (2) The present invention obtains the face videos of the instructor and students in real time through a camera to comprehensively record the teaching process; also analyzes the line-of-sight focus areas of the instructor and students through eye-tracking technology, and uses the clustering analysis method. Based on the Euclidean distance as the judgment basis, classifies the line-of-sight focus coordinates to form a set of line-of-sight focus areas. This process helps to understand and record the visual attention points of the instructor and students in the course; helps to discover problems in teaching, improve teaching quality and interaction effects. This comprehensive recording and analysis ability is of great significance for teaching research and teaching improvement. (3) Through the calculation and rating of the line-of-sight focus coincidence rate, the present invention can accurately and quantitatively evaluate the students' concentration in the course; the present invention also realizes a complete process from extracting the set of line-of-sight focus areas of the instructor and students, to constructing a trajectory line based on the time series and annotating the histogram, and then to calculating the coincidence rate and rating; its advantages are that on the one hand, it realizes the fine quantification of the students' concentration, avoiding the ambiguity of traditional subjective evaluation; on the other hand, based on the dynamic analysis of the time series, it can more truly reflect the changes in the students' attention at different stages of the course; in addition, by comparing the coincidence rate with the preset standard to obtain the concentration score, it provides an objective basis for teaching effect evaluation, and at the same time facilitates the operator to flexibly adjust the rating standard according to the actual situation, with strong practicability and adaptability. (4) Through the comparison and analysis of the histograms of the line-of-sight focus areas of the instructor and students, the present invention accurately judges whether the students have abnormal behaviors, and timely reminds the students by means of sound or highlighting, which helps the students to maintain concentration, improve classroom participation and learning effects; at the same time, the system can summarize information such as the students' concentration scores, the number of abnormal behaviors and timestamps to generate a course learning report and feedback it to the instructor, enabling the instructor to comprehensively understand the learning performance and state of each student, facilitating the instructor to adjust teaching strategies targeted, optimize the teaching process, improve the overall teaching quality and efficiency, promote the personalized development of teaching, and realize a more efficient interactive teaching mode. Brief Description of the Drawings
[0016] The present invention will be further described below with reference to the accompanying drawings.
[0017] Figure 1 is a schematic structural diagram of the system of the present invention; Figure 2It is a schematic flowchart of the method described in Embodiment 2 of the present invention; Figure 3 It is a schematic flowchart of the method described in Embodiment 3 of the present invention. Detailed implementation manners
[0018] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.
[0019] Embodiment 1 A multimedia interactive teaching machine control system, as Figure 1 shown, specifically includes the following: The applicable target object of this system is during the teaching process in the computer room, where the instructor teaches the students in the way of screen sharing. During this process, when the instructor focuses on explaining a knowledge point or making an electronic whiteboard note, the focus of the instructor's line of sight will be on a specific area (on the instructor's machine) on the display screen; and when the students are listening attentively, they will focus their line of sight on the specific area (on the students' machines) described by the instructor. If the students are not listening attentively, their line of sight will not converge on the specific area described by the instructor. Based on this characteristic, this embodiment gives a specific architecture of a multimedia interactive teaching machine control system to discuss this characteristic in detail; A data acquisition module, as the medium for the interaction between this system and the reality (interacting with the teaching machine, and the teaching machine includes the students' machines and the instructor's machine), interacts with the teaching machine through a wireless network or a wired network (depending on the actual network connection situation in the computer room); Specifically, the data acquisition module mainly obtains the teaching videos on the instructor's machine and the students' machines of the instructor and the students during the entire course, or the instructor's notes. Because during the teaching process in the computer room, the instructor shares the video image of the instructor's machine to the students' machines through the way of screen sharing, so only one copy of the teaching video needs to be obtained; It should be noted here that the teaching machine includes the students' machines and the instructor's machine; the configurations of the students' machines and the instructor's machine are the same; both the instructor's machine and the students' machines have numbers; the same configurations of the instructor's machine and the students' machines include cameras, displays, hosts, communication protocols, etc.; Both the student machines and the teaching machines are equipped with cameras. The cameras are integrated into the displays and have fixed positions. The cameras will capture the face videos of students or instructors in real time during classes (only for monitoring the learning status of students during classes, and the face videos are saved encrypted. After the learning status of students during classes is judged, they will be deleted to prevent data leakage). The data acquisition module obtains the face videos and teaching videos, and transmits the obtained face videos and teaching videos to the data storage module for storage operations through the system bus.
[0020] The data analysis module contains several computing units for processing the analysis and calculation steps described in this solution. The data analysis module will extract the face videos of instructors and students collected by the data acquisition module from the data storage module and perform time alignment processing. Through time alignment processing, it is possible to accurately observe the real-time reactions of students when the instructor is explaining a certain knowledge point or teaching link, so as to determine the real-time learning status of students. Then, the line-of-sight focus of the instructor at each moment is extracted from the face video of the instructor after alignment processing and mapped onto the instructor machine to be fitted into an instructor line-of-sight focus area. The instructor line-of-sight focus area is the area where the line-of-sight focus of the instructor gathers within a period of time. For example, if when the instructor is explaining a certain knowledge point, the instructor line-of-sight focus area is concentrated on specific courseware content, textbook paragraphs, or a certain part of the blackboard writing, then it can be determined that these contents are the parts that the instructor considers to be emphasized or key. It should be explained here that in the actual use of the method of mapping the instructor's line of sight focus to the instructor's machine, it is necessary to combine technologies such as face recognition technology and line of sight tracking technology, and use computer vision technology and image processing algorithms to analyze the instructor's face video; determine the position and orientation of the instructor's eyes through facial feature point detection technology, and further accurately determine the line of sight direction in combination with head pose estimation; then, according to the relative position relationship between the line of sight direction and the instructor's machine screen, map the line of sight focus to the corresponding coordinate position on the instructor's machine screen; perform such processing at each moment to obtain the line of sight focus coordinates of the instructor at each moment; finally, aggregate and fit all the line of sight focus coordinates within a period of time, and use methods such as clustering algorithms to aggregate the similar line of sight focus coordinates together to form an area, that is, the instructor's line of sight focus area (it should be explained that there are already pioneer examples in this regard in the prior art, and the determination of the line of sight focus and the line of sight focus area has been realized. For example, AU Optronics integrates cutting-edge sensing technology in the Micro-LED in-vehicle display, dynamically captures the driver's head pose and line of sight direction through a camera, and combines the position parameters of the in-vehicle screen to realize the real-time mapping of the line of sight focus; Apple Vision Pro: through the camera system combined with LiDAR and the R1 chip, it can capture the user's head and eye movements in real time, and output the line of sight vector through the Heterogeneous Neural Network (HNN), etc.).
[0021] Next, continue to obtain the student's line of sight focus area from the student's face video after time alignment processing, and compare it with the instructor's line of sight focus area within the same period of time. If there is an overlap, it can be determined that the student is listening during the instructor's lecture. If there is no overlapping part at all, it can be determined that the student has abnormal behavior, that is, not listening. By calculating the overlap rate between the instructor's line of sight focus area and the student's line of sight focus area, the student's concentration is evaluated based on the overlap rate.
[0022] Interaction feedback module. This module complements the data analysis module and generates corresponding processing measures based on the results of data processing in the data analysis module. If there is no overlap between any student line of sight focus area associated with the student and the determined instructor line of sight focus area in the data analysis module, that is, after determining that the student has abnormal behavior, the student's machine used by the student is used to remind the student to pay attention to the lecture in the form of a sound reminder; If there is an overlapping part after comparing the student's line of sight focus area obtained from the student's face video with the instructor's line of sight focus area within the same period of time, but the overlapping part is small, it can be determined that although the student is listening to the lecture, the listening quality is not high, and there may be a situation of attention transfer. At this time, the instructor's line of sight focus area is highlighted on the student's machine by means of high-brightness rendering to remind the student to listen to the lecture normally; At this point, the feedback processing of the students' listening behavior in the course is completed. When the course is over, this module will generate a course learning report associated with the student based on the data processing results of the data analysis module. The learning report includes: the student's concentration score in this course, the number of abnormal behaviors, and the start and end timestamps associated with each abnormal behavior; Feedback each student's learning report to the instructor, and send each student's learning report to the corresponding student for the student to review; The teacher can clearly understand the listening status of each student through this report, so that the teacher can adjust the teaching method in a targeted manner. For example, if most students show abnormal behavior when explaining a certain knowledge point, the teacher can reflect on the shortcomings of the teaching content or method here, and optimize the explanation ideas and adopt more vivid teaching methods in the next class, such as adding cases and interactive links, so as to improve the teaching effect; After receiving the learning report, students can clearly see their concentration and distraction performance in the course, and focus on reviewing the knowledge points corresponding to the periods when abnormal behaviors occurred to fill in learning gaps.
[0023] The data storage module includes several databases for storing any calculation step and analysis step described in this solution, and is also used to store the calculation result of any calculation step and the analysis result of the analysis step in this solution.
[0024] This embodiment introduces a multimedia interactive teaching machine control system architecture for monitoring the listening status of students based on the overlap of the instructor's and students' visual focus areas in a computer room teaching scenario; its purpose is to achieve accurate monitoring and feedback of students' listening behavior, obtain relevant video data on the instructor's machine and the student's machine through the data acquisition module, and perform time alignment, visual focus area extraction and comparison and other processing through the data analysis module. The interactive feedback module takes corresponding measures based on the analysis results, such as sound reminders, highlighting, etc., and finally generates a course learning report to feed back to the instructor and the student; it is intended to help the instructor understand the students' listening status in time so as to adjust the teaching method, and at the same time let the students know their own learning situation and optimize their learning behavior, so as to improve the overall teaching quality and learning effect, and realize efficient interaction and personalized management of the teaching process. The entire system design takes into account data security and privacy protection, and emphasizes the encrypted storage of face videos and the deletion processing after use.
[0025] Example 2 This embodiment discloses a method for determining the instructor's sight focus area and the trainee's sight focus area based on the embodiment 1, such as Figure 2 As shown, the specific steps include: By processing the data of the student machine and the instructor machine through the data acquisition module described in Embodiment 1, we can obtain the instructor's face video and the student's face video of the instructor and the student during the course; Next, perform relevant data processing. After the data processing, it is necessary to identify the instructor and the students in any course. In this embodiment, any course is analyzed to obtain the instructor of this course, denoted as , and determine the attending instructor The total number of students in this course of, denoted as ; As described in Embodiment 1, it can be known that the student machine has a number. The students in this course are sorted in descending order according to the number, and finally The student sequence associated with the students is obtained, expressed as: ; The purpose of this data processing process is to better organize and understand the personnel composition of the course.
[0026] Next, the instructor's face video collected by the cameras equipped on the teaching machine and the student machine is denoted as . It should be noted that the step of obtaining the instructor's face video is carried out in real time. That is to say, the obtained instructor's face video Changes over time and the video data volume increases because more video frames (video frames captured by the camera) are included. Similarly, the student's associated student face video is also obtained in real time.
[0027] Next, process all the student face videos of the students in the collected student sequence . Sort all the collected student face videos of the students in the order of the student sequence . After the sorting process, a student face video sequence is obtained, expressed as ; Then, determine any one student from the student sequence for example processing. Determine the student as the target for example processing. The remaining students are all processed according to the steps of processing the student . Then, extract the student face video associated with the student from the student face video sequence , denoted as , where i is a counting index with a value range of 1 to .
[0028] Taking the teaching and the instructor associated instructor face video as an example for processing. Similarly, for the student The associated student face video Perform synchronization processing, and the processing steps are the same; From the above content, the instructor face video can be obtained is composed of several frames of images. Therefore, first, extract each frame of image from the instructor face video That is, extract all the video frames in the instructor face video and represent the instructor face video in the form of a video frame sequence according to the order of the video frames, and obtain the instructor face video frame sequence, denoted as: , where is the total number of video frames in the instructor face video , is not a fixed value and increases with time.
[0029] At this time, it is also necessary to extract the teaching video after the instructor face video is aligned in time, that is, the picture displayed on the instructor machine screen, and record the extracted teaching video as ; Next, construct a two-dimensional coordinate system with the picture of the teaching video . This step is also to prepare for mapping the instructor's line-of-sight focus to the two-dimensional coordinate system to obtain the instructor's line-of-sight focus area; The specific method for constructing the two-dimensional coordinate system is as follows: Take the first pixel point in the lower left corner of the picture of the teaching video as the origin of the two-dimensional coordinate system, the lower boundary of the picture of the teaching video as the horizontal axis of the two-dimensional coordinate system, and the left boundary of the picture of the teaching video as the vertical axis of the two-dimensional coordinate system to construct the two-dimensional coordinate system associated with the teaching video , and denote this two-dimensional coordinate system as (that is, the two-dimensional coordinate system associated with the video pictures of the instructor machine and the student machine).
[0030] Starting from the first video frame in the instructor face video frame sequence of the instructor described in Embodiment 1 , use face recognition technology and eye movement tracking technology to extract the instructor's line-of-sight focus in the first video frame , and map it to the constructed two-dimensional coordinate system . By extracting the two-dimensional coordinates of the points mapped to the two-dimensional coordinate system , obtain the instructor's face video frame sequence The first video frame in is located in the teaching video The coordinates of the line of sight focus in the picture are marked as ; then, continue to process backward until , that is, the last video frame in the face video frame sequence Finally, obtain the number of video frames associated with the face video frame sequence in the two-dimensional coordinate system on the number of line of sight focus coordinates, that is, the instructor The line of sight focus of the instructor in the instructor's face video frame sequence is located in the teaching video The coordinates of the line of sight focus in the picture are sorted in the order of acquisition to obtain the set of line of sight focus coordinates , expressed as: .
[0031] Next, it is necessary to perform clustering analysis on the set of line of sight focus coordinates associated with the instructor From the determined set of line of sight focus coordinates : Obtain the first line of sight focus coordinate as a reference, and the reference is the basis for comparing subsequent line of sight focus coordinates; Then, obtain the next line of sight focus coordinate adjacent to the line of sight focus coordinate from the set of line of sight focus coordinates, that is, the line of sight focus coordinate , and extract the horizontal and vertical coordinates of the line of sight focus coordinate and the horizontal and vertical coordinates of the line of sight focus coordinate respectively. Calculate the Euclidean distance between the line of sight focus coordinate and the line of sight focus coordinate through the Euclidean distance formula, and mark the calculated result as: ; ; Compare the Euclidean distance between the line of sight focus coordinate and the preset Euclidean distance threshold by the operator in combination with the actual situation for comparison and judgment; If the Euclidean distance is less than or equal to the Euclidean distance threshold , then the line of sight focus coordinate and the line of sight focus coordinate Perform clustering with the (reference), and group them together. Then continue to obtain the next line-of-sight focus coordinate backward. Here, it is to obtain the line-of-sight focus coordinate , and repeat the steps of Euclidean distance comparison and clustering; If the Euclidean distance between two line-of-sight focus coordinates is greater than the Euclidean distance threshold during the comparison of Euclidean distances , then perform the following operations; Example: If during the above determination of the Euclidean distance between the line-of-sight focus coordinate and the line-of-sight focus coordinate , it is determined that the Euclidean distance is greater than the Euclidean distance threshold , then continue to judge the next adjacent line-of-sight focus coordinate backward, that is, judge the Euclidean distance between the line-of-sight focus coordinate and the line-of-sight focus coordinate . If f line-of-sight focus coordinates are continuously extracted, and the Euclidean distances between these f line-of-sight focus coordinates and the line-of-sight focus coordinate are all greater than the Euclidean distance threshold , then it is determined that the clustering of these f line-of-sight focus coordinates and the line-of-sight focus coordinate fails and they cannot be grouped together. At this time, discard the line-of-sight focus coordinate and cancel the reference identity of the line-of-sight focus coordinate ; And use the first line-of-sight focus coordinate among the f line-of-sight focus coordinates as the new reference. In this example, use the line-of-sight focus coordinate as the new reference; Based on the determined new reference: the line-of-sight focus coordinate , continue to repeat the above clustering judgment operation until the last line-of-sight focus coordinate in the line-of-sight focus coordinate set : is reached and the operation ends. Here, f is a value preset by the operator to prevent frequent replacement of the reference due to an occasional outlier coordinate, improving the stability of clustering.
[0032] During the clustering operation, if the total number of line-of-sight focus coordinates including the reference in the same group exceeds n, it means that the clustering of this group passes (the line-of-sight focus coordinates in this group can be n + 1, or n + 2, as long as it is greater than n). Then use a fixed circle to cover as many line-of-sight focus coordinates in this group as possible, and perform a fitting operation to obtain a circular area, and use this area as the instructor's line-of-sight focus area associated with the successfully clustered line-of-sight focus coordinates of this group, that is, the instructor An associated instructor's line-of-sight focus area, extract all the line-of-sight focus coordinates in this group in the order of the time line. Use the time stamp of the first line-of-sight focus coordinate as the start time stamp, and then extract the time stamp of the last line-of-sight focus coordinate in this group as the end time stamp. And use the start time stamp and the end time stamp as the start time stamp and the end time stamp of this instructor's line-of-sight focus area. Wherein, the value of n is preset by the operator, and the radius of the fixed circle is greater than the Euclidean distance threshold , and the specific value is determined by the operator according to the actual situation. Since the size of the fixed circle is fixed, the number of pixel points in the covered area is regarded as fixed. For the convenience of subsequent calculations, the number of pixel points in the area covered by the fixed circle is denoted as ; Then extract the next adjacent line-of-sight focus coordinate in the line-of-sight focus coordinate set : of the last line-of-sight focus coordinate associated with this instructor's line-of-sight focus area as a new benchmark, and repeat the above steps to continue to obtain the instructor associated instructor's line-of-sight focus area until the line-of-sight focus coordinate set : The last line-of-sight focus coordinate in , ending the clustering operation.
[0033] Denote the first instructor's line-of-sight focus area associated with the instructor as , and the subsequent instructor's line-of-sight focus areas are arranged in order backward. For example, the second instructor's line-of-sight focus area is denoted as .
[0034] Finally, summarize all the instructor's line-of-sight focus areas obtained in the above steps in the order of the time line associated with the instructor, and denote it as the instructor's line-of-sight focus area set, expressed as: , where represents the total number of the instructor's line-of-sight focus areas on the instructor's machine (teaching video) of the instructor in the course; So far, the processing of the instructor's line-of-sight focus area of the instructor ends. The obtained instructor's line-of-sight focus area set associated with the instructor , and obtain the student's line-of-sight focus area set associated with the student in this course and the student's line-of-sight focus area sets associated with other students respectively according to the above method.
[0035] This embodiment aims to process the face video data collected from the instructor machine and the trainee machine, extract the line-of-sight focus information of the instructor and the trainee, and map it to the two-dimensional coordinate system of the teaching video. By using face recognition and eye movement tracking technologies, the line-of-sight focus coordinates in each video frame are obtained, and then clustering analysis is performed to form the line-of-sight focus area sets of the instructor and the trainee. This method ultimately realizes the quantification and visualization of the attention distribution of the instructor and the trainee during the teaching process.
[0036] Embodiment 3 Based on Embodiment 1 and Embodiment 2, this embodiment further discloses a method for evaluating the concentration of a trainee based on the coincidence rate between the instructor's line-of-sight focus area and the trainee's line-of-sight focus area, and reminding the trainee in real time during the evaluation process, as Figure 3 shown, which specifically includes the following steps: According to the final processing in Embodiment 2, the instructor associated instructor line-of-sight focus area set can be obtained. Then, extract the trainee sequence and for any trainee in the course, extract the trainee line-of-sight focus area set associated with the trainee in the course: , where represents the total number of trainee line-of-sight focus areas on the trainee machine for trainee in the course. In this embodiment, trainee is used as an example for processing, and the remaining trainees are processed in the same way as trainee .
[0037] Determine the start timestamp and end timestamp of each instructor line-of-sight focus area in the instructor associated instructor line-of-sight focus area set . Similarly, extract the start timestamp and end timestamp of each trainee line-of-sight focus area from the trainee associated trainee line-of-sight focus area set ; Then construct a time series trajectory line , and the time series trajectory line has a time span from the start time of the current course to the end time of the current course; Then mark each instructor line-of-sight focus area in the instructor line-of-sight focus area set on the time series trajectory line in the form of a histogram column; and mark each trainee line-of-sight focus area in the trainee line-of-sight focus area set on the time series trajectory line It is marked in the form of a histogram above; the height of the histogram is constant and has no impact on this solution.
[0038] Instructor Any histogram mapped by the set of instructor's line-of-sight focus areas and the student The set of student's line-of-sight focus areas on the time-series trajectory line Any histogram mapped above may have overlapping parts. If an overlap occurs, then the start timestamp and end timestamp on the time-series trajectory line of the overlapping part are extracted, forming a timestamp interval, denoted as .
[0039] If there is any instructor The histogram mapped by any instructor's line-of-sight focus area associated with and the student If there is no overlap between the histograms of any student's line-of-sight focus area associated with, then it is determined that the student has abnormal behavior within the start timestamp and end timestamp associated with the instructor's line-of-sight focus area, and a sound reminder is sent to the student through the student machine used by the student to the student.
[0040] Instructor The histogram mapped by the set of instructor's line-of-sight focus areas and the student The set of student's line-of-sight focus areas on the time-series trajectory line The overlapping part of the histograms mapped above indicates that within the timestamp interval associated with the overlapping part, the instructor has an instructor's line-of-sight focus area, denoted as and the student has a student's line-of-sight focus area, denoted as However, it cannot be proven that and are the same area (or and There is an overlapping part between them), that is, it cannot be determined whether the student is listening attentively while the instructor is concentrating on teaching. Therefore, continue to extract the two-dimensional coordinates of the pixel points in the instructor's line-of-sight focus area associated with the instructor and the two-dimensional coordinates of the pixel points in the student's line-of-sight focus area associated with the student and judge the instructor's line-of-sight focus area and the student's line-of-sight focus area Whether there are overlapping parts. If the two-dimensional coordinates are the same, they overlap; if the two-dimensional coordinates are different, they do not overlap.
[0041] Extract the area of the instructor's line of sight focus and the area of the student's line of sight focus The number of pixel points in the overlapping part is denoted as ; Use to calculate the overlapping rate associated with the area of the instructor's line of sight focus and the area of the student's line of sight focus ; ; Then extract the overlapping rate rating standard preset by the operator: , and compare the calculated overlapping rate with the overlapping rate rating standard to determine the student's associated overlapping rate rating within the timestamp interval (the overlapping rate rating also reflects the student's learning state and learning concentration within the timestamp interval ). Among them, , , are all determined by the operator according to the actual situation, and is the minimum overlapping rate threshold for judging that the student is listening to the class normally. If the student has an overlapping rate less than within the timestamp interval is considered to be in an abnormal listening state at present. At this time, it is necessary to remind the student by highlighting the area of the instructor's line of sight focus, informing the student to pay attention to the class.
[0042] By repeating the above steps until the end of this course, the overlapping rate ratings of the student in this course can be determined, and the scores associated with different overlapping rate ratings are extracted, and the concentration score of the student in this course is calculated by the accumulation method to evaluate the concentration of the student .
[0043] Among them, the scores corresponding to different overlapping rate ratings in the overlapping rate rating standard are as follows: excellent corresponds to a score of 20; good corresponds to a score of 15; medium corresponds to a score of 10; poor corresponds to a score of 5.
[0044] The higher the concentration score, the higher the concentration of the student in the course; conversely, the lower the concentration score, the lower the concentration of the student in the course.
[0045] The instructor can set the standard attention score associated with this course by themselves and directly obtain all students with attention scores lower than the standard attention score and students with attention scores higher than or equal to the standard attention score after the course, which is convenient for the instructor to directly understand the students' listening situation.
[0046] In this embodiment, the coincidence rate is calculated by analyzing the timestamps and pixel coordinates of the line-of-sight focus areas of the instructor and students, and the attention level of the students during the course is quantified into specific scores; a time-series trajectory line is constructed to mark the line-of-sight focus areas of the instructor and students. Whether the students are listening attentively is judged by the timestamp interval and pixel coordinates of the overlapping part and rated according to the preset standard, and corresponding reminders are given to the students; finally, the attention scores of the students are comprehensively calculated, and the instructor can understand the students' listening situation based on this, including the list of students with attention scores lower than or meeting the standard attention score; this embodiment improves the automation and accuracy of teaching evaluation and provides real-time feedback for the teaching process.
[0047] Some of the data in the above-mentioned formulas are numerically calculated after removing their dimensions, and the content not described in detail in this specification belongs to the prior art well-known to those skilled in the art.
[0048] The above content is only an example and illustration of the present invention. Those skilled in the art of this technology can make various modifications or supplements to the described specific embodiments or use similar methods to replace them. As long as they do not deviate from the invention or exceed the scope defined by this claim book, they should fall within the protection scope of the present invention.
[0049] It should be stated that: all user data collected in this application are collected with the consent and authorization of the users. And the uses of the user data are all legal and compliant, and the use and processing of the user data comply with the relevant laws, regulations and standards of the relevant regions.
Claims
1. A multimedia interactive teaching machine control system, characterized in that The system includes the following: A data acquisition module that interacts with the teaching machine, obtains the instructor's face video and the student's face video of the instructor and the student, and transmits them to the data storage module for storage; A data analysis module that extracts the instructor's face video and the student's face video, performs time alignment, and determines the instructor's line-of-sight focus area associated with the instructor and the student's line-of-sight focus area associated with the student; Extract the start timestamp and end timestamp of the instructor's line-of-sight focus area, and obtain the student's line-of-sight focus area within the start timestamp and end timestamp; If the acquisition is successful, calculate the coincidence rate between the instructor's line-of-sight focus area and the student's line-of-sight focus area and evaluate the student's concentration; If the acquisition fails, it is regarded as an abnormal behavior of the student and recorded; An interactive feedback module that, if the student has an abnormal behavior, reminds the student through sound; If the coincidence rate is lower than the minimum coincidence rate threshold, highlight and render the instructor's line-of-sight focus area to remind the student; Generate a course learning report of the student and feedback it to the instructor after the course ends.
2. The control system of a multimedia interactive teaching machine according to claim 1, wherein The system also includes a data storage module that stores any calculation step and analysis step in this solution, and stores the calculation result of any calculation step and the analysis result of the analysis step in this solution.
3. The control system of a multimedia interactive teaching machine according to claim 1, wherein The data acquisition module interacts with the teaching machine in real time; The teaching machine includes a student machine and an instructor machine; The configurations of the student machine and the instructor machine are the same; Both the instructor machine and the student machine have numbers; The configuration includes a camera, a display, a host, and a communication protocol; The camera is integrated on the display and the position is fixed; The instructor machine shares the teaching video and the instructor's blackboard writing to the student machine.
4. The control system of a multimedia interactive teaching machine according to claim 3, wherein The specific method for the data acquisition module to obtain the instructor's face video and the student's face video of the instructor and the student is as follows: Identify the instructor in the course, marked as ; Determine the total number of students in the course, denoted as ; Sort the students in the order of the instructor aircraft numbers to obtain the student sequence ; Obtain instructor Instructor's face video in the course ; Obtain the trainee sequence All the trainee face videos of the trainees in it, and sort them in the order of the trainee sequence to obtain the trainee face video sequence .
5. The control system of a multimedia interactive teaching machine according to claim 4, wherein The specific method for the data analysis module to determine the instructor's line-of-sight focus area associated with the instructor and the student's line-of-sight focus area associated with the student is as follows: Determine any trainee in the trainee sequence , extract the associated trainee face video , where i is a counting index, and its value range is from 1 to ; Split the instructor face video according to the order of video frames into a sequence of instructor face video frames , where is the total number of video frames in Extract the teaching video with the instructor's face again The teaching video after time alignment processing ; Taking the first pixel point in the lower left corner of the screen as the origin, the lower boundary as the horizontal axis, and the left boundary as the vertical axis to construct the associated two-dimensional coordinate system, denoted as , which is also the two-dimensional coordinate system associated with the instructor aircraft display screen; Extract video frames , obtain The line of sight focus in and map it to to obtain the mapped point and extract the two-dimensional coordinates of the mapped point as The line of sight focus in is located at the line of sight focus coordinates in ; Repeat the above steps to obtain the line-of-sight foci of all video frames in the instructor's face video frame sequence are located at the set of line-of-sight focus coordinates in ; Pair Perform clustering to determine In All the instructor's line-of-sight focus areas associated therein are summarized in chronological order to obtain The set of instructor's line-of-sight focus areas associated , where Indicates The total number of the instructor's line-of-sight focus areas associated in the course; Similarly, the student and the student's line-of-sight focus areas and the set of the student's line-of-sight focus areas associated with other students during the course are obtained.
6. The control system of a multimedia interactive teaching machine according to claim 5, wherein, The data analysis module clusters to determine All specific ways of the instructor's line-of-sight focus areas associated in are as follows: Extract the set of line-of-sight focus coordinates The line-of-sight focus coordinates in , and use them as a reference, then extract backward , calculate , The Euclidean distance between ; Obtain the Euclidean distance threshold preset by the operator , if , then and are clustered as the same group, and continue to obtain the line-of-sight focus coordinates backward for clustering judgment with the benchmark; If , then extract for clustering judgment. If the Euclidean distances between f consecutively extracted line-of-sight focus coordinates and are all greater than , then discard , use as a reference, and continue to extract line-of-sight focus coordinates backward for clustering judgment, where f is a value preset by the operator; If the total number of line-of-sight focus coordinates in the same group, including the reference, exceeds n, it is regarded as successful clustering of the same group. A fixed circle is used to cover all the line-of-sight focus coordinates with successful clustering and fit them into a circular area, which is regarded as the instructor's line-of-sight focus area associated with the line-of-sight focus coordinates with successful clustering of the same group. The timestamp of the first line-of-sight focus coordinate in the same group is extracted as the start timestamp of the instructor's line-of-sight focus area, and then the timestamp of the last line-of-sight focus coordinate in the same group is extracted as the end timestamp of the instructor's line-of-sight focus area. Among them, the radius of the fixed circle is greater than the Euclidean distance threshold The value of n is preset by the operator; Determine the first instructor's line of sight focus area, denoted as , and continue to obtain the line of sight focus coordinates backward for clustering judgment until the line of sight focus coordinates that fail to cluster with the benchmark are obtained as the new benchmark. Repeat the above steps until the last line of sight focus coordinate in the set of line of sight focus coordinates is clustered to obtain All the instructor's line of sight focus areas associated in . The subsequent instructor's line of sight focus areas are recorded in numerical order.
7. The control system of a multimedia interactive teaching machine according to claim 6, characterized in that, The specific method for the data analysis module to calculate the coincidence rate between the instructor's line-of-sight focus area and the student's line-of-sight focus area and evaluate the student's concentration is as follows: S71. Extraction The set of instructor's line-of-sight focus areas associated ; S72. Extract the trainee The set of trainee's line of sight focus areas associated , where represents the total number of trainee's line of sight focus areas associated in the course; S73. Extraction The start timestamp and end timestamp of the line of sight focus area of all instructors Extract again from the start timestamp and the end timestamp of the line-of-sight focus area of all students; S74. Construct a time series trajectory line , The time span of which is from the start time to the end time of the course; For each instructor's line of sight focus area in , mark it in the form of a histogram on according to their respective associated start timestamps and end timestamps. S75. Then, mark the start timestamp and end timestamp associated with the line of sight focus area of each trainee in in the form of a histogram. S76. Extract from the time stamp interval when the histogram measured for any one instructor's line of sight focus area coincides with the histogram measured for any one student's line of sight focus area, denoted as ; S77. Then extract the instructor's line-of-sight focus area containing , and the trainee's line-of-sight focus area containing , and calculate the coincidence rate of the instructor's line-of-sight focus area and the trainee's line-of-sight focus area ; S78. Extract the coincidence rate rating standard preset by the operator: , and compare with the coincidence rate rating standard to obtain the coincidence rate rating associated within the timestamp interval . Among them, , , The specific values of are all determined by the operator in combination with the actual situation, and is the minimum coincidence rate threshold for judging that the trainee is attending class normally; Repeat the above steps to determine all the coincidence rate ratings in the course, calculate the concentration score based on the coincidence rate rating, and evaluate the concentration of the trainees.
8. A control system for a multimedia interactive teaching machine according to claim 7, characterized in that, The data analysis module calculates the concentration score based on the coincidence rate rating to evaluate the concentration of the students in the following specific way: In the coincidence rate rating standard, the score for excellent is 20; the score for good is 15; the score for medium is 10; the score for poor is 5; Obtain trainees Statistically analyze all the coincidence rate ratings associated in the course and the scores corresponding to different coincidence rate ratings in the coincidence rate rating standard The associated concentration scores; The concentration score is proportional to the student's concentration.
9. The control system of a multimedia interactive teaching machine according to claim 7, wherein In the interactive feedback module, the specific method for reminding the student includes the following steps: In the S76, if there is no overlap between the histogram mapped by any one of the instructor's line of sight focus areas associated with and the histogram of any one of the student's line of sight focus areas associated with it, it is determined that abnormal behavior occurs within the start timestamp and end timestamp associated with the instructor's line of sight focus area and a sound reminder is sent through the student machine used to send out a sound reminder. In step S77, if the coincidence rate is lower than the lowest coincidence rate threshold , the focus area of the instructor's line of sight is highlighted to remind the trainee.
10. A control system for a multimedia interactive teaching machine according to claim 9, characterized in that, In the interactive feedback module, the specific method for generating a course learning report of the student and feedbacking it to the instructor is as follows: Summarize students The concentration score in the course, the number of abnormal behaviors, and the students Generate the course learning report associated with each student based on the start timestamp and end timestamp associated with each abnormal behavior associated with the course in the course; Repeat the above steps to generate the course learning reports associated with each individual trainee in , and transmit them to the instructor's machine used by the instructor for the instructor to view.