A method, device, equipment, and storage medium for monitoring classroom attention.

By using a dual-camera system and a three-dimensional coordinate system to monitor the teacher's movement path and the students' line of sight, the problem of low accuracy in monitoring students' attention in the classroom is solved, and more accurate attention assessment is achieved.

CN120599540BActive Publication Date: 2026-05-05JIUJIANG DIGITAL IND DEV CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
JIUJIANG DIGITAL IND DEV CO LTD
Filing Date
2025-06-03
Publication Date
2026-05-05

AI Technical Summary

Technical Problem

The accuracy of current technologies for monitoring student attention in the classroom is low, which affects teaching effectiveness.

Method used

A dual-camera system is used, with the first camera positioned in front of the students for facial recognition and the second camera positioned above them for head recognition. Combined with a three-dimensional coordinate system, the system monitors the teacher's movement path and the students' line of sight, and assesses the students' attention through path similarity.

Benefits of technology

It improves the accuracy of monitoring student attention in the classroom, effectively identifying whether students are following the teacher's movements, thereby judging their concentration level.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120599540B_ABST
    Figure CN120599540B_ABST
Patent Text Reader

Abstract

Embodiments of the present application disclose a classroom attention monitoring method and device, equipment and a storage medium. In a teaching mode, a first video captured by a first camera capable of capturing a face and a second video captured by a second camera capable of capturing the overall appearance of a classroom are used to obtain the moving path of the head of a teacher and the line-of-sight path of the line of sight of each student, and the attention in the classroom is monitored through the similarity between the moving path and the line-of-sight path, which can effectively monitor whether the line of sight of the student in the teaching mode follows the movement of the teacher, so as to reflect the attention to the teaching content in the form of following the movement of the teacher, and effectively improve the accuracy of attention monitoring.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of educational information technology, and in particular to a method, device, equipment, and storage medium for monitoring classroom attention. Background Technology

[0002] In the classroom, the effectiveness of teaching is closely related to students' attention levels. Teachers not only have to deliver lectures but also constantly monitor students' attention, inevitably overlooking the attention spans of some students. Therefore, paying close attention to students' attention becomes crucial for improving teaching quality.

[0003] However, existing attention monitoring methods all suffer from low accuracy. Therefore, improving the accuracy of attention monitoring in the classroom is an urgent problem to be solved. Summary of the Invention

[0004] The main objective of this invention is to provide a method, device, equipment, and storage medium for monitoring classroom attention, which can solve the problem of low accuracy in monitoring students' attention in the classroom in the prior art.

[0005] To achieve the above objectives, the first aspect of the present invention provides a method for monitoring classroom attention, the method comprising:

[0006] If the current classroom environment is detected to have entered teaching mode, the system acquires a first video of the classroom captured by the first camera and a second video of the same time period captured by the second camera. The first camera is positioned in front of the student's seat in the classroom, and the second camera is positioned above the student's seat. The teaching mode refers to a classroom learning mode in which the teacher teaches knowledge points.

[0007] The first video is subjected to face recognition processing to obtain the face orientation angle of each student's face region in each video frame of the first video. The second video is subjected to head recognition processing to determine the first position of each student's head in the student desk and chair area at each time point of the second video frame, and the second position of the teacher's head in the lecture area at each time point of the second video frame. The first position and the second position both include the x value and y value in the first coordinate system, which is a three-dimensional coordinate system based on the three-dimensional space formed by the classroom.

[0008] The teacher's head movement path in the first coordinate system is determined based on the second position; the line of sight of each student in the first coordinate system is determined based on the face orientation angle, the first position, and the second position.

[0009] The attention of each student is monitored based on the similarity between the teacher's movement path and the student's gaze path.

[0010] Optionally, determining the movement path of the teacher's head in the first coordinate system based on the second position includes:

[0011] Obtain the height information of the instructor and the height information of the lecture area in the classroom;

[0012] Based on the height information of the instructor, the height information of the lecture area, and the second position, the head position information of the instructor in the first coordinate system at each time point is determined;

[0013] The teacher's movement path is formed by using the head position information at each time point in the first coordinate system.

[0014] Optionally, determining the line-of-sight path of each student in the first coordinate system based on the face orientation angle, the first position, and the second position includes:

[0015] Obtain the height of each student and the height of the desk and chair, and determine the height of each student's eye area when they are sitting down based on the height of each student and the height of the desk and chair.

[0016] Based on the height of each student's eye and the first position, determine the coordinates of each student's eye at each time point in the first coordinate system;

[0017] Based on the coordinates of each student's eye area in the first coordinate system at each time point, the face orientation angle, and the second position, determine the coordinates of each student's line of sight in the first coordinate system at each time point;

[0018] The line-of-sight path of each student is determined by using the coordinates of the line-of-sight point in the first coordinate system at each time point.

[0019] Optionally, determining the coordinates of each student's line-of-sight point in the first coordinate system at each time point based on the coordinates of each student's eye location in the first coordinate system, the face orientation angle, and the second position includes:

[0020] Based on the coordinates of the students' eye parts at the same time point, the face orientation angle determines the gaze vector of each student in the first coordinate system;

[0021] Determine a target plane in the first coordinate system that is parallel to the YZ axes of the first coordinate system and includes the second position;

[0022] The intersection of the line-of-sight vector and the target plane is taken as the coordinates of the student's line-of-sight point in the first coordinate system.

[0023] Optionally, the step of performing face recognition processing on the first video to obtain the face orientation angle of each student's face region in each frame of the first video includes:

[0024] The first video is traversed, and a face deep learning model is used to perform face recognition on the target video frames that are traversed, and the face region images contained in the target video frames are extracted.

[0025] The OpenCV cross-platform computer vision library is used to perform key point recognition on the face region images to determine the first set of coordinate points of the key points of each face region image in the camera coordinate system. The key point set includes: the center point of the left eye, the center point of the right eye, the center point of the chin, and the center of the eyebrows.

[0026] The first set of key points in each face region image is transformed to obtain a second set of coordinates in the second coordinate system.

[0027] The facial orientation angle of each facial region in the second set is calculated by taking the coordinates of the center point of the left eye, the center point of the right eye, the center point of the chin, and the center of the eyebrows. The directions of the second coordinate system are parallel to the directions of the first coordinate system.

[0028] Optionally, the calculation of the coordinates of the center point of the left eye, the center point of the right eye, the center point of the chin, and the center of the eyebrows of each face region in the second set to obtain the face orientation angle of each face region in the second coordinate system includes:

[0029] Determine the coordinates of the center point of the left eye and the center point of the right eye using a first line segment, and determine the coordinates of the center point of the chin and the center of the eyebrows using a second line segment;

[0030] Determine a first plane that is perpendicular to the first line segment and intersects the midpoint of the first line segment; determine a second plane that is perpendicular to the second line segment and includes the midpoint of the first line segment.

[0031] The three-dimensional vector formed by the lines intersecting the first plane and the second plane, starting from the midpoint of the first line segment and pointing outwards from the face, is taken as the face orientation angle.

[0032] Optionally, the step of monitoring the attention of each student based on the similarity between the teacher's movement path and the students' gaze paths includes:

[0033] Calculate the similarity between the movement path and the gaze path of each student to obtain the similarity value of each student's gaze tracking;

[0034] If the target similarity value is greater than or equal to a preset threshold, then the student's attention concentration corresponding to the target similarity is determined.

[0035] If the target similarity is less than a preset threshold, then it is determined that the student corresponding to the target similarity value is not focused.

[0036] To achieve the above objectives, a second aspect of the present invention provides a classroom attention monitoring device, the device comprising:

[0037] The acquisition module is used to acquire a first video of the classroom captured by a first camera and a second video of the same time period as the first video captured by a second camera if the current classroom environment is detected to have entered the teaching mode. The first camera is set in front of the student's seat in the classroom, and the second camera is located above the student's seat. The teaching mode refers to the classroom learning mode in which the teacher teaches knowledge points.

[0038] The first determining module is used to perform face recognition processing on the first video to obtain the face orientation angle of each student's face region in each video frame of the first video, and to perform head recognition processing on the second video to determine the first position of each student's head located in the student desk and chair area at each time point of the second video, and the second position of the head of the lecturer located in the lecture area at each time point of the second video. The first position and the second position both include x and y values ​​in a first coordinate system, which is a three-dimensional coordinate system based on the three-dimensional space formed by the classroom.

[0039] The second determining module is used to determine the movement path of the teacher's head in the first coordinate system based on the second position; and to determine the line of sight path of each student in the first coordinate system based on the face orientation angle, the first position, and the second position.

[0040] The monitoring module is used to monitor the attention of each student based on the similarity between the teacher's movement path and the students' gaze paths.

[0041] To achieve the above objectives, a third aspect of the present invention provides a computer-readable storage medium storing a computer program, which, when executed by a processor, causes the processor to perform the steps of the method provided in the first aspect.

[0042] To achieve the above objectives, a fourth aspect of the present invention provides a computer device including a memory and a processor, the memory storing a computer program that, when executed by the processor, causes the processor to perform the steps of the method provided in the first aspect.

[0043] The embodiments of the present invention have the following beneficial effects:

[0044] This invention provides a classroom attention monitoring method. In this method, if the current classroom environment is detected to have entered a teaching mode, a first video of the classroom captured by a first camera and a second video of the same time period captured by a second camera are acquired. The first camera is positioned in front of the students' seats, facing the students and able to capture their faces. The second camera is positioned above the students' seats, allowing it to capture a panoramic view of the classroom from a top-down angle. Furthermore, the teaching mode refers to a classroom learning mode where the teacher imparts knowledge points, indicating that students need to look up at the teacher and listen to the lecture. In this mode, students need to follow the teacher's position to better understand the content. Finally, facial recognition is performed on the first video. The process involves obtaining the facial orientation angles of each student's face in each video frame of the first video, performing head recognition processing on the second video, determining the first position of each student's head in the student seating area corresponding to the time point of each video frame in the second video, and the second position of the teacher's head in the lecturing area corresponding to the time point of each video frame in the second video. Both the first and second positions include x and y values ​​in a first coordinate system, which is a three-dimensional coordinate system based on the three-dimensional space formed by the classroom. The movement path of the teacher's head in the first coordinate system is determined based on the second position. Based on the facial orientation angle, the first position, and the second position, the gaze path of each student in the first coordinate system is determined. The attention of each student is monitored based on the similarity between the movement path and the gaze path. By using a first video captured by a first camera that can capture faces and a second video captured by a second camera that can capture the entire classroom in a teaching mode, the movement path of the teacher's head and the gaze paths of each student are obtained. Classroom attention is monitored by the similarity between the movement path and the gaze path. This method can effectively detect whether students' gazes follow the teacher's movements in a teaching mode, thus demonstrating their attention to the teaching content and improving the accuracy of attention monitoring. Attached Figure Description

[0045] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0046] in:

[0047] Figure 1 This is a flowchart illustrating the classroom attention monitoring method in an embodiment of the present invention;

[0048] Figure 2 This is a flowchart illustrating the method for determining the orientation angle of a person's face in an embodiment of the present invention.

[0049] Figure 3 This is a flowchart illustrating the method for determining the movement path of the teacher's head in the first coordinate system based on the second position in an embodiment of the present invention.

[0050] Figure 4 This is a schematic diagram of the program module of the classroom attention monitoring device in an embodiment of the present invention;

[0051] Figure 5 This is a structural block diagram of a computer device in an embodiment of the present invention. Detailed Implementation

[0052] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0053] To better understand the technical solution in this application, the application scenario is first introduced below. This scenario can be a school teaching scenario, applicable to daily teaching work in schools, such as classrooms for different grades in primary, middle, high school, or university. To monitor student attention in the classroom, cameras are installed, including a first camera and a second camera. The first camera is positioned in front of all student seats, meaning that students typically face the blackboard or display screen. To effectively capture student faces, the first camera can be placed in the area of ​​the blackboard or display screen or an adjacent area, allowing it to capture images of the classroom from the front. Furthermore, the number of first cameras in the classroom can be one or multiple. For example, the length of the plane containing the blackboard or display screen can be divided into two equal segments, defining the middle position and the two endpoints along the length. First cameras can be placed at the middle and endpoints, allowing the images captured by multiple first cameras to be combined for more complete image information of each student, thus improving facial recognition. Furthermore, the second camera can be located above the student seats, specifically on the classroom ceiling, allowing it to capture a panoramic view of the classroom from above. The classroom can be divided into two areas: a student desk and chair area and a lecture area. The second camera can be positioned in the center of the ceiling or in the center of the student desk and chair area. It should be noted that in practical applications, the positions and number of the first and second cameras can be set according to specific needs; no limitations are imposed here. Additionally, a control device is provided, which communicates with the first and second cameras to acquire the video footage captured by both cameras.

[0054] Please see Figure 1 The diagram below illustrates a classroom attention monitoring method according to an embodiment of this application. The method includes:

[0055] 101. If the current classroom environment is detected to have entered the teaching mode, the first video of the classroom captured by the first camera and the second video captured by the second camera at the same time period as the first video are obtained. The teaching mode refers to the classroom learning mode in which the teacher teaches knowledge points.

[0056] 102. Perform face recognition processing on the first video to obtain the face orientation angle of each student's face region in each video frame of the first video. Perform head recognition processing on the second video to determine the first position of each student's head in the student desk and chair area at each time point of the second video, and the second position of the teacher's head in the lecture area at each time point of the second video.

[0057] 103. Determine the movement path of the teacher's head in the first coordinate system based on the second position; determine the line of sight path of each student in the first coordinate system based on the face orientation angle, the first position, and the second position;

[0058] 104. Monitor the attention of each student based on the similarity between the teacher's movement path and the students' line of sight paths.

[0059] In this embodiment of the application, the above-mentioned classroom attention monitoring method can be implemented by a classroom attention monitoring device, which is a program module located in the control device. The control device can call the program module to implement the above-mentioned classroom attention monitoring method.

[0060] In this context, "classroom" typically refers to a period of 45 minutes during a lesson. To better monitor classroom attention, class time can be divided into teaching and non-teaching modes. The teaching mode involves the teacher lecturing on knowledge points. In this mode, students are expected to listen attentively and keep up with the teacher's pace. Students' gaze follows the teacher's movements. The inventors have discovered that listeners unconsciously fix their gaze on the lecturer when focusing on their content, which helps them better understand the lecture. Therefore, in the teaching mode, students' gaze follows the teacher's movements. The higher the similarity between the student's gaze path and the teacher's path, the more focused the student's attention. This application uses this principle to monitor student attention in the classroom. Furthermore, in the teaching mode, the teacher needs to be located within the lecture area for better monitoring of student attention. In addition, non-teaching mode refers to the mode in which the teacher does not need to teach knowledge points. For example, the teacher arranges for students to preview knowledge, do practice questions, take exams, etc. In these situations, students are usually in a daze. Other methods can be used to monitor students' attention in non-teaching mode, which will not be elaborated here.

[0061] It should be noted that in practical applications, the instructor can manually switch modes according to the class schedule. For example, after class begins, the instructor can activate the teaching mode by operating the control device, such as through physical or virtual buttons. Upon detecting this activation, the control device responds and monitors student attention using the method described in this embodiment. If the teaching mode has reached the duration of a full class period, the control device can automatically terminate the teaching mode. Alternatively, if the teaching mode has not reached the full duration, and a mode switching operation is detected, the device can switch from teaching mode to non-teaching mode in response to the operation. In practical applications, the activation and deactivation of the teaching mode can be configured as needed and are not limited here.

[0062] In this embodiment of the application, if the current classroom environment is detected to have entered the teaching mode, a first video of the classroom captured by the first camera and a second video captured by the second camera can be obtained. The start time of the first video is the same as the start time of the second video, and the end time of the first video is the same as the end time of the second video, so that the first video and the second video are videos of the same time period. Furthermore, in order to ensure that the first video and the second video are captured in the same time period, the time of the first camera and the second camera can be synchronized after each entry into the teaching mode.

[0063] After acquiring the first and second videos, facial recognition processing is performed on the first video to obtain the facial orientation angles of each student's face in each video frame of the first video. Head recognition processing is performed on the second video to determine the first position of each student's head in the student desk and chair area corresponding to the time point of each video frame of the second video, and the second position of the teacher's head in the lecture area corresponding to the time point of each video frame of the second video. The first video is mainly used to determine the facial orientation angles of the students in class, while the second video is used to locate the positions of the students and the teacher in the classroom. This position refers to the position in the first coordinate system. Both positions are located using the student's head and the teacher's head as the reference. In the teaching mode, the student's gaze follows the teacher and is usually looking at the teacher's head. Therefore, by locating the head, the teacher's movement path and the student's gaze path can be effectively determined.

[0064] Furthermore, both the first and second positions mentioned above include x and y values ​​in a first coordinate system. This first coordinate system is a three-dimensional coordinate system of the three-dimensional space formed by the classroom. In one feasible implementation, the three-dimensional coordinate system of the three-dimensional space formed by the classroom can be set with the lower left corner of the classroom as the origin, the direction from the lower left corner toward the ceiling as the Z-axis, the direction from the lower left corner toward the blackboard as the X-axis, and the direction perpendicular to the Z-axis and X-axis and passing through the classroom toward the outside of the classroom as the Y-axis. The first coordinate system mentioned later will all be based on this setting and will not be described in detail again.

[0065] It should be noted that each video frame has a corresponding time point. Since the first and second videos are from the same time period, the parameters obtained from the video frames at the same time point in both videos can be combined using the time point as a standard to obtain the required parameters. Furthermore, for the student, since the student's facial orientation angle can be obtained from the first video and the student's first position from the second video, in subsequent descriptions, when both facial orientation angle and first position are used, they refer to the facial orientation angle obtained from the video frames in the first video and the first position obtained from the video frames in the second video at the same time point, so that the parameters for the same student at that same time point can be obtained.

[0066] Please see Figure 2 Here is a flowchart of the method for determining the face orientation angle in this application, the method including:

[0067] 201. Traverse the first video, and use a face deep learning model to perform face recognition on the traversed target video frames to determine the face region images contained in the target video frames.

[0068] 202. Use the OpenCV cross-platform computer vision library to perform key point recognition on the face region images, and determine the first set of coordinate points of the key points of each face region image in the camera coordinate system. The key point set includes: the center point of the left eye, the center point of the right eye, the center point of the chin, and the center of the eyebrows.

[0069] 203. Perform coordinate transformation on the first set of key points of each face region image to obtain a second set of coordinates in the second coordinate system;

[0070] 204. Calculate the coordinates of the center point of the left eye, the center point of the right eye, the center point of the chin, and the center of the eyebrows of each face region in the second set to obtain the face orientation angle of each face region in the second coordinate system, wherein each direction of the second coordinate system is parallel to each direction of the first coordinate system.

[0071] In this embodiment of the application, it is necessary to determine the facial orientation angle of each student contained in each video frame of the first video. Specifically, each video frame in the first video is traversed in chronological order. For the target video frame that has been traversed, a deep learning model for face recognition is used to determine the face region images contained in the target video frame. For each face region image that has been determined, OpenCV (a cross-platform computer vision library) is used to perform key point recognition on each face region image to determine the first set of coordinate points of the key points of each face region image in the camera coordinate system. The key point set includes: the center point of the left eye, the center point of the right eye, the center point of the chin, and the center of the eyebrows.

[0072] OpenCV is a cross-platform computer vision library released under the BSD (Berkly Software Distribution, BSD License) license. It can run on various operating systems and can effectively achieve key point recognition.

[0073] Understandably, when performing facial recognition on the first video, the recognized facial features can be compared with the preset facial features of each student to determine the correspondence between each facial region image and the student, so that the facial orientation angle of each student in each video frame of the first video can be obtained.

[0074] Furthermore, considering that the change in the orientation angle of a face is relatively small in a short period of time, in order to reduce the amount of data processed and the occupation of system resources, face recognition can be performed first using a deep learning model to obtain a set of face region images of each student in the first video. This set contains the correspondence between the time point of the video frame and the face region image. If a certain video frame does not detect the face region image of student B, then the set of face region images of student B does not contain the correspondence between that video frame and the face region image. After obtaining the set of face region images of each student, image sampling is performed according to a preset step size, for example, sampling is performed at a step size of 3 seconds to obtain a set of sampled face region images. The method in this application is then completed using the set of sampled face region images to achieve the purpose of reducing the amount of data processed, reducing the occupation of system resources, and improving monitoring efficiency.

[0075] In this embodiment, by traversing the first video, a set of facial region images of each student in the first video can be obtained. Furthermore, key point recognition can be performed on each facial region image based on OpenCV to effectively obtain a first set of key points of each student in each video frame. It should be noted that the coordinates of the key points in the first set are determined based on the camera coordinate system. The origin of the camera coordinate system is the optical center of the first camera, the x-axis and y-axis are parallel to the x-axis and y-axis of the video frame, the z-axis is the optical axis of the first camera, which is perpendicular to the plane of the video frame, and the intersection of the optical axis and the image plane is the aforementioned origin, thus forming a rectangular coordinate system for the camera coordinate system.

[0076] Furthermore, to better determine the facial orientation angle of each student, the key points of the facial region images of each student in each video frame will be transformed to obtain a second set of coordinates for each facial region image in a second coordinate system. This second coordinate system is a three-dimensional coordinate system with the human body as the origin.

[0077] The second coordinate system can be set with the center point of the left eye in the face region image as the origin, with the x-axis parallel to the x-axis of the first coordinate system, the y-axis parallel to the y-axis of the first coordinate system, and the z-axis parallel to the z-axis of the first coordinate system. That is, all directions of the second coordinate system are parallel to all directions of the first coordinate system. It should be noted that this setting method allows the second coordinate system to be transformed into the first coordinate system, so that the movement path of the teacher and the line of sight of each student can be determined in the first coordinate system, so as to effectively monitor the students' attention.

[0078] It should be noted that in the embodiments of this application, the coordinates of the key points in the camera coordinate system are transformed to the coordinates in the second coordinate system. The transformation method is existing technology and will not be described in detail here.

[0079] In this embodiment of the invention, after transforming the first set of key points of each face region image to obtain the second set of coordinates in the second coordinate system, the coordinates of the left eye center point, right eye center point, chin center point and brow center of each face region image in the second set can be used to calculate the face orientation angle of each face region in the second coordinate system.

[0080] In this embodiment of the application, the facial orientation angle of each facial region in the second coordinate system can be calculated by using the following methods to calculate the coordinates of the center points of the left eye, right eye, chin, and eyebrows of each facial region in the second set. Specifically, this includes:

[0081] Step A: Determine the first line segment connecting the coordinates of the center point of the left eye and the center point of the right eye, and the second line segment connecting the coordinates of the center point of the chin and the center of the eyebrows;

[0082] Step B: Determine a first plane that is perpendicular to the first line segment and intersects the midpoint of the first line segment, and determine a second plane that is perpendicular to the second line segment and includes the midpoint of the first line segment.

[0083] Step C: The three-dimensional vector formed by the lines intersecting the first plane and the second plane, starting from the midpoint of the first line segment and pointing outwards from the face, is taken as the face orientation angle.

[0084] In this embodiment, after obtaining the second set, the second set contains subsets of key points for each face region in each video frame. One subset contains the coordinates of key points for a student's face image region in a video frame, and this subset includes the coordinates of the student's left eye center point, right eye center point, chin center point, and forehead center point. To obtain the student's facial orientation angle in the video frame, a first line segment can be formed by combining the coordinates of the left and right eye center points in the subset, and a second line segment can be formed by combining the coordinates of the chin center point and forehead center point in the subset. The two line segments are: the first line segment uses the center points of the left and right eyes to determine the angle of the student's head rotation, and the midpoint of the left and right eye centers (i.e., the midpoint of the first line segment) is used to determine the starting point of the face orientation angle for a more accurate result. The second line segment uses the center point of the chin and the center of the eyebrows. These two points can be used to determine whether the student has raised their chin, which can effectively determine the angle at which the student is looking up. Since the angle at which the student looks up is consistent with the angle at which the student's gaze is upward, the second line segment can be used to determine the angle at which the gaze is upward. Specifically, a first plane is determined that is perpendicular to the first line segment and intersects the midpoint of the first line segment. The angle of this first plane represents the student's head rotation angle. A second plane is determined that is perpendicular to the second line segment and includes the midpoint of the first line segment. The angle formed by this second plane represents the student's head tilt angle. In this way, the line where the first and second planes intersect represents the direction of the gaze. Specifically, the midpoint of the first line segment can be taken as the starting point of the face orientation angle, and this starting point is on the line where the first and second planes intersect. The three-dimensional vector formed by the line where the first and second planes intersect, pointing outwards from the face, is taken as the face orientation angle.

[0085] In this embodiment of the invention, by identifying the center point of the left eye, the center point of the right eye, the center point of the chin, and the center of the eyebrows in the student's face region image, it is possible to determine the student's head rotation angle based on the center point of the left eye and the center point of the right eye, and to determine the student's upward gaze angle based on the center point of the chin and the center of the eyebrows. Furthermore, the student's face orientation angle can be effectively determined based on the formed first plane and second plane, thereby improving the accuracy of the face orientation angle determination.

[0086] In this embodiment of the invention, after performing face recognition processing on the first video to obtain the face orientation angle of each student in each video frame, the second video also needs to be processed. Since the second video is captured by a second camera located on the ceiling of the classroom and capable of capturing the entire classroom from a top-down angle, head recognition processing will be performed on the second video. The main function of the second video is to locate the positions of students and teachers in the first coordinate system. After head recognition processing, the head located in the student desk and chair area is taken as the student's head, and the first position of each student's head in each video frame of the second video is determined. The head located in the lecture area is identified as the teacher's head, and the second position of the teacher's head corresponding to the time point in each video frame of the second video is determined. Furthermore, the movement path of the teacher's head in the first coordinate system is determined based on the second position, and the line of sight of the students in the first coordinate system is determined based on the face orientation angle, the first position, and the second position.

[0087] In the embodiments of this application, the first position and the second position are both the pixel positions of the student's and teacher's head in the video frame converted to the positions in the XY axis of the first coordinate system.

[0088] To obtain the first and second positions, the system will acquire the preset length and width information of the classroom, as well as the length information of the lecture area and the length information of the student desk and chair area in the actual scene. First, a head recognition model based on deep learning algorithm will be used to perform head recognition processing on each video frame of the second video to determine the head image region of each head in each video frame of the second video. Based on the ratio formed by the length information of the lecture area and the length information of the desk and chair area, each video frame will be divided into regions. The head located in the lecture area will be identified as the head of the lecturer, and the head located in the student desk and chair area will be identified as the head of the student. Furthermore, the second position of the lecturer's head and the first position of each student's head will be determined.

[0089] The second position of the instructor's head can be determined as follows:

[0090] First, the image of the teacher's head region in each video frame of the second video is determined. The smallest pixel in the X-axis direction of this head region image in the image coordinate system is determined as the calibration point of the teacher's head. It can be understood that, in the teaching mode, the teacher usually faces the students while lecturing, and the students look at the teacher's head. Therefore, the pixel with the smallest x-value is closest to the student's focus when looking at the teacher. That is, using this pixel can improve the accuracy of determining the teacher's movement path. The image coordinate system refers to the coordinate system formed by taking the lower left corner of the classroom in the captured video frame as the origin, the length direction of the classroom as the X-axis, and the width direction of the classroom as the Y-axis.

[0091] Secondly, after obtaining the teacher's pixel A (i.e., the pixel with the minimum x-value mentioned above) in each video frame of the second video, the distance between the two endpoints parallel to the X-axis and containing pixel A will be determined in each video frame. One endpoint, B, is located on the wall in front of the classroom in the video frame, and the other endpoint, C, is located on the wall behind the classroom in the video frame. The first length of the line segment formed by B and C, and the second length of the line segment formed by pixel A and C will be determined. The ratio formed by the second length and the first length will be determined, and this ratio will be multiplied by the actual length of the classroom in the first coordinate system to obtain the teacher's x-value in the first coordinate system. To achieve the goal of transforming the teacher's position from the X-axis in the video frame to the X-axis position in the classroom, similarly, the distance between the other two endpoints parallel to the Y-axis and including pixel A can be determined. Point E is located on the left wall of the classroom in the video frame, and point F is located on the right wall of the classroom in the video frame. The third length of the line segment formed by EF and the fourth length of the line segment formed by point E and pixel A are determined. The ratio of the fourth length to the third length is multiplied by the actual width of the classroom in the first coordinate system to obtain the y-value of the teacher in the first coordinate system. The obtained x-value and y-value constitute the second position of the teacher's head in the first coordinate system.

[0092] Furthermore, the method for determining the first position of each student's head may include:

[0093] First, taking the determination of the first position of a single student's head as an example, we can determine the image of the student's head region in each video frame of the second video. We then determine the largest pixel in the X-axis direction of this head region image as the calibration point of the student's head. Using the pixel G corresponding to this calibration point, we determine the student's first position, i.e., the distance between two endpoints parallel to the X-axis and containing pixel G. One endpoint, K, is located on the wall in front of the classroom in the video frame, and the other, L, is located on the wall behind the classroom in the video frame. We determine the fifth length of the line segment formed by K and L, and the sixth length of the line segment formed by pixel G and L. We then determine the ratio between the sixth and fifth lengths and multiply this ratio by... The actual length of the classroom in the first coordinate system is used to obtain the student's x-value in the first coordinate system, thus realizing the purpose of transforming the student's position from the X-axis position in the video frame to the X-axis position in the classroom. Similarly, the distance between the other two endpoints parallel to the Y-axis and including pixel point G can be determined. A point M is located on the left wall of the classroom in the video frame, and a point N is located on the right wall of the classroom in the video frame. The seventh length of the line segment formed by MN and the eighth length of the line segment formed by point M and pixel point G are determined. The ratio of the eighth length to the seventh length is multiplied by the actual width of the classroom in the first coordinate system to obtain the student's y-value in the first coordinate system. The obtained x-value and y-value constitute the first position of the student's head in the first coordinate system.

[0094] In this embodiment of the invention, after obtaining the face orientation angle and first position of each student and the second position of the teacher through the method described above, the movement path of the teacher's head and the line of sight of the students will be further obtained.

[0095] Please see Figure 3 This is a flowchart illustrating a method for determining the movement path of the teacher's head in a first coordinate system based on a second position, according to an embodiment of the present invention. The method includes:

[0096] Step 301: Obtain the height information of the instructor and the height information of the lecture area in the classroom;

[0097] Step 302: Based on the height information of the instructor, the height information of the lecture area, and the second position, determine the head position information of the instructor in the first coordinate system at each time point;

[0098] Step 303: Use the head position information at each time point in the first coordinate system to form the movement path of the instructor.

[0099] In this embodiment of the application, the height information of the teacher, the height information of the teaching area of ​​the classroom, and the length and width information of the classroom are first obtained. The system has the height information of the teacher, the height information of the teaching area of ​​the classroom, and the length and width information of the classroom preset in advance. The name of the teacher in the current classroom can be determined by facial recognition, and the corresponding height information can be obtained by searching the preset name and height information correspondence.

[0100] After obtaining the height information of the lecturer and the height information of the classroom's lecture area, the head position information of the lecturer in the first coordinate system at each time point can be determined based on the lecturer's height information, the height information of the lecture area, and the second position.

[0101] In this embodiment, the second position already includes the x and y values ​​of the teacher's head. Therefore, using the teacher's height information and the height information of the classroom's lecturing area, it is necessary to determine the z value of the teacher's head in the first coordinate system to determine the teacher's head position information. Specifically, the teacher's height information and the height information of the lecturing area can be used as the z value of the teacher's head position information. For example, if the teacher's height is 1.6m and the height information of the lecturing area is 0, then the z value of the teacher's head position information is determined to be 1.6m. Through the above method, the teacher's head position information at each time point can be effectively obtained.

[0102] In this embodiment of the invention, after obtaining the head position information of the instructor at each time point in the first coordinate system, the instructor's movement path is formed using this head position information. Specifically, since each time point corresponds to a video frame, the head position information of multiple video frames per second can be deduplicated. The average value of the remaining head position information after deduplication is calculated, and the average head position information is used as the head position information for that single second. In this way, the instructor's head position information in seconds is obtained, and the instructor's movement path is obtained by connecting the head position information in chronological order.

[0103] In this embodiment of the application, the line of sight of each student in the first coordinate system will also be determined based on the face orientation angle of each student, the first position and the second position of each student, specifically including the following steps:

[0104] Step D: Obtain the height of each student and the height of the desk and chair, and determine the height of each student's eye area when they are sitting down based on the height of each student and the height of the desk and chair.

[0105] Step E: Determine the coordinates of each student's eye position in the first coordinate system based on the height of each student's eye position and the first position;

[0106] Step F: Based on the coordinates of each student's eye location in the first coordinate system, the face orientation angle, and the second position, determine the coordinates of each student's line of sight in the first coordinate system at each time point;

[0107] Step G: Determine the line-of-sight path of each student using the coordinates of the line-of-sight points in the first coordinate system at each time point.

[0108] In this embodiment of the application, the height of each student and the height of the desks and chairs used are preset in the system, so that the height of the student sitting in front of the desk can be determined according to the student's height and the height of the desk and chairs according to the preset algorithm. Since the difference in the height of the eyes of different students from the top of their heads is small, the height of the student's eyes when sitting can be obtained by statistically averaging. Specifically, the height of the student sitting in front of the desk can be subtracted from the above average value, and the resulting value is the height of the student's eyes when sitting.

[0109] Furthermore, based on the height of each student's eye area and their initial position, the coordinates of each student's eye area in the first coordinate system are determined. The initial position already includes the student's x and y values ​​in the first coordinate system. The height of the student's eye area can be used as their z value in the first coordinate system to obtain the student's eye coordinates. After obtaining these coordinates, the gaze vector of each student at each time point in the first coordinate system can be determined based on the coordinates of their eye area and their face orientation angle. Specifically, combining these coordinates with the face orientation angle yields the student's gaze vector in the first coordinate system. Since the face orientation angle uses a second coordinate system, and the directions of its axes are the same and parallel to those of the first coordinate system, the components of the face orientation angle on the X, Y, and Z axes can be added to the x, y, and z values ​​of the student's eye area coordinates to obtain the aforementioned gaze vector.

[0110] After obtaining the aforementioned gaze vectors, the student's gaze vector at the same time point corresponding to the same video frame, along with the teacher's second position, will be used to determine the coordinates of the student's gaze point in the first coordinate system at that time point. These coordinates represent the student's gaze path. Specifically, a target plane parallel to the YZ axes and containing the second position can be defined in the first coordinate system. The intersection of the student's gaze vector and the target plane will be used as the coordinates of the student's gaze point. Furthermore, connecting the coordinates of each student's gaze points at each time point will yield the gaze path for each student.

[0111] Based on this, after obtaining the face orientation angle, first position, and second position, the movement path of the teacher's head in the first coordinate system can be determined according to the second position. Based on the face orientation angle, first position, and second position, the gaze path of each student in the first coordinate system can be determined. The similarity between the teacher's movement path and each student's gaze path can be calculated to obtain the similarity between the teacher's movement path and each student's gaze path. This similarity is used to determine each student's attention. Specifically, a preset threshold for similarity can be set, and the similarity between the teacher's movement path and each student's gaze path can be calculated to obtain the similarity of each student's gaze following. If the target similarity value is greater than or equal to the preset threshold, the student corresponding to the target similarity is determined to be focused. This target similarity is the similarity of any one of the aforementioned students. If the target similarity value is greater than the preset threshold, the student corresponding to the target similarity is determined to be unfocused. For example, if the similarity between the teacher's movement path and the student A's line of sight is greater than or equal to the preset threshold, then student A is determined to be in a state of focused attention. If the similarity between the teacher's movement path and the student A's line of sight is less than the preset threshold, then student A is determined to be in a state of inattentiveness, thus enabling effective monitoring of students' attention.

[0112] In this embodiment of the application, by using a first video captured by a first camera capable of capturing faces and a second video captured by a second camera capable of capturing the entire classroom in teaching mode, the movement path of the teacher's head and the video paths of each student's gaze are obtained. Classroom attention is monitored by the similarity between the teacher's movement path and the students' gaze paths. Since students need to pay attention to the teacher's lecture content in teaching mode and follow the teacher's movements with their eyes to better listen to and understand the lecture content, the similarity between the teacher's movement path and the students' video paths can be used to monitor the students' attention, effectively improving the accuracy of student attention monitoring.

[0113] Please see Figure 4 The diagram below shows the structural structure of a program module for a classroom attention monitoring device according to an embodiment of this application. The device includes:

[0114] The acquisition module 401 is used to acquire a first video of the classroom captured by a first camera and a second video of the same time period as the first video captured by a second camera if the current classroom environment is detected to have entered the teaching mode. The first camera is set in front of the student's seat in the classroom, and the second camera is located above the student's seat. The teaching mode refers to the classroom learning mode in which the teacher teaches knowledge points.

[0115] The first determining module 402 is used to perform face recognition processing on the first video to obtain the face orientation angle of each student's face region in each video frame of the first video, and to perform head recognition processing on the second video to determine the first position of each student's head located in the student desk and chair area at the time point of each video frame of the second video, and the second position of the head of the lecturer located in the lecture area at the time point of each video frame of the second video. The first position and the second position both include x and y values ​​in a first coordinate system, which is a three-dimensional coordinate system based on the three-dimensional space formed by the classroom.

[0116] The second determining module 403 is used to determine the movement path of the teacher's head in the first coordinate system based on the second position; and to determine the line of sight path of each student in the first coordinate system based on the face orientation angle, the first position, and the second position.

[0117] The monitoring module 404 is used to monitor the attention of each student based on the similarity between the teacher's movement path and the students' line of sight paths.

[0118] It should be noted that the specific implementation methods of each module in the device of this application can refer to the content of each step in the aforementioned classroom attention monitoring method, and will not be repeated here.

[0119] In this embodiment of the application, by using a first video captured by a first camera capable of capturing faces and a second video captured by a second camera capable of capturing the entire classroom in teaching mode, the movement path of the teacher's head and the video paths of each student's gaze are obtained. Classroom attention is monitored by the similarity between the teacher's movement path and the students' gaze paths. Since students need to pay attention to the teacher's lecture content in teaching mode and follow the teacher's movements with their eyes to better listen to and understand the lecture content, the similarity between the teacher's movement path and the students' video paths can be used to monitor the students' attention, effectively improving the accuracy of student attention monitoring.

[0120] Figure 5 An internal structural diagram of a computer device in one embodiment is shown. This computer device can specifically be a terminal or a server. Figure 5As shown, the computer device includes a processor, memory, and network interface connected via a system bus. The memory includes a non-volatile storage medium and internal memory. The non-volatile storage medium stores an operating system and may also store a computer program. When executed by the processor, this computer program causes the processor to perform the steps in the above-described method embodiments. The internal memory may also store a computer program, which, when executed by the processor, causes the processor to perform the steps in the above-described method embodiments. Those skilled in the art will understand that... Figure 5 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device to which the present application is applied. Specific computer devices may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.

[0121] In one embodiment, a computer device is provided, including a memory and a processor, the memory storing a computer program that, when executed by the processor, causes the processor to perform the following steps:

[0122] If the current classroom environment is detected to have entered teaching mode, the system acquires a first video of the classroom captured by the first camera and a second video of the same time period captured by the second camera. The first camera is positioned in front of the student's seat in the classroom, and the second camera is positioned above the student's seat. The teaching mode refers to a classroom learning mode in which the teacher teaches knowledge points.

[0123] The first video is subjected to face recognition processing to obtain the face orientation angle of each student's face region in each video frame of the first video. The second video is subjected to head recognition processing to determine the first position of each student's head in the student desk and chair area at each time point of the second video frame, and the second position of the teacher's head in the lecture area at each time point of the second video frame. The first position and the second position both include the x value and y value in the first coordinate system, which is a three-dimensional coordinate system based on the three-dimensional space formed by the classroom.

[0124] The teacher's head movement path in the first coordinate system is determined based on the second position; the line of sight of each student in the first coordinate system is determined based on the face orientation angle, the first position, and the second position.

[0125] The attention of each student is monitored based on the similarity between the teacher's movement path and the student's gaze path.

[0126] It should be noted that the specific implementation methods of the above steps can be found in the descriptions of each step in the aforementioned classroom attention monitoring method, and will not be repeated here.

[0127] In one embodiment, a computer-readable storage medium is provided storing a computer program that, when executed by a processor, causes the processor to perform the following steps:

[0128] If the current classroom environment is detected to have entered teaching mode, the system acquires a first video of the classroom captured by the first camera and a second video of the same time period captured by the second camera. The first camera is positioned in front of the student's seat in the classroom, and the second camera is positioned above the student's seat. The teaching mode refers to a classroom learning mode in which the teacher teaches knowledge points.

[0129] The first video is subjected to face recognition processing to obtain the face orientation angle of each student's face region in each video frame of the first video. The second video is subjected to head recognition processing to determine the first position of each student's head in the student desk and chair area at each time point of the second video frame, and the second position of the teacher's head in the lecture area at each time point of the second video frame. The first position and the second position both include the x value and y value in the first coordinate system, which is a three-dimensional coordinate system based on the three-dimensional space formed by the classroom.

[0130] The teacher's head movement path in the first coordinate system is determined based on the second position; the line of sight of each student in the first coordinate system is determined based on the face orientation angle, the first position, and the second position.

[0131] The attention of each student is monitored based on the similarity between the teacher's movement path and the student's gaze path.

[0132] It should be noted that the specific implementation methods of the above steps can be found in the descriptions of each step in the aforementioned classroom attention monitoring method, and will not be repeated here.

[0133] Those skilled in the art will understand that all or part of the processes in the above embodiments can be implemented by a computer program instructing related hardware. The program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments described above. Any references to memory, storage, databases, or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), dual data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), RAMbus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and RAMbus dynamic RAM (RDRAM), etc.

[0134] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0135] The embodiments described above are merely illustrative of several implementation methods of this application, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of this patent application. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this patent application should be determined by the appended claims.

Claims

1. A method for monitoring classroom attention, characterized in that, The method includes: If the current classroom environment is detected to have entered teaching mode, the first video of the classroom captured by the first camera and the second video captured by the second camera at the same time period as the first video are obtained. The first camera is set in front of the student's seat in the classroom, and the second camera is located above the student's seat. The teaching mode refers to the classroom learning mode in which the teacher teaches knowledge points, and in the teaching mode, the teacher is located in the teaching area. The first video is subjected to face recognition processing to obtain the face orientation angle of each student's face region in each video frame of the first video. The second video is subjected to head recognition processing to determine the first position of each student's head in the student desk and chair area at each time point of the second video frame, and the second position of the teacher's head in the lecture area at each time point of the second video frame. The first position and the second position both include the x value and y value in the first coordinate system, which is a three-dimensional coordinate system based on the three-dimensional space formed by the classroom. The teacher's head movement path in the first coordinate system is determined based on the second position; the line of sight of each student in the first coordinate system is determined based on the face orientation angle, the first position, and the second position. The attention of each student is monitored based on the similarity between the teacher's movement path and the student's line of sight. Determining the line of sight path of each student in the first coordinate system based on the face orientation angle, the first position, and the second position includes: Obtain the height of each student and the height of the desk and chair, and determine the height of each student's eye area when they are sitting down based on the height of each student and the height of the desk and chair. Based on the height of each student's eye and the first position, determine the coordinates of each student's eye at each time point in the first coordinate system; Based on the coordinates of each student's eye area in the first coordinate system at each time point, the face orientation angle, and the second position, determine the coordinates of each student's line of sight in the first coordinate system at each time point; The line-of-sight paths of each student are determined using the coordinates of their line-of-sight points at each time point in the first coordinate system. The step of determining the coordinates of each student's line of sight in the first coordinate system at each time point based on the coordinates of each student's eye location in the first coordinate system, the face orientation angle, and the second position includes: Based on the coordinates of the students' eye parts at the same time point, the face orientation angle determines the gaze vector of each student in the first coordinate system; Determine a target plane in the first coordinate system that is parallel to the YZ axes of the first coordinate system and includes the second position; The intersection of the line-of-sight vector and the target plane is taken as the coordinates of the student's line-of-sight point in the first coordinate system.

2. The method according to claim 1, characterized in that, Determining the movement path of the teacher's head in the first coordinate system based on the second position includes: Obtain the height information of the instructor and the height information of the lecture area in the classroom; Based on the height information of the instructor, the height information of the lecture area, and the second position, the head position information of the instructor in the first coordinate system at each time point is determined; The teacher's movement path is formed by using the head position information at each time point in the first coordinate system.

3. The method according to claim 1, characterized in that, The step of performing facial recognition processing on the first video to obtain the facial orientation angle of each student's face region in each frame of the first video includes: The first video is traversed, and a face deep learning model is used to perform face recognition on the target video frames that are traversed, and the face region images contained in the target video frames are extracted. The OpenCV cross-platform computer vision library is used to perform key point recognition on the face region images to determine the first set of coordinate points of the key points of each face region image in the camera coordinate system. The key point set includes: the center point of the left eye, the center point of the right eye, the center point of the chin, and the center of the eyebrows. The first set of key points of each face region image is transformed by coordinate transformation to obtain a second set of coordinates in a second coordinate system, which is a three-dimensional coordinate system with the human body as the origin. The facial orientation angle of each facial region in the second set is calculated by taking the coordinates of the center point of the left eye, the center point of the right eye, the center point of the chin, and the center of the eyebrows. The directions of the second coordinate system are parallel to the directions of the first coordinate system.

4. The method according to claim 3, characterized in that, The calculation of the coordinates of the center points of the left and right eyes, the center point of the chin, and the center of the eyebrows of each face region in the second set yields the face orientation angle of each face region in the second coordinate system, including: Determine the coordinates of the center point of the left eye and the center point of the right eye using a first line segment, and determine the coordinates of the center point of the chin and the center of the eyebrows using a second line segment; Determine a first plane that is perpendicular to the first line segment and intersects the midpoint of the first line segment; determine a second plane that is perpendicular to the second line segment and includes the midpoint of the first line segment. The three-dimensional vector formed by the lines intersecting the first plane and the second plane, starting from the midpoint of the first line segment and pointing outwards from the face, is taken as the face orientation angle.

5. The classroom attention monitoring method according to claim 1, characterized in that, The monitoring of each student's attention based on the similarity between the teacher's movement path and the students' gaze paths includes: Calculate the similarity between the movement path and the gaze path of each student to obtain the similarity value of each student's gaze tracking; If the target similarity value is greater than or equal to a preset threshold, then the student's attention concentration corresponding to the target similarity is determined. If the target similarity is less than a preset threshold, then it is determined that the student corresponding to the target similarity value is not focused.

6. A classroom attention monitoring device, characterized in that, The device includes: The acquisition module is used to acquire a first video of the classroom captured by a first camera and a second video of the same time period as the first video captured by a second camera if the current classroom environment is detected to have entered the teaching mode. The first camera is set in front of the student's seat in the classroom, and the second camera is located above the student's seat. The teaching mode refers to a classroom learning mode in which the teacher teaches knowledge points, and in the teaching mode, the teacher is located in the teaching area. The first determining module is used to perform face recognition processing on the first video to obtain the face orientation angle of each student's face region in each video frame of the first video, and to perform head recognition processing on the second video to determine the first position of each student's head located in the student desk and chair area at each time point of the second video, and the second position of the head of the lecturer located in the lecture area at each time point of the second video. The first position and the second position both include x and y values ​​in a first coordinate system, which is a three-dimensional coordinate system based on the three-dimensional space formed by the classroom. The second determining module is used to determine the movement path of the teacher's head in the first coordinate system based on the second position; and to determine the line of sight path of each student in the first coordinate system based on the face orientation angle, the first position, and the second position. The monitoring module is used to monitor the attention of each student based on the similarity between the teacher's movement path and the students' line of sight paths; Determining the line of sight path of each student in the first coordinate system based on the face orientation angle, the first position, and the second position includes: Obtain the height of each student and the height of the desk and chair, and determine the height of each student's eye area when they are sitting down based on the height of each student and the height of the desk and chair. Based on the height of each student's eye and the first position, determine the coordinates of each student's eye at each time point in the first coordinate system; Based on the coordinates of each student's eye area in the first coordinate system at each time point, the face orientation angle, and the second position, determine the coordinates of each student's line of sight in the first coordinate system at each time point; The line-of-sight paths of each student are determined using the coordinates of their line-of-sight points at each time point in the first coordinate system. The step of determining the coordinates of each student's line of sight in the first coordinate system at each time point based on the coordinates of each student's eye location in the first coordinate system, the face orientation angle, and the second position includes: Based on the coordinates of the students' eye parts at the same time point, the face orientation angle determines the gaze vector of each student in the first coordinate system; Determine a target plane in the first coordinate system that is parallel to the YZ axes of the first coordinate system and includes the second position; The intersection of the line-of-sight vector and the target plane is taken as the coordinates of the student's line-of-sight point in the first coordinate system.

7. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by a processor, the processor performs the steps of the method as described in any one of claims 1 to 6.

8. A computer device, comprising a memory and a processor, characterized in that, The memory stores a computer program that, when executed by the processor, causes the processor to perform the steps of the method as described in any one of claims 1 to 6.

Citation Information

Patent Citations

  • Method for evaluating class concentration degree of hearing-impaired students

    CN118430008A