Desktop learning concentration degree monitoring method and system

By collecting and analyzing learners' real-time facial and desktop images, combined with gaze tracing and item recognition technology, the problem that existing technology cannot monitor students' concentration on desktop items is solved, effectively monitoring and recording of learners' attention areas and items, and improving learning efficiency.

CN120029447APending Publication Date: 2025-05-23张芷铉
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411902416.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-23
Publication Date
2025-05-23

AI Technical Summary

Technical Problem

The existing technology is difficult to effectively monitor students' concentration on desktop items during the learning process, and it is impossible to identify and record the areas of attention and items of attention for students.

Method used

By collecting real-time facial images and desktop images of learners, using gaze tracing technology to obtain the human eye position and line of sight angle angle, combining desktop item recognition technology, calculate the overlap between the gaze angle area and the position of the item, judge the items that learners are concerned about, and count the time and type of items.

Benefits of technology

It realizes monitoring and recording of areas and items of attention on the desktop where the learner's sight is on, which can remind learners to maintain concentration and improve learning efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120029447A_ABST
    Figure CN120029447A_ABST
Patent Text Reader

Abstract

The invention provides a desktop learning concentration degree monitoring method and system, and the method comprises the steps: collecting a real-time face image of a learner and a desktop real-time image, obtaining the position and sight angle of the eyes of the learner in the image based on the real-time face image of the learner, determining the type and position of a desktop object based on the desktop real-time image, and carrying out the monitoring of the desktop learning concentration degree. The method comprises the following steps: calculating the overlap ratio of a sight angle area of a learner on a desktop and desktop article positions, judging articles concerned by the learner, counting the types and time of the articles concerned by the learner, regularly recording face images of the learner and images and data of the articles concerned by the learner, and generating traceable record data; according to the invention, the sight tracking data and the intelligent image object recognition data are fused to obtain the desktop attention area and object information of the sight of the learner, the time when the sight of the learner stays on the desktop and the classification and time data of objects concerned on the desktop can be monitored and recorded, and a signal is generated according to the data to remind the learner to keep concentration.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer vision technology, and in particular to a desktop learning concentration monitoring method and system. Background Art

[0002] In traditional learning, students mostly study at home on their own. The students’ concentration determines their learning situation and the consolidation of classroom knowledge to a certain extent, and most of it depends on the students’ self-consciousness. Therefore, students are prone to fatigue or lack of concentration, which will greatly affect the quality and effect of self-study. Therefore, the method of detecting students’ concentration is very important for self-study.

[0003] There are already some technologies that monitor students' attention, such as online learning concentration monitoring and analysis systems that use eye tracking technology, learning process recording systems, and smart desk lamps or desks that use eye tracking or human posture recognition.

[0004] However, the online learning concentration monitoring and analysis system based on eye tracking technology detects the learner's focus on the content displayed on the computer screen, not the learner's focus on the items on the desktop, and does not identify the content on the desktop. In addition, the existing learning process recording system only records the learner's facial image and desktop image during the learning process, and does not perform eye tracking or determine the learner's eye focus area and focus items. At the same time, the existing smart table lamps or desks with eye tracking or body posture recognition are designed to adjust lighting or remind sitting posture, and do not determine the learner's eye focus area and focus items. Summary of the invention

[0005] 1. Technical issues to be resolved

[0006] In view of the deficiencies in the prior art, the present invention provides a desktop learning concentration monitoring method and system to solve the problems raised in the above background technology.

[0007] (II) Technical solution

[0008] To achieve the above objectives, the present invention is implemented through the following technical solutions: A desktop learning concentration monitoring method comprises the following steps:

[0009] Collect learners' real-time facial images and desktop images;

[0010] Based on the learner's real-time facial image, the position and sight angle of the learner's eyes in the image are obtained, and based on the real-time desktop image, the type and position of desktop objects are determined;

[0011] Calculate the overlap between the learner's sight angle on the desktop and the location of the desktop objects to determine the objects that the learner is paying attention to;

[0012] Count the types of objects that learners pay attention to and the time they pay attention to them, and regularly record learners' facial images and the images and data of the objects they pay attention to, generating traceable record data;

[0013] When it is determined that the learner's attention to non-learning objects exceeds a threshold, or the time when the learner's eyes are away from the desk exceeds a threshold, an alarm is issued.

[0014] As a preferred embodiment of the present invention, the calculating the overlap between the learner's sight angle area on the desktop and the location of the desktop objects and determining the objects that the learner is paying attention to includes:

[0015] From the learner's real-time facial image, obtain the eye position data and sight angle in the image;

[0016] Among them, in the collected learner's face image, the eye tracking technology is used to obtain the human eye position data (x, y) and the eye angle in the image, and the eye angle includes the deflection angle α, the pitch angle β and the roll angle θ, and the x-axis, y-axis and z-axis are the coordinate system in the front image;

[0017] Periodically identify the type and location of objects in the desktop real-time image to determine the type of objects and their corresponding locations;

[0018] Fuse the gaze tracking data with the desktop object data to determine the objects on the desktop that the learner is focusing on;

[0019] Count the time data of learners' eye focus and record the images and data.

[0020] As a preferred embodiment of this embodiment, the fusing of the gaze tracking data and the desktop object data to determine the desktop object that the learner is paying attention to includes:

[0021] Calculate ROI_Y using the eye position data y, the sight pitch angle β, and the sight roll angle θ;

[0022] Physical height of human eye Z = Z0-y*K1*coa(A)

[0023] Physical pitch angle of human eye β′=β-A

[0024] ROI_Y can be obtained as follows: ROI_Y = Y0 - Z*tan(β′);

[0025] Calculate ROI_X using the eye position data x, sight deflection angle α, and sight roll angle θ:

[0026] First calculate the X-axis coordinate X0 of the intersection of the vertical line of the human eye center and the XOY plane;

[0027]

[0028] ROI_X=(Y0-ROI_Y)*tan(α)-X0

[0029] The sight rolling angle θ reflects the degree of head tilt of the learner and can be used to fine-tune ROI_X;

[0030] Based on the determined (ROI_X, ROI_Y), a square of a certain width is used to determine the focus area, the overlapping area S between the focus area and the object image is calculated, and a threshold of the overlapping area is set. When S>threshold, it is determined that the learner is focusing on the object.

[0031] As a preferred embodiment of this invention, the method of collecting time data of learners' sight attention and recording images and data includes:

[0032] The desktop items are classified into "learning" and "non-learning" categories. For example, "electronic digital products" such as mobile phones and "food and beverages" such as snacks are classified into the "non-learning" category; the line of sight out of the image range of the top camera is defined as "line of sight out of the desktop";

[0033] Count the time t1 when learners "look away from the desk" and the time t2 when they focus on "non-learning" items;

[0034] The learner's real-time facial image and desktop real-time image as well as the above statistical data are recorded at regular intervals to generate a traceable record.

[0035] As a preferred embodiment of this embodiment, when it is determined that the learner's attention time on non-learning objects exceeds a threshold, or the time when the learner's sight is out of the desk range exceeds a threshold, issuing an alarm includes:

[0036] Set the reminder threshold to determine whether the "eyes away from the desk" time t1 or the time of focusing on "non-learning" objects t2 exceeds the threshold in the recent cycle. If the threshold is exceeded, an alarm is triggered.

[0037] As a preferred embodiment of this invention, a desktop learning concentration monitoring system is used to implement the above-mentioned desktop learning concentration monitoring method, and is characterized by comprising:

[0038] An image acquisition unit, used to acquire real-time facial images and desktop images of learners;

[0039] A main controller, used for connecting to an image acquisition unit;

[0040] The main controller is also used to obtain the position and sight angle of the learner's eyes in the image based on the learner's real-time facial image, and determine the type and position of desktop objects based on the desktop real-time image;

[0041] Calculate the overlap between the learner's sight angle on the desktop and the location of the desktop objects to determine the objects that the learner is paying attention to;

[0042] Count the types of objects that learners pay attention to and the time they pay attention to them, and regularly record learners' facial images and the images and data of the objects they pay attention to, generating traceable record data;

[0043] An alarm device, used to warn learners through sound and light information, and controlled by the main controller;

[0044] The terminal is configured in the electronic device App. When the electronic device App is started, it accesses the main controller through the Internet to view the real-time image captured by the image acquisition unit, view the learning attention statistics, and receive reminder information.

[0045] As a preferred embodiment of the present invention, the image acquisition unit includes a top camera and a front camera. The top camera is used to capture images of desktop objects, and its pointing axis is on the same plane as the front camera. The front camera maintains a certain elevation angle to capture the learner's face, and its pointing axis is on the same plane as the top camera.

[0046] As a preferred embodiment of the present invention, the step of calculating the overlap between the learner's sight angle area on the desktop and the location of the desktop items to determine the items that the learner is paying attention to includes:

[0047] From the learner's real-time facial image, obtain the eye position data and sight angle in the image;

[0048] Among them, in the collected learner's face image, the eye tracking technology is used to obtain the human eye position data (x, y) and the eye angle in the image, and the eye angle includes the deflection angle α, the pitch angle β and the roll angle θ, and the x-axis, y-axis and z-axis are the coordinate system in the front image;

[0049] Periodically identify the type and location of objects in the desktop real-time image to determine the type of objects and their corresponding locations;

[0050] Fuse the gaze tracking data with the desktop object data to determine the objects on the desktop that the learner is focusing on;

[0051] Count the time data of learners' eye focus and record the images and data.

[0052] As a preferred embodiment of this embodiment, the fusing of the gaze tracking data and the desktop object data to determine the desktop object that the learner is paying attention to includes:

[0053] Calculate ROI_Y using the eye position data y, the sight pitch angle β, and the sight roll angle θ;

[0054] Physical height of human eye Z = Z0-y*K1*cos(A)

[0055] Physical pitch angle of human eye β′=β-A

[0056] ROI_Y can be obtained as follows: ROI_Y = TO-Z*tan(β′);

[0057] Calculate ROI_X using the eye position data x, sight deflection angle α, and sight roll angle θ:

[0058] First calculate the X-axis coordinate X0 of the intersection of the vertical line of the human eye center and the XOY plane;

[0059]

[0060] ROI_X=(Y0-ROI_Y)*tan(α)-X0

[0061] The sight rolling angle θ reflects the degree of head tilt of the learner and can be used to fine-tune ROI_X;

[0062] Based on the determined (ROI_X, ROI_Y), a square of a certain width is used to determine the focus area, the overlapping area S between the focus area and the object image is calculated, and a threshold of the overlapping area is set. When S>threshold, it is determined that the learner is focusing on the object.

[0063] As a preferred embodiment of this invention, the method of collecting time data of learners' sight attention and recording images and data includes:

[0064] The desktop items are classified into "learning" and "non-learning" categories. For example, "electronic digital products" such as mobile phones and "food and beverages" such as snacks are classified into the "non-learning" category; the line of sight out of the image range of the top camera is defined as "line of sight out of the desktop";

[0065] Count the time t1 when learners "look away from the desk" and the time t2 when they focus on "non-learning" items;

[0066] The learner's real-time facial image and desktop real-time image as well as the above statistical data are recorded regularly to generate a traceable record.

[0067] (III) Beneficial effects

[0068] The present invention provides a desktop learning concentration monitoring method and system, which has the following beneficial effects: by fusing the gaze tracking data and the intelligent image recognition data, the learner's gaze focus area and object information on the desktop are obtained, and the learner's gaze stay time on the desktop, the classification and time data of the desktop objects paid attention to can be monitored and recorded, and a signal is generated based on these data to remind the learner to maintain concentration. BRIEF DESCRIPTION OF THE DRAWINGS

[0069] Figure 1 This is a flow chart of the desktop learning concentration monitoring method of the present invention;

[0070] Figure 2 The obtained sight tracking data and the schematic diagram thereof of the present invention;

[0071] Figure 3 It is a schematic diagram of calculating ROI_Y of the present invention;

[0072] Figure 4 It is a schematic diagram of calculating X0 of the present invention;

[0073] Figure 5 It is a schematic diagram of calculating ROI_X of the present invention;

[0074] Figure 6 A schematic diagram for describing the area of ​​interest of the present invention;

[0075] Figure 7 A schematic diagram describing the physical coordinate system calibrated for the system of the present invention;

[0076] Figure 8 A schematic diagram of initialization parameters for system calibration of the present invention;

[0077] Figure 9 This is a block diagram of the desktop learning concentration monitoring system of the present invention. DETAILED DESCRIPTION

[0078] The embodiments of the present invention are described in detail below, examples of which are shown in the accompanying drawings, wherein the same or similar reference numerals throughout represent the same or similar elements or elements having the same or similar functions. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the present invention, and cannot be understood as limiting the present invention.

[0079] The disclosure below provides many different embodiments or examples to realize different structures of the present invention. In order to simplify the disclosure of the present invention, the parts and settings of specific examples are described below. Of course, they are only examples, and the purpose is not to limit the present invention. In addition, the present invention can repeat reference numbers and / or reference letters in different examples, and this repetition is for the purpose of simplicity and clarity, which itself does not indicate the relationship between the various embodiments and / or settings discussed. In addition, the present invention provides various specific examples of processes and materials, but those of ordinary skill in the art can be aware of the application of other processes and / or the use of other materials.

[0080] like Figure 1 As shown, an embodiment of the present invention provides a desktop learning concentration monitoring method, comprising the following steps:

[0081] S1: Collect learners’ real-time facial images and desktop images;

[0082] S2: Based on the learner's real-time facial image, the learner's eye position and sight angle in the image are obtained, and based on the real-time desktop image, the type and position of desktop objects are determined;

[0083] S3: Calculate the overlap between the learner's sight angle area on the desktop and the location of the desktop objects to determine the objects that the learner is paying attention to;

[0084] S4: Count the types of objects that learners pay attention to and the time they pay attention to them, and regularly record the learners' facial images and the images and data of the objects they pay attention to, to generate traceable record data;

[0085] S5: When it is determined that the learner's attention time on non-learning objects exceeds a threshold, or the time when the learner's sight is away from the desk exceeds a threshold, an alarm is issued.

[0086] Furthermore, the overlap between the learner's sight angle area on the desktop and the location of the desktop objects is calculated to determine whether the objects that the learner is paying attention to include:

[0087] S21: Obtain eye position data and sight angle in the image from the learner’s real-time facial image;

[0088] Among them, in the collected learner's face image, the eye tracking technology is used to obtain the eye position data (x, y) and the eye angle in the image. The eye angle includes the deflection angle α, the pitch angle β and the roll angle θ. The x-axis, y-axis and z-axis are the coordinate system in the front image;

[0089] like Figure 2 As shown in the figure, from the learner's face image taken by the front camera, the eye tracking technology is used to obtain the eye position data (x, y) and the eye angle in the image. The eye angle includes the deflection angle α, the pitch angle β and the roll angle θ. The x-axis, y-axis and z-axis are the coordinate system in the front image.

[0090] If the front camera is a depth camera, the distance data L′ between the center of the human eye and the front camera can also be obtained.

[0091] S22: periodically identifying the type and location of the objects in the desktop real-time image to determine the type of the objects and the corresponding location;

[0092] S23: Fuse the gaze tracking data and the desktop object data to determine the objects on the desktop that the learner is focusing on;

[0093] S24: Count the time data of the learner's eye focus and record the images and data.

[0094] Furthermore, by integrating the gaze tracking data and the desktop object data, it is determined that the objects on the desktop that the learner focuses on include:

[0095] like Figure 3As shown, ROI_Y is calculated using the eye position data y, the sight pitch angle β, and the sight roll angle θ;

[0096] Physical height of human eye Z = Z0-y*K1*cos(A)

[0097] Physical pitch angle of human eye β′=β-A

[0098] ROI_Y can be obtained as follows: ROI_Y = YO-Z*tan(β′)

[0099] If the front camera is a depth camera, the angle B between L and the optical axis is calculated from the depth information, and YO in the above formula can be corrected:

[0100] YO = L*COS(A+B);

[0101] like Figure 4 As shown, ROI_X is calculated using the eye position data x, the sight deflection angle α, and the sight rolling angle θ:

[0102] First calculate the X-axis coordinate X0 of the intersection of the vertical line of the human eye center and the XOY plane;

[0103]

[0104] Looking down from the Z axis, we can get Figure 5 , from the figure we can get:

[0105] ROI_X=(Y0-ROI_Y)*tan(α)-X0

[0106] The sight rolling angle θ reflects the degree of head tilt of the learner and can be used to fine-tune ROI_X;

[0107] Among them, Figure 6 As shown, if the front camera is a depth camera, the distance data L′ between the center of the human eye and the front camera can be obtained, and the human eye position data x has been obtained in advance. The distance L between the center point of the row where the center point of the human eye is located and the front camera can be obtained by calculation:

[0108]

[0109] Based on the determined (ROI_X, ROI_Y), a square of a certain width is used to determine the focus area, the overlapping area S between the focus area and the object image is calculated, and a threshold of the overlapping area is set. When S>threshold, it is determined that the learner is focusing on the object.

[0110] Furthermore, the time data of learners' eye gaze attention is counted and images and data are recorded, including:

[0111] The desktop items are classified into "learning" and "non-learning" categories. For example, "electronic digital products" such as mobile phones and "food and beverages" such as snacks are classified into the "non-learning" category; the line of sight out of the image range of the top camera is defined as "line of sight out of the desktop";

[0112] Count the time t1 when learners "look away from the desk" and the time t2 when they focus on "non-learning" items;

[0113] The learner's real-time facial image and desktop real-time image as well as the above statistical data are recorded at regular intervals to generate a traceable record.

[0114] Furthermore, when it is determined that the learner's attention to non-learning objects exceeds a threshold, or the time when the learner's sight is out of the desk range exceeds a threshold, an alarm is issued including:

[0115] Set the reminder threshold to determine whether the "eyes away from the desk" time t1 or the "non-learning" object attention time t2 has exceeded the threshold in the recent period. If the threshold is exceeded, an alarm is triggered.

[0116] The desktop learning concentration monitoring method of this embodiment integrates the gaze tracking data and the intelligent image recognition data to obtain the learner's gaze focus area and object information on the desktop. It can monitor and record the time the learner's gaze stays on the desktop, the classification and time data of the desktop objects paid attention to, and generate signals based on these data to remind the learner to maintain concentration.

[0117] This embodiment also provides a desktop learning concentration monitoring system, which is used to implement the above-mentioned desktop learning concentration monitoring method, and is characterized by comprising:

[0118] An image acquisition unit, used to acquire real-time facial images and desktop images of learners;

[0119] A main controller, used for connecting to an image acquisition unit;

[0120] The main controller is also used to obtain the position and sight angle of the learner's eyes in the image based on the learner's real-time facial image, and determine the type and position of desktop objects based on the desktop real-time image;

[0121] Calculate the overlap between the learner's sight angle on the desktop and the location of the desktop objects to determine the objects that the learner is paying attention to;

[0122] Count the types of objects that learners pay attention to and the time they pay attention to them, and regularly record learners' facial images and the images and data of the objects they pay attention to, generating traceable record data;

[0123] An alarm device, used to warn learners through sound and light information, and controlled by the main controller;

[0124] The terminal is configured in the electronic device App. When the electronic device App is started, it accesses the main controller through the Internet to view the real-time image captured by the image acquisition unit, view the learning attention statistics, and receive reminder information.

[0125] Among them Figure 7 and Figure 8 As shown, the system of the present application needs to calibrate the physical coordinate system and initialize parameters before implementing the desktop learning concentration monitoring method:

[0126] The physical coordinate system of the system calibration is as follows Figure 7 As shown, the front camera is located at the origin of the XYZ coordinate system and has an elevation angle A; the top camera is directly above the Y axis; the axis of the top camera and the axis of the front camera are both in the YOZ plane; the green dotted box is the field of view image of the front camera; the blue dotted box is the field of view image of the top camera.

[0127] Initialization parameters are obtained through camera calibration procedures and measurements. The measurement method can be a ruler or the information of the depth camera. During calibration, the learner sits upright in front of the desktop, keeps a normal learning distance from the desktop, and the center of the face is located at the 0 position of the X axis. The elevation angle of the front camera is adjusted so that the center of the eye is located at the center of the front image, such as Figure 7 , Figure 8 As shown. The initialization parameters include:

[0128] (1) Front camera image pixel width W1, pixel height H1, and the proportional coefficient k1 of each pixel to the physical coordinate distance unit;

[0129] (2) The pixel width W2 of the top camera image, the pixel height H2, and the proportionality coefficient k2 between each pixel and the physical coordinate distance unit;

[0130] (3) The elevation angle of the front camera is A, and the angle between the front camera image and the Z axis is also A;

[0131] (4) Y-axis coordinate YC of the center position of the overhead camera image

[0132] (5) Y-axis distance Y0 between the learner’s eye and the front camera;

[0133] (6) Z-axis height Z0 of the highest point of the front camera image.

[0134] Among the above parameters, Y0 and k2 will change as the learner moves forward and backward. If the front camera is a depth camera, the distance data measured by the depth camera can be used for adjustment.

[0135] Furthermore, the image acquisition unit includes a top camera and a front camera. The top camera is used to capture images of desktop objects, and its pointing axis is in the same plane as the front camera. The front camera maintains a certain elevation angle to capture the learner's face, and its pointing axis is in the same plane as the top camera.

[0136] Furthermore, the overlap between the learner's sight angle area on the desktop and the location of the desktop objects is calculated to determine whether the objects that the learner is paying attention to include:

[0137] From the learner's real-time facial image, obtain the eye position data and sight angle in the image;

[0138] Among them, in the collected learner's face image, the eye tracking technology is used to obtain the eye position data (x, y) and the eye angle in the image. The eye angle includes the deflection angle α, the pitch angle β and the roll angle θ. The x-axis, y-axis and z-axis are the coordinate system in the front image;

[0139] Periodically identify the type and location of objects in the desktop real-time image to determine the type of objects and their corresponding locations;

[0140] Fuse the gaze tracking data with the desktop object data to determine the objects on the desktop that the learner is focusing on;

[0141] Count the time data of learners' eye focus and record the images and data.

[0142] Furthermore, by integrating the gaze tracking data and the desktop object data, it is determined that the objects on the desktop that the learner focuses on include:

[0143] Calculate ROI_Y using the eye position data y, the sight pitch angle β, and the sight roll angle θ;

[0144] Physical height of human eye Z = Z0-y*K1*cos(A)

[0145] Physical pitch angle of human eye β′=β-A

[0146] ROI_Y can be obtained as follows: ROI_Y = TO-Z*tan(β′);

[0147] Calculate ROI_X using the eye position data x, sight deflection angle α, and sight roll angle θ:

[0148] First calculate the X-axis coordinate X0 of the intersection of the vertical line of the human eye center and the XOY plane;

[0149]

[0150] ROI_X=(Y0-ROI_Y)*tan(α)-X0

[0151] The sight rolling angle θ reflects the degree of head tilt of the learner and can be used to fine-tune ROI_X;

[0152] Based on the determined (ROI_X, ROI_Y), a square of a certain width is used to determine the focus area, the overlapping area S between the focus area and the object image is calculated, and a threshold of the overlapping area is set. When S>threshold, it is determined that the learner is focusing on the object.

[0153] Furthermore, the time data of learners' eye gaze attention is counted and images and data are recorded, including:

[0154] The desktop items are classified into "learning" and "non-learning" categories. For example, "electronic digital products" such as mobile phones and "food and beverages" such as snacks are classified into the "non-learning" category; the line of sight out of the image range of the top camera is defined as "line of sight out of the desktop";

[0155] Count the time t1 when learners "look away from the desk" and the time t2 when they focus on "non-learning" items;

[0156] The learner's real-time facial image and desktop real-time image as well as the above statistical data are recorded at regular intervals to generate a traceable record.

[0157] Although embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions and variations may be made to the embodiments without departing from the principles and spirit of the present invention, and that the scope of the present invention is defined by the appended claims and their equivalents.

Claims

1. A desktop learning concentration monitoring method, characterized in that: The following steps are involved: Collect learners' real-time facial images and desktop images; Based on the learner's real-time facial image, the position and sight angle of the learner's eyes in the image are obtained, and based on the real-time desktop image, the type and position of desktop objects are determined; Calculate the overlap between the learner's sight angle on the desktop and the location of the desktop objects to determine the objects that the learner is paying attention to; Count the types of objects that learners pay attention to and the time they pay attention to them, and regularly record learners' facial images and the images and data of the objects they pay attention to, generating traceable record data; When it is determined that the learner's attention to non-learning objects exceeds a threshold, or the time when the learner's eyes are away from the desk exceeds a threshold, an alarm is issued.

2. A desktop learning concentration monitoring method according to claim 1, characterized in that: The calculation of the overlap between the learner's sight angle area on the desktop and the location of the desktop items and the determination of the items that the learner is paying attention to include: From the learner's real-time facial image, obtain the eye position data and sight angle in the image; Among them, in the collected learner's face image, the eye tracking technology is used to obtain the human eye position data (x, y) and the eye angle in the image, and the eye angle includes the deflection angle α, the pitch angle β and the roll angle θ, and the x-axis, y-axis and z-axis are the coordinate system in the front image; Periodically identify the type and location of objects in the desktop real-time image to determine the type of objects and their corresponding locations; Fuse the gaze tracking data with the desktop object data to determine the objects on the desktop that the learner is focusing on; Count the time data of learners' eye focus and record the images and data.

3. A desktop learning concentration monitoring method according to claim 2, characterized in that: The fusion of the eye tracking data and the desktop object data to determine the desktop object that the learner is focusing on includes: Calculate ROI_Y using the eye position data y, the sight pitch angle β, and the sight roll angle θ; Physical height of human eye Z = Z0-y*K1*cos(A) Physical pitch angle of human eye β′=β-A ROI_Y can be obtained as follows: ROI_Y = Y0 - Z*tan(β′); Calculate ROI_X using the eye position data x, sight deflection angle α, and sight roll angle θ: First calculate the X-axis coordinate X0 of the intersection of the vertical line of the human eye center and the XOY plane; ROI_X=(YO-ROI_Y)*tan(α)-X0 The sight rolling angle θ reflects the degree of head tilt of the learner and can be used to fine-tune ROI_X; Based on the determined (ROI_X, ROI_Y), a square of a certain width is used to determine the focus area, the overlapping area S between the focus area and the object image is calculated, and a threshold of the overlapping area is set. When S>threshold, it is determined that the learner is focusing on the object.

4. A desktop learning concentration monitoring method according to claim 2, characterized in that: The method of collecting the time data of the learner's sight attention and recording the images and data includes: The desktop items are classified into "learning" and "non-learning" categories. For example, "electronic digital products" such as mobile phones and "food and beverages" such as snacks are classified as "non-learning"; the line of sight out of the image range of the top camera is defined as "line of sight out of the desktop"; Count the time t1 when learners "look away from the desk" and the time t2 when they focus on "non-learning" items; The learner's real-time facial image and desktop real-time image as well as the above statistical data are recorded at regular intervals to generate a traceable record.

5. A desktop learning concentration monitoring method according to claim 4, characterized in that: When it is determined that the learner's attention to non-learning objects exceeds a threshold, or the time of sight away from the desk exceeds a threshold, issuing an alarm includes: Set the reminder threshold to determine whether the "eyes away from the desk" time t1 or the "non-learning" object focus time t2 exceeds the threshold in the recent cycle. If the threshold is exceeded, an alarm is triggered.

6. A desktop learning concentration monitoring system, used to implement the desktop learning concentration monitoring method according to any one of claims 1 to 5, characterized in that: include: An image acquisition unit, used to acquire real-time facial images and desktop images of learners; A main controller, used for connecting to an image acquisition unit; The main controller is also used to obtain the position and sight angle of the learner's eyes in the image based on the learner's real-time facial image, and determine the type and position of desktop objects based on the desktop real-time image; Calculate the overlap between the learner's sight angle on the desktop and the location of the desktop objects to determine the objects that the learner is paying attention to; Count the types of objects that learners pay attention to and the time they pay attention to them, and regularly record learners' facial images and the images and data of the objects they pay attention to, generating traceable record data; An alarm device, used to warn learners through sound and light information, and controlled by the main controller; The terminal is configured in the electronic device App. When the electronic device App is started, it accesses the main controller through the Internet to view the real-time image captured by the image acquisition unit, view the learning attention statistics, and receive reminder information.

7. A desktop learning concentration monitoring system according to claim 6, characterized in that: The image acquisition unit includes a top camera and a front camera. The top camera is used to capture images of desktop items, and its pointing axis is on the same plane as the front camera. The front camera maintains a certain elevation angle to capture the learner's face, and its pointing axis is on the same plane as the top camera.

8. A desktop learning concentration monitoring system according to claim 6, characterized in that: The calculation of the overlap between the learner's sight angle area on the desktop and the location of the desktop items and the determination of the items that the learner is paying attention to include: From the learner's real-time facial image, obtain the eye position data and sight angle in the image; Among them, in the collected learner's face image, the eye tracking technology is used to obtain the human eye position data (x, y) and the eye angle in the image, and the eye angle includes the deflection angle α, the pitch angle β and the roll angle θ, and the x-axis, y-axis and z-axis are the coordinate system in the front image; Periodically identify the type and location of objects in the desktop real-time image to determine the type of objects and their corresponding locations; Fuse the gaze tracking data with the desktop object data to determine the objects on the desktop that the learner is focusing on; Count the time data of learners' eye focus and record the images and data.

9. A desktop learning concentration monitoring system according to claim 8, characterized in that: The fusion of the eye tracking data and the desktop object data to determine the desktop object that the learner is focusing on includes: Calculate ROI_Y using the eye position data y, the sight pitch angle β, and the sight roll angle θ; Physical height of human eye Z = Z0-y*K1*cos(A) Physical pitch angle of human eye β′=β-A ROI_Y can be obtained as follows: ROI_Y = Y0 - Z*tan(β′); Calculate ROI_X using the eye position data x, sight deflection angle α, and sight roll angle θ: First calculate the X-axis coordinate X0 of the intersection of the vertical line of the human eye center and the XOY plane; ROI_X=(YO-ROI_Y)*tan(α)-X0 The sight rolling angle θ reflects the degree of head tilt of the learner and can be used to fine-tune ROI_X; Based on the determined (ROI_X, ROI_Y), a square of a certain width is used to determine the focus area, the overlapping area S between the focus area and the object image is calculated, and a threshold of the overlapping area is set. When S>threshold, it is determined that the learner is focusing on the object.

10. A desktop learning concentration monitoring system according to claim 8, characterized in that: The method of collecting the time data of the learner's sight attention and recording the images and data includes: The desktop items are classified into "learning" and "non-learning" categories. For example, "electronic digital products" such as mobile phones and "food and beverages" such as snacks are classified as "non-learning"; the line of sight out of the image range of the top camera is defined as "line of sight out of the desktop"; Count the time t1 when learners "look away from the desk" and the time t2 when they focus on "non-learning" items; The learner's real-time facial image and desktop real-time image as well as the above statistical data are recorded at regular intervals to generate a traceable record.