Learning monitoring method, device, storage medium and electronic equipment

By collecting facial images and interaction data of students during their homework in real time and combining them with facial feature and motion trajectory analysis, the problem of inaccurate learning status monitoring in existing technologies is solved, and accurate assessment of students' learning status and auxiliary improvement of efficiency are achieved.

CN120471744BActive Publication Date: 2025-10-03浙江海亮科技有限公司
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510941206.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-07-09
Publication Date
2025-10-03
Estimated Expiration
2045-07-09

AI Technical Summary

Technical Problem

In the existing technology, the student learning status monitoring method has a low ability to identify the user's true learning status and cannot accurately identify deceptive behaviors such as scripts and feints, resulting in inaccurate evaluation results and inability to effectively assist in improving learning efficiency.

Method used

By collecting facial image data and interaction data of students during their homework in real time, combining facial features and motion trajectory analysis, judging user behavior characteristics, integrating multi-dimensional data for weighted calculation, and evaluating students' concentration scores.

Benefits of technology

It achieves accurate monitoring of students’ learning status, avoids the influence of scripts and false actions, and improves the accuracy of assessment and auxiliary improvement of learning efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120471744B_ABST
    Figure CN120471744B_ABST
Patent Text Reader

Abstract

The present disclosure provides a learning monitoring method, device, storage medium, and electronic device. The method includes: in response to a target user submitting a current assignment, obtaining facial image data and interaction data corresponding to the target user, the facial image data and interaction data being the target user's user data during the current assignment; determining parameter information corresponding to at least one behavioral feature based on multiple feature points and interaction data corresponding to the facial image data, the at least one behavioral feature characterizing the target user as being absent-minded; performing weighted calculation on the parameter information corresponding to the at least one behavioral feature and the assignment accuracy to determine a concentration score, which is used to evaluate the target user's learning status. The method disclosed herein can integrate multi-source data for analysis and evaluation, improve the accuracy of the assessment of the user's learning status, prevent cheating from affecting the assessment results, and further improve the user's user experience.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of computer technology, and in particular to a learning monitoring method, device, storage medium, and electronic device. Background Art

[0002] When students use related applications to do homework, various types of user data may be generated. Related technologies need to monitor students' behavior based on the generated user data in order to evaluate and improve students' learning status. Summary of the Invention

[0003] The present disclosure provides a learning monitoring method, device, storage medium and electronic device.

[0004] According to a first aspect of the present disclosure, a learning monitoring method is provided, which includes: in response to a target user submitting a current assignment, obtaining facial image data and interaction data corresponding to the target user, the interaction data including screenshot data, mouse click data, and assignment accuracy, and the facial image data and interaction data are user data of the target user during the current assignment; based on multiple feature points and interaction data corresponding to the facial image data, determining parameter information corresponding to at least one behavioral feature, the parameter information including the number and duration of the behavioral feature, and at least one behavioral feature characterizing that the target user is in an absent-minded state; performing weighted calculation on the parameter information and assignment accuracy corresponding to the at least one behavioral feature to determine a concentration score, and the concentration score is used to evaluate the learning status of the target user.

[0005] In some embodiments of the present disclosure, parameter information corresponding to at least one behavioral feature is determined based on multiple feature points and interaction data corresponding to facial image data, including: preprocessing the multiple feature points to obtain boundary range parameters, the boundary range parameters include multiple boundary point coordinates and center point coordinates; determining acceleration parameters based on the boundary range parameters; determining the number and duration of behavioral features corresponding to each behavioral feature of at least one behavioral feature based on the boundary range parameters, acceleration parameters, facial image data, multiple feature points, and at least one of the interaction data.

[0006] In some embodiments of the present disclosure, multiple feature points are preprocessed to obtain boundary range parameters, including: based on a preset screen size, normalizing the multiple feature points to determine an initial affine transformation matrix, the initial affine transformation matrix including an initial scaling ratio, an initial rotation angle, and an initial translation coordinate; based on the multiple feature points, the initial affine transformation matrix is ​​updated through a relative depth ratio model to obtain a target affine transformation matrix; based on the eye distance and mouth-eye height corresponding to the multiple feature points, the avatar data is determined through the target affine transformation matrix, the avatar data including face width and face height; based on the avatar data and the preset screen size, the coordinates of multiple boundary points and the center point coordinates are determined.

[0007] In some embodiments of the present disclosure, acceleration parameters are determined based on boundary range parameters, including: determining the coordinates of multiple avatar center points within a preset time period based on the boundary range parameters; calculating the displacement difference of the multiple avatar center point coordinates according to the preset time length to obtain the instantaneous acceleration corresponding to each preset time length in the preset time period; determining the acceleration parameters within the preset time period based on the instantaneous acceleration corresponding to each preset time length, the acceleration parameters including the acceleration mean, the acceleration standard deviation, and the proportion of low acceleration.

[0008] In some embodiments of the present disclosure, based on at least one of a boundary range parameter, an acceleration parameter, facial image data, a plurality of feature points, and interaction data, the number of behavioral feature times and duration corresponding to each behavioral feature in at least one behavioral feature are determined, including at least one of the following: when at least one of the eye aspect ratio and the acceleration standard deviation corresponding to the plurality of feature points meets a preset condition, the user state within the preset duration is determined to be a dozing state, and the number of dozing times is increased by a count value, and the dozing duration is increased by a preset duration; when the acceleration mean is less than the acceleration mean threshold, and the low acceleration ratio is greater than the low acceleration ratio threshold, the user state within the preset duration is determined to be a dazed state. , and increase the number of daze times by a count value, and increase the daze duration by a preset duration; when at least one boundary point coordinate in the current boundary range parameter is not in the historical reference range corresponding to the boundary point coordinate, determine that the user state within the preset duration is an unruly state, and increase the number of unruly times by a count value, and increase the unruly duration by a preset duration. The current boundary range parameter is obtained by filtering the boundary range parameters corresponding to the daze state and the dozing state in the boundary range parameter; according to the preset duration, when the text similarity between the target area in the screenshot data and the current job text data is less than or equal to the similarity threshold, increase the number of irrelevant times by a count value, and increase the irrelevant duration by a preset duration.

[0009] In some embodiments of the present disclosure, the preset conditions include: the eye aspect ratio is zero; the acceleration standard deviation is greater than the drowsiness threshold; the number of times the acceleration standard deviation reaches a peak within a preset time period exceeds a preset number.

[0010] In some embodiments of the present disclosure, the method further includes: if facial image data is not collected for a first period of time during the current operation, increasing the number of misalignment times by a count value and increasing the misalignment duration by the first period of time.

[0011] In some embodiments of the present disclosure, the method further includes any of the following: when there is temporal overlap between the mouse click data in the interaction data and the screenshot data, determining the corresponding area of ​​the mouse click data in the screenshot data as the target area; when there is no temporal overlap between the mouse click data in the interaction data and the screenshot data, based on the screenshot data, determining the area where the overlap between the boundary range corresponding to the multiple screen partitions determined by the partition model for the screenshot data and the boundary range parameters satisfies the overlap condition as the target area.

[0012] According to a second aspect of the present disclosure, a learning monitoring device is provided, characterized in that the device comprises:

[0013] An acquisition module is configured to acquire facial image data and interaction data corresponding to the target user in response to the target user submitting the current job, wherein the interaction data includes screenshot data, mouse click data, and job accuracy, and the facial image data and interaction data are user data of the target user during the current job;

[0014] a determination module configured to determine parameter information corresponding to at least one behavioral feature based on a plurality of feature points corresponding to the facial image data and the interaction data, the parameter information including the number and duration of the behavioral feature, wherein the at least one behavioral feature indicates that the target user is in an absentee state;

[0015] The evaluation module is used to perform weighted calculation on the parameter information and operation accuracy corresponding to at least one behavioral feature to determine the concentration score, which is used to evaluate the learning status of the target user.

[0016] According to a third aspect of the present disclosure, a computer-readable storage medium is provided, on which a computer program is stored, characterized in that when the computer program is executed by a processor, the method of the first aspect is implemented.

[0017] According to a fourth aspect of the present disclosure, an electronic device is provided, comprising a storage medium, a processor, and a computer program stored on the storage medium and executable on the processor, wherein the method of the first aspect is implemented when the processor executes the computer program.

[0018] The learning monitoring method disclosed herein integrates multi-dimensional data, including facial image feature points and interaction data between the target user and the system, to determine the frequency and duration of various behavioral characteristics that characterize the target user's slacking behavior. This is then used to calculate a weighted concentration score for the target user, allowing for assessment and intervention of the target user's learning status. This method improves the accuracy of the assessment of the user's learning status, prevents cheating from affecting the assessment results, and further enhances the user experience.

[0019] It should be understood that the content described in this section is not intended to identify the key or important features of the embodiments of the present application, nor is it intended to limit the scope of the present application. Other features of the present application will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] The accompanying drawings are provided to facilitate a better understanding of the present invention and do not constitute a limitation of the present disclosure.

[0021] Figure 1 A flow chart of a learning monitoring method provided by an embodiment of the present disclosure;

[0022] Figure 2 A schematic diagram of a process for preprocessing multiple feature points provided in an embodiment of the present disclosure;

[0023] Figure 3 A flowchart of another learning monitoring method provided by an embodiment of the present disclosure;

[0024] Figure 4A It is an overall flow chart of a learning monitoring method;

[0025] Figure 4B This is a schematic diagram of facial feature point positioning;

[0026] Figure 4C A flowchart for mapping the visible area;

[0027] Figure 4D Flowchart for desertion analysis;

[0028] Figure 4E Flowchart for correlation detection;

[0029] Figure 4F Flowchart for correlation detection;

[0030] Figure 5 A schematic diagram of the structure of a learning monitoring device provided in an embodiment of the present disclosure;

[0031] Figure 6 A schematic diagram of the hardware structure of an electronic device provided in an embodiment of the present disclosure. DETAILED DESCRIPTION

[0032] The following description of exemplary embodiments of the present disclosure is made in conjunction with the accompanying drawings, including various details of the embodiments of the present disclosure to facilitate understanding. These details should be considered as merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications may be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, for the sake of clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.

[0033] In related technologies, the monitoring method of user learning status has a low ability to identify the user's true learning status, and cannot accurately identify deceptive behaviors such as scripts and fake movements, and thus cannot accurately assist in improving the user's learning efficiency based on the concentration assessment results.

[0034] In order to solve the problems in the related art, the present disclosure proposes a learning monitoring method, which collects image data and interaction data of the user's operation process in real time, and comprehensively analyzes and judges the image data, angle data and operation accuracy during the operation process and after the operation is submitted, and combines facial features and motion trajectory analysis to judge the user's behavioral characteristics, and obtains the number and duration of various behavioral characteristics of slacking off, so as to fuse multi-dimensional data and obtain the user's concentration score through weighted calculation to accurately evaluate the learning status. Through the learning monitoring method disclosed in the present disclosure, the various behavioral characteristics of the user can be accurately identified, and deceptive behaviors such as scripts and fake actions can be avoided from affecting the judgment of behavioral characteristics, so that a more accurate concentration score can be obtained, and accurate monitoring of the user's learning status can be achieved, thereby helping the user to improve their learning efficiency and status.

[0035] The following describes the learning monitoring method, apparatus, electronic device, storage medium, and computer program product according to embodiments of the present disclosure with reference to the accompanying drawings.

[0036] The learning monitoring method disclosed herein is executed by a learning monitoring device, which includes an image acquisition module and a screen capture module, wherein the image acquisition module is used to capture the image in front of the screen in real time to obtain image data, and the screen capture module is used to take screenshots at a preset frequency to obtain screenshot data.

[0037] Specifically, the image acquisition module may obtain facial image data with a human face, and may also obtain image data without a human face when shooting. Among them, the facial image data with a human face is used for analysis by the determination module to determine the parameter information corresponding to the behavioral characteristics, and the image data without a human face will be judged as the user is in an improper state.

[0038] Specifically, the screenshot data obtained by the screenshot module is used by the determination module for analysis to determine parameter information corresponding to the behavior characteristics.

[0039] For example, the image acquisition module may be a camera in a tablet computer.

[0040] Figure 1 This is a flow chart of a learning monitoring method provided by an embodiment of the present disclosure. Figure 1 As shown, the method includes:

[0041] Step 101: In response to a target user submitting a current job, facial image data and interaction data corresponding to the target user are acquired.

[0042] In some embodiments, the interaction data includes screenshot data, mouse click data, operation accuracy, facial image data, and the interaction data is user data of the target user during the current operation.

[0043] In some embodiments, in response to a target user submitting a current job, facial image data and interaction data corresponding to the target user can be obtained after the target user submits the job. The facial image data is image data of the current target user in front of the screen captured by the image acquisition module, and the interaction data is data generated by the human-computer interaction between the target user and the system, such as the target user clicking a mouse anywhere on the screen, the target user switching screen display interfaces, or the accuracy of the job after the target user submits the job.

[0044] In some embodiments, the interaction data may also be data generated by the user sliding or clicking a finger on the screen to operate the system display interface.

[0045] In some embodiments, the screenshot data may be an image of the current screen captured at the moment the screenshot is triggered. The triggering of the screenshot may be to continuously detect whether the target user switches screens while performing the current task. If so, a screenshot of the current screen is taken every X seconds (e.g., 1 second), and the data is continuously reported and stored for subsequent processing.

[0046] In the embodiment of the present disclosure, user data is the data generated by all operations of the target user between the start time of opening the current job and the end time of submitting the job, including image data, mouse operation data, and the accuracy data of the system's judgment on the submitted job, etc.

[0047] Specifically, the facial image data may be data reflecting the facial image of the target user extracted at a preset frequency or a preset time interval.

[0048] For example, the student opens the corresponding designated homework, the tablet turns on the camera, and specifies the effective completion time range of the homework.

[0049] For example, in the scenario where a tablet opens a homework app, the system monitors the user's behavior in real time through the tablet camera and captures the user's facial image at a specific time point. In order to monitor the user's homework concentration, the system extracts several frames of images at fixed intervals for subsequent facial feature point analysis. For example, the system extracts images at 1:00, 1:05, 1:10, and 1:15, and names them as Figure (1), Figure (2), Figure (3), and Figure (4), respectively.

[0050] In some embodiments, the facial image data and interaction data may be obtained by sampling all behavioral data during the target user's job at fixed intervals in response to the target user submitting the current job, wherein the fixed interval may be every Δt seconds.

[0051] In the embodiment of the present disclosure, the mouse click data can be the system monitoring the mouse click event and recording the interaction coordinates. Got it.

[0052] In the embodiment of the present disclosure, the screenshot data can be a system Take a screenshot at any time and get the screenshot ; Perform noise reduction processing and use the bilateral filter algorithm to suppress noise while retaining edge features to obtain screenshot data.

[0053] For example, Take a screenshot at any time and process the data as follows: Tablet Screenshot + Screenshot Noise Reduction: Always perform tablet screen capture and image preprocessing: Screen capture: Get tablet screen screenshots ; Noise reduction processing: Bilateral filter algorithm is used to suppress noise while retaining edge features.

[0054] Specifically, the user data of the target user during the current operation may include facial image data, screenshot data, mouse click data, and operation accuracy.

[0055] Step 102: Determine parameter information corresponding to at least one behavioral feature based on a plurality of feature points corresponding to the facial image data and the interaction data.

[0056] In some embodiments, the parameter information includes the number and duration of the behavioral characteristics, and at least one behavioral characteristic indicates that the target user is in an absent-minded state.

[0057] In some embodiments, the multiple feature points corresponding to the facial image data can be obtained by performing facial feature point detection using the Dlib 68-point model.

[0058] Specifically, the Dlib library can be used to perform keypoint detection on the extracted facial image data. This model can detect 68 different facial keypoints, including characteristic points for multiple facial regions such as the eyes, eyebrows, nose, mouth, and chin. For each facial image in the facial image data, the Dlib 68-point model is used to extract the keypoints of the target user's face. 68 keypoints are extracted from the image. Each keypoint contains two coordinate values, X and Y, which represent the location of a specific facial part in the image.

[0059] Specifically, the Dlib 68-point model can be used to obtain the positioning information of multiple feature points in each frame of facial image data. The positioning information includes the horizontal and vertical coordinate values ​​of feature points such as the outer canthus of the left and right eyes, the center of the mouth, and the tip of the nose.

[0060] For example, Figure 4B As shown in the schematic diagram, in Dlib's 68-point model, points 37 to 42 are used to mark the boundary of the left eye, points 43 to 48 are used to mark the boundary of the right eye, points 49 to 68 are used to mark the boundary of the mouth, and points 28 to 36 are used to mark the boundary of the nose. Each point corresponds to a specific feature point of the user's eyes and mouth, which is used to monitor the state of the human eye. For each image, the system extracts eye (6 feature points for the left and right eyes), mouth (20 feature points), and nose (9 feature points) data. Data example: arrays of user eye behavior, mouth behavior, and nose tip behavior detection. If there is no eye, mouth, or nose tip data at a certain moment, the coordinates are set to: {"point": 37, "x": 0, "y":0}. The nose tip positioning coordinates are: .

[0061] Furthermore, by obtaining the locations of multiple feature points, the location of the midpoint of the line connecting the outer canthi of the left and right eyes can be calculated using the following formula: ,in, For the outer canthus of the left eye, The outer canthus of the right eye. By locating the left and right corners of the mouth, the center of the mouth can be calculated. The calculation formula is: ,in, For the left corner of the mouth, Right corner of mouth.

[0062] In an embodiment of the present disclosure, parameter information corresponding to at least one behavioral feature is determined based on multiple feature points and interaction data corresponding to facial image data, including: preprocessing multiple feature points to obtain boundary range parameters, the boundary range parameters include multiple boundary point coordinates and center point coordinates; based on multiple feature points and interaction data, acceleration parameters are determined; based on at least one of the boundary range parameters, acceleration parameters, facial image data, multiple feature points, and interaction data, the number of behavioral features and duration corresponding to each behavioral feature in at least one behavioral feature are determined.

[0063] In some embodiments, preprocessing multiple feature points to obtain boundary range parameters can be performed through coordinate normalization, and the screen direction is further determined to obtain the boundary range parameters through horizontal screen mapping matrix calculation or vertical screen mapping matrix calculation.

[0064] In some embodiments, the boundary range parameters obtained by preprocessing the feature points can be the dynamic activity range of the target user in front of the screen, that is, the multiple boundary point coordinates can include the left boundary point coordinates, the right boundary point coordinates, the upper boundary point coordinates, and the lower boundary point coordinates. The boundary range is the rectangular range corresponding to the upper left corner coordinate point and the lower right corner coordinate point, so that the dynamic activity range corresponding to each frame of facial image data can be obtained, that is, multiple boundary point coordinates and center point coordinates are obtained.

[0065] In the embodiment of the present disclosure, the method of preprocessing multiple feature points is not limited by the present disclosure. The image coordinate system and the screen coordinate system can be first unified, and then a mapping relationship is established from the pixel ratio in the image data to the screen ratio of the screen size. Based on the mapping relationship and the coordinates of the multiple feature points, the dynamic activity range of the target user can be obtained by calculation, that is, the coordinates of multiple boundary points and the center point coordinates are obtained.

[0066] In some embodiments, the method for preprocessing multiple feature points can also be to detect the physical distance of the target user relative to the screen through a detection device, and then update the mapping relationship obtained above based on the physical distance, and use the updated mapping relationship and the coordinates of multiple feature points to obtain the dynamic activity range of the target user by calculation, that is, to obtain multiple boundary point coordinates and center point coordinates.

[0067] In some embodiments, based on the boundary range parameters, the acceleration parameters are determined, including: based on the boundary range parameters, determining the coordinates of multiple avatar center points within a preset time period; according to the preset time length, calculating the displacement difference of the multiple avatar center point coordinates to obtain the instantaneous acceleration corresponding to each preset time length in the preset time period; based on the instantaneous acceleration corresponding to each preset time length, determining the acceleration parameters within the preset time period, the acceleration parameters including the acceleration mean, the acceleration standard deviation, and the proportion of low acceleration.

[0068] In an embodiment of the present disclosure, based on the boundary range parameters, the coordinates of multiple avatar center points within a preset time period are determined, and the coordinates of the avatar center point corresponding to each facial image in the facial image data can be determined based on the boundary range parameters corresponding to each facial image.

[0069] In an embodiment of the present disclosure, the preset time period can be a pre-set time window, and determining the coordinates of multiple avatar center points corresponding to the facial image data within the time window can be obtaining the facial image data within the time window, and then obtaining the multiple feature points corresponding to the facial image data within the time window, as well as obtaining the coordinates of the avatar center point corresponding to each facial image.

[0070] For example, Timeframes where intervals Time period, obtain continuous avatar center data:

[0071] .

[0072] In an embodiment of the present disclosure, the preset duration may be the time interval between two adjacent facial images.

[0073] In an embodiment of the present disclosure, the displacement difference of the center coordinates of multiple avatars is calculated according to the preset time length to obtain the instantaneous acceleration corresponding to each preset time length in the preset time period. The displacement difference corresponding to each preset time length in the preset time period can be calculated by the displacement difference.

[0074] Specifically, the displacement difference calculation can be based on the coordinates of the avatar center point corresponding to the start time of the preset duration and the coordinates of the avatar center point corresponding to the end time of the preset duration. By calculating the difference of the horizontal coordinate and the difference of the vertical coordinate respectively, the displacement difference corresponding to the preset duration is obtained, that is, including the difference of the horizontal coordinate and the difference of the vertical coordinate.

[0075] For example, displacement difference calculation (within the time window):

[0076]

[0077] Furthermore, the instantaneous acceleration corresponding to each preset time length is calculated by the displacement difference. The instantaneous speed corresponding to the two time points of each preset time length can be first calculated by the horizontal coordinate difference and the vertical coordinate difference, and then the instantaneous acceleration corresponding to the preset time length is calculated based on the instantaneous speed. The instantaneous acceleration includes the horizontal coordinate instantaneous acceleration and the vertical coordinate instantaneous acceleration.

[0078] For example, the instantaneous speed: ;

[0079] Instantaneous acceleration:

[0080] .

[0081] In some embodiments, based on the instantaneous acceleration corresponding to each preset time length, the acceleration parameters within the preset time period are determined. The horizontal coordinate instantaneous acceleration and the vertical coordinate instantaneous acceleration corresponding to each preset time length can be synthesized to obtain the acceleration scalar corresponding to each preset time length, and then based on the acceleration scalars corresponding to multiple preset time lengths within the preset time period, the acceleration mean and acceleration standard deviation within the preset time period are calculated.

[0082] Furthermore, based on the acceleration scalars corresponding to multiple preset durations within a preset time period, the proportion of acceleration scalars with acceleration values ​​less than a preset acceleration threshold among all acceleration scalars is determined as the low acceleration ratio. Thus, the acceleration parameters: acceleration mean, acceleration standard deviation, and low acceleration ratio are obtained.

[0083] In some embodiments, the preset acceleration threshold is a preset minimum acceleration value, which can be customized according to the scenario or needs, and its specific value is not limited in this disclosure.

[0084] For example, based on the instantaneous acceleration , the resultant acceleration scalar:

[0085] ;

[0086] Further calculations are performed to obtain the mean acceleration:

[0087] ;

[0088] Acceleration standard deviation:

[0089] ;

[0090] Low acceleration ratio: defines the acceleration threshold

[0091] .

[0092] In some embodiments, based on at least one of a boundary range parameter, an acceleration parameter, facial image data, a plurality of feature points, and interaction data, the number of behavioral feature times and duration corresponding to each behavioral feature in at least one behavioral feature are determined, including at least one of the following: when at least one of the eye aspect ratio and the acceleration standard deviation corresponding to the plurality of feature points meets a preset condition, the user state within the preset time period is determined to be a dozing state, and the number of dozing times is increased by a count value, and the dozing duration is increased by a preset time period; when the acceleration mean is less than the acceleration mean threshold, and the low acceleration ratio is greater than the low acceleration ratio threshold, the user state within the preset time period is determined to be a dazed state, and The number of daze times is increased by a count value, and the daze duration is increased by a preset duration; when the coordinates of at least one boundary point in the current boundary range parameters are not in the historical reference range corresponding to the boundary point coordinates, the user state within the preset duration is determined to be an unbehavioral state, and the number of unbehavioral times is increased by a count value, and the unbehavioral duration is increased by a preset duration. The current boundary range parameters are obtained by filtering the boundary range parameters corresponding to the daze state and the dozing state in the boundary range parameters; according to the preset duration, when the text similarity between the target area in the screenshot data and the current job text data is less than or equal to the similarity threshold, the number of irrelevant times is increased by a count value, and the irrelevant duration is increased by a preset duration.

[0093] In some embodiments, the target user may have multiple behavioral characteristics during the current task, such as at least two of the following: dozing, daydreaming, unruly, and irrelevant. Alternatively, the target user may have only one of the following behavioral characteristics:

[0094] 1. Dozing judgment:

[0095] In some embodiments, when at least one of the eye aspect ratio and acceleration standard deviation corresponding to multiple feature points meets preset conditions, the user status within the preset time period is determined to be a dozing state, and the number of dozes is increased by a count value, and the dozing time is increased by the preset time period.

[0096] In some embodiments, the preset conditions include: the eye aspect ratio is zero; the acceleration standard deviation is greater than the drowsiness threshold; the number of times the acceleration standard deviation reaches a peak within a preset time period exceeds a preset number.

[0097] Specifically, the eye aspect ratio can be calculated based on multiple feature points corresponding to each facial image of the target user over a preset time period. The eye aspect ratio reflects whether the target user has their eyes open. For example, when the eye aspect ratio approaches 0, it indicates that the target user has their eyes closed; otherwise, it indicates that their eyes are open.

[0098] For example, the eye aspect ratio is calculated based on the following formula:

[0099]

[0100] Among them, p1 - p6 are the six key points of a single eye (for example, 37->42 for the left eye and 43->48 for the right eye).

[0101] When eyes are closed: the EAR value approaches 0 (the distance between the upper and lower eyelids is close to 0).

[0102] In some embodiments, when at least one of the following conditions is met: the eye aspect ratio is zero, the acceleration standard deviation is greater than the dozing threshold, and the number of times the acceleration standard deviation reaches a peak within a preset time period exceeds a preset number, then the state of the target user within the preset time period is a dozing state.

[0103] Specifically, the dozing threshold and the preset number of times are pre-set values ​​and can be customized according to specific scenarios or needs, which is not limited by the present disclosure.

[0104] For example, the parameter definition: The fluctuation threshold for dozing; is the number of acceleration peaks detected within the time window; is the acceleration peak value threshold detected within the time window. The user is in a dozing state if any of the following three conditions are met:

[0105]

[0106]

[0107]

[0108] If the dozing state is met, the number of dozing times +1, the dozing time + .

[0109] 2. Determination of daze:

[0110] In some embodiments, when the acceleration mean is less than the acceleration mean threshold and the low acceleration ratio is greater than the low acceleration ratio threshold, the user state within the preset time period is determined to be a daze state, and the number of daze times is increased by a count value, and the daze time is increased by the preset time period.

[0111] In some embodiments, the acceleration mean threshold and the low acceleration ratio threshold are pre-set thresholds and can be customized according to the scenario or needs.

[0112] For example, the parameter definition of daze determination is: is the mean threshold of the dazed acceleration, is the low acceleration ratio threshold.

[0113] Specifically, if the acceleration average within the preset time period is less than the acceleration average threshold, and the low acceleration ratio is greater than the low acceleration ratio threshold, the user state of the target user within the preset time period is a dazed state.

[0114] For example, the user is in a daze state when the following two conditions are met at the same time:

[0115]

[0116] If the daze state is met, the number of dazes +1, the daze duration + .

[0117] 3. Improper judgment:

[0118] In some embodiments, when at least one boundary point coordinate in the current boundary range parameter is not within the historical reference range corresponding to the boundary point coordinate, the user state within the preset time period is determined to be an improper state, and the number of improper times is increased by a count value, and the improper time period is increased by a preset time period. The current boundary range parameter is obtained by filtering the boundary range parameters corresponding to the dazed state and the dozing state in the boundary range parameters.

[0119] Specifically, the current boundary range parameter is a parameter obtained by removing the boundary range parameters determined as the dozing state and the dazed state during the current work period.

[0120] In some embodiments, the historical reference range corresponding to the boundary point coordinates can be based on the historical records of the target user, and the historical dynamic activity range of the target user is obtained by calculation, which represents the activity range maintained by the target user for a long time while doing homework, and is used to compare and judge the activity range during the current homework period to determine whether the target user is within the normal activity range.

[0121] For example, the historical benchmark model is constructed by using the sliding window statistical method. Based on the user's past M operation data, the reasonable range of the left limit, the reasonable range of the right limit, the reasonable range of the upper limit, and the reasonable range of the lower limit are calculated. The following takes the calculation process of the left limit as an example:

[0122] Left boundary mean:

[0123]

[0124] Left boundary standard deviation:

[0125]

[0126] Reasonable range interval of the left boundary: define the standard deviation multiple coefficient ,like ;

[0127]

[0128] The same applies to the right, upper and lower boundaries.

[0129] Furthermore, when at least one boundary point coordinate in the current boundary range parameter is not within the historical reference range corresponding to the boundary point coordinate, the user status within the preset time length is determined to be an improper state. In other words, when all boundary point coordinates are within the corresponding historical reference range, the user status of the target user will be determined to be proper.

[0130] For example, every During the time period, check whether the left, right, upper and lower boundaries are within their respective reasonable ranges. If they are beyond the reasonable range, the number of times of misalignment will be +1, and the cumulative time of misalignment will be + .

[0131] 4. Irrelevant judgment:

[0132] In some embodiments, according to a preset duration, when the text similarity between the target area in the screenshot data and the current job text data is less than or equal to a similarity threshold, the number of irrelevant times is increased by a count value and the irrelevant duration is increased by a preset duration.

[0133] In some embodiments, irrelevant determination is triggered when there is screenshot data in the interactive data, that is, during the current operation, if the target user switches screens, the irrelevant determination process will be triggered.

[0134] Among them, coherence and similarity can be used interchangeably, and similarity, correlation and coherence can be used interchangeably.

[0135] In some embodiments, the similarity threshold is a preset minimum value of text similarity. If the text similarity is less than or equal to the minimum value, it indicates that the text relevance is low and the user status is determined to be irrelevant.

[0136] In some embodiments, the target area in the screenshot data may be obtained by partitioning the screenshot data and determining the area where the target user clicks the mouse as the target area, or determining each partition as the target area.

[0137] In some embodiments, the target area may be obtained by calculating the overlap between each partition and the dynamic activity range of the target user, and determining the area with the highest overlap as the target area.

[0138] In some embodiments, the current job text data may be obtained by performing feature extraction on the current job to obtain the job theme of the current job.

[0139] In some embodiments, text similarity calculation is performed on the current job text data and the data of the target area to obtain a similarity value. Specifically, the text similarity calculation method can adopt the text similarity method in the related art.

[0140] For example, text A is the text of the target area, and text B is the text of the current task. The vectors are obtained by using BERT semantic vectorization. , ; Calculate semantic similarity: Use the cosine similarity method to obtain:

[0141]

[0142]

[0143] Depend on It can be deduced that: Represents the dot product of vector A and vector B; and Represents vectors and The modulus (length);

[0144] If similarity ≥ threshold , it means that the text is relevant.

[0145] Every Time period, to determine the relevance. If it is determined to be relevant, mark the user as focusing on searching for job-related content within that time period. Otherwise, it is determined to be irrelevant, and the number of irrelevant searches + 1 and the irrelevant search duration + .

[0146] In the above embodiment, by performing the above determination on the target user, the number and duration of various behavioral characteristics of the target user are obtained.

[0147] In some embodiments, the method further includes: if facial image data is not collected for a first period of time during the current operation, increasing the number of misalignment times by a count value and increasing the misalignment duration by the first period of time.

[0148] Specifically, if there is a time period during the current operation when the user's facial image is not collected, it means that the target user is not in the area in front of the screen, and the user's status can be determined to be improper.

[0149] For example, a threshold time period is defined Seconds, judge continuous If the face data is not collected within seconds, it means the user arrive If you are not in front of the tablet at the time (threshold time period), the number of times you are not straightened up will be +1, and the duration of not straightening up will be + , repeat this step continuously while no job is submitted.

[0150] In some embodiments, a duration threshold may also be set. Furthermore, if the first duration is greater than or equal to the duration threshold, it will be judged as improper. Conversely, if the first duration is less than the duration threshold, it will not be judged as improper.

[0151] Step 103: Perform weighted calculation on the parameter information and the task accuracy corresponding to at least one behavioral feature to determine a concentration score.

[0152] In some embodiments, the concentration score is used to evaluate the learning status of the target user.

[0153] Specifically, the number of times of absent-mindedness and the duration of absent-mindedness are determined based on the number and duration of behavioral characteristics of at least one behavioral characteristic; the concentration score is determined based on the preset number weight, duration weight, homework accuracy weight, number of absent-mindedness, duration of absent-mindedness, and homework accuracy.

[0154] In some embodiments, the number of absenteeism is obtained by adding up the number of behavioral feature times corresponding to at least one behavioral feature obtained in step 102, and the total number is obtained as the number of absenteeism; the duration of absenteeism is obtained by adding up the behavioral feature duration corresponding to at least one behavioral feature, and the total duration is obtained as the absenteeism duration.

[0155] In some embodiments, the job accuracy is determined by the system when the target user submits the current job.

[0156] In some embodiments, the preset number weight, duration weight, and operation accuracy weight can be the weights of various pre-set parameters, which can be obtained by comprehensive calculation based on historical records, or customized according to scenarios or needs, which is not limited in this disclosure.

[0157] In some embodiments, the concentration score is determined based on the preset number of weights, duration weights, homework accuracy weights, number of absenteeism, duration of absenteeism, and homework accuracy. The concentration score can be calculated using the following formula:

[0158]

[0159] in, is the weight of the number of desertions, is the weight of the absenteeism duration, is the desertion threshold, ; is the start time of the current job, is the end time of the current job, is the accuracy weight of the assignment.

[0160] In some embodiments, the method further includes any of the following: when the concentration score is in the first interval, determining that the target user's learning status is highly focused; when the concentration score is in the second interval, determining that the target user's learning status is moderately focused; when the concentration score is in the third interval, determining that the target user's learning status needs improvement; when the concentration score is in the fourth interval, determining that the target user's learning status is low focused.

[0161] In some embodiments, the first interval, the second interval, the third interval, and the fourth interval can be customized according to the scenario or needs, for example, the first interval is 90-100 points, the second interval is 70-89 points, the third interval is 50-69 points, and the fourth interval is less than 50 points.

[0162] In some embodiments, after obtaining the target user's concentration score, the target user's concentration level can be obtained based on the above-mentioned interval, and then based on the level, the target user's learning monitoring report can be obtained, and the learning monitoring report can be submitted to the target user's supervisor to achieve the purpose of monitoring the target user's learning and realize accurate learning status assessment.

[0163] In some embodiments, the target user's learning status is highly focused, indicating that the target user's efficiency and self-management are excellent; the target user's learning status is moderately focused, indicating that the target user is occasionally distracted but it does not affect the learning results; the target user's learning status is needs improvement, indicating that the target user's distracted behavior is obvious and may affect the learning effect; the target user's learning status is low focused, indicating that timely intervention is needed for the target user to avoid the target user not focusing on learning, resulting in a decline in learning outcomes.

[0164] In summary, the technical solution provided by this disclosure integrates multi-dimensional data, including facial image feature points and interaction data between the target user and the system, to obtain multiple behavioral characteristics that characterize the frequency and duration of the target user's slacking behavior. This is then used to calculate the target user's focus score through weighted calculation, allowing for assessment and intervention of the target user's learning status. This method improves the accuracy of the assessment of the user's learning status, prevents cheating from affecting the assessment results, and further enhances the user experience.

[0165] As a possible implementation, Figure 2 The flowchart of preprocessing multiple feature points shown in FIG. 1 includes the following steps based on the above embodiment:

[0166] Step 201: Based on a preset screen size, multiple feature points are normalized to determine an initial affine transformation matrix.

[0167] In some embodiments, the initial affine transformation matrix includes an initial scaling factor, an initial rotation angle, and an initial translation coordinate.

[0168] In some embodiments, the preset screen size is the screen size of the current target user performing the current task, and the preset screen size is determined to a specific value after the device determines it. Specifically, the preset screen size includes screen height and screen width.

[0169] Specifically, based on the preset screen size, multiple feature points are normalized to determine the initial affine transformation matrix, which can be used to determine the scaling ratio, translation coordinates, and rotation angles of the feature points of the facial image data relative to the preset screen size.

[0170] In an embodiment of the present disclosure, the initial facial image and the plane where the screen is located are preset to be parallel, and the initial rotation angle is 0.

[0171] For example, the image coordinate system is a 2D (two-dimensional) coordinate system (unit: pixel) based on the camera imaging plane, where the origin is the upper left corner of the image (0,0), the X-axis is horizontal to the right, and the Y-axis is vertically downward; the screen coordinate system is the physical display area of ​​the tablet device (unit: mm / pixel density), where the origin is the center point of the screen (0,0), the X-axis is horizontal to the right, and the Y-axis is vertically upward (opposite to the Y-axis of the image coordinate system).

[0172] Furthermore, the reference point can be determined among multiple feature points, and the rule for selecting the reference point can be to use the midpoint of the line connecting the outer canthi of the two eyes in the image coordinate system as the main reference point, the midpoint of the corner of the mouth in the image coordinate system as the vertical reference point, and the tip of the nose in the image coordinate system as the auxiliary calibration point.

[0173] Specifically, the coordinates of the outer canthus of the left eye and the outer canthus of the right eye among multiple feature points are used to obtain the length (horizontal width) of the eye line (eye_distance, abbreviated as ED):

[0174]

[0175] in, The X-axis coordinate of the outer canthus of the left eye; The X-axis coordinate of the outer canthus of the right eye; The Y-axis coordinate of the outer canthus of the left eye; The Y-axis coordinate of the outer canthus of the right eye.

[0176] Through the coordinates of the outer canthus of the left eye, the outer canthus of the right eye, and the coordinates of the midpoint of the mouth corner among multiple feature points, the vertical distance (horizontal height) from the eye to the mouth is obtained (mouth_height for short) MH :

[0177]

[0178] in, Y-axis coordinate of the midpoint of the mouth corner; The Y-axis coordinate of the outer canthus of the left eye; The Y-axis coordinate of the outer canthus of the right eye.

[0179] Furthermore, based on the above eye line length, the vertical distance from the eyes to the mouth, and the preset screen size, the pixel ratio is converted to the screen ratio:

[0180] Horizontal scaling : ;

[0181] Vertical scaling : ;

[0182] in, : The length of the eye line (horizontal width) eye_distance; The vertical distance from the eyes to the mouth (horizontal height).

[0183] Furthermore, by using the coordinates of the screen center in the preset screen size and the coordinates of the midpoint of the line connecting the outer canthi of both eyes, we can determine the required translation distance and align the center point with the center of the screen:

[0184]

[0185]

[0186] in, X-axis coordinate of the midpoint of the line connecting the outer canthi of both eyes; The Y-axis coordinate of the midpoint of the line connecting the outer canthi of both eyes.

[0187] Based on the scaling ratio and translation distance obtained above, the initial affine transformation matrix is ​​obtained:

[0188] Markdown:

[0189] [Scale X, Rotate Angle, Move X]

[0190] [rotation angle, scale Y, move Y]

[0191] For example, substitute the above matrix values ​​into:

[0192] [ ,0, ]

[0193] [0, , ]

[0194] in, horizontal scaling; Vertical scaling factor.

[0195] In some embodiments, the learning monitoring device is not provided with a detection device, such as a TOF sensor for detecting the distance or depth between the face and the screen, and the initial affine transformation matrix is ​​directly used to determine the subsequent avatar data.

[0196] In some embodiments, the learning monitoring apparatus is provided with a detection device, and the initial affine transformation matrix can be dynamically updated according to the distance or depth detected by the detection device, and the updated matrix is ​​used to determine subsequent avatar data.

[0197] Step 202: Based on multiple feature points, the initial affine transformation matrix is ​​updated through a relative depth scale model to obtain a target affine transformation matrix.

[0198] In some embodiments, the data update of the initial affine transformation matrix using the relative depth scale model can be performed as follows:

[0199] Normalized reference factor construction: In the initial state (when the user is at the calibration distance), define the standardized reference factor:

[0200]

[0201] Among them, the visual depth of the nose tip: ;

[0202] Face visual width : The pixel distance between the detected left and right boundary points of the face (for example, point 2 and point 17) in the image.

[0203] Then, a dynamic depth ratio calculation is performed. When the user moves their head, the current depth is calculated in real time:

[0204]

[0205] Among them, the numerator: the current frame ratio, reflecting the projection characteristics of the head's immediate distance;

[0206] Denominator: Initial benchmark , used to offset inherent differences in individual facial dimensions;

[0207] Screen adaptation (screen_diagonal): Use the physical size of the screen diagonal (e.g., a 6.1-inch phone vs. a 27-inch monitor) to ensure the compensation is appropriate for the display ratio of different devices.

[0208] Stability control parameter (k): Suppresses high-frequency noise caused by camera jitter or face detection errors (usually 0.1-0.5); ensures smooth and stable cursor movement by reducing the step size.

[0209] In some embodiments, the target scaling ratio is adjusted according to the current depth, and the calculated current depth is dynamically injected into the initial affine transformation matrix to obtain the target affine transformation matrix.

[0210] For example, the calculated Dynamically inject the affine transformation matrix. Specific operations:

[0211] Matrix update rule:

[0212] Scaling compensation: adjust the scaling factor based on depth (Z value):

[0213] .

[0214] Step 203: Based on the eye distance and mouth-eye height corresponding to the multiple feature points, the head portrait data is determined through the target affine transformation matrix.

[0215] In some embodiments, the avatar data includes face width and face height.

[0216] In some embodiments, based on the eye distance and mouth-eye height corresponding to multiple feature points, the target affine transformation matrix can be used to obtain the actual displayed face width and face height of the avatar. The eye distance and mouth-eye height can be calculated based on the coordinates of the multiple feature points.

[0217] In some embodiments, the eye distance and mouth-eye height are converted into avatar data through a target affine transformation matrix.

[0218] Specifically, calculate the actual displayed width and height of the avatar: face width = original eye distance × horizontal scaling ratio × safety factor; face height = original mouth-eye height × vertical scaling ratio × safety factor. The safety factor can be customized between 1.2 and 1.5 based on the scenario or needs.

[0219] Furthermore, different values ​​may be set for the safety factor in the horizontal direction and the safety factor in the vertical direction.

[0220] For example: FW = ED × scale.x × 1.2 (a safety factor of 1.2 means a 20% safety margin); for example: FH = MH × scale.y × 1.5 (the safety factor is 1.5, which means leaving 50% more safety space).

[0221] Step 204 : Determine the coordinates of multiple boundary points and the center point based on the avatar data and the preset screen size.

[0222] In an embodiment of the present disclosure, multiple boundary point coordinates and center point coordinates are determined based on avatar data and a preset screen size, wherein the multiple boundary point coordinates include left boundary point coordinates, right boundary point coordinates, upper boundary point coordinates, and lower boundary point coordinates, and the center point coordinates are the avatar center point coordinates.

[0223] Specifically, the coordinates of the left and right boundary points are determined as follows: left limit = face width / 2, right limit = screen width - face width / 2; for example, ; .

[0224] The coordinates of the upper and lower boundary points are as follows: upper limit = face height / 2; lower limit = screen height - face height / 2. For example, ; .

[0225] Based on the above boundary points, the boundary range is obtained, that is, the dynamic boundary matrix E corresponding to the coordinate points of the upper left corner and the lower right corner:

[0226]

[0227] Dynamic limit coordinates are expressed as:

[0228] ;

[0229] Dynamic limit matrix area .

[0230] For example, the screen size is 1080×1920, and the affine transformation parameters are: scale x=2.0, scale y=1.8; safety factor: 1.2 for horizontal and 1.5 for vertical; face width = ED×2.0×1.2, ED=100px, then face width = 240px; face height = MH×1.8×1.5, MH=150px, face height = 405px; dynamic limit: the left and right movable range is: 240 / 2=120px, 1080-120=960px; the up and down movable range is 405 / 2=202.5px, 1920-2-2.5=1717.5px; which means that the center of the avatar can only move within the horizontal range of 120~960px and the vertical range of 202.5~1717.5px of the screen.

[0231] In some embodiments, in the embodiments of the present disclosure, the center point coordinates are determined based on the avatar data and the preset screen size. The horizontal coordinate value and the vertical coordinate value of the center point can be obtained by using the left limit, upper limit, face width, and face height obtained above.

[0232] Specifically, avatar center.x = left limit + face width / 2:

[0233]

[0234] Avatar center.y = upper limit + face height / 2:

[0235]

[0236] In the above embodiment, multiple feature points corresponding to the facial image data are normalized based on the preset screen size to obtain an affine matrix that converts the pixel size into the screen size. Then, based on the avatar data obtained by the affine transformation and the preset screen size, the boundaries of the target user's dynamic activity range and the coordinates of the avatar center point can be obtained for use in the process of determining behavioral characteristics. This allows the determination of behavioral characteristics to integrate multi-source data, avoid the impact of cheating and other behaviors on the accuracy of the determination, and improve the accuracy and effectiveness of the determination of the target user's behavioral characteristics.

[0237] As a possible implementation, Figure 3 The flowchart of another learning monitoring method shown in the figure, based on the above embodiment, includes any one of the following steps:

[0238] Step 301: When there is temporal overlap between the mouse click data in the interaction data and the screenshot data, a corresponding area of ​​the mouse click data in the screenshot data is determined as a target area.

[0239] In an embodiment of the present disclosure, if the target user has mouse click data and screenshot data at the same time in the interaction data, it means that the target user has clicked the screen at that moment, which can be understood as the target user's current focus area is in the mouse click area, and the mouse click area can be further determined as the target area for judging text relevance.

[0240] In some embodiments, multiple areas are pre-set in each frame of the screenshot data. If the mouse coordinates corresponding to the mouse click data at the same moment fall into any one of the multiple areas, the corresponding area is determined as the target area.

[0241] For example, judging Whether the user triggers a mouse click event at the moment. If a mouse click event is triggered, the collected mouse coordinates are determined through the spatial mapping relationship. Does it fall into the predefined An area in , inferring the page where the user's gaze is located .

[0242] Furthermore, the relevance between the assignment topic and the page topic where the sight is located is determined. Specifically, the assignment topic is semantically vectorized and combined with Figure 1 In step 102, the text similarity calculation is performed to determine whether the homework topic is Is there any relevance?

[0243] Step 302: When there is no temporal overlap between the mouse click data in the interactive data and the screenshot data, based on the screenshot data, an area where the overlap between the boundary range corresponding to the multiple screen partitions determined by the partition model for the screenshot data and the boundary range parameters satisfies the overlap condition is determined as the target area.

[0244] In some embodiments, the interactive data does not contain the screenshot data and the mouse click data at the same time, and further based on the screenshot data and the mouse click data Figure 2 The boundary range parameters obtained in the calculation are used to calculate the overlap to determine the target area.

[0245] For example, judging Whether the user triggers the mouse click event at the moment, if the mouse click event is not triggered, according to the dynamic boundary matrix obtained and the above The cross-sectional area is calculated by , and the total area of ​​each intersection area ... . Set the minimum cross-proportion threshold , filter out the values ​​greater than the threshold , sort from large to small and get the largest proportion , in order to determine the page where the user's sight is , get Corresponding layout theme

[0246] Specifically, the partition model can be a deep learning model LayoutLMv3, the input of which is multiple frames of screenshots in the screenshot data, and the output is the classification of multiple screen partitions and the upper left and lower right coordinate points of the matrix.

[0247] Furthermore, the overlap calculation is performed for the multiple screen partitions and the dynamic boundary area corresponding to the boundary range parameter, which may be to calculate the intersection area and the intersection area ratio of each screen partition and the dynamic boundary area.

[0248] For example, suppose there are 2 matrices and ,

[0249] Matrix non-overlap rule: Rectangle A is to the left of Rectangle B: ≤ ; Rectangle A is to the right of rectangle B: ≥ ; Rectangle A is above rectangle B: ≤ ; Rectangle A is below rectangle B: ≥ .

[0250] When the above conditions are not met, it proves that there is overlap between the matrices.

[0251] Cross-sectional area The calculation formula is as follows:

[0252]

[0253] is the value of the coordinate of the lower right corner of matrix A on the x-axis, and so on 、 、 、 、 、 、 .

[0254] Calculation of overlapping area ratio:

[0255]

[0256] is the area of ​​the dynamic limit matrix, is the calculated intersection area of ​​rectangle A and rectangle B.

[0257] Furthermore, for each screen partition, the text of each screen partition can be obtained through OCR image content recognition, and the text vector can be obtained through BERT, and finally the content of each screen partition can be obtained.

[0258] In some embodiments, a screen partition whose overlap with the dynamic boundary range satisfies an overlap condition is determined as a target area, wherein the overlap condition may be a highest overlap.

[0259] Furthermore, text relevance calculation is performed based on the text of the target area and the current job. For example, the screenshot data of the target area is first semantically vectorized, and the obtained semantic vector is used to calculate text similarity with the semantic vector of the current job to obtain a similarity value, which is then used in the coherence determination process.

[0260] In the above embodiment, by judging the temporal overlap of the mouse click data and the screenshot data, the target area for text similarity calculation can be determined. Different determination methods can be used for different situations to further improve the efficiency and accuracy of the judgment, thereby improving the efficiency and accuracy of the overall learning monitoring method.

[0261] In summary, this disclosure integrates multi-dimensional data, using facial image feature points and interaction data between the target user and the system to derive multiple behavioral characteristics that characterize the number and duration of the target user's slacking behavior. This is then used to calculate a weighted concentration score for the target user, allowing for assessment and intervention of the target user's learning status. This method improves the accuracy of the assessment of the user's learning status, prevents cheating from affecting the assessment results, and further enhances the user experience.

[0262] The following is a specific implementation of the learning monitoring method:

[0263] Figure 4A The following is the overall flow chart.

[0264] 1. Camera collects facial images

[0265] 1. Video frame extraction

[0266] In the scenario where the tablet opens the homework app, the system monitors the user's behavior in real time through the tablet camera and captures the user's facial image at a specific time point. In order to monitor the user's homework concentration, the system extracts several frames of images at fixed intervals for subsequent facial feature point analysis. In this embodiment, the system extracts images at 1:00, 1:05, 1:10, and 1:15, respectively, and names them as Figure (1), Figure (2), Figure (3), and Figure (4).

[0267] These extracted image frames contain the user's complete facial information, which is used for subsequent facial feature point positioning and eye data calculation.

[0268] 2. Facial feature point positioning using the Dlib 68-point model

[0269] Dlib is an open source deep learning library that is widely used in the field of facial feature point detection. Through the Dlib library, the system can detect key points on the extracted facial images. The model can detect 68 different facial key points, such as Figure 4B As shown, it contains feature points of multiple facial areas such as eyes, eyebrows, nose, mouth and chin.

[0270] In each image, the system uses the Dlib 68-point model to extract key points from the user's face. The system extracts 68 key points from the image. Each key point contains two coordinate values, X and Y, that represent the location of a specific facial part in the image.

[0271] 3. Positioning of human eye features and mouth feature points

[0272] In Dlib's 68-point model, points 37 to 42 mark the left eye boundary, points 43 to 48 mark the right eye boundary, points 49 to 68 mark the mouth boundary, and points 28 to 36 mark the nose boundary. Each point corresponds to a specific feature point of the user's eyes and mouth, used to monitor eye status.

[0273] For each image, the system extracts eye (6 feature points for the left and right eyes), mouth (20 feature points), and nose (9 feature points) data. For example, the data includes arrays of user eye behavior, mouth behavior, and nose tip behavior detection. If there is no eye, mouth, or nose tip data at a certain moment, the coordinates are set to: {"point": 37, "x": 0, "y": 0}.

[0274] 4. Positioning the outer canthi of the left and right eyes and the center of the mouth

[0275] The midpoint of the line connecting the outer canthi of the left and right eyes is located as shown in Table 1:

[0276]

[0277] The calculation formula is as follows:

[0278]

[0279] Nose tip positioning: Nose tip ( NC ) is index 34, and the coordinate system example (x, y) = ( NC.x , NC.y ).

[0280]

[0281] Mouth center positioning: left corner of mouth ( LM ), index number is 49, coordinate system example (x, y) = ( LM.x , LM.y ); right corner of mouth ( RM ), index number is 55, coordinate system example (x, y) = ( RM.x , RM.y ).

[0282] The calculation formula is as follows:

[0283]

[0284] 5. Visual area mapping

[0285] like Figure 4C The flowchart shown.

[0286] 5.1 Coordinate system definition and mapping relationship

[0287] (1) Coordinate system division

[0288] Image coordinate system: 2D coordinate system based on the camera imaging plane (unit: pixel);

[0289] Origin: the upper left corner of the image (0,0);

[0290] X axis: horizontal to the right, Y axis: vertically downward;

[0291] Screen coordinate system: the physical display area of ​​the tablet device (unit: mm / pixel density);

[0292] Origin: the center of the screen (0,0);

[0293] X-axis: horizontal to the right, Y-axis: vertically upward (opposite to the Y-axis direction of the image coordinate system).

[0294] (2) The rules for selecting benchmark points are shown in Table 2:

[0295]

[0296] 5.2 Matrix calculation process

[0297] 5.2.1 Establishing the Affine Transformation Matrix

[0298] 5.2.1.1 Measure dimensions (calculate reference parameters), as shown in Table 3:

[0299]

[0300] Get the length (horizontal width) of the eye line ED

[0301]

[0302] The X-axis coordinate of the outer canthus of the left eye;

[0303] The X-axis coordinate of the outer canthus of the right eye;

[0304] The Y-axis coordinate of the outer canthus of the left eye;

[0305] The Y-axis coordinate of the outer canthus of the right eye.

[0306] Calculate the vertical distance from the eyes to the mouth (horizontal height) mouth_height MH :

[0307]

[0308] in: Y-axis coordinate of the midpoint of the mouth corner;

[0309] The Y-axis coordinate of the outer canthus of the left eye;

[0310] The Y-axis coordinate of the outer canthus of the right eye.

[0311] 5.2.1.2 Scaling (normalization)

[0312] Pixel ratio converted to screen ratio:

[0313] Horizontal scaling

[0314]

[0315] Vertical scaling

[0316]

[0317] The length of the eye line (horizontal width) eye_distance;

[0318] The vertical distance from the eyes to the mouth (horizontal height) mouth_height.

[0319] 5.2.1.3 Finding the center (coordinate system alignment)

[0320] Calculate the distance to move (align the center point with the center of the screen)

[0321]

[0322]

[0323] in, X-axis coordinate of the midpoint of the line connecting the outer canthi of both eyes;

[0324] The Y-axis coordinate of the midpoint of the line connecting the outer canthi of both eyes.

[0325] 5.2.1.4 Generating Matrix

[0326] Markdown:

[0327] [Scale X, Rotate Angle, Move X]

[0328] [rotation angle, scale Y, move Y]

[0329] Substitute the above matrix values ​​into:

[0330] [ ,0, ]

[0331] [0, , ]

[0332] in, horizontal scaling;

[0333] Vertical scaling factor.

[0334] 5.2.2 Depth Compensation Calculation

[0335] To achieve a stable display effect when the head position changes, dynamic compensation is performed through a relative depth scale model. This is divided into the following three stages:

[0336] (1) Construction of normalized reference factor: In the initial state (when the user is at the calibration distance), define the normalized reference factor:

[0337]

[0338] Nose tip visual depth : The nose tip depth obtained by the TOF sensor, then: ;

[0339] Face visual width : The pixel distance between the detected left and right boundary points of the face (for example, point 2 and point 17) in the image.

[0340] (2) Dynamic depth ratio calculation: When the user moves his head, the current depth ratio is calculated in real time.

[0341]

[0342] Quantity explanation:

[0343] (a) Relative depth change term:

[0344] Numerator: current frame ratio, reflecting the projection characteristics of the head's immediate distance;

[0345] Denominator: Initial benchmark , used to offset inherent differences in individual facial dimensions;

[0346] Ratio significance: A value of 1.2 means the user moves forward 20%, and a value of 0.8 means the user moves back 20%;

[0347] (b) Screen adaptation (screen_diagonal)

[0348] Use the physical size of the screen diagonal (e.g., a 6.1-inch phone vs. a 27-inch monitor) to ensure the compensation is appropriate for the display ratio of the device.

[0349] Example: The same head movement is mapped to a small pixel displacement on a small screen and a large displacement on a large screen.

[0350] (c) Stability control parameter (k)

[0351] Suppress high-frequency noise caused by camera jitter or face detection errors (usually 0.1-0.5);

[0352] Ensure smooth and stable cursor movement by reducing the step size.

[0353] (3) Compensation execution and display optimization:

[0354] will be calculated Dynamically inject the affine transformation matrix:

[0355] Matrix update rule: The initial affine transformation matrix is:

[0356] Scaling compensation: According to the depth ( Z value) to adjust the scaling factor:

[0357]

[0358] 5.2.3 Dynamic Boundary Generation Algorithm

[0359] The edges of the screen include the range of motion and dynamic limits.

[0360] 5.2.3.1 Calculating size

[0361] Calculate the actual width and height of the avatar based on the parameters after affine transformation in step 5.2.1

[0362] Face width = original eye distance × horizontal scaling ratio × safety factor; for example: FW = ED × scale.x × 1.2 (leave 20% more safety margin);

[0363] Face height = original mouth and eye height × vertical scaling ratio × safety factor; for example: FH = MH × scale.y ×1.5 (leave 50% more safety space).

[0364] 5.2.3.2 Determining Boundaries

[0365] Calculate the movable limit position according to the screen size and avatar size:

[0366] (1) Left and right boundaries: left limit = face width / 2

[0367]

[0368] Right limit = screen width - face width / 2

[0369]

[0370] (2) Upper and lower boundaries: Upper limit = face height / 2

[0371]

[0372] Lower limit = screen height - face height / 2

[0373]

[0374] The dynamic boundary matrix E corresponds to the coordinate points of the upper left and lower right corners.

[0375]

[0376] Dynamic limit coordinate representation

[0377] Dynamic limit matrix area

[0378] (3) Process of obtaining the center point of the portrait: portrait center.x = left limit + face width / 2

[0379]

[0380] Avatar center.y = upper limit + face height / 2

[0381]

[0382] 6. Analysis of desertion

[0383] like Figure 4D As shown in the flowchart, the user starts doing homework, opens the homework app, and the tablet camera automatically turns on to record the start time of the homework. , capture a frame from the tablet camera every s seconds.

[0384] 6.1 Calculation process when the user is not in front of the tablet and in a reasonable area

[0385] Defining threshold time periods Seconds, judge continuous If the face data is not collected within seconds, it means the user arrive Time (Threshold Time Period) Not in front of the tablet, the number of times of misalignment +1, the duration of misalignment + , repeat this step continuously during the period when no job is submitted; otherwise, it means there is face data (abnormal frames are removed and then spliced), and the threshold time period is defined Seconds, combined with step 5 visual area mapping to obtain data (head portrait center, left and right boundaries and upper and lower boundaries), calculate whether the head portrait center is within the boundary range, if it continues If the center of the second profile picture is not within the boundary, it means that the user is in front of the tablet, but not straight; the number of times of not straightening is recorded +1, and the length of time of not straightening is + .

[0386] The center point of the avatar must be within the range, otherwise it is considered out of the boundary range.

[0387]

[0388] 6.2 Analysis of Absence After Submission

[0389] Automatic submission of jobs or manual submission of jobs by users, recording of job end time , get the job period -> All facial data are analyzed (frames not containing faces are removed and reassembled).

[0390] Timeframes where intervals At intervals (e.g., every 3 seconds), obtain data on the center of the avatar, its left and right boundaries, and its upper and lower boundaries through (visible area mapping in step 5).

[0391] Eye aspect ratio EAR calculation rules:

[0392]

[0393] Among them, p1 - p6 are the six key points of a single eye (for example, 37->42 for the left eye and 43->48 for the right eye).

[0394] When eyes are closed: the EAR value approaches 0 (the distance between the upper and lower eyelids is close to 0).

[0395] 6.2.3 Head Center (HC) Analysis

[0396] Timeframes where intervals Time period, obtain continuous avatar center data:

[0397]

[0398] 6.2.3.1 Instantaneous acceleration calculation

[0399] (1) Displacement difference calculation (within the time window)

[0400]

[0401] (2) Instantaneous speed

[0402]

[0403] (3) Instantaneous acceleration

[0404]

[0405] (4) Synthetic acceleration scalar

[0406]

[0407] 6.2.3.2 Stability index extraction

[0408] (1) Mean acceleration

[0409]

[0410] (2) Acceleration standard deviation

[0411]

[0412] (3) Low acceleration ratio, defining the acceleration threshold

[0413]

[0414] Sleepiness determination:

[0415] Parameter definition: The fluctuation threshold for dozing; is the number of acceleration peaks detected within the time window; is the acceleration peak value threshold detected within the time window. The user is in a dozing state if any of the following three conditions are met:

[0416] The user is in a dozing state and any of the following three conditions are met:

[0417]

[0418]

[0419]

[0420] If the dozing state is met, the number of dozing times + 1, the duration of dozing + .

[0421] Daze judgment:

[0422] Parameter definition: is the mean threshold of the dazed acceleration, is the low acceleration ratio threshold.

[0423] The user is in a daze state when both of the following conditions are met:

[0424]

[0425] If the daze state is met, the number of dazes + 1, the daze duration + .

[0426] 6.2.4 Comparative Analysis of Historical Data in the Visible Area

[0427] 6.2.4.1 Historical Sedimentation of User Normal Boundary Data

[0428] After the job is submitted, step 6.2.3 analyzes the avatar center, filtering out data that corresponds to the time period of daze and dozing. The user's normal data (including the avatar center, left and right boundaries, and upper and lower boundaries) is used as the user's historical reasonable behavior feature data;

[0429] Historical benchmark model construction: Using sliding window statistics method, based on the user's past M operation data.

[0430] Left boundary mean:

[0431]

[0432] Left boundary standard deviation:

[0433]

[0434] Reasonable range interval of the left boundary: define the standard deviation multiple coefficient ,like

[0435]

[0436] The same applies to the right, upper and lower boundaries.

[0437] 6.2.4.2 Calculation of statistical indicators

[0438] Every During the time period, check whether the left, right, upper and lower boundaries are within their respective reasonable ranges. If they are beyond the reasonable range, the number of times of misalignment will be +1, and the cumulative time of misalignment will be + .

[0439] 2. Application Switching and Job Correlation Detection Process

[0440] 2.1 Flowchart Figure 4E and Figure 4F The specific process is as follows.

[0441] 2.2 Pre-data acquisition calculation instructions

[0442] Continuous collection: After the user submits a job, the system collects seconds) to sample and analyze all behavioral data during their operation;

[0443] Trigger detection: The focus determination process is initiated only when the screen is detected to be switched to a non-operating application.

[0444] 2.2.1 Data source collection instructions:

[0445] Listen for mouse click events and record interaction coordinates ;

[0446] According to the above existing process, the dynamic limit matrix is ​​obtained and area .

[0447] 2.2.2 Instructions for using the deep learning model LayoutLMv3

[0448] enter ,Using the fine-tuned LayoutLMv3 model, we can obtain the classification and matrix of the region corresponding to the upper left and lower right coordinate points.

[0449] 2.2.2 Intersection area calculation to illustrate the areas of different themes

[0450] Suppose there are 2 matrices:

[0451] ;

[0452] ,

[0453] Existing matrix non-overlap rules:

[0454] Rectangle A is to the left of rectangle B: ≤ ;

[0455] Rectangle A is to the right of rectangle B: ≥ ;

[0456] Rectangle A is above rectangle B: ≤ ;

[0457] Rectangle A is below rectangle B: ≥ .

[0458] When the above conditions are not met, it proves that there is overlap between the matrices.

[0459] Cross-sectional area The calculation formula is as follows:

[0460]

[0461] is the value of the coordinate of the lower right corner of matrix A on the x-axis, and so on 、 、 、 、 、 、 .

[0462] Calculation of overlapping area ratio:

[0463]

[0464] is the area of ​​the dynamic limit matrix mentioned above, is the calculated cross-area.

[0465] 2.2.3 Text Similarity Calculation

[0466] Suppose there are two texts, text A and text B, and the vectors are obtained by BERT semantic vectorization respectively. , ;

[0467] Calculate semantic similarity: use cosine similarity method to get . ,

[0468] Depend on It can be deduced that: Represents the dot product of vector A and vector B.

[0469] and Represents vectors and The modulus (length);

[0470] If similarity ≥ threshold , it means that the text is relevant.

[0471] 2.3 Single Screenshot Analysis Process

[0472] by Take a screenshot at any time and process the relevant data according to the following steps

[0473] 2.3.1 Tablet Screenshot + Screenshot Noise Reduction

[0474] exist Always perform tablet screen capture and image preprocessing:

[0475] Screenshot Capture: Get tablet screenshots ;

[0476] Noise reduction: Bilateral filter algorithm is used to suppress noise while retaining edge features.

[0477] 2.3.2 Screenshot layout partition

[0478] Take a screenshot of the deep learning model LayoutLMv3 from step 2.2.2 Partition and get each area The upper left and lower right corner coordinate point data.

[0479]

[0480] Among them, bbox is the corresponding area coordinate, and label is the corresponding category.

[0481] 2.3.3 Screen Layout Partition Cropping + Regional OCR Content Extraction

[0482] according to Coordinate pair Cut and get Sub-images and perform OCR image content recognition on each sub-image to obtain the corresponding text , get text vectors through BERT , and finally get each Content .

[0483]

[0484] 2.3.4 Determining the User's Viewpoint

[0485] judge Whether the user triggers the mouse click event at the moment:

[0486] If a mouse click event is triggered, the mouse coordinates collected in 2.2.1 are determined through the spatial mapping relationship. Does it fall into the predefined An area in , inferring the page where the user's gaze is located .

[0487] If the mouse click event is not triggered, the dynamic boundary matrix obtained according to 2.2.2 and the above The cross-sectional area is calculated by , and the total area of ​​each intersection area ... .

[0488] Set the minimum cross-proportion threshold , filter out the values ​​greater than the threshold , sort from large to small and get the largest proportion , in order to determine the page where the user's sight is Get Corresponding layout theme .

[0489] 2.3.4 Determining the relevance between the assignment topic and the focus page topic

[0490] The assignment topic is semantically vectorized and combined with the text similarity calculation in 2.2.3 to determine whether the assignment topic is related to Is there any relevance?

[0491] 2.4 Overall calculation and analysis process

[0492] Every Time period, combined with the above 2.3 single screenshot analysis process to determine the relevance, if it is determined to be relevant, mark the user as focused on searching for job-related content in that time period, otherwise it is determined to be irrelevant and record the number of irrelevant searches + 1, the irrelevant search duration + .

[0493] 3. Comprehensive evaluation calculation

[0494] Parameter definition: is the weight of the number of desertions, is the weight of the absenteeism duration, is the desertion threshold, ; is the start time of the current job, is the end time of the current job, is the accuracy weight of the assignment.

[0495]

[0496] Level of concentration:

[0497] Levels are divided according to scores:

[0498] 90-100 points: Highly focused (excellent efficiency and self-management);

[0499] 70-89 points: Moderate concentration (occasionally distracted but not affecting results);

[0500] 50-69 points: Needs improvement (distractive behavior is obvious and may affect learning effect);

[0501] <50 points: Low concentration (immediate intervention required).

[0502] In summary, the above solution has the following beneficial effects:

[0503] 1. By integrating multi-source data such as mouse operation, camera visual analysis, eye behavior tracking, and motion trajectory smoothness, a multi-dimensional learning behavior monitoring system is constructed;

[0504] 2. Use Dlib's 68-point facial feature recognition model to accurately capture the dynamic changes of key feature points such as the user's eyes and mouth;

[0505] 3. Use motion trajectory analysis algorithms (including instantaneous acceleration calculation and stability index extraction) to effectively identify real learning behavior;

[0506] 4. Multimodal data fusion technology ensures that the system can distinguish between real learning behavior and cheating behavior (such as script operations, fake movements, etc.).

[0507] 5. There have been significant improvements in accuracy, anti-cheating capabilities and user experience, providing reliable technical support for online education quality assurance.

[0508] Corresponding to the above-mentioned learning monitoring method, the present invention also provides a learning monitoring device. Since the device embodiment of the present invention corresponds to the above-mentioned method embodiment, details not disclosed in the device embodiment can be referred to the above-mentioned method embodiment and will not be repeated in this invention.

[0509] Figure 5 A structural diagram of a learning monitoring device provided by an embodiment of the present disclosure is shown in FIG. Figure 5 As shown, the device includes:

[0510] Acquisition module 510, for acquiring facial image data and interaction data corresponding to the target user in response to the target user submitting the current job, wherein the interaction data includes screenshot data, mouse click data, and job accuracy. The facial image data and interaction data are user data of the target user during the current job;

[0511] A determination module 520 is configured to determine parameter information corresponding to at least one behavioral feature based on the plurality of feature points corresponding to the facial image data and the interaction data, the parameter information including the number and duration of the behavioral feature, wherein the at least one behavioral feature indicates that the target user is in an absentee state;

[0512] The evaluation module 530 is used to perform weighted calculation on the parameter information and the homework accuracy corresponding to at least one behavioral feature to determine a concentration score, which is used to evaluate the learning status of the target user.

[0513] In some embodiments of the present disclosure, the determination module is also used to preprocess multiple feature points to obtain boundary range parameters, which include multiple boundary point coordinates and center point coordinates; determine acceleration parameters based on the boundary range parameters; and determine the number and duration of behavioral features corresponding to each behavioral feature in at least one behavioral feature based on the boundary range parameters, acceleration parameters, facial image data, multiple feature points, and at least one item of interaction data.

[0514] In some embodiments of the present disclosure, the determination module is also used to normalize multiple feature points based on a preset screen size to determine an initial affine transformation matrix, where the initial affine transformation matrix includes an initial scaling ratio, an initial rotation angle, and an initial translation coordinate; based on multiple feature points, the initial affine transformation matrix is ​​updated through a relative depth ratio model to obtain a target affine transformation matrix; based on the eye distance and mouth-eye height corresponding to the multiple feature points, the avatar data is determined through the target affine transformation matrix, where the avatar data includes face width and face height; based on the avatar data and the preset screen size, the coordinates of multiple boundary points and the center point coordinates are determined.

[0515] In some embodiments of the present disclosure, the determination module is also used to determine the coordinates of multiple avatar center points within a preset time period based on the boundary range parameters; calculate the displacement difference of the multiple avatar center point coordinates according to the preset time length, and obtain the instantaneous acceleration corresponding to each preset time length in the preset time period; based on the instantaneous acceleration corresponding to each preset time length, determine the acceleration parameters within the preset time period, and the acceleration parameters include the acceleration mean, acceleration standard deviation, and low acceleration ratio.

[0516] In some embodiments of the present disclosure, the determination module is further configured to determine that the user state within a preset time period is a dozing state, and increase the number of dozing times by a count value, and increase the dozing time by a preset time period, when at least one of the eye aspect ratio and acceleration standard deviation corresponding to multiple feature points meets a preset condition; when the acceleration mean is less than the acceleration mean threshold and the low acceleration ratio is greater than the low acceleration ratio threshold, determine that the user state within the preset time period is a daze state, and increase the number of daze times by a count value, and increase the daze time by a preset time period; in the current boundary range parameter When the coordinates of at least one boundary point are not within the historical reference range corresponding to the boundary point coordinates, the user state within the preset time period is determined to be an improper state, and the number of improper times is increased by a count value, and the improper time period is increased by the preset time period. The current boundary range parameter is obtained by filtering the boundary range parameters corresponding to the dazed state and the dozing state in the boundary range parameters; according to the preset time period, when the text similarity between the target area in the screenshot data and the current job text data is less than or equal to the similarity threshold, the number of irrelevant times is increased by a count value, and the irrelevant time period is increased by the preset time period.

[0517] In some embodiments of the present disclosure, the preset conditions include: the eye aspect ratio is zero; the acceleration standard deviation is greater than the drowsiness threshold; the number of times the acceleration standard deviation reaches a peak within a preset time period exceeds a preset number.

[0518] In some embodiments of the present disclosure, the determination module is further configured to increase the number of misalignment times by a count value and increase the misalignment duration by a first duration if no facial image data is collected for a first duration during the current operation.

[0519] In some embodiments of the present disclosure, the evaluation module is also used to determine the corresponding area of ​​the mouse click data in the screenshot data as the target area when there is a time overlap between the mouse click data in the interaction data and the screenshot data; when there is no time overlap between the mouse click data in the interaction data and the screenshot data, based on the screenshot data, determine the area where the overlap of the boundary range corresponding to the multiple screen partitions determined by the partition model of the screenshot data and the boundary range parameters meets the overlap condition as the target area.

[0520] It should be noted that the above explanation of the method embodiment is also applicable to the device of this embodiment, and the principles are the same, which is not limited in this embodiment.

[0521] Based on the above Figures 1 to 3 The method shown in FIG. 1 is a method for performing the above-mentioned steps. Accordingly, this embodiment further provides a computer program product, including a computer program, which implements the above-mentioned steps when executed by a processor. Figures 1 to 3 The method shown.

[0522] Based on the above Figures 1 to 3 The method shown in FIG. 1 is a method for performing the above-mentioned steps. Accordingly, this embodiment further provides a computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, the computer program can realize the above-mentioned steps. Figures 1 to 3 The method shown.

[0523] Based on this understanding, the technical solution of the present application can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (which can be a CD-ROM, USB flash drive, mobile hard disk, etc.), and includes a number of instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to execute the methods of various implementation scenarios of the present application.

[0524] like Figure 6 FIG. 1 is a schematic diagram of the hardware structure of an electronic device of the present invention, comprising:

[0525] at least one processor 601; and,

[0526] A memory 602 in communication with at least one of the processors 601; wherein,

[0527] The memory 602 stores instructions that can be executed by at least one of the processors. The instructions are executed by at least one of the processors to enable the at least one of the processors to perform the learning monitoring method described above.

[0528] Figure 6 A processor 601 is taken as an example.

[0529] The electronic device may further include an input device 603 and a display device 604 .

[0530] The processor 601, the memory 602, the input device 603 and the display device 604 may be connected via a bus or other means, with the bus connection being used as an example in the figure.

[0531] The memory 602 is a non-volatile computer-readable storage medium that can be used to store non-volatile software programs, non-volatile computer executable programs, and modules, such as the program instructions / modules corresponding to the review content generation method in the embodiment of the present application, for example, Figure 1 4 . The processor 601 executes the non-volatile software programs, instructions and modules stored in the memory 602 to perform various functional applications and data processing, that is, to implement the learning monitoring method in the above embodiment.

[0532] The memory 602 may include a program storage area and a data storage area, wherein the program storage area may store an operating system, an application program required for at least one function; the data storage area may store data created according to the use of the review content generation method, etc. In addition, the memory 602 may include a high-speed random access memory, and may also include a non-volatile memory, such as at least one disk storage device, a flash memory device, or other non-volatile solid-state storage device. In some embodiments, the memory 602 may optionally include a memory remotely arranged relative to the processor 601, and these remote memories may be connected to the device for executing the review content generation method via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.

[0533] The input device 603 can receive user clicks and generate signal inputs related to user settings and function controls of the review content generation method. The display device 604 can include a display device such as a display screen.

[0534] The one or more modules are stored in the memory 602 and, when executed by the one or more processors 601 , execute the learning monitoring method in any of the above method embodiments.

[0535] Optionally, the physical device may also include a user interface, a network interface, a camera, a radio frequency (RF) circuit, a sensor, an audio circuit, a Wi-Fi module, and the like. The user interface may include a display screen and an input unit such as a keyboard. Optional user interfaces may also include a USB interface and a card reader interface. Optionally, the network interface may include a standard wired interface or a wireless interface (such as a Wi-Fi interface).

[0536] Those skilled in the art will understand that the above-mentioned physical device structure provided in this embodiment does not constitute a limitation on the physical device, and may include more or fewer components, or a combination of certain components, or different component arrangements.

[0537] The storage medium may also include an operating system and a network communication module. The operating system is a program that manages the hardware and software resources of the physical device, supporting the execution of information processing programs and other software and / or programs. The network communication module is used to enable communication between components within the storage medium, as well as with other hardware and software within the physical information processing device.

[0538] Through the description of the above implementation methods, those skilled in the art can clearly understand that this application can be implemented by means of software plus the necessary general hardware platform, or by hardware. By applying the solution of this embodiment, compared with the current existing technology, this embodiment optimizes the original code of the first object to obtain the corrected code corresponding to the original code; determines the code difference between the original code and the corrected code; based on the code difference, determines the corrected code level of the corrected code, and pushes the corrected code level to the first object. The code quality assessment based on the code difference reduces the influence of subjective factors and improves the accuracy and efficiency of the code quality assessment.

[0539] It should be noted that, in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprises," "comprising," or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, article, or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or device. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of other identical elements in the process, method, article, or device comprising the element.

[0540] The foregoing is merely a list of specific embodiments of the present application, intended to enable those skilled in the art to understand and implement the present application. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present application. Therefore, the present application is not limited to the embodiments described herein, but is intended to conform to the broadest scope consistent with the principles and novel features of the present application.

Claims

1. A learning monitoring method, characterized in that: The method comprises: In response to a target user submitting a current job, acquiring facial image data and interaction data corresponding to the target user, the interaction data including screenshot data, mouse click data, and job accuracy, the facial image data and interaction data being user data of the target user during the current job; Determining parameter information corresponding to at least one behavioral feature based on the plurality of feature points corresponding to the facial image data and the interaction data, the parameter information including the number of behavioral features and duration, the at least one behavioral feature indicating that the target user is in an absentee state; Performing weighted calculation on the parameter information corresponding to the at least one behavioral feature and the homework accuracy to determine a concentration score, wherein the concentration score is used to evaluate the learning status of the target user; The determining of parameter information corresponding to at least one behavioral feature based on the plurality of feature points corresponding to the facial image data and the interaction data includes: Preprocessing the multiple feature points to obtain boundary range parameters, where the boundary range parameters include multiple boundary point coordinates and a center point coordinate; the boundary range parameters are the dynamic activity range of the target user in front of the screen; determining an acceleration parameter based on the boundary range parameter; Determining the number of behavioral features and the duration of each behavioral feature in the at least one behavioral feature based on at least one of the boundary range parameter, the acceleration parameter, the facial image data, the plurality of feature points, and the interaction data; The preprocessing of the plurality of feature points to obtain boundary range parameters includes: Based on a preset screen size, normalize the multiple feature points to determine an initial affine transformation matrix, where the initial affine transformation matrix includes an initial scaling ratio, an initial rotation angle, and an initial translation coordinate; Based on the multiple feature points, the initial affine transformation matrix is ​​updated using a relative depth scale model to obtain a target affine transformation matrix; Based on the eye distance and mouth-eye height corresponding to the multiple feature points, determining the head portrait data through the target affine transformation matrix, the head portrait data including face width and face height; Determining the coordinates of the plurality of boundary points and the coordinates of the center point based on the avatar data and the preset screen size; The determining of the number of behavioral features and the duration of each behavioral feature in the at least one behavioral feature based on at least one of the boundary range parameter, the acceleration parameter, the facial image data, the plurality of feature points, and the interaction data includes: According to the preset time length, when the text similarity between the target area in the screenshot data and the current job text data is less than or equal to the similarity threshold, the number of irrelevant times is increased by a count value, and the irrelevant time length is increased by the preset time length; The method further comprises: In a case where the mouse click data in the interaction data and the screenshot data overlap in time, determining a corresponding area of ​​the mouse click data in the screenshot data as the target area; In the case that there is no time overlap between the mouse click data in the interactive data and the screenshot data, based on the screenshot data, the area where the overlap between the multiple screen partitions determined by the partition model for the screenshot data and the boundary range corresponding to the boundary range parameters meets the overlap condition is determined as the target area.

2. The method according to claim 1, characterized in that The determining of the acceleration parameter based on the boundary range parameter includes: Based on the boundary range parameters, determining the coordinates of multiple avatar center points within a preset time period; Calculating the displacement difference of the center point coordinates of the multiple portraits according to the preset time length to obtain the instantaneous acceleration corresponding to each preset time length in the preset time period; Based on the instantaneous acceleration corresponding to each preset time length, the acceleration parameters within the preset time period are determined, where the acceleration parameters include the acceleration mean, the acceleration standard deviation, and the low acceleration ratio.

3. The method according to claim 2, characterized in that The determining, based on at least one of the boundary range parameter, the acceleration parameter, the facial image data, the plurality of feature points, and the interaction data, of the number of behavioral feature times and duration corresponding to each of the at least one behavioral feature, further includes at least one of the following: If at least one of the eye aspect ratios and the acceleration standard deviation corresponding to the multiple feature points meets a preset condition, determining that the user state within the preset time period is a dozing state, increasing the number of dozing times by a count value, and increasing the dozing time by the preset time period; If the acceleration mean is less than the acceleration mean threshold and the low acceleration ratio is greater than the low acceleration ratio threshold, determine that the user state within the preset time period is a daze state, increase the number of daze times by one, and increase the daze duration by the preset time period; When at least one boundary point coordinate in the current boundary range parameter is not within the historical reference range corresponding to the boundary point coordinate, the user state within the preset time period is determined to be an improper state, and the number of improper times is increased by a count value, and the improper time period is increased by the preset time period. The current boundary range parameter is obtained by filtering the boundary range parameters corresponding to the dazed state and the dozing state in the boundary range parameters.

4. The method according to claim 3, characterized in that The preset conditions include: the eye aspect ratio is zero; the acceleration standard deviation is greater than the dozing threshold; the number of times the acceleration standard deviation reaches a peak within a preset time period exceeds a preset number; The method further includes: if the facial image data is not collected for a first period of time during the current operation, increasing the number of times of misalignment by a count value and increasing the misalignment duration by the first period of time.

5. A learning monitoring device, characterized in that: The device comprises: an acquisition module, configured to acquire facial image data and interaction data corresponding to the target user in response to the target user submitting the current job, the interaction data including screenshot data, mouse click data, and job accuracy, the facial image data and interaction data being user data of the target user during the current job; a determination module configured to determine parameter information corresponding to at least one behavioral feature based on a plurality of feature points corresponding to the facial image data and the interaction data, the parameter information including the number and duration of the behavioral feature, wherein the at least one behavioral feature indicates that the target user is in an absentee state; an evaluation module, configured to perform weighted calculation on the parameter information corresponding to the at least one behavioral feature and the homework accuracy rate to determine a concentration score, wherein the concentration score is used to evaluate the learning status of the target user; The determining of parameter information corresponding to at least one behavioral feature based on the plurality of feature points corresponding to the facial image data and the interaction data includes: Preprocessing the multiple feature points to obtain boundary range parameters, where the boundary range parameters include multiple boundary point coordinates and a center point coordinate; the boundary range parameters are the dynamic activity range of the target user in front of the screen; determining an acceleration parameter based on the boundary range parameter; Determining the number of behavioral features and the duration of each behavioral feature in the at least one behavioral feature based on at least one of the boundary range parameter, the acceleration parameter, the facial image data, the plurality of feature points, and the interaction data; The preprocessing of the plurality of feature points to obtain boundary range parameters includes: Based on a preset screen size, normalize the multiple feature points to determine an initial affine transformation matrix, where the initial affine transformation matrix includes an initial scaling ratio, an initial rotation angle, and an initial translation coordinate; Based on the multiple feature points, the initial affine transformation matrix is ​​updated using a relative depth scale model to obtain a target affine transformation matrix; Based on the eye distance and mouth-eye height corresponding to the multiple feature points, determining the head portrait data through the target affine transformation matrix, the head portrait data including face width and face height; Determining the coordinates of the plurality of boundary points and the coordinates of the center point based on the avatar data and the preset screen size; The determining of the number of behavioral features and the duration of each behavioral feature in the at least one behavioral feature based on at least one of the boundary range parameter, the acceleration parameter, the facial image data, the plurality of feature points, and the interaction data includes: According to the preset time length, when the text similarity between the target area in the screenshot data and the current job text data is less than or equal to the similarity threshold, the number of irrelevant times is increased by a count value, and the irrelevant time length is increased by the preset time length; Wherein, the device further comprises: a first target area determination module, configured to, when there is temporal overlap between the mouse click data in the interaction data and the screenshot data, determine a corresponding area in the screenshot data of the mouse click data as the target area; The second target area determination module is used to determine, based on the screenshot data, a region where the overlap between the multiple screen partitions determined for the screenshot data through the partition model and the boundary range corresponding to the boundary range parameters meets the overlap condition as the target area when there is no time overlap between the mouse click data in the interaction data and the screenshot data.

6. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method according to any one of claims 1 to 4 is implemented.

7. An electronic device comprising a storage medium, a processor, and a computer program stored in the storage medium and executable on the processor, wherein: When the processor executes the computer program, the method according to any one of claims 1 to 4 is implemented.

Citation Information

Patent Citations

  • Fatigue driving judging method based on wearable devices

    CN109984762A

  • Online classroom student concentration evaluation method and system based on multi-feature fusion

    CN114663734A

  • Remote intelligent monitoring method based on picture similarity comparison

    CN115393615A