A method, system and memory for assessing concentration

By setting up abnormal behavior and attention scoring models, video frames of surveillance recordings are analyzed, calculated, and converted into normal distribution intervals. This solves the problems of time-consuming, costly, and lack of standards in existing technologies for classroom attention evaluation, and enables real-time and accurate evaluation of students' attention.

CN116189036BActive Publication Date: 2026-01-02DATA SPACE RES INST

Patent Information

Application Number
CN202211663909.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-12-23
Publication Date
2026-01-02
Estimated Expiration
2042-12-23

AI Technical Summary

Technical Problem

Existing methods for assessing classroom attention are time-consuming, costly, lack unified standards and real-time performance, and machine vision-based detection technologies lack a scientific evaluation system, making it impossible to accurately determine students' attention levels.

Method used

By setting up abnormal behavior and attention scoring models, video frames of surveillance recordings are analyzed to calculate the attention scores of the monitored objects and convert them into normal distribution intervals. Combined with get out of class start and end discrimination methods, real-time norms are generated in real time, which are suitable for classroom video analysis in different scenarios.

Benefits of technology

It enables objective and accurate evaluation of students' concentration, dynamically reflects changes in concentration, is applicable to various scenarios, reduces manual observation and data processing work, improves evaluation efficiency, and reduces sampling and experimental errors of traditional methods.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116189036B_ABST
    Figure CN116189036B_ABST
Patent Text Reader

Abstract

The present application relates to the field of classroom attention evaluation, especially to a concentration evaluation method, system and memory. The present application firstly judges the abnormal behavior of the monitoring object based on the video frame, and brings the abnormal behavior into the set concentration evaluation model to calculate the concentration score of the monitoring object; then the concentration scores of the monitoring objects in the same group are converted into normal distribution, and the concentration level of the monitoring object is judged according to the position of each concentration score on the normal distribution. The present application realizes the calculation of the concentration score of the monitoring object in different environments; then the monitoring objects which need to be uniformly evaluated and compared with each other are analyzed together according to the concentration score, so as to obtain the uniform evaluation of the concentration of the monitoring objects in different scenes, so that the concentration evaluation method is more widely applicable, and the evaluation result is more reliable.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of classroom attention evaluation, in particular to a concentration evaluation method, system and memory. BACKGROUND

[0002] Classroom concentration refers to the degree to which students concentrate their attention on classroom content during class. Attention is the direction and concentration of individual psychological activity to a certain object, and is the basis of cognitive activity. Attention is closely related to working memory, information processing speed, etc., and can predict intelligence and academic performance to a certain extent. Students with high attention levels are less likely to be disturbed by irrelevant information, while students with learning difficulties often have low attention levels. According to the diagnostic criteria given by the American Diagnostic and Statistical Manual of Mental Disorders, Fifth Edition (DSM-5), long-term inattention in children may be related to ADHD (Attention Deficit Hyperactivity Disorder).

[0003] Traditional concentration and attention measurement methods are mostly in the form of questionnaires and scales, which require students, parents or teachers to fill in. This method has many problems: first, the filling in of paper quality forms requires time and manpower, and it is difficult to conduct large sample testing; second, it is not suitable for repeated testing and cannot reflect the dynamic changes in students' concentration levels; third, students' answers will be affected by social desirability, and the results lack validity. Existing concentration detection technology based on machine vision can collect and identify students' concentration behavior indicators, but the evaluation of the results is too subjective, lacks a unified measurement standard, and cannot determine the level of students' concentration, which lacks practical guiding significance for parents and schools. SUMMARY

[0004] In order to solve the defects of the attention evaluation means in the prior art, the present application provides a concentration evaluation method, which is more widely applicable and has more reliable evaluation results.

[0005] The concentration evaluation method provided by the present application comprises the following steps:

[0006] S1, setting an abnormal behavior and concentration score model, video frame analysis is performed on the monitoring video, and the concentration score of the monitoring object is calculated in combination with the concentration score model;

[0007] The concentration score model is:

[0008]

[0009] Wherein, C n is the concentration score of the monitoring object n, I is the number of abnormal behaviors, i is the ordinal number, 1≦i≦I; a i represents the weight of the i-th abnormal behavior, t (i,n)denotes the duration of the n-th monitoring object showing the i-th abnormal behavior, T (i,n) denotes t (i,n) corresponding total monitoring duration;

[0010] S2, calculate the mean μ and the standard deviation σ of the concentration scores of the monitoring objects in the same group, convert the concentration scores of the monitoring objects in the same group into a normal distribution, and divide the normal distribution into a plurality of normal distribution intervals, wherein the normal distribution intervals correspond to the set concentration levels one by one; the monitoring objects in the same group refer to all the monitoring objects that need to be evaluated together;

[0011] S3, determine the concentration level of the monitoring objects in the group according to the position of the normal distribution interval corresponding to the concentration score of each monitoring object.

[0012] Preferably, for judging the concentration of students in the classroom, first, the class-on and class-off are distinguished in step S1, and then the concentration score of the monitoring object is calculated based on the video frames of the class-on time and the concentration score model; the class-on and class-off distinguishing includes the following steps:

[0013] SB1, extract the video frame at time t from the monitoring video, detect the video frame using a target detection model, obtain the students in the video frame, and set a monitoring box, wherein the monitoring box and the position of the students in the video frame correspond one by one;

[0014] SB2, calculate the connected area s(t) of the monitoring box and the effective area S(t) of the monitoring area at time t according to the video frame; the effective area S(t) of the monitoring area is the circumscribed rectangle area surrounding all the monitoring boxes;

[0015] SB3, determine whether S(t) / S0≧T2 is established; if yes, let λ=λ1 and d(t)=s(t) / S(t); if not, let λ=λ2 and d(t)=S(t) / S0; λ is a set parameter, λ1 and λ2 are both set constants, and 0<λ1<λ2<1; T2 is a set second threshold value;

[0016] SB4, calculate the corrected personnel density d(t)', d(t)'=λd(t-1)'+(1-λ)d(t); determine whether d(t)'≧T1 is established; if yes, determine that time t is in the class-on state; otherwise, determine that time t is in the class-off state; T1 is a set first threshold value.

[0017] Preferably, 0.3≦T1<T2≦0.8.

[0018] Preferably, S2 specifically includes the following steps:

[0019] S21, calculate the mean μ and the standard deviation σ of the concentration scores of the monitoring objects in the same group, and convert the concentration scores of the monitoring objects in the same group into a normal distribution;

[0020]

[0021]

[0022] wherein M is a set of same-group monitoring objects, i.e. a set of all monitoring objects that are sorted together for the concentration score, |M| represents the number of monitoring objects in the set M;

[0023] S22, dividing the normal distribution into five normal distribution intervals: [0, μ-1.8σ), [μ-1.8σ, μ-0.6σ), [μ-0.6σ, μ+0.6σ), [μ+0.6σ, μ+1.8σ), [μ+1.8σ, 100); the distribution of the five normal distribution intervals corresponds to the concentration level “low”, “lower”, “medium”, “higher” and “high”.

[0024] Preferably, the camera cruising track contains K cruising points, the camera time of each cruising point is t0, the monitoring period is T0=Q×K×t0, and K n represents a set of cruising points, when the camera is at any one of the K n cruising points, the monitoring object n is within the monitoring range of the camera;

[0025] For the monitoring object n, the calculation formula of t i and T i in a monitoring period is as follows:

[0026]

[0027] r (q,k,n,j,i) represents the number of continuous video frames in which the monitoring object n is in the i-th abnormal behavior, which is photographed by the camera for the j-th time at the k-th cruising point in the q-th cruising; J (q,k,n,i) represents the number of times in which the monitoring object n is in the i-th abnormal behavior, which is photographed by the camera at the k-th cruising point in the q-th cruising; m represents the number of continuous video frames in 1 second, represents the rounding up of r (q,k,n,j,i) / m;

[0028]

[0029] y (q,k,n,i) represents a binary number, when the monitoring object n is in the i-th abnormal behavior, which is monitored by the camera for the q-th time at the k-th cruising point, y (q,k,n,i) takes 1; otherwise, y (q,k,n,i) takes 0.

[0030] Preferably, the abnormal behavior includes head abnormal behavior, and the judgment condition of the head abnormal behavior is that the head posture yaw angle is greater than or equal to a set yaw threshold value or the head posture pitch angle is greater than or equal to a set pitch threshold value; the head posture yaw angle and the head posture pitch angle of the monitoring object are calculated by the following formulas:

[0031] δ x = γ x - (θ h - 180°) ; δ y = θ v + γ y

[0032]

[0033] δ x represents the horizontal error of the head posture rotation angle, and δ y represents the vertical error of the head posture rotation angle.

[0034] θ h represents the horizontal angle of the center point of the camera shooting direction, and θ v represents the vertical angle of the center point of the camera shooting direction; γ x represents the actual horizontal angle difference between the target point m and the center point o on the video frame, and γ y represents the actual vertical angle difference between the target point m and the center point o; the target point m is the center point of the target face on the video frame, and the target face is the face of the monitoring object.

[0035] f h represents the initial value of the head posture yaw angle obtained by analyzing the face in the video frame through the head posture algorithm, and f v represents the initial value of the head posture pitch angle obtained by analyzing the face in the video frame through the head posture algorithm.

[0036] Preferably:

[0037] When x < w / 2, γ x = a x - β x

[0038] When x ≥ w / 2, γ x = a x + β x

[0039] γ y = a y + β y

[0040]

[0041]

[0042] w and h represent the length and height of the bounding rectangle of the video frame respectively, (x, y) and (w / 2, h / 2) represent the coordinates of the target point m and the center point o on the video frame respectively, a x represents the horizontal angle error of the target point m to the center point o on the video frame, a y represents the vertical angle error of the target point m to the center point o on the video frame;

[0043] β x represents the error compensation of the target point m (x, y) in the horizontal direction, β y represents the error compensation of the target point m (x, y) in the vertical direction;

[0044] k h represents the error value of the camera field of view angle and the video frame field of view angle in the horizontal direction; k v represents the error value of the camera field of view angle and the video frame field of view angle in the vertical direction; k dh represents the error component of the diagonal direction field of view angle of the video frame in the horizontal direction; k dv represents the error component of the diagonal direction field of view angle of the video frame in the vertical direction.

[0045] Preferably, S1 is followed by S4: sorting the concentration scores of the same group of monitoring objects from high to low, and the concentration score sorting of the monitoring object n is recorded as R n , and calculating the percentile position of the monitoring object; the percentile position of the monitoring object n is recorded as P n ;

[0046] P n = 100 - (100R n - 50) / |M|

[0047] Wherein, M is a set of monitoring objects of the same group, that is, a set of all monitoring objects which are sorted by concentration scores together, and |M| represents the number of monitoring objects in the set M.

[0048] The application also provides a concentration evaluation system and a memory, which provide a carrier for the above-mentioned concentration evaluation method.

[0049] The application provides a concentration evaluation system, which comprises a storage module and a processor, the storage module stores a computer program, the processor is connected with the storage module, and the processor is used for executing the computer program to realize the concentration evaluation method.

[0050] The application provides a memory, which is characterized by storing a computer program, and the computer program is used for realizing the concentration evaluation method when being executed.

[0051] The present application has the advantages of:

[0052] (1) The concentration evaluation method provided by the present application collects students' concentration data through big data, establishes real-time norms (average value and standard deviation of samples), and compares students' concentration with the concentration levels of other students in the same time period in real time. The position of the student in the grade group objectively and accurately evaluates the student's concentration in a certain period. This method not only overcomes the shortcomings of traditional scale establishment, such as high cost and lack of timeliness due to the inability to update the norm, but also improves the existing video collection technology which lacks a scientific evaluation system. Through this method, the school and parents can push the student's concentration in real time, which helps parents and schools to understand the student's concentration and the changes in the student's psychological state, so as to help parents better understand the student's psychological state and help schools better and more targeted teaching design and intervention. The present application calculates the concentration score based on the monitoring video in real time, dynamically generates real-time norms, and can effectively avoid the disadvantages of traditional norms that are not updated in time and cannot reflect the real position of students' concentration level in the same group.

[0053] (2) The present application mainly proposes a method for real-time judgment of class start and end state based on classroom video frames from the perspective of classroom video analysis. The method can be organically combined with student attendance systems, student classroom behavior analysis systems and other application systems.

[0054] (3) The class start and end discrimination method used in the present application calculates the personnel density based on the connected area s(t) of the student monitoring box in the video frame and the effective area S(t) of the monitoring area. The present application does not need to know the total number of people in the scene when calculating the personnel density, so it has better scene adaptability. The method of calculating personnel density based on the connected area s(t) of the student monitoring box and the effective area S(t) of the monitoring area proposed by the present application can be used not only in the conventional classroom teaching scene, but also in dance classrooms, drama classrooms and movie theaters for real-time personnel density analysis and calculation. The class start and end discrimination method proposed by the present application can effectively and timely discriminate class start and end according to the changes in classroom video content, providing a stable state signal for real-time data collection and subsequent data analysis. The revised personnel density obtained by updating the momentum personnel density in the present application can stably discriminate class start and end, significantly improving the stability and robustness of the personnel density index.

[0055] (4) The concentration score model adopted by the present application can evaluate the concentration of the monitored object in combination with various abnormal behaviors, and the evaluation result has high reliability, which is beneficial to improve the concentration evaluation efficiency and reduce a large amount of manual observation and data arrangement work. The weight of the abnormal behavior in the present application is set by an expert, and the abnormal behavior recognition can be obtained through a machine learning recognition model. Through the innovative fusion of expert evaluation and machine learning, the concentration is quickly and accurately evaluated. The present application can evaluate the concentration of each monitored object based on the continuous video frames parsed from the monitoring video, avoiding the sampling error caused by the traditional method and the experimental error caused by directly collecting data from a non-real environment.

[0056] (5) The abnormal behavior in the present application can be flexibly set, and the concentration score model is constructed based on the deduction principle (deduction of abnormal behavior score). The concentration monitoring method in the present application is more in line with the concentration evaluation in the actual classroom scene, and is also applicable to various different scenes, and has a wide application range.

[0057] (6) In the present application, when analyzing the head abnormal behavior, the center point of the target face on the video frame is taken as the target point, the head posture turning angle error is calculated according to the relative position of the target point and the center point of the video frame, the error is calculated for each target face, and the error is corrected according to the position of each target face in the video frame. The deflection angle of the target face in the camera monitoring picture can be accurately calculated, and more accurate head posture recognition results can be obtained when the head posture is analyzed based on this, and the accuracy of head abnormal behavior judgment is ensured.

[0058] (7) In the present application, the inconsistency between the camera field of view angle and the video frame field of view angle is fully considered. Firstly, the position of the target point is compensated for error, and then the angle difference between the target point and the center point is calculated according to the compensated target point, and then the head posture turning angle is corrected, so as to ensure the accuracy of error calculation, and further ensure the accuracy of head posture recognition. BRIEF DESCRIPTION OF DRAWINGS

[0059] Figure 1 It is a flow chart for the class starting and ending discrimination method;

[0060] Figure 2 It is a schematic diagram of the monitoring frame connected area and the monitoring area effective area;

[0061] Figure 3 It is a schematic diagram of the camera cruising point shown in the embodiment;

[0062] Figure 4 It is a schematic diagram of the student position distribution in the embodiment;

[0063] Figure 5 It is a head posture yaw angle and a head posture pitch angle calculation flow chart;

[0064] Figure 6 A head pose algorithm flowchart.

[0065] Figure 7 A concentration evaluation method flowchart. DETAILED DESCRIPTION

[0066] The concentration evaluation method provided by the application first judges the abnormal behavior of the monitored object based on the video frame, and brings the abnormal behavior into a set concentration evaluation model to calculate the concentration score of the monitored object; then the concentration scores of the monitored objects in the same group are converted into a normal distribution, and the concentration level of the monitored object is judged according to the position of each concentration score on the normal distribution.

[0067] In the application, first, frame extraction is performed based on the monitoring video of the space, that is, the monitoring area, so as to realize the monitoring of the abnormal behavior of each monitored object and the calculation of the concentration score, and the calculation of the concentration score of the monitored object in different environments is realized; then the monitored objects that need to be uniformly evaluated and compared with each other are analyzed together according to the concentration score, so as to obtain the uniform evaluation of the concentration of the monitored objects in different scenes, so that the application scope of the concentration evaluation method is wider, and the evaluation result is more reliable.

[0068] When the application is applied to evaluate the concentration of students, the concentration score can only be calculated based on the performance of the students in the classroom, so first, the video frames in the class time need to be screened for abnormal behavior analysis.

[0069] A class start and end discrimination method

[0070] The class start and end discrimination method provided by the embodiment combines area calculation with personnel density, and judges whether the class is in session according to the area calculation, so as to be more suitable for different classroom situations.

[0071] In the embodiment, first, the monitoring video of the monitoring area is acquired, frame extraction is performed on the monitoring video, and one frame of video frame at each time is acquired for analysis.

[0072] In the embodiment, first, the video frame is detected by a target detection model to obtain the students in the video frame, and a monitoring box is set for the detected students, the monitoring box and the position of the detected students in the video frame are one-to-one corresponding. The connected area of the monitoring box is denoted as s(t), and the area of the circumscribed rectangle of all monitoring boxes is denoted as the effective area S(t) of the monitoring area. The connected area s(t) of the monitoring box can be obtained through OpenCV software.

[0073] The effective area S(t) of the monitoring area is the area of the circumscribed rectangle of all monitoring boxes, and specifically, the monitoring box can be set as a rectangular box, and then:

[0074] S(t) = W' x H'

[0075] W' = max(x 12 ,x 22 ,…,x i2 ,…,x I2 )-min(x 11 ,x 21 ,…,x i1 ,…,x I1 )

[0076] H' = max(y 12 ,y 22 ,…,y i2 ,…,y I2 )-min(y 11 ,y 21 ,…,y i1 ,…,y I1 )

[0077] Wherein, W' and H' are length and width of effective area of monitoring area respectively.

[0078] (x i1 ,y i1 ) is the coordinate of the upper left corner of the i-th monitoring frame on the video frame, (x i2 ,y i2 ) is the coordinate of the lower right corner of the i-th monitoring frame on the video frame; 1 ≦ i ≦ I, I is the number of students detected in the video frame, that is, the number of monitoring frames.

[0079] Referring to Figure 1 , the class start and end determination method proposed in the embodiment includes the following steps SB1-SB4:

[0080] SB1, extract the video frame at time t from the monitoring video of the monitoring area, detect the video frame by using the target detection model, obtain the students in the video frame, and set the monitoring frame, which is one-to-one corresponding to the position of the students in the video frame. The target detection model is a common technical means in the field of vision. Combined with Figure 2 It can be seen that, by analyzing the video frame by the target detection model, even if students 1, 5, 11 and 12 are missed, the real-time personnel density obtained by the final calculation is still guaranteed because the present application is based on area calculation of personnel density, thereby ensuring the reliability of the final class start and end determination.

[0081] SB2, calculate the connected area s(t) of the monitoring frame and the effective area S(t) of the monitoring area at time t according to the video frame, as Figure 2 shown.

[0082] SB3. Determine whether S(t) / S0≧T2 holds true; if yes, let λ=λ1, d(t)=s(t) / S(t); if no, let λ=λ2, d(t)=S(t) / S0; S0 represents the area of ​​the video frame, specifically S0=W×H, where W and H are the length and width of the video frame, respectively. λ2>λ1.

[0083] SB4. Calculate the corrected personnel density d(t)', d(t)'=λd(t-1)'+(1-λ)d(t); determine whether d(t)'≧T1 holds true; if yes, then determine that time t is the state of class; otherwise, determine that time t is the state of get out of class dismissal.

[0084] A focus rating model

[0085] This invention calculates the attention score of the monitored object by combining a attention scoring model with specified abnormal behaviors; the higher the attention score, the higher the attention of the monitored object.

[0086] The focus rating model is as follows:

[0087]

[0088] Among them, C n Let I be the focus score of the monitored object n, i be the number of abnormal behaviors, and i be the ordinal number, where 1 ≤ i ≤ I; a i t represents the weight of the i-th abnormal behavior. (i,n) T represents the duration of the i-th abnormal behavior exhibited by the monitored object n. (i,n) Indicates t (i,n) The corresponding total monitoring duration. In this implementation, a i Manually labeled.

[0089]

[0090] r (q,k,n,j,i) K represents the number of consecutive video frames captured by the camera at the k-th patrol point during the q-th patrol, showing the monitored object n exhibiting the i-th abnormal behavior; Q represents the maximum number of patrols, i.e., the number of times the camera cycles within a monitoring period; one cycle of the camera refers to the trajectory of the camera passing through each patrol point exactly once; K n This represents the set of cruise points, when the camera is at point K. n At any of the patrol points, the monitored object n is within the camera's monitoring range;

[0091] J (q,k,n,i) This represents the number of times the camera captures the monitored object n exhibiting the i-th abnormal behavior at the k-th patrol point during the q-th patrol; m represents the number of consecutive video frames per second. Indicates that for r (q,k,n,j,i)rounded up.

[0092]

[0093] y (q,k,n,i) indicates a binary number, when the camera monitors the monitoring object n to perform the i-th abnormal behavior at the k-th cruise point in the q-th cruise, then y (q,k,n,i) takes 1; otherwise, y (q,k,n,i) takes 0.

[0094] Embodiment 1: t i and T i Calculation process example

[0095] Assume that students in a classroom are divided into three rows, the cruise points 1-6 of the camera are as shown in Figure 3 , the path of one cruise is "1-2-3-4-5-6" or "6-5-4-3-2-1". The camera stays at each cruise point for t0, a monitoring period contains Q cruises, and the camera takes m frames of images per second;

[0096] In this embodiment, K = 6, and it is assumed that:

[0097] t0 = 1 minute, Q = 2, m = 30, and n = 2.

[0098] The monitoring period T0 = Q x K x t0 = 12 minutes.

[0099] Assume that the shooting areas corresponding to the cruise points 1, 2, 3, and 4 are as shown in Figure 4 the circles A1, A2, A3, and A4.

[0100] In a monitoring period, the camera shooting time at each cruise point is as follows:

[0101] Cruise point 1: 1st minute + 12th minute

[0102] Cruise point 2: 2nd minute + 112th minute

[0103] Cruise point 3: 3rd minute + 10th minute

[0104] Cruise point 4: 4th minute + 9th minute

[0105] Cruise point 5: 5th minute + 8th minute

[0106] Cruise point 6: 6th minute + 7th minute

[0107] Referring to Figure 4, the position B1 of student A is only in the shooting area corresponding to the cruise point 1, the position B12 of student B is in the shooting area corresponding to the cruise point 1 and the cruise point 2, and the position B234 of student C is in the shooting area corresponding to the cruise point 2, the cruise point 3 and the cruise point 4. The yawning monitoring results of students A, B and C by the camera in the monitoring period are shown in Table 1.

[0108] Table 1: Yawning monitoring statistics of A, B and C

[0109]

[0110] In this embodiment, let r (q,k,n,j,哈欠) , let.

[0111] It can be known from Table 1 that:

[0112] K 甲 ={cruise point 1};

[0113] r (q=1,k=1,甲,j=1,哈欠) =10 frames; r (q=1,k=1,甲,j=2,哈欠) =21 frames; r (q=1,k=1,甲,j=3,哈欠) =6 frames;

[0114] r (q=2,k=1,甲,j=1,哈欠) =10 frames; r (q=2,k=1,甲,j=2,哈欠) =21 frames; r (q=2,k=1,甲,j=3,哈欠) =6 frames;

[0115] K 乙 ={cruise point 1; cruise point 2};

[0116] r (q=1,k=1,乙,j=1,哈欠) =21 frames; r (q=1,k=1,乙,j=2,哈欠) =6 frames;

[0117] r (q=1,k=2,乙,j=1,哈欠) =25 frames; r (q=1,k=2,乙,j=2,哈欠) =31 frames;

[0118] r (q=2,k=1,乙,j=1,哈欠) =20 frames; r (q=2,k=2,乙,j=1,哈欠) =6 frames;

[0119] K 丙 ={cruise point 2; cruise point 3; cruise point 4}.

[0120] r (q=1,k=2,丙,j=1,哈欠) =25 frames; r (q=1,k=2,丙,j=2,哈欠) =31 frames;

[0121] r (q=1,k=3,丙,j=1,哈欠) =11 frames; r (q=1,k=4,丙,j=1,哈欠) =26 frames;

[0122] r (q=2,k=2,丙,j=1,哈欠) = 6 frames; r (q=2,k=3,丙,j=1,哈欠) = 31 frames; r (q=2,k=4,丙,j=1,哈欠) = 6 frames;

[0123] It can be seen that, in the current monitoring period, there are:

[0124]

[0125]

[0126] For student B, there are:

[0127]

[0128]

[0129] For student C, there are:

[0130]

[0131]

[0132] In this embodiment, when the concentration score model is used to calculate the concentration of the monitoring object, the monitoring video is first obtained, the video frame is analyzed, and the duration t (i,n) and the corresponding total monitoring time T (i,n) of each abnormal behavior of each monitoring object are obtained; then the video frame analysis data t (i,n) and T (i,n) of the monitoring object are substituted into the concentration score model to obtain the concentration score of each monitoring object.

[0133] The abnormal behavior can be set as head abnormal behavior, yawning, expression abnormality, and sitting posture abnormality.

[0134] The abnormal behavior can be identified by a recognition model obtained by machine learning. For example, a neural network recognition model can be constructed to identify the input monitoring video as a continuous video frame output, and the recognition model is also used to label whether each video frame contains the corresponding abnormal behavior. In this way, the t (i,n) and T (i,n) corresponding to the abnormal behavior can be calculated according to the analysis of the recognition model on the monitoring video.

[0135] The abnormal behavior can also be identified by setting the corresponding judgment condition, and after analyzing the monitoring video into continuous video frames, the judgment condition is combined with the monitoring object on each video frame to determine whether the corresponding abnormal behavior exists.

[0136] A head abnormal behavior identification method

[0137] In the embodiment, a head abnormal behavior judgment condition is set, and when the target face on the video frame meets the judgment condition, it is judged that the monitoring object on the video frame has head abnormal behavior. The target face is the face of the monitoring object.

[0138] The judgment condition of the abnormal behavior is that the head posture yaw angle is greater than or equal to a set yaw threshold, or the head posture pitch angle is greater than or equal to a set pitch threshold.

[0139] The existing head posture algorithm is used to analyze the captured image to obtain the head posture rotation angle of the target face in the captured image. The head posture rotation angle includes a head posture yaw angle f h and a head posture pitch angle f v . When analyzing the head posture rotation angle, the existing head posture algorithm only considers the captured image and does not consider the camera distortion, so there is an error in the finally obtained head posture rotation angle.

[0140] In the embodiment, when calculating the head posture yaw angle f and the head posture pitch angle f , the influence of camera distortion on the video frame is considered, so that the horizontal error δ x and the vertical error δ y of the head posture rotation angle caused by camera distortion are calculated, and the head posture rotation angle obtained by the head posture algorithm is corrected by combining the horizontal error δ x and the vertical error δ y of the head posture rotation angle, so as to ensure the accuracy of the head posture rotation angle recognition.

[0141] Referring to Figure 5 , the head posture correction method based on camera distortion compensation in the embodiment specifically includes the following steps SA1-SA3.

[0142] SA1, obtaining the horizontal field of view angle fov h , the vertical field of view angle fov v , the component fov dh of the diagonal direction field of view angle in the horizontal direction, and the component fov dv of the diagonal direction field of view angle in the vertical direction of the video frame; obtaining the horizontal field of view angle FOV h and the vertical field of view angle FOV v of the camera providing the video frame; and the video frame is obtained by analyzing the monitoring video from the camera.

[0143] SA2, analyzing the video frame by the head posture algorithm to obtain the head posture yaw angle f h and the head posture pitch angle f v of the target face; obtaining the coordinates of the target point m and the center point o on the video frame, and calculating the head posture rotation angle horizontal error δx and vertical error δ y The target point is the center point of the target face on the video frame, and specifically, the center point of the circumscribed rectangle or regular polygon of the target face on the video frame can be selected as the target point.

[0144] SA3, combined with f h , f v , δ x and δ y Calculate the corrected head posture rotation angle; the corrected head posture rotation angle includes the corrected head posture yaw angle and the corrected head posture pitch angle

[0145]

[0146] represents the corrected head posture yaw angle, that is, the horizontal rotation angle; represents the corrected head posture pitch angle, that is, the vertical rotation angle; f h represents the head posture yaw angle obtained by analyzing the face in the video frame through the head posture algorithm, f v represents the head posture pitch angle obtained by analyzing the face in the video frame through the head posture algorithm.

[0147] Head posture rotation angle horizontal error δ x and vertical error δ y The calculation formula is:

[0148] δ x = γ x - (θ h - 180°) (2-1)

[0149] δ y = θ v + γ y (2-2)

[0150] θ h represents the horizontal angle of the center point of the camera shooting direction, θ v represents the vertical angle of the center point of the camera shooting direction; γ x represents the actual horizontal angle difference between the target point m and the center point o on the video frame, γ y represents the actual vertical angle difference between the target point m and the center point o; the target point m is the center point of the target face on the video frame.

[0151] When x < w / 2, then: γ x = a x - β x

[0152] When x >= w / 2, then: γ x = a x + β x

[0153] γ y = a y + β y

[0154] a x denotes the horizontal angular error of the target point m to the center point o on the video frame, a y denotes the vertical angular error of the target point m to the center point o on the video frame; β x denotes the error compensation of the target point m in the horizontal direction, β y denotes the error compensation of the target point m in the vertical direction.

[0155]

[0156]

[0157] w and h denote the length and height of the circumscribed rectangle of the video frame, respectively, (x, y) and (w / 2, h / 2) denote the coordinates of the target point m and the center point o on the video frame, respectively; fov h denotes the horizontal field of view of the video frame, fov v denotes the vertical field of view of the video frame.

[0158]

[0159]

[0160] k h denotes the error value of the camera field of view and the video frame field of view in the horizontal direction; k v denotes the error value of the camera field of view and the video frame field of view in the vertical direction; k dh denotes the error component of the diagonal direction field of view of the video frame in the horizontal direction; k dv denotes the error component of the diagonal direction field of view of the video frame in the vertical direction.

[0161] k h = (FOV h - fov h ) / 2

[0162] k v = (FOV v - fov v ) / 2

[0163] k dh = (fov dh - fovh ) / 2-k h

[0164] k dv =k v -(FOV v -fov dv ) / 2

[0165] fov h denotes the horizontal field of view of the video frame, fov v denotes the vertical field of view of the video frame, fov dh denotes the component of the diagonal direction field of view of the video frame in the horizontal direction, fov dv denotes the component of the diagonal direction field of view of the video frame in the vertical direction; FOV h denotes the horizontal field of view of the camera, FOV v denotes the vertical field of view of the camera; FOV h and FOV v are intrinsic parameters of the camera, which can be read from the camera manual.

[0166] Head pose algorithm

[0167] There are many head pose algorithms in the prior art, among which, neural network is commonly used to implement the head pose algorithm. In the embodiment, a head pose recognition model is obtained through autonomous learning of a machine model, so as to analyze the video frame through the head pose recognition model and obtain the head pose rotation angle of each face.

[0168] With reference to Figure 6 , in the embodiment, the steps of implementing the head pose algorithm are as follows.

[0169] SC1, constructing a neural network and labeling samples; the samples are video frames, the labeling labels of the samples are the head pose rotation angles of each face in the video frames, the input of the neural network is the video frame, and the output of the neural network is the head pose rotation angle of each face in the video frame;

[0170] SC2, learning the labeled samples by the neural network to train the network parameters, and taking the trained neural network as a head pose recognition model;

[0171] SC3, inputting the target video frame into the head pose recognition model to obtain the head pose rotation angle of each face in the target video frame output by the head pose recognition model, i.e. the head pose yaw angle f h and the head pose pitch angle f v of the face.

[0172] In the embodiment, the labeling labels of the samples are artificially labeled.

[0173] Specifically, a plurality of labeled samples are constructed in SC1, and the labeled samples are divided into training samples and test samples; in SC2, the neural network is iterated multiple times, each time the neural network learns a specified number of training samples to update the network parameters, and then calculates the head posture angle of each face in the set number of test samples as the model label of the test sample through the updated neural network, and calculates the model label accuracy of the neural network according to the difference between the model label of the test sample and the labeled label; if the model label accuracy of the neural network is lower than the set value, the neural network learns new training samples; in this way, through parameter iteration, until the model label accuracy of the neural network reaches the set value, the neural network is used as a head posture recognition model.

[0174] A concentration evaluation method

[0175] Reference Figure 7 The concentration evaluation method provided by the embodiment comprises the following steps S1-S3.

[0176] S1, set the abnormal behavior and the concentration score model, and analyze the video frames of the monitoring video during the class time, and calculate the concentration score of the monitoring object by combining the concentration score model.

[0177] S2, calculate the average value μ and the standard deviation σ of the concentration scores of the same group of monitoring objects, convert the concentration scores of the same group of monitoring objects into a normal distribution, and divide the normal distribution into a plurality of normal distribution intervals, and the normal distribution interval corresponds to the set concentration level one by one.

[0178] The same group of monitoring objects refers to all monitoring objects that need to be evaluated together; for example, when the method is applied to a school, the monitoring video is obtained based on the class to calculate the concentration scores of each student as the monitoring object, and then the students of the same grade in the school are taken as the same group of monitoring objects for unified analysis.

[0179] In this step, the concentration scores of the same group of monitoring objects can be divided into five normal distribution intervals: [0, μ-1.8σ), [μ-1.8σ, μ-0.6σ), [μ-0.6σ, μ+0.6σ), [μ+0.6σ, μ+1.8σ), [μ+1.8σ, 100); the five normal distribution intervals correspond to the concentration levels “low”, “lower”, “medium”, “higher” and “high”.

[0180] Wherein, μ and σ are the mean and standard deviation of the normal distribution, respectively:

[0181]

[0182]

[0183] M is a set of monitoring objects in the same group, that is, a set of all monitoring objects for which concentration score ranking is performed together, and |M| represents the number of monitoring objects in the set M.

[0184] S3, determining the concentration level of the monitoring object in the group according to the position of the normal distribution interval corresponding to the concentration score of each monitoring object, that is, obtaining whether the concentration score of the monitoring object is "low", "lower", "medium", "higher" or "high".

[0185] S4, ranking the concentration scores of the monitoring objects in the same group from high to low, and the concentration score ranking of the monitoring object n is denoted as R n , calculating the percentile position of the monitoring object; the percentile position of the monitoring object n is denoted as P n , P n =100-(100R n -50) / |M|.

[0186] The percentile position P n calculated in the embodiment conforms to the form of psychological assessment and can be directly used for psychological assessment.

[0187] Example 2: Concentration assessment in a classroom of a school

[0188] In this embodiment, the concentration of students in a classroom of a school is assessed by using the concentration assessment method provided by the present application. In this embodiment, when the class start and end time is determined, the parameters are set as: λ1=0.95; λ2=0.98, T1=0.4, T2=0.5.

[0189] In this embodiment, the abnormal behaviors are set as: head abnormal behavior, yawning, expression abnormality and sitting posture abnormality. Among them, the head abnormal behavior is determined by the head abnormal behavior recognition method described above; the yawning, expression abnormality and sitting posture abnormality are determined by the neural network-based recognition model of machine learning.

[0190] In this embodiment, the concentration level of the students is assessed by a plurality of subject teachers and class leaders, that is, "low", "lower", "medium", "higher" or "high" as the actual concentration level.

[0191] In this embodiment, the number of students tested is 311, which belongs to 6 classes. First, the concentration scores of the students in each class are calculated based on the monitoring video of the class time, and then the concentration scores of the students in the 6 classes are converted into a normal distribution. According to the interval position of the concentration score of each student on the normal distribution, the concentration level of the student is determined, that is:

[0192] If the concentration score is in the interval [0, μ-1.8σ), the concentration level of the student is evaluated as "low";

[0193] If the concentration score is in the interval [μ-1.8σ, μ-0.6σ), the concentration level evaluation value of the student is "low";

[0194] If the concentration score is in the interval [μ-0.6σ, μ+0.6σ), the concentration level evaluation value of the student is "medium";

[0195] If the concentration score is in the interval [μ+0.6σ, μ+1.8σ), the concentration level evaluation value of the student is "high";

[0196] If the concentration score is in the interval [μ+1.8σ, 100), the concentration level evaluation value of the student is "very high".

[0197] In this embodiment, the final statistical result is that the concentration level evaluation values of 308 persons are consistent with the actual concentration levels. It can be seen that the concentration evaluation result obtained by the present application is highly consistent with the actual performance.

[0198] The above is only a preferred embodiment of the present application, and is not used to limit the present application. Any modification, equivalent replacement and improvement within the spirit and principle of the present application shall be included in the protection scope of the present application.

Claims

1. A concentration evaluation method characterized by, Comprise the following steps: S1, set abnormal behavior and concentration score model, video frame analysis is carried out to monitoring video, and the concentration score of monitoring object is calculated in combination with concentration score model; The concentration score model is: wherein C n is the concentration of the monitoring object n, I is the number of abnormal behaviors, i is the ordinal number, 1 ≦ i ≦ I; a i represents the weight of the i-th abnormal behavior, t (i,n) represents the duration of the i-th abnormal behavior performed by the monitoring object n, T (i,n) represents t (i,n) corresponds to the total monitoring duration; S2, the average value μ and the standard deviation σ of the concentration score of the same group of monitoring objects are calculated, the concentration score of the same group of monitoring objects is converted into normal distribution, and the normal distribution is divided into a plurality of normal distribution intervals, and the normal distribution interval corresponds to the set concentration level one by one;The same group of monitoring objects refers to all monitoring objects that need to be evaluated together; S3, the concentration level of the monitoring object in the group is judged according to the position of the normal distribution interval corresponding to the concentration score of each monitoring object.

2. The concentration evaluation method according to claim 1, wherein For judging the concentration of students in class, step S1 first distinguishes between class and non-class based on video frames, and then calculates the concentration score of monitoring objects based on the video frames of the class time combined with the concentration score model;The class and non-class discrimination comprises the following steps: SB1, the video frame at time t is extracted from the monitoring video, the video frame is detected by using the target detection model, the students in the video frame are obtained, and the monitoring box is set, the monitoring box and the position of the student in the video frame correspond one by one; SB2, the connected area s(t) of the monitoring box and the effective area S(t) of the monitoring area at time t are calculated according to the video frame;The effective area S(t) of the monitoring area is the circumscribed rectangle area surrounding all monitoring boxes; SB3, judge whether S(t) / S0≧T2 is established;Yes, let λ=λ1, d(t)=s(t) / S(t);No, let λ=λ2, d(t)=S(t) / S0;λ is a set parameter, λ1 and λ2 are both set constants, 0<λ1<λ2<1;T2 is a set second threshold;S0 represents the area of the video frame; SB4, calculate the corrected personnel density d(t)',d(t)'=λd(t-1)'+(1-λ)d(t);Judge whether d(t)'≧T1 is established;Yes, it is judged that t time is class state;Otherwise, it is judged that t time is non-class state;T1 is a set first threshold.

3. The concentration evaluation method according to claim 2, wherein 0.3≦T1<T2≦0.

8.

4. The concentration evaluation method according to claim 1, wherein S2 specifically comprises the following steps: S21, the average value μ and the standard deviation σ of the concentration score of the same group of monitoring objects are calculated, and the concentration score of the same group of monitoring objects is converted into normal distribution; Wherein, M is the set of the same group of monitoring objects, that is, the set of all monitoring objects for concentration score sorting together, |M| represents the number of monitoring objects in set M; S22, the normal distribution is divided into five normal distribution intervals: [0, μ-1.8σ), [μ-1.8σ, μ-0.6σ), [μ-0.6σ, μ+0.6σ), [μ+0.6σ, μ+1.8σ), [μ+1.8σ, 100);The five normal distribution intervals correspond to concentration levels "low", "lower", "medium", "higher" and "high" respectively.

5. The concentration evaluation method according to claim 1, wherein Let the camera cruise track contain K cruise points, the camera time of each cruise point is t0, the monitoring period is T0=Q×K×t0, let K n represent a set of cruise points, when the camera is at any one of the K n cruise points, the monitoring object n is within the monitoring range of the camera. For a monitoring object n, its t i and T i The calculation formula is as follows: r (q,k,n,j,i) represents the number of continuous video frames in which the monitoring object n is in the i-th abnormal behavior at the j-th time of shooting at the k-th cruise point in the q-th cruise of the camera; J (q,k,n,i) represents the number of times in which the monitoring object n is in the i-th abnormal behavior at the k-th cruise point in the q-th cruise of the camera; m represents the number of continuous video frames in 1 second, represents the number of times in which the monitoring object n is in the i-th abnormal behavior at the k-th cruise point in the q-th cruise of the camera; m represents the number of continuous video frames in 1 second, (q,k,n,j,i) is rounded up to m. y (q,k,n,i) y represents a binary number, when the camera monitors the monitoring object n to perform the i-th abnormal behavior at the k-th cruise point in the q-th cruise, then y (q,k,n,i) takes 1; otherwise, y (q,k,n,i) takes 0.

6. The concentration evaluation method according to claim 5, wherein The abnormal behavior includes head abnormal behavior, and a judgment condition of the head abnormal behavior is that a head posture yaw angle is greater than or equal to a set yaw threshold value or a head posture pitch angle is greater than or equal to a set pitch threshold value; the head posture yaw angle and the head posture pitch angle of the monitoring object are calculated by the following formulas: delta x = gamma x - (theta h - 180°); delta y = theta v + gamma y δ x horizontal error of the head pose rotation angle, δ y vertical error of the head pose rotation angle; θ h denotes the horizontal angle of the center point of the camera shooting direction, θ v denotes the vertical angle of the center point of the camera shooting direction; γ x denotes the actual horizontal angle difference between the target point m and the center point o on the video frame, γ y denotes the actual vertical angle difference between the target point m and the center point o; the target point m is the center point of the target face on the video frame, and the target face is the face of the monitoring object; f h represents a yaw angle initial value of a head pose obtained by analyzing a face in a video frame through a head pose algorithm, f v represents a pitch angle initial value of a head pose obtained by analyzing a face in a video frame through a head pose algorithm.

7. The concentration evaluation method of claim 6, characterized in that: When x < w / 2, then: γ x = a x - β x When x >= w / 2, then: γ x = a x + β x gamma y = a y + beta y w and h represent the length and height of the bounding rectangle of the video frame, respectively, (x, y) and (w / 2, h / 2) represent the coordinates of the target point m and the center point o on the video frame, respectively, a x represents the horizontal angle error of the target point m to the center point o on the video frame, a y represents the vertical angle error of the target point m to the center point o on the video frame; β x denotes the error compensation of the target point m(x,y) in the horizontal direction, β y denotes the error compensation of the target point m(x,y) in the vertical direction; k h represents the error value of the camera field of view and the video frame field of view in the horizontal direction;k v represents the error value of the camera field of view and the video frame field of view in the vertical direction;k dh represents the error component of the diagonal direction field of view of the video frame in the horizontal direction;k dv represents the error component of the diagonal direction field of view of the video frame in the vertical direction.

8. The concentration evaluation method according to claim 1, wherein S1 is followed by S4: ranking the concentration scores of the monitoring objects in the same group from high to low, the concentration score of monitoring object n is ranked as R n , and calculating the percentile position of the monitoring object; the percentile position of monitoring object n is denoted as P n . P n = 100 - (100R n - 50) / |M| Wherein, M is a set of same-group monitoring objects, i.e. a set of all monitoring objects which are sorted together for the concentration score, |M| represents the number of monitoring objects in the set M.

9. A concentration assessment system characterized by comprising: The application further provides a computer readable storage medium, which stores a computer program.

10. A memory, comprising: The application further provides a computer readable storage medium, which stores a computer program.

Citation Information

Patent Citations

  • Classroom concentration degree detection method and device

    CN111931585A

  • Online classroom student concentration evaluation method and system based on multi-feature fusion

    CN114663734A

Cited By

  • Student classroom concentration intelligent evaluation system and method based on multi-modal data

    CN122336858A