Sight tracking network-based crowd concentration real-time evaluation method

By using the eye-tracking network Cognition_Net, combined with a multi-scale feature fusion module and a bidirectional long short-term memory network, a non-intrusive, objective, and real-time assessment of the attention level of multiple people was achieved, solving the problem of accurate evaluation of teaching level in large communication venues.

CN121482852APending Publication Date: 2026-02-06CHONGQING LIANGJIANG JIANIAN HEALTH EXAMINATION CENT CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511616265.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-11-06
Publication Date
2026-02-06

AI Technical Summary

Technical Problem

Existing attention detection methods cannot achieve non-invasive, objective, and real-time assessment of the attention levels of multiple people, especially lacking accurate methods for evaluating teaching performance in large communication settings.

Method used

A gaze-tracking network-based approach is adopted, which uses the lightweight gaze-tracking network Cognition_Net, combined with a multi-scale feature fusion module and a bidirectional long short-term memory network, to capture crowd videos using a camera, perform face detection and attention signal analysis, and calculate the group's level of focus.

Benefits of technology

It enables objective and real-time assessment of the attention levels of multiple people, provides accurate evaluation of teaching quality, and reduces subjective arbitrariness.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121482852A_ABST
    Figure CN121482852A_ABST
Patent Text Reader

Abstract

The invention discloses a crowd concentration degree real-time evaluation method based on a sight tracking network. The method comprises the following steps: establishing a light-weight sight tracking network CognittionNet; collecting a video I (t) containing a crowd, and performing face detection Fi and storing; sending each Fi into a sight tracking network CognittionNet to obtain an attention signal Li (t); counting the standard deviation sigmai of Li (t) every 30 seconds to obtain the concentration degree sigmamean of one group; therefore, the objective evaluation of the attention levels of multiple persons is realized by only adopting one camera.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application belongs to the field of deep learning, and particularly relates to a crowd concentration real-time evaluation method based on a gaze tracking network. BACKGROUND

[0002] Attention is a kind of mental activity of human beings, and is a direction and concentration to a certain object; attention is the ability of directing and concentrating the mental activity to a certain object.

[0003] The attention detection method includes a questionnaire method, a method based on electroencephalogram signals and a method based on virtual reality, and these methods cannot be quantified, have high hardware requirements, have complex hardware requirements, and there is no non-invasive, objective and real-time detection method which is not affected by the head posture.

[0004] At present, in some large communication places such as a classroom or a meeting, the presentation level of a presenter is related to the attraction to the audience, but the existing evaluation method often has subjective randomness. SUMMARY

[0005] Figure 1 The flowchart of the patent; Figure 2 The device diagram of the detection system; Figure 3 The overall structure of the gaze tracking network; Figure 4 The wavelet attention feature fusion module (WA) is introduced; Figure 5 The attention signal; Figure 6 The multi-person attention detection; Figure 7 The multi-person attention signal. The purpose of the application is to realize the objective detection of the multi-person attention level by using one camera, so as to provide an accurate evaluation method for the presentation level. The application will be described in detail below in combination with the drawings.

[0006] The crowd concentration real-time evaluation method based on the gaze tracking network, the flow is as shown in Figure 1 The following steps are included: Step one, a light gaze tracking network Cognition_Net is established; Step two, a video collection containing a crowd, face detection F i and storage are performed; Step three, each F i is sent into the gaze tracking network Cognition_Net to obtain an attention signal Li (t) ; Step 4: Statistical analysis every 30 seconds L i (t) Standard deviation σ i To obtain the level of focus of a group. σ mean .

[0007] Check the system's hardware and software, such as Figure 2 As shown, the setup includes a PC and a webcam. The PC hardware consists of an Intel(R) Core(TM) i9-9900K, an NVIDIA GTX1080Ti, 16GB of DDR4 RAM, and a 500GB hard drive; the software environment is an Ubuntu 16.04 operating system, Python 3.7.4, and the PyTorch 1.4.0 deep learning framework. The webcam operates at a frame rate of 60fps with a resolution of 1280*720.

[0008] The camera was placed on the top edge of the computer monitor, and the participants looked at the screen from a distance of about 60cm, in a natural sitting posture.

[0009] The following explains the specific implementation of the four steps: Step one involves building a lightweight eye-tracking network, Cognition_Net, which specifically employs the following two steps: 1) A multi-scale feature fusion module (EfficientMLCA) and a bidirectional long short-term memory network (Bi-LSTM) are embedded into the mid-to-high-level feature outputs of the original backbone network ResNeSt26d, and a wavelet attention feature fusion module (WA) is introduced to form the Cognition_Net network. The overall network structure is as follows: Figure 3 As shown, a wavelet attention feature fusion module (WA) is introduced, as follows: Figure 4 As shown; 2) The Cognition_Net network was trained using the ETH-XGaze dataset.

[0010] Step two involves video capture and face detection of a crowd. F i And store it, specifically using the following steps: 1) For an input video containing a crowd I(t) The face of each person is detected using a non-convex mixture norm error coding algorithm. F i ; 2) These F i Store as a list { F 1,F 2 ,F 3 ,…}

[0011] In step three, for each F i The signal is fed into the gaze-tracking network Cognition_Net to obtain attention signals. L i (t) The following methods are specifically adopted: Face image F i Input Cognition_Net, output gaze direction { pitch i (t), yaw i (t)}, calculate according to the following formula A i (t) :

[0012] in, a and r Influenced by factors such as the focal length, working distance, aberrations, and spherical aberration of the imaging system, here a Take 1.2, r Take 0.6.

[0013] After investigation A i (t) Kalman filter (noise amplitude estimated at 3.5*10) -4 )get B i (t) After investigation B i (t) Attention signal is obtained through wavelet smoothing (Mallat wavelet decomposition, 7 decomposition levels, wavelet basis coif5). L i (t) The result is as follows: Figure 5 As shown.

[0014] Statistics are collected every 30 seconds in step four. L i (t) Standard deviation σ i To obtain the level of focus of a group. σ mean Multiple people captured from a single camera and facial recognition, such as... Figure 6 As shown, the attention signals of multiple people are as follows Figure 7 As shown.

[0015] The specific steps are as follows: 1) Statistics every 30 seconds L i (t) Standard deviation σ i ; 2) Adopt σ i To represent each person's attention level; 3) Calculate multiple people within a group σ i mean σ mean ; 4) Utilize σ mean To measure the level of focus of a group.

Claims

1. A method for real-time assessment of crowd attention based on eye-tracking networks, characterized in that, Includes the following steps: Step 1: Build a lightweight eye-tracking network, Cognition_Net; Step two involves video capture and face detection of a crowd. F i And store; Step 3, for each F i The signal is fed into the gaze-tracking network Cognition_Net to obtain attention signals. L i (t) ; Step 4: Statistical analysis every 30 seconds L i (t) Standard deviation σ i To obtain the level of focus of a group. σ mean .

2. The method for real-time assessment of crowd attention based on eye-tracking networks according to claim 1, characterized in that, Step one involves building a lightweight eye-tracking network, Cognition_Net, which specifically follows these steps: Step 1: Embed a multi-scale feature fusion module (EfficientMLCA) and a bidirectional long short-term memory network (Bi-LSTM) into the mid-to-high-level feature outputs of the original backbone network ResNeSt26d, and introduce a wavelet attention feature fusion module (WA) to form the Cognition_Net network. Step 2: Train the Cognition_Net network using the ETH-XGaze dataset.

3. The method for real-time assessment of crowd attention based on eye-tracking networks according to claim 1, characterized in that, Step two involves video capture and face detection of a crowd. F i And store it, specifically using the following steps: Step 1: Input a video containing a crowd. I(t) Detect each person's face F i ; Step two, take these F i Store as a list { F 1, F 2, F 3,…}.

4. The method for real-time assessment of crowd attention based on eye-tracking networks according to claim 1, characterized in that, In step three, for each F i The signal is fed into the gaze-tracking network Cognition_Net to obtain attention signals. L i (t) The following methods are specifically adopted: Face image F i Input Cognition_Net, output gaze direction { Pitch i (t), yaw i (t) }, calculate according to the following formula A i (t) : After investigation A i (t) Filtering B i (t) After investigation B i (t) Smoothing to obtain attention signals L i (t) .

5. A method for real-time assessment of crowd attention based on an eye-tracking network according to claim 1, characterized in that, Step 4: Statistics every 30 seconds L i (t) Standard deviation σ i To obtain the level of focus of a group. σ mean The specific steps are as follows: Step 1: Statistical analysis every 30 seconds L i (t) Standard deviation σ i ; Step two, using σ i To represent each person's attention level; Step 3: Calculate the number of people within the group σ i The mean σ mean ; Step 4, using σ mean To measure the level of focus of a group.