Digital teaching interaction method

By installing tracking sensor hardware in the classroom to analyze students' eye movement status in real time, the problem of single cognitive state evaluation dimensions in the digital teaching system is solved, and multi-dimensional quantitative analysis of students' cognitive state and improvement of teaching interactivity is achieved.

CN120199121AInactive Publication Date: 2025-06-24CHANGCHUN INST OF ELECTRONIC TECH
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510663412.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-22
Publication Date
2025-06-24
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The existing digital teaching system has the problem of single cognitive state assessment dimensions and lack of quantitative analysis of students' current status.

Method used

By installing dual-mode camera arrays and 9-axis inertial sensor hardware in the classroom, students' eye movement status are tracked in real time and uploading data to a cloud server for in-depth analysis to obtain students' classroom concentration.

Benefits of technology

It provides multi-dimensional quantitative analysis of students' cognitive status, helps teachers better understand students' learning status, enhance teaching interactivity, improve students' learning experience, and provide objective data support for teaching evaluation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120199121A_ABST
    Figure CN120199121A_ABST
Patent Text Reader

Abstract

The invention discloses a digital teaching interaction method, which comprises the following steps: installing tracking sensor hardware in a classroom, and tracking the eyeball motion state of each student in real time in the classroom process; uploading the collected data to a cloud server in real time, and deeply analyzing the classroom concentration degree of each student in real time by a distributed computing unit embedded in the cloud server; acquiring the classroom concentration degree of all the students according to the classroom concentration degree of each student acquired in real time; the cloud server wirelessly transmits the classroom concentration degree of all the students to the classroom display terminal in real time, and a teacher performs teaching reminding according to the real-time classroom concentration degree displayed in the teacher display terminal, so that the real-time classroom concentration degree is kept at a certain threshold value, and deep semi-automatic interaction is completed.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of digital education, and in particular to a digital teaching interaction method. Background Art

[0002] With the rapid development of Internet and multimedia technologies, digital teaching resources have been widely applied. However, existing digital teaching methods often suffer from problems such as insufficient interactivity and poor user experience, resulting in limited teaching effects.

[0003] For example: In the existing classroom digital teaching process, the evaluation dimension of the cognitive state is single. Traditional digital systems only rely on the correct rate of answering questions and lack quantitative analysis of the current state. Summary of the Invention

[0004] The purpose of this part is to outline some aspects of the embodiments of the present invention and briefly introduce some preferred embodiments. Simplifications or omissions may be made in this part, as well as in the abstract and title of the present application, to avoid obscuring the purpose of this part, the abstract, and the title. Such simplifications or omissions shall not be used to limit the scope of the present invention.

[0005] In view of the problems existing in the above-mentioned existing digital teaching systems, the present invention is proposed.

[0006] Therefore, the technical problem solved by the present invention is to solve the problem that the existing digital teaching system has a single evaluation dimension of the cognitive state and lacks quantitative analysis of the current state.

[0007] To solve the above technical problem, the present invention provides the following technical solution: A digital teaching interaction method, comprising the following steps: S1: Install tracking sensor hardware in the classroom to track the eye movement states of each student during the class in real time; wherein, the installed tracking sensor hardware is: a dual-mode camera array, including a visible light camera, a near-infrared camera, a 9-axis inertial sensor, and a distributed computing unit; S2: Upload the collected data to the cloud server in real time, and the distributed computing unit embedded in the cloud server deeply analyzes the class concentration of each student in real time; S3: Obtain the class concentration of all students based on the class concentration of each student obtained in real time; S4: The cloud server wirelessly transmits the class concentration of all students to the classroom display terminal in real time, and the teacher gives teaching reminders according to the real-time class concentration displayed on the teacher display terminal to keep it at a certain threshold, completing deep semi-automatic interaction.

[0008] As a preferred embodiment of the digital teaching interaction method of the present invention, the specific steps for real-time tracking of the eye movement states of each student during the class are as follows: H1: After real-time image acquisition, perform image preprocessing and output the preprocessed image; H2: Based on the preprocessed image, extract pupil features through a pupil feature extraction network; H3: Based on the preprocessed image, perform gaze estimation through a gaze estimation model; H4: Based on the preprocessed image, obtain temporal attention and spatial attention; H5: Establish a concentration evaluation model for feature extraction to obtain the proportion of the main fixation residence time, the gaze transfer frequency, and the pupil diameter change rate.

[0009] As a preferred embodiment of the digital teaching interaction method of the present invention, in step H1, the specific steps for preprocessing the acquired image are as follows: Q1: Perform multispectral image fusion: Input the RGB image (denoted as I rgb ), and the IR image (denoted as I ir ) for fusion. The specific fusion output formula is as follows:

[0010]

[0011] where (x, y) are the image pixel coordinates; 0.6 and 0.4 are the weighting coefficients; Y is the grayscale image output after fusion; Q2: Perform adaptive illumination compensation on the grayscale image output after fusion; Q3: Perform dynamic ROI extraction on the compensated grayscale image and output the face bounding box coordinates and the eye sub-region coordinates.

[0012] where:

[0013] Face bounding box coordinates:

[0014] ;

[0015] Eye sub-region coordinates:

[0016] .

[0017] As a preferred embodiment of the digital teaching interaction method of the present invention, the cloud server performs real-time in-depth analysis of the class concentration of each student, and the steps for synchronously obtaining the class concentration of all students are as follows: E1: Obtain the proportion of the main fixation residence time, the gaze transfer frequency, and the pupil diameter change rate of each student; E2: Establish a concentration evaluation model, input each feature, and output the class concentration of the corresponding student; E3: Real-time and uniformly obtain the class concentration of each student; E4: Use the deviation formula to obtain the concentration deviation between each student and the standard, and obtain the maximum concentration deviation as the class concentration of the overall students.

[0018] As a preferred embodiment of the digital teaching interaction method of the present invention, specifically: the established concentration evaluation model is as follows:

[0019]

[0020] Among them, ε is the classroom concentration; α is the proportion of the fixation time of the main fixation point, and its unit is %; β is the saccade frequency, and its unit is Hz; γ is the pupil diameter change rate; 1.28, -0.44, and -ln2 are all adjustment constants.

[0021] As a preferred embodiment of the digital teaching interaction method of the present invention, specifically: the defined concentration standard is 1; then, the concentration deviation between each student and the standard is obtained by using the deviation formula, and the maximum concentration deviation is expressed as the classroom concentration of the whole students. The specific basis is as follows:

[0022]

[0023] Among them, ε1 is the classroom concentration of the first student; ε2 is the classroom concentration of the second student; ε r is the classroom concentration of the rth student; ‖‖2 is the function expression of the second-order norm.

[0024] As a preferred embodiment of the digital teaching interaction method of the present invention, specifically: the threshold in S4 is set to 0.64 or 0.641.

[0025] The present invention provides a digital teaching interaction method, which has the following beneficial effects:

[0026] It provides a quantitative analysis of the cognitive state of students, exceeding the traditional single-dimensional evaluation method;

[0027] Through real-time tracking and evaluation, it helps teachers better understand the learning status of students, so as to carry out effective teaching interventions;

[0028] It enhances the interactivity of teaching and improves the learning experience of students;

[0029] It provides objective data support for teaching evaluation, which helps to improve the teaching quality and effect. BRIEF DESCRIPTION OF THE DRAWINGS

[0030] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings. Among them:

[0031] Figure 1This is the overall method flowchart of the digital teaching interaction method provided by the present invention.

[0032] Figure 2 This is the method flowchart for real-time tracking of the eye movement states of each student during the classroom process provided by the present invention.

[0033] Figure 3 This is the method flowchart for preprocessing the collected images provided by the present invention.

[0034] Figure 4 This is the method flowchart for the cloud server to perform real-time in-depth analysis of the classroom concentration of each student and synchronously obtain the classroom concentration of all students provided by the present invention.

[0035] Figure 5 This is the overall algorithm flow of the dynamic ROI extraction technology provided by the present invention. Detailed implementation manners

[0036] To make the above objects, features, and advantages of the present invention more obvious and understandable, the following will describe the detailed implementation manners of the present invention with reference to the accompanying drawings of the specification. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0037] In the existing classroom digital teaching process, the cognitive state evaluation dimension is single. Traditional digital systems only rely on the answering accuracy rate and lack quantitative analysis of the current state.

[0038] Therefore, please refer to Figure 1 , the present invention provides a digital teaching interaction method, including the following steps:

[0039] S1: Install tracking sensor hardware in the classroom to real-time track the eye movement states of each student during the classroom process; among them, the installed tracking sensor hardware is: a dual-mode camera array, including a visible light camera (1920x1080@60fps) and a near-infrared camera (850nm LED + filter), a 9-axis inertial sensor (head pose tracking), and a distributed computing unit (NVIDIA Jetson Xavier);

[0040] Among them:

[0041] 1. Technical specifications of the dual-mode camera array

[0042] 1.1 Visible light camera (RGB module)

[0043] Sensor model: Sony IMX477 (1 / 2.3" back-illuminated CMOS);

[0044] Optical parameters:

[0045] Focal length: f = 8 mm (adjustable focus range ±15%);

[0046] F-number: F / 2.0 (low-light optimized);

[0047] Field of view: 78° × 64° (diagonal 94°);

[0048] Electronic characteristics:

[0049] Quantum efficiency: 62% @ 550 nm;

[0050] Full well capacity: 10400 e − ;

[0051] Read noise: 2.3 e − / pix;

[0052] 1.2 Near-infrared camera (NIR module)

[0053] Light source configuration:

[0054] Wavelength: 850 nm (VCSEL array, compliant with IEC 60825-1 Class 1 eye-safe standard);

[0055] Irradiance: 1.5 mW / cm2 @ 50 cm;

[0056] Filter characteristics:

[0057] Bandpass range: 840 - 860 nm;

[0058] Cut-off depth: OD6 @ 400 - 800 nm;

[0059] Transmittance: >92% @ 850 nm;

[0060] Sensor model: OmniVision OV9282 (global shutter);

[0061] Resolution: 1280 × 800 @ 120 fps;

[0062] Near-infrared sensitivity: 0.38 lux (typical value);

[0063] 1.3 Dual-mode synchronization mechanism

[0064] Hardware synchronization: μs-level trigger synchronization implemented through FPGA;

[0065] Exposure time alignment error: <50 μs;

[0066] Frame buffer queue depth: 8 frames (to prevent data loss);

[0067] Optical baseline: The distance between the two cameras d = 65 mm;

[0068] Depth resolution: ΔZ = (Z 2 / fd) × Δx;

[0069] At a distance of 2 m, ΔZ ≈ 15 mm;

[0070] 2. Detailed Explanation of 9-Axis Inertial Measurement Unit (IMU)

[0071] 2.1 Sensor composition, as shown in Table 1 below:

[0072] Table 1

[0073] 2.2 Sensor fusion algorithm

[0074] Where:

[0075] , θ, ψ: Roll / pitch / yaw angles;

[0076] m e : Geomagnetic field vector (about 50 μT);

[0077] 2.3 Performance indicators

[0078] Static error of attitude angle: <0.5° RMS;

[0079] Dynamic delay: <5 ms @ 200 Hz output rate;

[0080] 3. Technical parameters of the distributed computing unit

[0081] 3.1 NVIDIA Jetson Xavier NX

[0082] Computing core:

[0083] 6-core NVIDIA Carmel ARMv8.2 CPU @ 1.9 GHz;

[0084] 384-core Volta GPU + 48 Tensor Cores;

[0085] Memory architecture:

[0086] 8GB LPDDR4x @ 51.2 GB / s;

[0087] 16GB eMMC 5.1 storage;

[0088] Interface capabilities:

[0089] 2x MIPI CSI-2 (supporting 6 camera access);

[0090] 1x GbE + 2x USB3.1;

[0091] 3.2 Real-time computing performance, as shown in Table 2 below:

[0092] Table 2 Task type Algorithm Latency Power consumption Image preprocessing CLAHE + ROI extraction 8 ms / frame 3.2W Pupil tracking PupilNet inference 11 ms / frame 4.8W Gaze estimation GazeNet inference 14 ms / frame 5.1W

[0093] 3.3 Thermal design

[0094] Thermal design power (TDP): 15W (passive cooling mode);

[0095] Operating temperature range: -25°C to 80°C (meeting classroom environment requirements);

[0096] 4. System integration solution

[0097] 4.1 Mechanical structure design

[0098] Camera installation:

[0099] Pitch angle adjustment range: -30° to +45°;

[0100] Pan-tilt stepper motor accuracy: 0.072° / step;

[0101] Electromagnetic compatibility:

[0102] Pass FCC Part 15 Class B certification;

[0103] Radiated emission limit: <40 dBμV / m @ 3m;

[0104] 4.2 Power management

[0105] Input voltage: 12V DC (PoE power supply compatible);

[0106] Power consumption distribution:

[0107] Camera array: 7.2W (RGB: 4W, NIR: 3.2W);

[0108] IMU module: 0.3W;

[0109] Computing unit: dynamically adjustable (5 - 15W);

[0110] 5. Classroom deployment parameters, as shown in Table 3 below:

[0111] Table 3 Parameter Typical value Description Installation height 2.8-3.2m Suspended from the classroom ceiling Coverage 6×8m² Maximum monitoring area of a single device Student density ≤ 30 people Ensure unobstructed line of sight Lighting conditions 300 - 1500 lux Support automatic gain control

[0112] This hardware configuration can achieve sub - centimeter line - of - sight tracking accuracy in a typical classroom environment.

[0113] It should be noted that the series of tracking sensors installed in the classroom are all creative applications of existing conventional electronic devices, and no redundant elaboration will be made here.

[0114] S2: Upload the collected data to the cloud server in real - time, and the distributed computing unit embedded in the cloud server deeply analyzes the classroom concentration of each student in real - time;

[0115] S3: Obtain the classroom concentration of all students based on the classroom concentration of each student obtained in real - time;

[0116] S4: The cloud server wirelessly transmits the classroom concentration of all students to the classroom display terminal in real - time. The teacher makes teaching reminders based on the real - time classroom concentration displayed on the teacher's display terminal to keep it within a certain threshold, completing deep semi - automatic interaction.

[0117] Furthermore, referring to Figure 2 , the specific steps for real - time tracking of the eye movement state of each student during the classroom process are as follows:

[0118] H1: After real - time image acquisition, perform image pre - processing and output the pre - processed image;

[0119] H2: Based on the pre - processed image, extract pupil features through a pupil feature extraction network;

[0120] H3: Based on the pre - processed image, perform line - of - sight estimation through a line - of - sight estimation model;

[0121] H4: Based on the pre - processed image, obtain temporal attention and spatial attention;

[0122] H5: Establish a concentration evaluation model for feature extraction to obtain the proportion of the dwell time of the main fixation point, the frequency of line - of - sight transfer, and the change rate of pupil diameter.

[0123] Furthermore, referring to Figure 3 , in step H1, the specific steps for pre - processing the collected image are as follows:

[0124] Q1: Perform multi - spectral image fusion: Input the RGB image (denoted as I rgb ) and the IR image (denoted as I ir ) for fusion. Among them, the specific fusion output formula is:

[0125]

[0126] Among them, (x, y) are the pixel coordinates of the image; 0.6 and 0.4 are the weighting coefficients; Y is the grayscale image output after fusion;

[0127] Q2: Perform adaptive illumination compensation on the grayscale image output after fusion;

[0128] It should be noted that the specific details of the adaptive illumination compensation are described as follows:

[0129] 1. Algorithm Definition and Input / Output

[0130] Input image: Y(x, y), grayscale image, size W×H, pixel value range [0, L−1] (L = 256 for 8-bit image);

[0131] Output image: I enhanced (x, y), the enhanced grayscale image;

[0132] Mathematical expression:

[0133]

[0134] 2. Explanation of Key Parameters See Table 4

[0135] Table 4 Parameter Symbol representation Value Physical meaning Contrast limit threshold <![CDATA[C clip > 2.0 Control the maximum height of the histogram bin to prevent noise amplification Block size <![CDATA[T w ×T h > 8×8 Divide the image into 8×8 = 64 local regions

[0136] 3. Block Processing Flow

[0137] Step 1: Image Blocking Divide the image Y into N T non-overlapping Tiles:

[0138] Number of Tiles:

[0139]

[0140] Tile Coordinates: The range of the (i, j)-th Tile:

[0141]

[0142] Step 2: Local Histogram Calculation Calculate the grayscale histogram for each Tile:

[0143]

[0144] k ∈ [0, L−1]: Gray level;

[0145] δ(.): Indicator function (takes 1 when the condition is true, otherwise takes 0);

[0146] Step 3: Contrast Limitation

[0147] Calculate the maximum allowable bin height:

[0148]

[0149] (Assume the Tile size is 128×128, and it needs to be adjusted when the original image size is not 1024×1024);

[0150] Crop and redistribute: For each histogram H i,j :

[0151]

[0152] Step 4: Histogram equalization Calculate the cumulative distribution function (CDF) for each Tile:

[0153]

[0154] Normalized mapping function:

[0155]

[0156] CDF min =min k CDF i,j (k);

[0157] Step 5: Bilinear interpolation fusion For each pixel (x, y), find the nearest 4 Tile center points (i, j), (i+1, j), (i, j+1), (i+1, j+1), and calculate the weights:

[0158]

[0159] 4. Parameter impact analysis

[0160] Clip Limit = 2.0:

[0161] Physical meaning: The upper limit of the allowable bin height is 3 times the average bin height (1 + C clip = 3);

[0162] Effect:

[0163] When C clip = 0, it degrades to ordinary histogram equalization;

[0164] C clip = 2.0 can effectively suppress noise while retaining details;

[0165] Tile = 8×8:

[0166] Computational amount: Each Tile processes 128×128 pixels (when the original image is 1024×1024);

[0167] Trade-off:

[0168] Increasing the Tile size → reducing the local contrast enhancement effect;

[0169] Reducing the Tile size → increasing the risk of block artifacts;

[0170] 5. The complete table of mathematical symbols is shown in Table 5 below: Table 5 Symbol Meaning Unit / Range Y(x, y) Input image grayscale value Integer ∈ [0, 255] <![CDATA[T w , T h > Width and height of the Tile Pixels (here 128) <![CDATA[H i,j (k)]]> Histogram of the (i, j)-th Tile <![CDATA[count ∈ [0, T w ·T h > <![CDATA[C clip > Contrast limit coefficient Dimensionless (here 2.0) <![CDATA[H max > Upper limit of the histogram bin height Count (here 192) <![CDATA[f i,j (k)]]> Gray mapping function Integer ∈ [0, 255]

[0171] Through the above processing flow, CLAHE significantly improves the local contrast while suppressing noise, and is particularly suitable for the problem of blurred eye features caused by uneven illumination in the classroom environment. The bilinear interpolation strategy effectively eliminates the mutation at the Tile boundary, ensuring the spatial continuity of the enhanced image.

[0172] Q3: Perform dynamic ROI extraction on the compensated grayscale image, and output the face bounding box coordinates and eye sub-region coordinates;

[0173] Where:

[0174] Face bounding box coordinates:

[0175] ;

[0176] Eye sub-region coordinates:

[0177] 。

[0178] Among them, in the face detection results output by YOLOv5s, the bounding box coordinates are represented by the top-left origin + width and height, and the specific definition is as follows:

[0179]

[0180] = (top-left x coordinate, top-left y coordinate, width, height)

[0181] Specifically, please refer to Table 6 below: (unit: pixel)

[0182] Table 6 Symbol Mathematical definition Illustration description Value range Calculation example <![CDATA[x face > Position of the upper left corner of the bounding box on the x-axis of the image Offset from the left edge of the image to the right <![CDATA[0≤x face <W img > <![CDATA[If the image width W img = 1920, then x face = 300 represents 300 pixels from the left edge]]> <![CDATA[y face > Position of the upper left corner of the bounding box on the y-axis of the image Offset from the upper edge of the image downwards <![CDATA[0≤y face <H img > <![CDATA[If the image height H img = 1080, then y face = 150 means 150 pixels from the upper edge]]> <![CDATA[w face > Horizontal span of the bounding box Width of the face region <![CDATA[w face >0]]> <![CDATA[w face = 200 indicates that the face width is 200 pixels]]> <![CDATA[h face > Vertical span of the bounding box Height of the face region <![CDATA[h face >0]]> <![CDATA[h face = 300 indicates that the face height is 300 pixels]]>

[0183] Specifically, please refer to Figure 5 , for an overview of the algorithm flow of the dynamic ROI extraction technology.

[0184] Where:

[0185] ① YOLOv5s model structure:

[0186] # yolov5s.yaml

[0187] backbone:

[0188] - [-1, 1, Conv, [64, 6, 2, 2]] # 0-P1 / 2

[0189] - [-1, 1, Conv, [128, 3, 2]] # 1-P2 / 4

[0190] - [-1, 3, C3,

[128] ] # 2

[0191] - [-1, 1, Conv, [256, 3, 2]] # 3-P3 / 8

[0192] head:

[0193] - [-1, 1, Conv, [256, 3, 2]] # 18-P5 / 32

[0194] - [[17, 20, 23], 1, Detect, [1]] # Detection head

[0195] Key parameters:

[0196] Input size: 640×640×3 (BGR format);

[0197] Output dimension: 3×(5 + 1) (3 scales, 1 class);

[0198] Each detection box: (x center , y center , w box , h box , c onf );

[0199] ② Post - processing of face detection

[0200] Step 1: Coordinate denormalization

[0201] x det , y det , w det , h det : Normalized coordinates output by the model (range [0, 1]);

[0202] W orig ×H orig : Original image resolution (e.g., 1920×1080);

[0203] Step 2: Non - maximum suppression (NMS)

[0204] Intersection over Union threshold: IoU thres = 0.5;

[0205] Confidence threshold: conf thres = 0.7;

[0206] NMS formula:

[0207]

[0208] ③Eye sub-region localization

[0209] Anatomical localization formula:

[0210]

[0211] Coefficient source:

[0212] 0.3: Horizontally, the eyes are located at 30% of the face width;

[0213] 0.25: Vertically, the eyes are located at 25% of the face height;

[0214] Error tolerance: Δx, Δy ≤ 3 pixels

[0215] Error compensation technology:

[0216] Sub-pixel interpolation:

[0217]

[0218] Optical flow tracking:

[0219] Use the Lucas-Kanade algorithm to calculate the offset between adjacent frames:

[0220]

[0221] where J is the image gradient matrix and ΔI is the luminance change;

[0222] ④Dynamic adjustment strategy

[0223] Lighting adaption: <

[0224] Adjust the search range according to the average image luminance:

[0225]

[0226] μ lum : The average gray value of the eye ROI region;

[0227] Multi-object processing:

[0228] Prioritize the detected multiple faces:

[0229]

[0230] The specific symbols in the above operation process are shown in Table 7:

[0231] Table 7 Symbol Meaning Unit / Range <![CDATA[x face > x-coordinate of the upper left corner of the face box Pixels ∈ [0, W{orig}] <![CDATA[w face > Width of the face box Pixels > 0 Δx Horizontal positioning error Pixels ≤ 3 J Image gradient matrix <![CDATA[R 2×n > <![CDATA[IoU thres > Overlap threshold Dimensionless [0, 1]

[0232] The performance indicators of the above model are shown in Table 8 below:

[0233] Table 8 Index Laboratory environment Classroom environment Face detection FPS 58 42 Localization accuracy (px) 1.2±0.5 2.3±1.1 Extreme angle detection rate 92% 85% Multi-face processing latency 8 ms 12 ms

[0234] Through the combination of anatomical prior and sub-pixel compensation, this model achieves a positioning error of <3 pixels at 1080p resolution, meeting the requirements of real-time eye tracking in classroom scenarios.

[0235] It should be noted that:

[0236] I. Based on the preprocessed image, pupil feature extraction is performed through a pupil feature extraction network, including the following technologies:

[0237] 1. Input data specification

[0238] Input image:

[0239] Preprocessed eye region image (grayscale image);

[0240] Size: W×H = 128×128 pixels;

[0241] Pixel value range: I(x, y) ∈ [0, 255] (0 = pure black, 255 = pure white);

[0242] Symbol explanation:

[0243] x: Image column coordinate (horizontal direction), value range [0, 127];

[0244] y: Image row coordinate (vertical direction), value range [0, 127];

[0245] 2. Network architecture decomposition

[0246] 2.1 Backbone network (MobileNetV3-small)

[0247] Function: Extract basic features such as pupil edges and textures in the image;

[0248] Key parameters:

[0249] α = 0.75: Width factor, controlling the number of channels in each layer to 75% of the original version (for example, the original 64 channels → 48);

[0250] Bottleneck structure:

[0251] Input → 1×1 convolution (dimensionality increase) → Depthwise separable convolution (feature extraction) → SE module (channel attention) → 1×1 convolution (dimensionality reduction);

[0252] SE module formula:

[0253]

[0254] GAP: Global average pooling, compressing the feature map to 1×1×C;

[0255] W1 ∈ R C / r×C, W2 ∈ R C×C / r: Weights of the fully connected layer (r = 4 is the compression ratio);

[0256] σ: Sigmoid function, outputting channel weights in [0, 1];

[0257] 2.2 Multi-scale feature pyramid (FPN)

[0258] Function: Fuse features of different scales, taking into account both local details and global semantics;

[0259] Hierarchical operations:

[0260] P5 layer (high-level semantics): Perform 1×1 convolution on the output of the backbone network and upsample to 16×16;

[0261] P4 layer (mid-level features): Add P5 and the mid-level features of the backbone network, and output 16×16 after convolution;

[0262] P3 layer (low-level details): Upsample P4 to 32×32 and fuse it with the lower-level features;

[0263] Fusion formula:

[0264]

[0265] P3 up : Low-level features after upsampling;

[0266] P5 down : High-level features after downsampling;

[0267] 2.3 Dual-branch output head

[0268] Coordinate regression branch:

[0269] Structure: 3×3 convolution (512 → 64 channels) → fully connected layer (64 → 2) → Sigmoid activation;

[0270] Output: Normalized pupil center coordinates (x p , y p ) ∈ [0, 1] 2

[0271] Decoding formula:

[0272]

[0273] W = 128, H = 128: Width and height of the input image;

[0274] Ellipse parameter branch:

[0275] Structure: 3×3 convolution (512→64 channels) → fully connected layer (64→3) → Sigmoid activation;

[0276] Output: (a′, b′, θ′) ∈ [0, 1] 3

[0277] Physical parameter decoding:

[0278]

[0279] 3. Loss function design

[0280] 3.1 Coordinate error (MSE Loss)

[0281] Formula:

[0282]

[0283] Symbol explanation:

[0284] N: Batch size (usually 32);

[0285] (x gt , y gt ): True pupil center coordinates (manually annotated);

[0286] 3.2 Shape error (IoU Loss)

[0287] Calculation steps:

[0288] Ellipse discretization: Convert the predicted ellipse and the true ellipse into 100 boundary points;

[0289] Generate mask: Create a binary image, with 1 inside the ellipse region and 0 outside;

[0290] IoU calculation:

[0291]

[0292] M pred : Predicted ellipse mask

[0293] M gt : True elliptical mask

[0294] Loss value:

[0295]

[0296] 3.3 Regularization terms

[0297] Physical constraint terms:

[0298]

[0299] Parameter explanation:

[0300] λ1 = 0.1: Penalize the case where the minor axis exceeds the major axis;

[0301] λ2 = 0.05: Penalize the rotation angle deviation from the historical mean θ prior ;

[0302] θ prior : Rotation angle calculated by moving average (dynamically updated);

[0303] 3.4 Total loss function

[0304]

[0305] Weight design: Focus on coordinate accuracy, supplemented by shape and physical constraints;

[0306] 4. Dynamic post - processing and calibration

[0307] 4.1 Sub - pixel localization optimization

[0308] Bilinear interpolation:

[0309]

[0310] Function: Interpolate at non - integer coordinates with an accuracy of 0.1 pixel;

[0311] 4.2 Optical flow tracking compensation

[0312] Lucas - Kanade algorithm:

[0313]

[0314] Symbol explanation:

[0315] J: Image gradient matrix (J = [∂I / ∂x, ∂I / ∂y]);

[0316] ΔI: Gray - level change vector between adjacent frames;

[0317] 4.3 Online calibration mechanism

[0318] Trigger condition: the coordinate fluctuation is less than 2 pixels for 10 consecutive frames;

[0319] Parameter update rule:

[0320]

[0321] Function: slowly adapt to individual differences (such as pupil size changes);

[0322] 5. The key parameters are as shown in Table 9

[0323] Table 9 Parameter Symbol Value Physical meaning Input size W×H 128×128 Eye region resolution Width factor α 0.75 Network channel scaling ratio Regularization term weight <![CDATA[λ1, λ2]]> 0.1,0.05 Constraint the physical rationality of the ellipse Learning rate η 5e-4 Parameter update step size Moving average coefficient - 0.9 / 0.8 Historical parameter decay rate

[0324] 6. Performance optimization strategy

[0325] TensorRT acceleration: convert the model to FP16 precision, and the inference speed is increased by 2.1 times;

[0326] Adaptive resolution: dynamically adjust the input size (128×128 or 64×64) according to the pupil movement speed

[0327] Heterogeneous computing:

[0328] The GPU processes convolution operations;

[0329] The CPU processes the post-processing logic (coordinate conversion, optical flow tracking).

[0330] II. Based on the preprocessed image, perform gaze estimation through a gaze estimation model, including the following technologies:

[0331] 1. The composition of the input data is shown in Table 10

[0332] Table 10 Input item Symbol Dimension / Format Physical meaning Eye ROI image <![CDATA[I eye > 128×128×1 Preprocessed eye grayscale image Head pose quaternion <![CDATA[q=(q w , q x , q y , q z )]]> 4D vector Head rotation state (unit quaternion) Scene depth map D(x, y) 640×480×1 Distance from the camera to the object (unit: meter)

[0333] 2. The full process of gaze vector calculation

[0334] 2.1 From the eye coordinate system to the camera coordinate system

[0335] 3D coordinates of the pupil center:

[0336]

[0337] Parameter explanation:

[0338] (x p , y p ): Normalized coordinates of the pupil center (from the output of PupilNet);

[0339] (cx , c y ): Coordinates of the optical center of the camera (calibration parameter, unit: pixel);

[0340] (d x , d y ): Physical size of a single pixel (e.g., 3.4 μm);

[0341] (f x , f y ): Focal length of the camera (unit: mm);

[0342] 2.2 Head Pose Compensation

[0343] Quaternion to Rotation Matrix:

[0344]

[0345] Symbol Explanation:

[0346] q w : Real part of the quaternion, q x , q y , q z : Imaginary part;

[0347] R ∈ R 3×3 : Head rotation matrix;

[0348] World Coordinate System Transformation:

[0349]

[0350] t head : Head translation vector (estimated by IMU integration);

[0351] 2.3 Depth Adaptive Calibration

[0352] Calculation of the End Point of the Line of Sight:

[0353]

[0354] Parameter Explanation:

[0355] (x g , y g ): Intersection coordinates of the line of sight vector and the depth map;

[0356] D(x g , y g ): Depth value at the intersection point (unit: meter);

[0357] 3. Neural Network Architecture (GazeNet)

[0358] 3.1 Model Structure

[0359] def GazeNet():

[0360] input_img = Input(shape=(128,128,1)) # Eye image

[0361] input_pose = Input(shape=(4,)) # Quaternion

[0362] # Image feature extraction

[0363] x = Conv2D(32, (3,3), activation='relu')(input_img)

[0364] x = MaxPooling2D(2,2)(x) # 64x64

[0365] x = Conv2D(64, (3,3), activation='relu')(x)

[0366] x = MaxPooling2D(2,2)(x) # 32x32

[0367] x = Flatten()(x) # 32x32x64 → 65536 dimensions

[0368] # Pose feature fusion

[0369] pose_feat = Dense(64, activation='relu')(input_pose)

[0370] x = Concatenate()([x, pose_feat]) # 65536 + 64 = 65600 dimensions

[0371] # Regression output

[0372] output = Dense(3, activation='linear')(x) # Output 3D gaze vector

[0373] return Model(inputs=[input_img, input_pose], outputs=output)

[0374] 3.2 Key parameters are shown in Table 11

[0375] Table 11 Parameter Value Description Number of convolutional kernels 32, 64 Number of feature maps extracted per layer Pooling size 2×2 Downsampling rate Number of fully connected units 64 Head pose feature dimension Output dimension 3 <![CDATA[Viewing vector (g x , g y , g z )]]>

[0376] 4. Loss Function Design

[0377] 4.1 Angular Loss

[0378]

[0379] Geometric Interpretation: The angle between the predicted line of sight and the true line of sight (unit: radian)

[0380] 4.2 Magnitude Loss

[0381]

[0382] Function: To constrain the length accuracy of the line of sight vector

[0383] 4.3 Temporal Smoothness

[0384]

[0385] Physical Meaning: To suppress mutations between adjacent frames (second-order difference constraint)

[0386] 4.4 Total Loss Function

[0387]

[0388] 5. Real-Time Optimization Techniques

[0389] 5.1 Dynamic Weight Adjustment

[0390] Depth Confidence Weight:

[0391]

[0392] ω: Head angular velocity (unit: degrees / second)

[0393] Function: When the head rotates rapidly, reduce the reliability weight of the depth map

[0394] 5.2 Caching Prediction Mechanism

[0395] Formula:

[0396]

[0397] Function: To maintain the smoothness of the line of sight during network latency

[0398] 6. The symbol meanings are shown in Table 12

[0399] Table 12 Symbol Meaning Unit / Range q Head pose quaternion Unit quaternion (∥q∥ = 1) R(q) Rotation matrix 3×3 orthogonal matrix <![CDATA[v eye > Eye coordinate system vector 3D vector (dimensionless) D(x, y) Depth value Meter (m) <![CDATA[L angle > Angle error loss Radian (rad) <![CDATA[g pred > Predicted gaze vector 3D vector (world coordinate system)

[0400] 7. The performance indicators are shown in Table 13

[0401] Table 13 Index Laboratory environment Classroom environment Angle error 0.8°±0.3° 1.5°±0.7° Inference speed 18 ms / frame 25 ms / frame Multi-person tracking 12 people @ 30 fps 8 people @ 30 fps Power consumption 5.2W 6.8W

[0402] 8. Calibration process

[0403] Nine-point calibration method:

[0404] The user sequentially gazes at 9 points with known positions on the screen;

[0405] Record the corresponding gaze vectors {g i} and screen coordinates {(u i , v i )};

[0406] Affine transformation solution:

[0407]

[0408] Least squares method to solve the parameter matrix:

[0409]

[0410] Through the combination of multi-sensor data fusion and deep learning, the model has achieved centimeter-level gaze tracking accuracy in a complex classroom environment, providing reliable basic data for attention analysis.

[0411] III. Based on the preprocessed images, obtain temporal attention and spatial attention, including the following technologies:

[0412] 1. Input data specification

[0413] Input sequence:

[0414] The preprocessed continuous frame image sequence {I t}, t=1 T where T = 10 represents a 10-second time window (assuming a frame rate of 10 fps, a total of 100 frames);

[0415] The size of each frame of the image: 128×128 pixels, and the gray value range is [0, 255];

[0416] Pupil coordinate sequence: {(xt, yt)}t=1T (from the output of PupilNet)

[0417] 2. Temporal attention mechanism

[0418] 2.1 LSTM temporal modeling

[0419] Network structure:

[0420] model = Sequential()

[0421] model.add(LSTM(units=128, input_shape=(T, 2))) # Input: T time steps, each step with 2 features (x, y)

[0422] model.add(Dense(64, activation='relu'))

[0423] model.add(Dense(T, activation='softmax')) # Output: Attention weights for each time step

[0424] Explanation of symbols:

[0425] T = 10: Time window length (seconds);

[0426] units = 128: Number of neurons in the LSTM hidden layer;

[0427] Input dimension: (None, 10, 2) (batch size × time steps × features);

[0428] 2.2 Attention weight calculation

[0429] LSTM hidden state:

[0430]

[0431] h t ∈R 128 : Hidden state vector at time step t;

[0432] Attention score:

[0433]

[0434] W a ∈R 128 ,b a ∈R: Learnable parameters;

[0435] Normalized weight:

[0436]

[0437] α t : Attention weight at time step t, ∑α t = 1;

[0438] 2.3 Temporal feature aggregation

[0439] Weighted feature vector:

[0440]

[0441] Physical meaning: highlighting the line-of-sight position at important moments (such as periods of long fixation);

[0442] 2.4 Key parameters, see Table 14

[0443] Table 14 Parameter Value Function Time window T 10s Balance short-term dynamics and long-term patterns Number of LSTM layers 1 Prevent overfitting Dropout rate 0.2 Enhance generalization ability

[0444] 3. Spatial attention mechanism

[0445] 3.1 Heatmap generation

[0446] Gaussian kernel diffusion:

[0447]

[0448] σ = 15: controlling the radius of the attention area (pixels);

[0449] Output: single-frame heatmap M t ∈ [0, 1] 128×128 ;

[0450] 3.2 Temporal weighted aggregation

[0451] Cumulative heatmap:

[0452]

[0453] α t : from the temporal attention weight;

[0454] 3.3 Spatial weight calculation

[0455] Region segmentation: dividing the image into an 8×8 grid, with each grid being 16×16 pixels;

[0456] Region weight:

[0457]

[0458] G ij : the (i, j)th grid region, i, j ∈ [0, 7];

[0459] 3.4 Key parameters, see Table 15

[0460] Table 15 Parameter Value Function Gaussian kernel σ 15 pixels Control the influence range of a single fixation Grid size 8×8 Balance localization accuracy and computational complexity

[0461] 4. Spatiotemporal attention joint analysis

[0462] 4.1 Attention score calculation

[0463] Time concentration:

[0464]

[0465] The smaller the entropy value, the more concentrated the time attention;

[0466] Spatial concentration:

[0467]

[0468] The larger the value, the more focused the spatial attention;

[0469] 4.2 Attention state determination

[0470]

[0471] 4.3 Visual output

[0472] Time attention curve: Plot the variation of α t with time;

[0473] Spatial heat map: Overlay M total on the original image;

[0474] 5. Symbols are shown in Table 16

[0475] Table 16 Symbol Meaning Unit / Range T Time window length Seconds (default 10 seconds) <![CDATA[α t > Time attention weight <![CDATA[[0, 1], Σα t = 1]]> <![CDATA[β ij > Spatial grid weight <![CDATA[[0, 1], ∑β ij = 1]]> σ Standard deviation of Gaussian kernel Pixels (default 15) <![CDATA[S time > Time attention entropy value Unitless (the smaller the value, the more concentrated)

[0476] 6. Practical application examples

[0477] Scenario: A student is solving a math problem in class

[0478] Time attention:

[0479] In the first 3 seconds, α t is relatively high → Reading the question;

[0480] In the middle 5 seconds, α t is stable → Continuously calculating;

[0481] In the last 2 seconds, α t suddenly increases → Checking the answer;

[0482] Spatial attention:

[0483] β 6,2 = 0.7 (corresponding to the area of the scratch paper);

[0484] β 3,5 = 0.1 (occasionally glancing at the clock);

[0485] Judgment result: Time entropy S time = 1.1, spatial peak S space = 0.7 → High attention state;

[0486] Through the joint analysis of the spatio-temporal attention mechanism, the system can accurately identify the attention focus and its stability of students, providing a quantitative basis for teaching evaluation.

[0487] IV. Feature extraction technology of the concentration evaluation model is described as follows:

[0488] 1. Proportion of the dwell time of the main fixation point ((T main ))

[0489] 1.1 Definitions and formulas

[0490]

[0491] Explanation of symbols:

[0492] ((x t , y t )): The line-of-sight coordinates of the t-th frame (unit: pixel);

[0493] ((x main , y main )): The coordinates of the main fixation point (determined by the clustering algorithm, such as K means );

[0494] (r): The effective radius (default 50 pixels, corresponding to an angular view of about 2°);

[0495] (δ): The indicator function (1 when the condition is satisfied, otherwise 0)

[0496] (T total ): The total number of time steps (e.g., 10 seconds @ 10 fps → 100 frames);

[0497] 1.2 Calculation steps

[0498] Determine the main fixation point: Perform DBSCAN clustering on all fixation points ({(x t , y t ))} within 10 seconds, and take the center of the largest cluster:

[0499]

[0500] (C): The set of frame indices of the largest cluster;

[0501] Statistical dwell time: Traverse all frames. If the distance between the current coordinates and the main fixation point is ≤ (r), then increment the counter by 1;

[0502] Example calculation: If 75 frames satisfy the condition within 10 seconds, then:

[0503]

[0504] 2. Saccade Frequency (F transition )

[0505] 2.1 Definitions and Formulas

[0506]

[0507] Explanation of Symbols:

[0508] Saccade: A rapid movement of the line of sight where the speed of the line of sight exceeds the threshold (30° / second);

[0509] (T total ) : Total time (seconds)

[0510] 2.2 Saccade Detection Algorithm

[0511] Speed Calculation: The speed of the line of sight movement between adjacent frames:

[0512]

[0513] (Δt = 0.1) seconds (frame interval of 10fps);

[0514] Angle Conversion: Convert to angular velocity according to camera parameters (such as focal length (f = 8mm), pixel size (d = 3.4μm)):

[0515]

[0516] Event Determination:

[0517]

[0518] Example Calculation: If 12 saccades are detected within 10 seconds:

[0519]

[0520] 3. Pupil Diameter Change Rate ((ΔD pupil ))

[0521] 3.1 Definitions and Formulas

[0522]

[0523] Explanation of Symbols:

[0524] (a t ) : Pupil diameter of the (t)-th frame (unit: pixel);

[0525] (a t ={a' t + b' t)} / 2 ) (average of the major axis (a' t ) and the minor axis (b' t ))

[0526] 3.2 Calculation Steps

[0527] Data Smoothing: Apply moving average filtering to the original pupil diameter sequence ({a t}):

[0528]

[0529] Window size = 5 frames (0.5 seconds) to suppress high-frequency noise;

[0530] Extreme Value Detection:

[0531] (max(a t )): The maximum pupil diameter within the time window;

[0532] (min(a t )): The minimum pupil diameter within the time window;

[0533] Example Calculation: If the pupil diameter fluctuates between 45 and 55 pixels within 10 seconds, the average = 50 pixels:

[0534]

[0535] 4. Feature Symbols are shown in Table 17

[0536] Table 17 Symbol Meaning Unit / Range <![CDATA[T main > Percentage of fixation time at the main fixation point [0,1] r Effective fixation radius Pixels (default 50) <![CDATA[F transition > Saccade frequency Times per second (Hz) <![CDATA[v angle > Saccade angular velocity Degrees per second (° / s) <![CDATA[ΔD pupil > Pupil diameter change rate Unitless (ratio)

[0537] 5. Dynamic Threshold Adjustment Rule

[0538] Automatically adjust the determination threshold according to the course type:

[0539] Table 18 Course type <![CDATA[T main Threshold]]> <![CDATA[F transition Threshold]]> Lectures 0.7 0.5Hz Lab sessions 0.6 0.8Hz Discussion sessions 0.5 1.2Hz

[0540] Determination Logic:

[0541]

[0542] 6. Practical Application Example

[0543] Scenario: A student is taking a 30-minute math exam:

[0544] Main Fixation Point: The test paper question area (coordinates (320, 240), radius 50 pixels);

[0545] Dwell Time: The fixation point falls in this area within 25 minutes → (T main= 25 / 30 ≈ 0.83 );

[0546] Number of saccades: 18 rapid eye movements detected → (F transition}= 18 / 30 = 0.6Hz);

[0547] Pupil change: diameter fluctuates from 3.2mm to 3.8mm → (ΔD pupil = (3.8 - 3.2) / 3.5 ≈ 0.17 );

[0548] Judgment result:

[0549] (T main =0.83 > 0.7), (F transition =0.6 < 0.7) → High concentration level.

[0550] Furthermore, referring to Figure 4 , the cloud server performs real-time in-depth analysis of the classroom concentration levels of each student, and synchronously obtains the classroom concentration levels of all students, including the following steps:

[0551] E1: Obtain the proportion of the dwell time of the main fixation point, the saccade frequency, and the pupil diameter change rate for each student;

[0552] E2: Establish a concentration evaluation model, input various features, and output the classroom concentration level of the corresponding student;

[0553] E3: Real-time and uniformly obtain the classroom concentration levels of each student;

[0554] E4: Use the deviation formula to obtain the concentration deviation between each student and the standard, and obtain the maximum concentration deviation, which is expressed as the classroom concentration level of the overall students.

[0555] Furthermore, the established concentration evaluation model is specifically:

[0556]

[0557] where ε is the classroom concentration level; α is the proportion of the dwell time of the main fixation point, %; β is the saccade frequency, Hz; γ is the pupil diameter change rate; 1.28, -0.44, and -ln2 are all adjustment constants.

[0558] Furthermore, define the concentration standard as 1;

[0559] Then, use the deviation formula to obtain the concentration deviation between each student and the standard, and obtain the maximum concentration deviation, which is expressed as the classroom concentration level of the overall students, specifically according to the following formula:

[0560]

[0561] Among them, ε1 is the classroom concentration of the first student; ε2 is the classroom concentration of the second student; ε r is the classroom concentration of the r-th student.

[0562] Specifically, the threshold value in S4 is set to 0.64 or 0.641.

[0563] In order to verify the technical effects of the present invention, the following simulation experiment is now carried out:

[0564] Purpose of the experiment

[0565] Verify the effectiveness of the digital teaching interaction method in improving teaching interactivity and students' learning experience.

[0566] Experimental design

[0567] Experimental environment setting: Install the tracking sensor hardware described in the present invention in a standard classroom.

[0568] Selection of experimental objects: Select students of different grades and subjects as experimental objects.

[0569] Collection of experimental data: During the normal teaching process, collect the eye movement state data of students in real time.

[0570] Experimental steps

[0571] Installation and debugging of hardware: Install and debug hardware devices such as dual-mode camera arrays and 9-axis inertial sensors to ensure the accuracy and real-time of data collection.

[0572] Data collection: During normal teaching activities, collect the eye movement data of each student, including pupil characteristics, line of sight direction, etc.

[0573] Data processing and analysis: Analyze the collected data, and use preprocessing and feature extraction technologies to obtain relevant feature values.

[0574] Concentration evaluation: Use the established concentration evaluation model to calculate the classroom concentration of each student.

[0575] Data recording: Record the concentration data of each student and compare it with the results of other evaluation methods.

[0576] Table of partial experimental verification data

[0577] Table 19 Student ID Percentage of fixation time at the main fixation point (%) Saccade frequency (Hz) Pupil diameter change rate 1 75 0.5 0.03 2 68 0.6 0.02 3 80 0.4 0.04 4 60 0.7 0.01 5 72 0.5 0.03

[0578] Note:

[0579] Proportion of the residence time of the main fixation point: The proportion of the total time that the student focuses on the main teaching content in the total observation time.

[0580] Gaze transfer frequency: The number of times the gaze transfers from one area to another within a unit of time.

[0581] Pupil diameter change rate: An indicator reflecting the changes in students' emotions and attention.

[0582] Classroom concentration: The value of students' classroom concentration calculated according to the model.

[0583] The present invention provides a digital teaching interaction method, which has the following beneficial effects:

[0584] It provides a quantitative analysis of students' cognitive states, going beyond traditional single-dimensional assessment methods;

[0585] Through real-time tracking and evaluation, it helps teachers better understand students' learning states, thereby enabling effective teaching interventions;

[0586] It enhances the interactivity of teaching and improves students' learning experiences;

[0587] It provides objective data support for teaching evaluation, contributing to the improvement of teaching quality and effects.

[0588] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical solutions of the present invention can be modified or equivalently replaced without departing from the spirit and scope of the technical solutions of the present invention, and they should all be covered within the scope of the claims of the present invention.

Claims

1. A digital teaching interaction method, characterized in that, It includes the following steps: S1: Install tracking sensor hardware in the classroom to track the eye movement status of each student in real time during the class; among them, the installed tracking sensor hardware is: a dual-mode camera array, including a visible light camera, a near-infrared camera, a 9-axis inertial sensor, and a distributed computing unit; S2: Upload the collected data to the cloud server in real time, and the distributed computing unit embedded in the cloud server deeply analyzes the class concentration of each student in real time; S3: Obtain the class concentration of all students based on the class concentration of each student obtained in real time; S4: The cloud server wirelessly transmits the class concentration of all students to the classroom display terminal in real time, and the teacher gives teaching reminders according to the real-time class concentration displayed on the teacher's display terminal to keep it at a certain threshold, completing deep semi-automatic interaction.

2. The digital teaching interaction method according to claim 1, wherein Specifically, tracking the eye movement status of each student in real time during the class includes the following steps: H1: Collect images in real time and perform image preprocessing, and output the preprocessed images; H2: Based on the preprocessed images, extract pupil features through a pupil feature extraction network; H3: Based on the preprocessed images, estimate the line of sight through a line of sight estimation model; H4: Based on the preprocessed images, obtain temporal attention and spatial attention; H5: Establish a concentration evaluation model for feature extraction, and obtain the proportion of the main fixation residence time, the line of sight transfer frequency, and the pupil diameter change rate.

3. The digital teaching interaction method according to claim 2, wherein In step H1, the preprocessing of the collected images specifically includes the following steps: Q1: Perform multispectral image fusion: Input the RGB image (denoted as I rgb ), and the IR image (denoted as I ir ), and perform fusion. Among them, the specific fusion output formula is: ; Among them, (x, y) are the pixel coordinates of the image; 0.6 and 0.4 are weighting coefficients; Y is the grayscale image output after fusion; Q2: Perform adaptive illumination compensation on the fused grayscale image output; Q3: Extract the dynamic ROI of the compensated grayscale image, and output the face bounding box coordinates and the eye sub-region coordinates; Among them: Face bounding box coordinates: ; Eye sub-region coordinates: 。 4. The digital teaching interaction method according to claim 3, wherein The cloud server deeply analyzes the class concentration of each student in real time, and synchronously obtaining the class concentration of all students includes the following steps: E1: Obtain the proportion of the main fixation residence time, the line of sight transfer frequency, and the pupil diameter change rate of each student; E2: Establish a concentration evaluation model, input each feature, and output the class concentration of the corresponding student; E3: Uniformly obtain the class concentration of each student in real time; E4: Use the deviation formula to obtain the concentration deviation between each student and the standard, and obtain the maximum concentration deviation as the class concentration of the whole students.

5. The digital teaching interaction method according to claim 4, wherein The established concentration evaluation model is specifically: ; Among them, ε is the class concentration; α is the proportion of the main fixation residence time, and its unit is %; β is the line of sight transfer frequency, and its unit is Hz; γ is the pupil diameter change rate; 1.28, -0.44, and -ln2 are all adjustment constants.

6. The digital teaching interaction method according to claim 5, characterized in that: Define the concentration standard as 1; Then, using the deviation formula to obtain the concentration deviation between each student and the standard, and obtaining the maximum concentration deviation as the class concentration of the whole students is specifically based on the following formula: ; Among them, ε1 is the classroom concentration of the first student; ε2 is the classroom concentration of the second student; ε r is the classroom concentration of the r-th student; ‖‖2 is the functional expression of the second-order norm.

7. The digital teaching interaction method according to claim 6, wherein: The threshold in S4 is set to 0.64 or 0.641.

Citation Information

Patent Citations

  • Algorithm for judging degree of concentration of students in class

    CN108021893A

  • A student attention analysis system and a judgment method thereof

    CN109086676A

  • Eye movement tracking sliding input method and system, intelligent terminal and eye movement tracking device

    CN114546102A

  • Intelligent teaching system

    CN115019570A

  • Online teaching platform capable of realizing student concentration degree detection

    CN115937923A