Image recognition-based classroom teaching student attention analysis system and method

By using an image recognition-based system to collect feature points and posture shifts of students' eyes with infrared cameras and analyzing eye movements and facial amplitudes, the real-time and accuracy problems of classroom attention monitoring in existing technologies are solved, enabling dynamic tracking and accurate identification of students' attention.

CN121096030BActive Publication Date: 2026-07-24WUHAN TIANTIAN INTERACTIVE TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
WUHAN TIANTIAN INTERACTIVE TECH CO LTD
Filing Date
2025-10-30
Publication Date
2026-07-24

AI Technical Summary

Technical Problem

Existing technologies cannot effectively monitor students' classroom attention in real time, have difficulty identifying subtle movements, lack real-time data collection, and lack continuous quantitative analysis of attention fluctuations, resulting in delayed classroom feedback.

Method used

An image recognition-based system is used to collect the coordinates of feature points of both eyes through an infrared wide dynamic range camera. The system analyzes eye movement spatial feature groups, image quality composition rate, posture shift and facial amplitude. Combined with multidimensional data and physiological signals, key data segments are selected for attention determination.

Benefits of technology

It enables dynamic tracking of students' attention fluctuations, improves the accuracy and real-time response capability of abnormal state identification, enhances the accuracy and timeliness of monitoring, and supports continuous and reliable analysis of individual behavioral states.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121096030B_ABST
    Figure CN121096030B_ABST
Patent Text Reader

Abstract

The application relates to the technical field of attention analysis, in particular to a classroom teaching student attention analysis system and method based on image recognition, which comprises an eye movement space fusion module, an image source weight evaluation module, a posture offset recognition module, a face amplitude analysis module and an attention judgment module. The application realizes dynamic tracking of student behavior micro-variation and attention fluctuation in a classroom scene through multi-dimensional image data and physiological signal collaborative analysis, comprehensive binocular space feature optimization, bone dynamic proportion comparison and face amplitude difference, signal quality distribution and weight normalization, improves the accuracy and real-time response capability of abnormal state recognition through data screening and sensitive index extraction, guarantees monitoring accuracy under different angles, complex illumination and various actions, enhances timely capture of group differences and individual mutations, and provides continuous and reliable data support for individual behavior state changes in the classroom.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of attention analysis technology, and in particular to a system and method for analyzing student attention in classroom teaching based on image recognition. Background Technology

[0002] Attention analysis involves monitoring and analyzing an individual's attention state in a specific environment. It primarily utilizes sensor data acquisition, physiological signal detection, behavioral feature extraction, and visual behavior analysis to obtain and determine information such as the subject's level of concentration, trends, and behavioral characteristics. This field is widely applied in education, healthcare, human-computer interaction, and driving safety. Traditional classroom student attention analysis systems rely on teachers' experience observing student behavior or using basic video recording playback to make subjective judgments or simple qualitative statistics about student attention. Common methods include manually observing students' eye direction, facial expressions, body movements, and conducting post-class questionnaires.

[0003] Existing technologies rely on subjective observation or basic playback, which cannot achieve dynamic capture of student behavior details. The observation process is easily affected by the limited perspective and changes in lighting, making it difficult to identify some subtle movements. The data collection process lacks real-time performance, and questionnaire feedback is affected by subjective judgment, failing to accurately reflect the student's real-time state. There is a lack of effective mechanisms to deal with sudden behavioral changes or fluctuations in attention, resulting in delayed classroom feedback, limited ability to identify abnormal behavior, and difficulty in supporting continuous quantitative analysis of student status. Summary of the Invention

[0004] To address the problems existing in the prior art, embodiments of the present invention provide a system and method for analyzing student attention in classroom teaching based on image recognition. The technical solution is as follows: The system includes: The eye-tracking spatial fusion module is used to analyze the coordinates of binocular feature points collected by infrared wide dynamic range cameras based on the camera equipment set up in the classroom, determine the relationship between the angle from the tip of the nose to the lens and the facial reference line, screen for abnormal data, and obtain eye-tracking spatial feature groups. The image source weight evaluation module is used to compare the signal intensity of each acquisition source based on the eye movement spatial feature group, analyze the sharpness index of the different acquisition sources, calculate the structure ratio, and obtain the image quality composition rate. The posture offset recognition module is used to calculate the spatial length from the head to the shoulder based on the image quality composition rate, analyze the distance from the shoulder to the hand, determine data segments that continuously deviate from the structural interval, filter key skeleton segments, and obtain the dynamic offset amplitude of the skeleton. The facial amplitude analysis module is used to analyze the amplitude changes of the corners of the mouth, the wings of the nose, the jawline and the eye socket area based on the dynamic offset amplitude of the skeleton, to determine the interval differences with other students, to screen key data of fluctuation amplitude, and to obtain facial dynamic fluctuation index. The attention determination module is used to set weights for each item based on the facial dynamic fluctuation index, standardize each data item, calculate the attention fluctuation discrimination result by combining the weights.

[0005] On one hand, the eye-tracking spatial fusion module includes: The coordinate extraction submodule is used to acquire binocular image frames captured by an infrared wide dynamic range camera based on the camera equipment set up in the classroom, analyze the eye feature points in each frame, calculate and record the spatial coordinates of the left and right eyes, and generate a set of binocular spatial coordinates. The image screening submodule is used to analyze the brightness and image clarity of the pupil region in the image frame based on the binocular spatial coordinate set, determine whether it meets the preset standard, filter out coordinate data that does not meet the requirements, calculate the spatial angle between the tip of the nose and the lens, determine the degree of matching between the angle and the facial feature reference line, remove frame data that deviates from the normal, and obtain the usable coordinate range of the eye. The distribution fusion submodule is used to analyze the relative interval between coordinate points based on the available coordinate range of the eye, determine the distribution of dense areas, and perform balance adjustment according to the area density to optimize the coordinate distribution structure and obtain the eye movement spatial feature group.

[0006] On one hand, the image source weight evaluation module includes: The signal strength comparison submodule is used to analyze the signal strength information acquired by each image acquisition source based on the eye movement spatial feature group, determine the distribution state of the signal strength of each source, compare the signal performance of different image sources in the same time period, filter the acquisition source data with insufficient signal strength, and obtain the signal strength distribution result. The sharpness analysis submodule is used to analyze the image sharpness parameters of each image acquisition source based on the signal intensity distribution results, compare the sharpness differences between acquisition sources, screen image sources with weak sharpness performance, and generate sharpness difference data of acquisition sources. The visible range screening submodule is used to detect the degree of matching between the camera angle and facial orientation of each image source based on the clarity difference data of the acquired sources, determine whether it meets the visible range standard, identify the data sources that meet the conditions, and obtain the image quality composition rate.

[0007] On one hand, the attitude offset recognition module includes: The spatial length calculation submodule is used to obtain the spatial positioning points of the head and shoulders based on the image quality composition rate, calculate the three-dimensional straight-line distance between the two points, collect the coordinate data of key nodes of the shoulder and hand, calculate the spatial length between the shoulder and hand, and generate a set of bone segment distance parameters. The proportional change judgment submodule is used to call the set of bone segment distance parameters, compare the length ratio from head to shoulder with that from shoulder to hand, analyze the changes in proportional change in continuous frames, determine whether the proportional change continues to exceed the structural interval reference, and obtain the structural proportional fluctuation indicator. The offset segment screening submodule is used to detect the spatial position and time sequence of continuous skeletal movements in the proportional deviation segment based on the structural proportional fluctuation indicator, determine the offset amplitude of each frame in the skeleton sequence, filter data segments with stable structural deviation trends and concentrated movement spans, and obtain the dynamic offset amplitude of the skeleton.

[0008] On one hand, the facial amplitude analysis module includes: The regional amplitude calculation submodule is used to analyze the coordinate positions of key points in the corners of the mouth, the wings of the nose, the jawline and the eye socket region in continuous images based on the dynamic offset amplitude of the skeleton, calculate the displacement change range of each key point in the image sequence, determine the contour deformation features in each region, and obtain the key change range of the face. The stretching change recognition submodule is used to detect the lateral stretching path of the corner of the mouth region in continuous frames based on the key facial change range, calculate the continuous change of the distance between the corners of the mouth in the key time period, identify the stretching direction and change trend, and obtain the dynamic change amount of the corners of the mouth. The amplitude fluctuation screening submodule is used to compare the range of mouth corner changes of each student in the same time period based on the dynamic change of mouth corner, determine whether the individual mouth corner fluctuation is different from the group distribution, screen data segments with abnormal fluctuation amplitude changes, and obtain facial dynamic fluctuation index.

[0009] On one hand, the attention determination module includes: The feature normalization processing submodule is used to call the eye movement spatial feature group and the skeleton dynamic offset amplitude based on the facial dynamic fluctuation index, detect the value range of each dimension of data in each feature, calculate its relative distribution interval in the current sample, and obtain the normalized feature combination data. The combined weight setting submodule is used to set the data proportion weights of eye movement, skeleton and facial features based on the normalized feature combination data, allocate the proportion of each type of feature participating in the index construction, calculate the weighted feature results under the weights according to the categories, and generate the feature weighted mapping results. The status indicator generation submodule is used to call the feature weighted mapping result, determine the numerical fluctuation trend of the weighted data in each time period, identify the range of continuous rise or fall, extract the time feature composed of the fluctuation direction and the slope of change, and obtain the attention fluctuation discrimination result.

[0010] On the one hand, the signal strength refers to the signal reliability of the face or eye images captured by each camera, and the structure ratio refers to the proportional distribution of signal strength and clarity index among all acquisition sources.

[0011] On the other hand, a method for analyzing student attention in classroom teaching based on image recognition is provided, including the following steps: S1: Based on the camera equipment set up in the classroom, analyze the coordinates of the feature points of both eyes collected by the infrared wide dynamic range camera, determine the relationship between the angle from the tip of the nose to the lens and the facial reference line, screen for abnormal data, and obtain the eye movement spatial feature group. S2: Based on the eye-tracking spatial feature group, compare the signal intensity of each acquisition source, analyze the sharpness index of the difference acquisition source, calculate the structure ratio, and obtain the image quality composition rate. S3: Based on the image quality composition rate, calculate the spatial length from head to shoulder, analyze the distance from shoulder to hand, determine data segments that continuously deviate from the structural interval, filter key skeleton segments, and obtain the dynamic offset amplitude of the skeleton. S4: Based on the dynamic offset amplitude of the skeleton, analyze the amplitude changes of the corners of the mouth, the wings of the nose, the jawline and the eye socket area, determine the interval differences with other students, screen key data of fluctuation amplitude, and obtain facial dynamic fluctuation index. S5: Based on the facial dynamic fluctuation index, set the weights for each item, standardize the data for each item, calculate the attention fluctuation discrimination result by combining the weights.

[0012] The beneficial effects of the technical solutions provided in the embodiments of the present invention include at least the following: By combining multidimensional image data with physiological signals for collaborative analysis, and integrating binocular spatial feature optimization, dynamic proportion comparison of bones, and facial amplitude differences, along with signal quality distribution and weight normalization, we can dynamically track subtle changes in student behavior and fluctuations in attention in classroom scenarios. Through data filtering and extraction of sensitive indicators, we can improve the accuracy of abnormal state identification and real-time response capabilities, ensure monitoring accuracy under different angles, complex lighting, and diverse actions, enhance the timely capture of group differences and individual mutations, and provide continuous and reliable data support for changes in individual behavioral states in the classroom. Attached Figure Description

[0013] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0014] Figure 1 This is a schematic diagram of the system of the present invention; Figure 2 This is a flowchart of the eye-tracking spatial fusion module of the present invention; Figure 3 This is a flowchart of the image source weight evaluation module of the present invention; Figure 4 This is a flowchart of the attitude offset recognition module of the present invention; Figure 5 This is a flowchart of the facial amplitude analysis module of the present invention; Figure 6 This is a flowchart of the attention determination module of the present invention. Detailed Implementation

[0015] The technical solution of the present invention will now be described with reference to the accompanying drawings.

[0016] In embodiments of the present invention, words such as "exemplarily," "for example," etc., are used to indicate that something is an example, illustration, or description. Any embodiment or design described as "exemplary" in the present invention should not be construed as being more preferred or advantageous than other embodiments or designs. Specifically, the use of the word "exemplary" is intended to present the concept in a concrete manner. Furthermore, in embodiments of the present invention, the meaning expressed by "and / or" can be both, or either one.

[0017] In the embodiments of this invention, the terms "image" and "picture" may sometimes be used interchangeably. It should be noted that, without emphasizing the distinction between them, they convey the same meaning. Similarly, the terms "of," "corresponding (relevant)," and "corresponding" may sometimes be used interchangeably. It should be noted that, without emphasizing the distinction between them, they convey the same meaning.

[0018] In this embodiment of the invention, sometimes a subscript such as W1 may be written in a non-subscript form such as W1. When the difference is not emphasized, the meaning they express is the same.

[0019] To make the technical problems, technical solutions and advantages of the present invention clearer, a detailed description will be given below in conjunction with the accompanying drawings and specific embodiments.

[0020] This invention provides a student attention analysis system for classroom teaching based on image recognition, such as... Figure 1 As shown, the system includes: The eye-tracking spatial fusion module is used to analyze the spatial coordinates of binocular feature points acquired by infrared wide dynamic range cameras based on the camera equipment set up in the classroom. It compares the relative differences between pupil signal intensity and image sharpness index, determines the correspondence between the angle from the tip of the nose to the lens and the facial feature reference line, screens data where the signal intensity and sharpness index have lower limits, optimizes the remaining binocular coordinates, balances the coordinate space distribution, and obtains the eye-tracking spatial feature group. The image source weight evaluation module is used to compare the signal intensity levels of each image acquisition source based on eye-tracking spatial feature groups, analyze the differences in sharpness index between different acquisition sources, determine the correspondence between camera angle and face orientation, filter data that meet the visible range standard, calculate the structural proportion of signal intensity and sharpness index in the acquisition source, and obtain the image quality composition rate. The posture shift recognition module is used to calculate the spatial length from the head to the shoulder in the skeleton recognition unit based on the image quality composition rate, analyze the spatial distance from the shoulder to the hand, compare the proportional change between the two lengths, determine the data segments whose proportions continuously deviate from the standard structure range, filter the skeleton segments with key shift degrees, and obtain the dynamic shift amplitude of the skeleton. The facial amplitude analysis module is used to analyze the amplitude changes of the corners of the mouth, nose, jawline and eye socket area in multiple frames of images based on the dynamic offset amplitude of the skeleton, calculate the change process of the stretching amplitude of the corners of the mouth, determine the interval difference of the student's corners of the mouth amplitude with other students in the same time period, screen data with large fluctuations in the corners of the mouth amplitude, and obtain the facial dynamic fluctuation index. The attention determination module is used to determine attention fluctuations based on facial dynamic fluctuation indicators, call eye movement spatial feature groups and skeleton dynamic offset amplitude, set the weight of each item of eye movement, skeleton and face, perform standard normalization processing on each item of data in turn, and calculate each item of data by weighted combination to obtain the attention fluctuation determination result.

[0021] The eye-tracking spatial feature set includes the distribution of eye gaze intersections, eye-tracking behavior classification, and spatial aggregation characteristics. The image quality composition rate includes the proportion of high signal-to-noise ratio, statistics of effective acquisition time periods, and image source classification labels. The skeleton dynamic offset amplitude includes action posture type, skeleton offset level, and dynamic persistence label. The facial dynamic fluctuation index includes expression change category, micro-expression persistence statistics, and facial state label. The attention fluctuation discrimination result includes discrimination level, stage trend label, and attention transfer node.

[0022] In the eye-tracking spatial fusion module, the infrared wide dynamic range camera refers to an infrared imaging camera that can automatically adjust exposure and capture clear images under different lighting conditions. It is mainly used to collect facial and eye images of students in the classroom from the front and side. The binocular feature point spatial coordinates refer to the specific position coordinates of the student's left and right eyes in three-dimensional space, reflecting the actual orientation and gaze position of the eyeballs. The pupil signal intensity refers to the measurement results of the visibility, contrast, or recognition confidence of the pupil area in the image, which is often used to judge the effectiveness of eye-tracking detection. The image sharpness index is a quantitative measure of the image quality captured by the camera, which usually reflects whether there are problems such as blurriness, out-of-focus, or low resolution in the image. The facial feature reference line refers to the spatial benchmark set based on the facial geometry (such as the eye-nose-mouth alignment line), which is used to judge whether the head orientation, facial orientation, etc. are reasonable. Data with lower limit conditions refers to data that is too low in signal intensity or sharpness index during the data screening process, which is considered invalid data. The remaining binocular coordinates refer to the effective set of binocular spatial coordinates used for analysis and subsequent fusion after removing invalid data.

[0023] In the image source weight evaluation module, signal strength level refers to the signal reliability of the face or eye images captured by each camera, used to measure whether the image can support accurate feature point detection; sharpness index difference refers to the comparison of image sharpness index between different acquisition sources, used to filter image sources with better image quality; corresponding situation refers to the matching degree between the camera angle and the actual orientation of the student's face, used to exclude low-usability angles such as side faces and back faces; visible range standard refers to the set qualified ranges for angle, distance, field of view, etc. between the camera and the student's face, and exceeding these ranges is considered unusable; structure proportion refers to the proportional distribution of signal strength and sharpness index among all acquisition sources, used for subsequent weighted calculation.

[0024] In the posture deviation recognition module, the skeleton recognition unit refers to the hardware or algorithm component that uses image recognition technology to detect and locate key skeletal points such as the student's head, shoulders, and hands in real time; the proportional change refers to the changes in the two sets of spatial lengths, the head-shoulder distance and the shoulder-hand distance, and their ratio, used to determine the state related to attention, such as sitting posture and hand movements; the standard structural interval refers to the reasonable range of the head-shoulder and shoulder-hand spatial ratios measured based on normal human sitting posture and gestures, and continuous deviations indicate abnormalities; the degree of deviation key refers to the skeletal movement segment where, after judgment, the proportional change is found to be significantly outside the standard structural interval and lasts for a long time; the skeletal segment refers to the action sequence containing several continuous skeletal data points within a sampling period, used for time series analysis.

[0025] In the facial amplitude analysis module, amplitude change refers to the change in the geometric position or deformation amplitude of key points in the corners of the mouth, the wings of the nose, the jawline, and the eye socket area over time in an image sequence; the stretching amplitude change process refers to the trend of change formed by continuously calculating the distance between key points of the corners of the mouth when the corners of the mouth open and close, purse, etc. at different times; interval difference refers to the comparison between the range of facial parameter changes of this student in a certain period of time and the range of facial parameter changes of other students in the classroom in the same period of time; data with large fluctuation amplitude refers to data segments with drastic fluctuations in the values ​​of facial parameter changes within a certain period of time, which are often used to identify abnormal expressions or fluctuations in attention.

[0026] In the attention assessment module, each weight refers to the weight assigned to the assessment results of the three types of data: eye movement, skeleton, and face, which is used to distinguish importance when making a comprehensive score; standard normalization processing refers to adjusting feature data from different sources to the same numerical range through mathematical methods, which facilitates weighted combination and direct comparison; weighted combination calculation refers to multiplying each feature data by its corresponding weight and then summing them to obtain the attention state index.

[0027] like Figure 2 As shown, the eye-tracking spatial fusion module includes: The coordinate extraction submodule is used to acquire binocular image frames captured by an infrared wide dynamic range camera based on the camera equipment set up in the classroom, analyze the eye feature points in each frame, calculate and record the spatial coordinates of the left and right eyes, and generate a set of binocular spatial coordinates by combining the image frames with temporal numbering. Eye images are captured by infrared wide dynamic range cameras placed in different corners of the classroom. Each camera captures the student's facial image at a frequency of 30 frames per second. The pixel set of the binocular region in each frame is selected, and edge recognition and grayscale contrast segmentation are performed on the left and right eye image regions in turn to extract the eyeball boundary and locate the center point of the pupil. The pixel mapping is converted into spatial coordinates. If the student is located 1.8 meters above the ground and 2 meters horizontally in front of the camera, it can be converted into three-dimensional coordinate data at a distance of 0.5 meters from the lens through the principle of binocular imaging. According to the acquisition order of the image frames, each set of left and right eye spatial coordinates is assigned a time number, which starts from frame number 1 and increases frame by frame. Each set of data includes a time number, a triplet of left eye spatial coordinates and a triplet of right eye spatial coordinates. Approximately 300 sets of data will be recorded in a 10-second classroom segment. Each set of data is stored in a structured list to establish a set of binocular spatial coordinates.

[0028] The image screening submodule is used to analyze the brightness and image clarity of the pupil region in the image frame based on the binocular spatial coordinate set, determine whether it meets the preset standard, filter out coordinate data that does not meet the requirements, calculate the spatial angle between the tip of the nose and the lens, determine the degree of matching between the angle and the facial feature reference line, remove frame data that deviates from the normal, and obtain the usable coordinate range of the eye. Based on frame-by-frame analysis of the binocular spatial coordinate set, the brightness value and image sharpness parameters of the pupil region in the corresponding image are analyzed. The brightness value is calculated by averaging the gray pixels in the pupil region. If the value is less than 50, it is considered that the image sharpness in the dark area is low. The image sharpness parameter is calculated based on the image edge sharpness and contour contrast. If the value is lower than the set benchmark value of 60, it is judged as a blurry image. Images with frame numbers 124 and 167 are judged as low brightness and blurry images and are removed. At the same time, by extracting the spatial vector of the nose tip position coordinate and the lens center point in the image, and combining it with the position of facial feature lines, it is determined whether the angle between the two is between 5° and 20°. If the angle exceeds this range, it is considered that the student's head is tilted too much, and the coordinate data of frame numbers 158 and 159 are excluded. Finally, only the frame coordinates with a brightness above 50, a sharpness above 60, and an angle between the nose tip and the lens within the range of 5° to 20° are retained to form the usable coordinate range of the eyes.

[0029] The distribution fusion submodule is used to calculate the spatial distribution density of coordinate points based on the available coordinate range of the eye, analyze the relative interval between coordinate points, determine the distribution of dense areas, and perform balance adjustment according to the area density to optimize the coordinate distribution structure and obtain the eye movement spatial feature group. The spatial distribution of all coordinate points is statistically analyzed, and the coordinate values ​​of each point on the X, Y, and Z axes are recorded. The distance between any two points in three-dimensional space is calculated, and a point-to-point spacing matrix is ​​generated. The average spacing and density clustering of all points are classified and judged. If the number of spatially clustered points in a certain area exceeds 8 and the average spacing is less than 0.1 meters, it is identified as a high-density area. Four such clustered areas were identified in the frame data of the front row of the classroom. These areas are fused and adjusted by shifting neighboring points towards the center point and replacing extreme values ​​with local average values. If the original coordinates are (1.21, 0.84, 1.48) and (1.22, 0.83, 1.47), they are unified into (1.215, 0.835, 1.475) after fusion, ensuring that the overall distribution in space is balanced. Isolated points or duplicate recorded points in areas with low density are removed. Finally, the adjusted point coordinate set is output, and a complete eye-tracking spatial feature set is constructed.

[0030] like Figure 3 As shown, the image source weight evaluation module includes: The signal strength comparison submodule is used to analyze the signal strength information acquired by each image acquisition source based on eye-tracking spatial feature groups, determine the distribution state of the signal strength of each source, compare the signal performance of different image sources in the same time period, filter the acquisition source data with insufficient signal strength, and call the remaining image frames for processing to obtain the signal strength distribution results. By reading signal strength data sequences from different image acquisition sources, the signal level of each camera within the same time period is statistically analyzed. The average signal strength values ​​of sources A, B, and C are extracted within a 10-second window, recorded as 82 for source A, 76 for source B, and 48 for source C. Each source's signal strength value is compared to a preset benchmark value of 60. If the signal strength is lower than this benchmark value, it is marked as a source with insufficient signal. The fluctuation amplitude of each source within each second is compared, and periods with amplitudes exceeding 15 are defined as unstable signal intervals. When analyzing the distribution of each source, the signal strength within consecutive frames is... Frames with an intensity greater than 80 are considered high-intensity frames, and those less than 60 are considered low-intensity frames. Signal stability is determined by calculating the proportion of high-intensity frames in each acquisition source. If the proportion exceeds 70%, it is defined as a stable source. In a classroom environment, source A is facing the student's face, source B is located to the right rear, and source C is located at a far angle. Source C has a weaker signal due to angle obstruction and distance attenuation. Therefore, all frame numbers corresponding to source C are marked and removed. At the same time, the remaining frames of sources A and B are synchronized and aligned. Time consistency is ensured by matching frame numbers. The signal values ​​of the two sources are averaged and smoothed to generate the signal intensity distribution results for that time period.

[0031] The sharpness analysis submodule is used to analyze the image sharpness parameters of each image acquisition source based on the signal strength distribution results, compare the sharpness differences between acquisition sources, screen out image sources with weak sharpness performance, retain high-quality image data, and call this part of the data to generate sharpness difference data of acquisition sources. After removing sources with insufficient signal, the image sharpness parameters of the retained acquisition sources were calculated frame by frame. The edge sharpness and detail contrast values ​​of each frame were extracted, and the average sharpness was calculated for each acquisition source. The average sharpness of source A was 78, and that of source B was 68. These were compared with the sharpness benchmark value of 65. It was determined that the sharpness performance of source A was higher than the benchmark range, while source B was in the critical range. The frame number range of source B with a sharpness below 65 was marked on the sharpness distribution curve, and these frames were removed. It was also found that the proportion of frames with a sharpness range of 70 to 85 in source A reached 80%, which was determined to be a high-quality image source. In this process, the image resolution information and brightness balance value of sources A and B at the same time were also compared. If the brightness value of source A in a certain frame was higher than 120 and the contrast range was between 40 and 60, it was retained; otherwise, it was discarded. This process was performed between the 35th and 45th seconds of the classroom video, and finally, the sharpness difference data of the acquisition sources composed of high-quality image frames was obtained.

[0032] The visible range screening submodule is used to detect the degree of matching between the camera angle and face orientation of each image source based on the difference in sharpness of the acquired sources, determine whether it meets the visible range standard, filter the data sources that meet the conditions, and statistically analyze the structural proportion of signal strength and sharpness parameters of the acquired sources to obtain the image quality composition rate. Based on the difference in image sharpness from the collected sources, the camera angles of each image source were matched and detected with the students' facial orientation. The camera angle parameters and facial orientation vectors of each frame were read, and the visible range was determined by comparing the angle between the two in three-dimensional coordinates. If the angle was between 10 and 25 degrees, it was considered to be within the visible range. If the angle exceeded 30 degrees or was less than 5 degrees, it was marked as an invalid frame. In this example, the average angle of source A was 18 degrees and the average angle of source B was 28 degrees. It was determined that some frames of source B exceeded the standard range and their corresponding frames were removed. Subsequently, the signal strength and sharpness parameters of the remaining frames were statistically analyzed. The ratio of the number of frames with signal strength in the range of 70 to 90 to the total number of frames was calculated, and the proportion of frames with sharpness in the range of 65 to 85 was also calculated. If both ratios were above 60%, they were classified into the high-quality structure ratio area. In the classroom monitoring segment, the signal strength frame ratio of source A was 75%, and the sharpness frame ratio was 82%. After summarizing, the structural ratio results of the signal strength and sharpness parameters of each collected source were formed, and the image quality composition rate was obtained.

[0033] like Figure 4 As shown, the attitude shift recognition module includes: The spatial length calculation submodule is used to obtain the spatial positioning points of the head and shoulders in the skeleton recognition unit based on the image quality composition rate, calculate the three-dimensional straight-line distance between the two points, collect the coordinate data of key nodes of the shoulder and hand, calculate the spatial length between the shoulder and hand, and generate a set of skeleton segment distance parameters. By calling the coordinates of key skeleton nodes extracted by the skeleton recognition unit, the coordinates of the head and shoulder positioning points are first read. This process is repeated for each frame. In frame 120, the head coordinates are (0.52, 1.45, 1), and the shoulder coordinates are (0.52, 1.25, 1). By comparing the differences in the spatial values ​​of the three coordinate components, the vertical straight-line length between the head and shoulder is determined to be 0.20 meters. Subsequently, the coordinates of the left shoulder and left hand, and the right shoulder and right hand within the same frame are acquired and processed to extract the three-dimensional spatial positions of the shoulder and hand. In frame 120, the left shoulder coordinates are (0.42, 1.25, 1), and the left hand coordinates are (0.22, 1.05, 1.05, 1.2 ... (1.05), the arm spatial length corresponding to the difference between the three axes was measured to be 0.287 meters. If the coordinates on the right side show the right shoulder as (0.62, 1.25, 1) and the right hand as (0.80, 1.1, 1.1), then the right arm segment distance is calculated to be 0.255 meters. The left and right arm distances are recorded separately, and the head-shoulder distance and the shoulder-hand distances on both sides are labeled as H_C and C_H, respectively, forming three sets of skeletal segment distance parameters. Based on this, a set of skeletal segment distance parameters is established. Each item in the set includes the time frame number, the spatial length from the head to the shoulder, the spatial length from the left shoulder to the left hand, and the spatial length from the right shoulder to the right hand. For a video clip of 30 frames per second, 150 sets of segment distance data can be formed within 5 seconds. A complete set of skeletal segment distance parameters is formed by traversing the data.

[0034] The proportional change judgment submodule is used to call the set of bone segment distance parameters, compare the length ratio from head to shoulder with that from shoulder to hand, analyze the changes in proportional change in continuous frames, determine whether the proportional change continues to exceed the structural interval benchmark, filter out data segments without fluctuation characteristics, and obtain structural proportional fluctuation indicators. For each set of data, the ratio of head-to-shoulder distance and shoulder-to-hand distance is calculated. After calculation, the ratio results are compared with the structural interval benchmark. The structural interval benchmark is set to a range of 0.65 to 0.85 based on the statistical data of human skeletal structure in a seated posture. If the head-to-shoulder length is 0.20 meters and the shoulder-to-hand length is 0.26 meters, then the ratio value of the current frame is 0.77, which meets the standard range. If the head-to-shoulder length is measured to be 0.18 meters and the shoulder-to-hand length is 0.30 meters in frame number 132, the ratio is 0.60, which is lower than the lower limit of the standard, then it is marked as a structural offset frame. The ratio value of each frame is continuously statistically analyzed for 5 consecutive frames. If the ratio of 3 out of 5 consecutive frames is... If the value exceeds the standard range, it is judged as a continuous deviation from the structural state. A structural fluctuation label is added to such frame sequences. After traversing and processing the full frame data, the proportional fluctuation judgment logic is applied to the time axis. Time intervals with continuous offset segments are marked and extracted. Frame sequences with proportional values ​​that are always stable between 0.70 and 0.80 without significant changes are judged as data segments without fluctuation characteristics and are filtered out. In the classroom monitoring video, the proportional value fluctuates between 0.78 and 0.81 in the range of 200 to 220, without significant deviation from the structural baseline range. This time period is judged as a stable posture segment and removed from the results. Finally, the structural proportional fluctuation label is obtained.

[0035] The offset segment screening submodule is used to detect the spatial position and time sequence of continuous skeletal movements in the proportional deviation segment based on the structural proportional fluctuation indicator, determine the offset amplitude of each frame in the skeleton sequence, filter data segments with stable structural deviation trends and concentrated movement spans, and obtain the dynamic offset amplitude of the skeleton. The formula for determining the offset magnitude of each frame in the skeleton sequence is: ; Calculate the skeleton offset amplitude parameter, filter data segments with stable structural deviation trends and concentrated movement spans to obtain the dynamic offset amplitude of the skeleton. Representing the Frame skeleton offset magnitude parameter, This represents the number of students identified. Representing the Frame number The distance between a student's head and shoulders. Representing the The average distance from head to shoulder for each student within this paragraph. Representing the Frame number The distance between a student's shoulder and hand. Representing the The average distance from shoulder to hand for each student within this paragraph. Representing the The parameters representing the proportion of signal strength and sharpness index of each student in the current frame image quality composition rate; The skeleton offset magnitude parameter is calculated based on the relative changes between the spatial distance from the head to the shoulder and the spatial distance from the shoulder to the hand of each student detected in each frame and their respective average spatial distances, combined with the image quality composition rate weight of each student in that frame. It is used to quantitatively reflect the overall degree to which all students in the current frame deviate from their normal state in terms of spatial posture structure.

[0036] First, obtain the key skeleton coordinates of each student in frame f, and determine the spatial distance from their head to their shoulder. Spacing between shoulder and hand Measurements were collected, taking the 10th frame of a certain segment as an example, and the data collected from students S01, S02, and S03 were collected respectively. The collected measurements were 23.5cm, 25.1cm, and 22.7cm. The measurements were 41.8cm, 43.2cm, and 39.6cm; the average head-to-shoulder distance for the three students mentioned above was within this paragraph. The average shoulder-to-hand distances were 24.0cm, 24.5cm, and 23.0cm, respectively. The measurements were 42.0cm, 42.5cm, and 40.0cm respectively. Based on the results of the image quality composition rate assessment module, the normalized proportions of the image signal strength and sharpness index for the three students were calculated. The values ​​are 0.31, 0.36, and 0.33 respectively. During execution, the relative offset ratio of the frame structure for each student is calculated sequentially. For S01, we have: ; ; Then its combined offset contribution term is: ; Similarly, for S02: ; ; Its weighted offset is: ; S03 calculates to: ; ; Its combined offset is: ; In summary, substitute the contributions of the three students in this frame into the execution formula: ; This value This indicates the offset tendency of the current frame within the overall student skeleton structure. This value is then used as the skeleton offset magnitude parameter for the 10th frame. The above calculation is repeated for consecutive frames, and the results are then compared. Perform timing recording; if three or more frames within a five-frame consecutive interval satisfy the condition... In cases where the system identifies segments with stable structural deviations, it uses frame time-series indexing to pinpoint and encapsulate the start and end positions of these segments, forming data segments corresponding to the dynamic offset amplitude of the skeleton. The formula calculates the difference between the structural change ratios of the head-shoulder and shoulder-hand segments, eliminating absolute distance measurement errors caused by individual body size differences, and also introduces an image composition rate weighting factor. The contribution of samples with different image quality is suppressed or enhanced, and high-confidence samples are given greater influence weight in the weighted summation. This avoids the situation where low-quality frames dominate the overall parameter judgment process, and obtains more representative frame-level offset parameters for accurately defining the pose deviation region.

[0037] like Figure 5 As shown, the facial amplitude analysis module includes: The regional amplitude measurement submodule is used to analyze the coordinate positions of key points in the corners of the mouth, nose, jawline and eye socket region in continuous images based on the dynamic offset amplitude of the skeleton, calculate the displacement change range of each key point in the image sequence, determine the contour deformation features in each region, and obtain the key change range of the face. The structural point coordinates of key facial regions are extracted frame by frame from the frame sequence. First, the spatial positions of the outer corner of the mouth, the edge of the nasal ala, the lowest point of the jawline, and the upper edge of the eye socket are identified in each frame. The image coordinates of these points are mapped to a three-dimensional coordinate system, and the X, Y, and Z values ​​of the corresponding points are recorded. Each frame is recorded once and a timestamp is added. For example, in image frame number 125, the coordinates of the left corner of the mouth are (0.54, 1.38, 1.12), and the coordinates of the right corner of the mouth are (0.66, 1.38, 1.12). A temporal coordinate vector sequence is established for these points within 120 consecutive frames. Then, the change in the coordinate position of the same feature point in adjacent frames is calculated sequentially. The coordinate change distance of each point is calculated in the three axes, and the squared differences are summed and averaged. The displacement range of the point throughout the entire time period is obtained. For example, the displacement of the point on the right side of the nose in the X-axis direction between frame numbers 120 and 150 is 0.014 meters, the displacement in the Y-axis direction is 0.009 meters, and the displacement in the Z-axis direction is 0.003 meters. The overall displacement range of the point is the combination of the changes in these three directions. Then, the deformation degree of the four groups of points, namely the corner of the mouth, the wing of the nose, the jaw, and the eye socket, is classified and evaluated. The amplitude benchmark value is set to 0.015 meters. If the maximum displacement range of the key point in a certain area exceeds this value, it is determined that there is a significant contour change in the area. Otherwise, it is marked as a low amplitude area. For example, the maximum displacement of the right corner of the mouth in the 10-second segment is detected to reach 0.022 meters, which is higher than the benchmark value. Therefore, the area is included in the facial key change range.

[0038] The stretching change recognition submodule is used to detect the lateral stretching path of the corner of the mouth region in continuous frames based on key facial change intervals, calculate the continuous change of the distance between the corners of the mouth within key time periods, identify the stretching direction and change trend, analyze the temporal pattern of the corner of the mouth deformation, and obtain the dynamic change amount of the corner of the mouth. The lateral coordinate data of the corner of the mouth region is extracted from consecutive frames. In each frame, the X-axis coordinates of the left and right corners of the mouth are obtained and used to calculate the lateral distance between the corners. Then, a sequence graph of the distance between the corners of the mouth changing over time is plotted in the consecutive frames. In a 5-second sequence, 30 frames of data are collected per second, resulting in 150 sets of distance values. The two time points with the maximum and minimum stretching of the corners of the mouth are extracted from this data, and the change in distance is calculated. If the initial value is 0.115 meters and the maximum value is 0.136 meters, the change is 0.021 meters. If this change is within 0.0... A distance of 18 meters or more is considered a moderate stretching range. The slope sign and value during the change process are detected. If the distance between the corners of the mouth shows a linear upward trend within 15 consecutive frames, it is determined to be an outward stretching state. If it shows a decreasing trend in subsequent consecutive frames, it indicates that the corners of the mouth are changing to a closing trend. Combining the trend label and amplitude change data, the stretching direction of the corners of the mouth in this period is identified as "open first and then close". The sequence is classified as a time segment of an alternating opening and closing pattern. Finally, three descriptive labels of change amplitude, direction and trend are added to each segment of corner change data, and the dynamic change amount of the corners of the mouth is output.

[0039] The amplitude fluctuation screening submodule is used to compare the range of mouth corner changes of each student in the same time period based on the dynamic changes of mouth corner, to determine whether the individual mouth corner fluctuations are different from the group distribution, to screen data segments with abnormal fluctuation amplitude changes, and to obtain facial dynamic fluctuation indicators. The changes in the corners of the mouths of multiple students within the same time period are compared. The maximum and minimum distances between the corners of the mouths of all students and their corresponding ranges of change are extracted from the sampling data of 30 frames per second. The average value and fluctuation range of the changes in the corners of the mouths of each student are calculated. For example, in the time period from frame 320 to frame 350, the change in the corners of the mouths of student A is 0.023 meters, student B is 0.011 meters, and student C is 0.019 meters. First, the average value of the changes in the corners of the mouths of all students in this time period is calculated to be 0.017 meters. Then, the anomaly judgment standard is set as the average value ± 0. If a student's change in the corner of their mouth is higher than 0.023 meters or lower than 0.011 meters, they are considered to have an abnormal fluctuation range. If student A's value is higher than the upper limit during this period, it is marked as a high fluctuation segment. At the same time, it is analyzed whether the 20 frames before and after this segment are also in the abnormal range. If the abnormality lasts for more than 15 frames, it is classified as a continuous fluctuation state, and a dynamic offset label is added to this segment of data. Finally, the abnormal fluctuation value, the corresponding student number, the start and end frame numbers, and the fluctuation duration are organized into a structured record as a facial dynamic fluctuation index.

[0040] like Figure 6 As shown, the attention determination module includes: The feature normalization processing submodule is used to call the eye-tracking spatial feature group and the skeleton dynamic offset amplitude based on the facial dynamic fluctuation index, detect the value range of each dimension of data in each feature, calculate its relative distribution interval in the current sample, and uniformly map it to the standard ratio space to obtain normalized feature combination data. The eye-tracking spatial feature set and the skeleton dynamic offset amplitude were used to simultaneously extract and process three types of feature data. First, samples were traversed for the three dimensions of eye-tracking spatial feature set: gaze intersection distribution, eye-tracking behavior classification, and spatial aggregation characteristics. In each frame, the X and Y coordinate ranges of the student's gaze point were counted and the range was recorded. If the range was 0.38 meters in the horizontal direction and 0.14 meters in the vertical direction, it indicated that the feature value range for that dimension was within the above range. Next, the shoulder and hand offset amplitudes in the frame sequence were extracted for the three dimensions of action posture type, skeleton offset level, and dynamic persistence marker in the skeleton dynamic offset amplitude. If the maximum offset value was 0.32 meters and the minimum offset value was 0.08 meters in a certain time period, the value range for that dimension was 0.24 meters. Subsequently, the facial... The dynamic fluctuation index extracts data on facial expression change categories, micro-expression persistence statistics, and facial state labels to identify changes in mouth corner stretching and jaw contour. The maximum and minimum change values ​​are recorded as 0.018 meters and 0.005 meters, respectively, with a difference of 0.013 meters. The original values ​​of each dimension are then linearly mapped according to their respective maximum and minimum value ranges, transforming the original values ​​into a standard ratio space between 0 and 1. For example, if the change value of the mouth corner in a certain frame is 0.011 meters, its normalized value is 0.011 minus 0.005 divided by 0.013, resulting in 0.46. After completing the normalization process for all dimensions of data, a set of standardized feature combination data is formed. The nine normalized feature values ​​of each frame are combined and stored as a unified feature vector, ultimately yielding the normalized feature combination data.

[0041] The combined weight setting submodule is used to set the data proportion weight of eye movement, skeleton and facial features based on the normalized feature combination data, allocate the proportion of each type of feature to participate in the index construction, calculate the weighted feature results under the weights according to the categories, and generate the feature weighted mapping results. Structural weights were assigned to three feature categories: eye-tracking, skeleton, and face. Initial weights were set based on the participation of each feature category in the teacher's teaching scenario, with eye-tracking weighted at 0.4, skeleton at 0.3, and face at 0.3. Subsequently, the weighted sum was calculated by multiplying the normalized feature value of each dimension in each frame by its corresponding category weight. For example, in frame 150, the normalized values ​​of the three eye-tracking features were 0.60, 0.75, and 0.50, so the weighted sum was 0.60 + 0.75 + 0.50 + 0.4 + 3. The weighted result for eye movement features is 0.37, the weighted result for the three skeleton features is 0.20, 0.30, and 0.25, with a weighted result of 0.075, and the weighted result for the three facial features is 0.48, 0.52, and 0.43, with a weighted result of 0.143. The total weighted value is the sum of the three, which is 0.588. This value represents the overall weighted output value of the current frame. This operation is performed on all data frames, forming a time-ordered feature weighted mapping sequence. Each frame is accompanied by a weighted value label, and this sequence is the feature weighted mapping result.

[0042] The status indicator generation submodule is used to call the feature weighted mapping results, determine the numerical fluctuation trend of the weighted data in each time period, identify the range of continuous rise or fall, extract the time features composed of the fluctuation direction and the slope of change, and obtain the attention fluctuation discrimination result. A sliding window trend determination analysis is performed on the weighted value sequence. Each analysis window consists of 5 frames. The difference between the first and last weighted values ​​and the slope are calculated within each 5-frame window. If the weighted value increases from 0.48 to 0.63 within a window, the trend is determined to be "rising," with a difference of 0.15 and a positive slope. The starting frame number, ending frame number, slope value, and direction of change are recorded. If the trend is upward for three consecutive windows, the current state is determined to be "continuously rising." If the weighted value decreases from 0.63 to 0 within the subsequent two windows... If the value drops from 0.55 to 0.46, the trend changes to "decreasing" and the slope becomes negative. The turning point frame is marked as the inflection point of the fluctuation and the node is recorded. At the same time, if the weighted value fluctuation amplitude in a frame exceeds 0.25, it is judged as a frame of violent fluctuation. This standard is determined by multiplying the average fluctuation value of historical samples by 0.12. If there are three or more violent fluctuation frames in a sequence, it is marked as an unstable state segment. The direction, duration of each fluctuation interval, maximum slope and fluctuation node list are output and integrated to form the attention fluctuation discrimination result.

[0043] A method for analyzing student attention in classroom teaching based on image recognition includes the following steps: S1: Based on the camera equipment set up in the classroom, analyze the coordinates of the feature points of both eyes collected by the infrared wide dynamic range camera, determine the relationship between the angle from the tip of the nose to the lens and the facial reference line, screen for abnormal data, and obtain the eye movement spatial feature group. S2: Based on the eye-tracking spatial feature group, compare the signal intensity of each acquisition source, analyze the sharpness index of the difference acquisition source, calculate the structure ratio, and obtain the image quality composition rate. S3: Based on the image quality composition rate, calculate the spatial length from head to shoulder, analyze the distance from shoulder to hand, identify data segments that continuously deviate from the structural interval, filter key skeleton segments, and obtain the dynamic offset amplitude of the skeleton. S4: Based on the dynamic offset amplitude of the skeleton, analyze the amplitude changes of the corners of the mouth, the wings of the nose, the jawline and the eye socket area, determine the interval differences with other students, screen key data of fluctuation amplitude, and obtain facial dynamic fluctuation index. S5: Based on the facial dynamic fluctuation index, set various weights, standardize the data for each item, and calculate the attention fluctuation discrimination result according to the weight combination.

[0044] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in the present invention should be included within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.

Claims

1. A classroom teaching student attention analysis system based on image recognition, characterized in that, include: The eye-tracking spatial fusion module is used to analyze the coordinates of binocular feature points collected by infrared wide dynamic range cameras based on the camera equipment set up in the classroom, determine the relationship between the angle from the tip of the nose to the lens and the facial reference line, screen for abnormal data, and obtain eye-tracking spatial feature groups. The image source weight evaluation module is used to compare the signal intensity of each acquisition source based on the eye movement spatial feature group, analyze the sharpness index of the different acquisition sources, calculate the structure ratio, and obtain the image quality composition rate. The posture offset recognition module is used to calculate the spatial length from the head to the shoulder based on the image quality composition rate, analyze the distance from the shoulder to the hand, determine data segments that continuously deviate from the structural interval, filter key skeleton segments, and obtain the dynamic offset amplitude of the skeleton. The facial amplitude analysis module is used to analyze the amplitude changes of the corners of the mouth, the wings of the nose, the jawline and the eye socket area based on the dynamic offset amplitude of the skeleton, to determine the interval differences with other students, to screen key data of fluctuation amplitude, and to obtain facial dynamic fluctuation index. The attention determination module is used to set various weights based on the facial dynamic fluctuation index, standardize each data standard, calculate the attention fluctuation discrimination result by combining the weights; The attitude shift recognition module includes: The spatial length calculation submodule is used to obtain the spatial positioning points of the head and shoulders based on the image quality composition rate, calculate the three-dimensional straight-line distance between the two points, collect the coordinate data of key nodes of the shoulder and hand, calculate the spatial length between the shoulder and hand, and generate a set of bone segment distance parameters. The proportional change judgment submodule is used to call the set of bone segment distance parameters, compare the length ratio from head to shoulder with that from shoulder to hand, analyze the changes in proportional change in continuous frames, determine whether the proportional change continues to exceed the structural interval reference, and obtain the structural proportional fluctuation indicator. The offset segment screening submodule is used to detect the spatial position and time sequence of continuous skeletal movements in the proportional deviation segment based on the structural proportional fluctuation indicator, determine the offset amplitude of each frame in the skeleton sequence, filter data segments with stable structural deviation trends and concentrated movement spans, and obtain the dynamic offset amplitude of the skeleton. The eye movement spatial feature set includes the distribution of eye gaze intersections, eye movement behavior classification, and spatial aggregation characteristics. The image quality composition rate includes the proportion of high signal-to-noise ratio, statistics of effective acquisition time periods, and image source classification labels. The skeleton dynamic offset amplitude includes action posture type, skeleton offset level, and dynamic persistence marker. The facial dynamic fluctuation index includes expression change category, micro-expression persistence statistics, and facial state label. The attention fluctuation discrimination result includes discrimination level, stage trend marker, and attention transfer node. The signal strength refers to the reliability of the signal in the face or eye images captured by each camera, and the structure ratio refers to the proportional distribution of signal strength and sharpness index among all acquisition sources.

2. The image recognition-based student attention analysis system for classroom teaching according to claim 1, characterized in that, The eye-tracking spatial fusion module includes: The coordinate extraction submodule is used to acquire binocular image frames captured by an infrared wide dynamic range camera based on the camera equipment set up in the classroom, analyze the eye feature points in each frame, calculate and record the spatial coordinates of the left and right eyes, and generate a set of binocular spatial coordinates. The image screening submodule is used to analyze the brightness and image clarity of the pupil region in the image frame based on the binocular spatial coordinate set, determine whether it meets the preset standard, filter out coordinate data that does not meet the requirements, calculate the spatial angle between the tip of the nose and the lens, determine the degree of matching between the angle and the facial feature reference line, remove frame data that deviates from the normal, and obtain the usable coordinate range of the eye. The distribution fusion submodule is used to analyze the relative interval between coordinate points based on the available coordinate range of the eye, determine the distribution of dense areas, and perform balance adjustment according to the area density to optimize the coordinate distribution structure and obtain the eye movement spatial feature group.

3. The image recognition-based student attention analysis system for classroom teaching according to claim 1, characterized in that, The image source weight evaluation module includes: The signal strength comparison submodule is used to analyze the signal strength information acquired by each image acquisition source based on the eye movement spatial feature group, determine the distribution state of the signal strength of each source, compare the signal performance of different image sources in the same time period, filter the acquisition source data with insufficient signal strength, and obtain the signal strength distribution result. The sharpness analysis submodule is used to analyze the image sharpness parameters of each image acquisition source based on the signal intensity distribution results, compare the sharpness differences between acquisition sources, screen image sources with weak sharpness performance, and generate sharpness difference data of acquisition sources. The visible range screening submodule is used to detect the degree of matching between the camera angle and facial orientation of each image source based on the clarity difference data of the acquired sources, determine whether it meets the visible range standard, identify the data sources that meet the conditions, and obtain the image quality composition rate.

4. The image recognition-based student attention analysis system for classroom teaching according to claim 1, characterized in that, The facial amplitude analysis module includes: The regional amplitude calculation submodule is used to analyze the coordinate positions of key points in the corners of the mouth, the wings of the nose, the jawline and the eye socket region in continuous images based on the dynamic offset amplitude of the skeleton, calculate the displacement change range of each key point in the image sequence, determine the contour deformation features in each region, and obtain the key change range of the face. The stretching change recognition submodule is used to detect the lateral stretching path of the corner of the mouth region in continuous frames based on the key facial change range, calculate the continuous change of the distance between the corners of the mouth in the key time period, identify the stretching direction and change trend, and obtain the dynamic change amount of the corners of the mouth. The amplitude fluctuation screening submodule is used to compare the range of mouth corner changes of each student in the same time period based on the dynamic change of mouth corner, determine whether the individual mouth corner fluctuation is different from the group distribution, screen data segments with abnormal fluctuation amplitude changes, and obtain facial dynamic fluctuation index.

5. The image recognition-based student attention analysis system for classroom teaching according to claim 1, characterized in that, The attention determination module includes: The feature normalization processing submodule is used to call the eye movement spatial feature group and the skeleton dynamic offset amplitude based on the facial dynamic fluctuation index, detect the value range of each dimension of data in each feature, calculate its relative distribution interval in the current sample, and obtain the normalized feature combination data. The combined weight setting submodule is used to set the data proportion weights of eye movement, skeleton and facial features based on the normalized feature combination data, allocate the proportion of each type of feature participating in the index construction, calculate the weighted feature results under the weights according to the categories, and generate the feature weighted mapping results. The status indicator generation submodule is used to call the feature weighted mapping result, determine the numerical fluctuation trend of the weighted data in each time period, identify the range of continuous rise or fall, extract the time feature composed of the fluctuation direction and the slope of change, and obtain the attention fluctuation discrimination result.

6. A method for analyzing student attention in classroom teaching based on image recognition, wherein the method is applied to the image recognition-based student attention analysis system for classroom teaching as described in any one of claims 1-5, characterized in that, Includes the following steps: S1: Based on the camera equipment set up in the classroom, analyze the coordinates of the feature points of both eyes collected by the infrared wide dynamic range camera, determine the relationship between the angle from the tip of the nose to the lens and the facial reference line, screen for abnormal data, and obtain the eye movement spatial feature group. S2: Based on the eye-tracking spatial feature group, compare the signal intensity of each acquisition source, analyze the sharpness index of the difference acquisition source, calculate the structure ratio, and obtain the image quality composition rate. S3: Based on the image quality composition rate, calculate the spatial length from head to shoulder, analyze the distance from shoulder to hand, determine data segments that continuously deviate from the structural interval, filter key skeleton segments, and obtain the dynamic offset amplitude of the skeleton. S4: Based on the dynamic offset amplitude of the skeleton, analyze the amplitude changes of the corners of the mouth, the wings of the nose, the jawline and the eye socket area, determine the interval differences with other students, screen key data of fluctuation amplitude, and obtain facial dynamic fluctuation index. S5: Based on the facial dynamic fluctuation index, set the weights for each item, standardize the data for each item, calculate the attention fluctuation discrimination result by combining the weights.