Student physical and mental state evaluation method and system based on behavior-physiological synergistic compensation

CN122657977BActive Publication Date: 2026-09-25XUZHOU MEDICAL UNIVERSITY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202611150851.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-07-31
Publication Date
2026-09-25
Estimated Expiration
2046-07-31

AI Technical Summary

Technical Problem

这种信噪比的剧烈波动使得生理指标提取的可靠性大幅下降,无法满足对学生心理压力及情绪波动进行精准长周期预警的准临床级要求

Benefits of technology

[0019]1、本发明实现了基于身心一致性的精准评价,有效破解了传统课堂评价中的假装听课难题;相比传统仅依赖抬头率等物理姿态判断专注度、极易受目光呆滞、思维游离等假性专注行为误导的监测手段,通过引入心率变异性等生理指标作为判定认知负荷的客观标准,利用生理唤醒度与物理姿态的双重验证,能够精准辨识隐性走神状态。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122657977B_ABST
    Figure CN122657977B_ABST
Patent Text Reader

Abstract

The application discloses a student physical and mental state evaluation method and system based on behavior-physiological collaborative compensation, relates to the cross technical field of computer vision, medical informatics and intelligent education, and the method comprises the following steps: determining a physiological perception area and a behavior perception area based on a preprocessed video sequence; reconstructing an original blood volume pulsation signal and calculating a heart rate variability index, extracting a hand writing frequency feature, and generating a behavior feature; constructing a phase-energy two-dimensional verification mechanism according to the original blood volume pulsation signal and the hand writing frequency feature, and generating a physiological feature; fusing the behavior feature and the physiological feature, performing multi-level physical and mental state determination based on the fusion result, combining a preset individual physiological benchmark image to calibrate the determination result, and generating a student physical and mental state evaluation result. The application realizes accurate evaluation based on physical and mental consistency, and effectively solves the problem of pretending to listen to a class in traditional classroom evaluation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the interdisciplinary fields of computer vision, medical informatics, and smart education. Specifically, it relates to a method and system for evaluating the physical and mental state of students based on behavior-physiological synergistic compensation. Background Technology

[0002] The focus of educational evaluation has shifted from a single summative assessment to real-time, dynamic monitoring of students' classroom learning process, mental health, and physical workload. However, traditional classroom evaluation models heavily rely on manual observation by teachers or assistant observers. This model faces multiple challenges in practical application: First, manual observation is highly subjective and experience-dependent, making it difficult to develop standardized evaluation rubrics; second, limited by human physiological field of vision and attention allocation, teachers cannot simultaneously conduct long-term, comprehensive behavioral recordings of all students in the classroom while completing their teaching tasks; the most critical issue is that manual methods can only capture macroscopic external movements, completely failing to perceive subtle fluctuations in students' internal physiological parameters, such as heart rate, psychological stress, and the balance of the autonomic nervous system, resulting in a significant lag in the perception of sub-optimal mental health or deep fatigue.

[0003] In the field of automatic monitoring technology based on computer vision algorithms, while existing deep learning models have made some progress in macro-level indicators such as student attendance statistics and group head-up rate detection, their recognition granularity remains at the level of coarse posture classification. In real dynamic teaching scenarios, students in a head-down posture exhibit extremely high semantic ambiguity: existing monitoring solutions, relying solely on low-dimensional head tilt angles or torso posture features, struggle to accurately distinguish whether students are actively taking notes, improperly operating mobile devices, or experiencing fatigue or even subtle inattentiveness due to cognitive overload. Due to a lack of in-depth analysis of hand micro-movement features, pen tip interaction trajectories, and their energy distribution, existing behavior recognition technologies often suffer from extremely high false positive rates when assessing students' true engagement and learning depth. This lack of recognition dimensions prevents the system from providing education administrators with refined behavioral analysis data that has substantial reference value.

[0004] In the field of physiological signal and psychological state assessment, current mainstream solutions still focus on contact-based wearable devices such as smart bracelets and chest-strap ECG recorders. While these devices offer high measurement accuracy in laboratory environments, their widespread adoption in classrooms faces significant engineering obstacles. These include student resistance due to privacy concerns and the feeling of foreign objects, comfort issues with prolonged wear, and the frequent charging and high maintenance costs associated with managing large amounts of hardware. In recent years, non-contact sensing technologies such as remote photoplethysmography (rPPG) have shown great potential, but their performance outside of controlled laboratory environments has been subpar. In dynamic classroom settings, frequent student movements, such as significant handwriting tremors, head turns, or shifts during conversations, introduce severe motion artifacts, easily masking subtle blood oxygen metabolism characteristics. This drastic fluctuation in signal-to-noise ratio significantly reduces the reliability of physiological indicator extraction, failing to meet the near-clinical requirements for accurate, long-term early warning of student psychological stress and emotional fluctuations.

[0005] No effective solutions have yet been proposed to address the problems in the relevant technologies. Summary of the Invention

[0006] In response to the problems in related technologies, this invention proposes a method and system for evaluating the physical and mental state of students based on behavioral-physiological synergistic compensation, so as to overcome the aforementioned technical problems existing in the existing related technologies.

[0007] Therefore, the specific technical solution adopted by the present invention is as follows:

[0008] According to one aspect of the present invention, a method for evaluating the physical and mental state of students based on behavioral-physiological synergistic compensation is provided, the method comprising:

[0009] S1. Preprocess the acquired video sequence, determine the physiological perception region and the behavior perception region based on the preprocessed video sequence, and perform optical flow compensation and encapsulation processing on the physiological perception region and the behavior perception region respectively.

[0010] S2. Based on the physiological sensing region after optical flow compensation, the original blood volume pulsation signal is reconstructed and the heart rate variability index is calculated. Simultaneously, hand writing frequency features are extracted from the encapsulated behavior sensing region to generate behavior features.

[0011] S3. Based on the original blood volume pulsation signal and hand writing frequency characteristics, a phase-energy dual-dimensional verification mechanism is constructed, and the verification mechanism is used to distinguish between real motion artifacts and false motion artifacts to generate physiological characteristics.

[0012] S4. Integrate behavioral and physiological characteristics, perform multi-level physical and mental state assessment based on the integration results, and calibrate the assessment results in conjunction with the preset individual physiological benchmark profile to generate student physical and mental state evaluation results.

[0013] According to another aspect of the present invention, a student physical and mental state evaluation system based on behavior-physiological synergistic compensation is also provided, the system comprising:

[0014] The perception region determination module is used to preprocess the acquired video sequence, determine the physiological perception region and the behavior perception region based on the preprocessed video sequence, and perform optical flow compensation and encapsulation processing on the physiological perception region and the behavior perception region respectively.

[0015] The behavioral feature generation module is used to reconstruct the original blood volume pulsation signal and calculate the heart rate variability index based on the physiological sensing area after optical flow compensation, and simultaneously extract the hand writing frequency features from the encapsulated behavioral sensing area to generate behavioral features.

[0016] The physiological feature generation module is used to construct a phase-energy dual-dimensional verification mechanism based on the original blood volume pulsation signal and hand writing frequency characteristics, and to use the verification mechanism to distinguish between real motion artifacts and false motion artifacts to generate physiological features.

[0017] The evaluation result acquisition module is used to integrate behavioral and physiological characteristics, perform multi-level physical and mental state assessments based on the integrated results, and calibrate the assessment results in conjunction with preset individual physiological benchmark profiles to generate student physical and mental state evaluation results.

[0018] The beneficial effects of this invention are as follows:

[0019] 1. This invention achieves accurate evaluation based on the consistency of mind and body, effectively solving the problem of pretending to listen in traditional classroom evaluation. Compared with traditional monitoring methods that rely solely on physical posture such as head-up rate to judge focus and are easily misled by false focus behaviors such as blank stares and wandering thoughts, this invention introduces physiological indicators such as heart rate variability as an objective standard for judging cognitive load. By using the dual verification of physiological arousal and physical posture, it can accurately identify hidden inattentive states.

[0020] 2. This invention proposes a behavior-driven physiological signal compensation logic. By introducing a phase-energy dual-dimensional triggering mechanism, it effectively reduces the frequency of false touches in the system at rest or under slight movement. At the same time, in the process of using filters to correct physiological signals in reverse, harmonic suppression constraints are introduced. While filtering out strong motion artifacts, the detailed features of physiological signals are fully preserved. This ensures that the system can still obtain high-confidence physiological indicators in dynamic teaching activities, greatly improving the industrial-grade feasibility of non-contact sensing technology.

[0021] 3. This invention delves into the subtle periodic differences in hand movements, providing teachers with accurate student behavior profiles to facilitate differentiated teaching and refined management; by calculating the percentage deviation of physiological indicators relative to the resting baseline, it can detect potential stress overload, anxiety, or depression risks in students at an early stage; this non-invasive and non-intrusive screening method greatly reduces the cost of early warning of campus psychological crises and the psychological resistance of students, and has extremely high social security value.

[0022] 4. This invention adopts a purely visual, non-contact perception solution, which does not require students to wear any hardware sensors. Moreover, the algorithm logic supports lightweight deployment at the edge, which significantly reduces the construction and maintenance costs of smart classrooms while protecting student privacy. It provides a practical and feasible technical path for large-scale digital evaluation of student performance and mental health surveys.

[0023] 5. By introducing adaptive filtering, this invention can improve the positioning accuracy of key points in the edge area of ​​the classroom. By performing white balance calibration and orthogonal projection, it effectively removes the interference of ambient light fluctuations, so that the heart rate extraction under complex lighting conditions is stably controlled within the preset range. In addition, by using a failover backup mechanism, it ensures the continuity of signal extraction when students turn their heads at large angles or are partially occluded, thereby improving the accuracy of multi-target detection. Attached Figure Description

[0024] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0025] Figure 1 This is a flowchart according to an embodiment of the present invention;

[0026] Figure 2 This is a principle block diagram according to an embodiment of the present invention;

[0027] Figure 3 This is a system overall flowchart according to an embodiment of the present invention;

[0028] Figure 4 This is a schematic diagram of dynamic locking of the ROI region according to an embodiment of the present invention;

[0029] Figure 5 This is a logic diagram for recognizing and determining the frequency of hand movements according to an embodiment of the present invention;

[0030] Figure 6 This is a multi-level state determination logic diagram based on physical-psychological consistency verification according to an embodiment of the present invention;

[0031] Figure 7 This is a diagram of the improved 3D-ResNet behavioral feature extraction architecture according to an embodiment of the present invention;

[0032] Figure 8 This is a flowchart of rPPG signal purification based on orthogonal projection and zero-phase filtering according to an embodiment of the present invention;

[0033] Figure 9 This is a schematic diagram of an RLS adaptive noise cancellation logic based on behavior feedback driven according to an embodiment of the present invention;

[0034] Figure 10 This is a logic diagram of dynamic calibration and deviation quantification based on individual physiological benchmarks according to an embodiment of the present invention.

[0035] In the picture:

[0036] 1. Perception area determination module; 2. Behavioral feature generation module; 3. Physiological feature generation module; 4. Evaluation result acquisition module. Detailed Implementation

[0037] To further illustrate the various embodiments, the present invention provides accompanying drawings, which are part of the disclosure of the present invention. These drawings are mainly used to illustrate the embodiments and can be used in conjunction with the relevant descriptions in the specification to explain the operating principles of the embodiments. With reference to these drawings, those skilled in the art should be able to understand other possible implementation methods and the advantages of the present invention.

[0038] According to embodiments of the present invention, a method and system for evaluating the physical and mental state of students based on behavioral-physiological synergistic compensation are provided.

[0039] The present invention will now be further described in conjunction with the accompanying drawings and specific embodiments, such as... Figure 1 and Figure 3 As shown, according to an embodiment of the present invention, a student physical and mental state evaluation method based on behavior-physiological synergistic compensation includes:

[0040] S1. Preprocess the acquired video sequence, determine the physiological perception region and the behavior perception region based on the preprocessed video sequence, and perform optical flow compensation and encapsulation processing on the physiological perception region and the behavior perception region respectively.

[0041] In this optional embodiment, S1 includes:

[0042] S11. Obtain the video sequence of the classroom scene, and perform lens distortion correction, automatic gain control and white balance processing on the video sequence in sequence to complete the preprocessing of the video sequence and generate standardized video frames.

[0043] S12. Input the standardized video frames into the preset lightweight target detection network to locate the student target box and obtain the student target image. Input the student target image into the preset human pose estimation model and output the coordinates of the skeletal key points by combining the heat map regression algorithm.

[0044] S13. Extract facial key point coordinates from skeletal key point coordinates, extract facial sub-regions based on facial key point coordinates, and perform occlusion judgment on facial sub-regions to obtain physiological perception regions; wherein, the physiological perception regions include the main physiological perception region and the backup physiological perception region.

[0045] In this optional embodiment, S13 includes:

[0046] S131. Based on the coordinates of the skeletal key points, the facial key point coordinates are obtained using face alignment technology, and the facial key point coordinates are mapped to a preset three-dimensional deformable model to construct a three-dimensional local topological structure.

[0047] S132. Based on the three-dimensional local topology, extract the grid vertex coordinates of the facial sub-region, and use the inverse perspective transformation to project the grid vertex coordinates to generate the main physiological perception region; wherein, the facial sub-region includes the forehead center sub-region and the bilateral cheek sub-regions.

[0048] S133. Calculate the confidence score of facial key points and use the local binary mode operator to analyze the skin texture consistency to obtain the occlusion weight of each facial sub-region.

[0049] S134. Compare the occlusion weight with a preset threshold. If the occlusion weight exceeds the preset threshold, trigger the failover backup mechanism and obtain the backup physiological perception area as the physiological perception area. If the occlusion weight does not exceed the preset threshold, use the main physiological perception area as the physiological perception area.

[0050] In this optional embodiment, triggering the failover backup mechanism and obtaining the backup physiological sensing region as the physiological sensing region includes:

[0051] Based on the coordinates of facial key points and mesh vertices, a pose estimation algorithm and a preset camera intrinsic parameter matrix are used to obtain the three-dimensional rotation matrix of the target head. The three-dimensional rotation matrix is ​​decomposed into trigonometric functions to obtain the Euler angles of the target head in three-dimensional space. The Euler angles include pitch, yaw and roll angles. Based on the Euler angles, the area behind the ear or the skin area on the side of the neck is dynamically located through spatial geometric coordinate transformation to obtain a backup physiological perception area, which is then used as the physiological perception area.

[0052] S14. Based on the coordinates of the key points of the skeleton, construct the hand-object interaction area, use the semantic recognition algorithm to capture the geometric edge features of the hand-object interaction area, and perform perspective transformation correction, resampling and normalization processing on the geometric edge features in sequence to generate the behavior perception area.

[0053] S15. Perform optical flow compensation on the physiological perception region using the dense optical flow algorithm, and encapsulate the behavior perception region using a preset deep learning framework.

[0054] It should be noted that the process involves acquiring classroom video sequences containing multiple target objects using an infrared-enhanced high-definition camera and preprocessing them: lens distortion correction is performed, and iterative refinement based on higher-order Taylor expansion (using Newton's iteration method) and residual lookup table (Residual LUT) compensation is used to eliminate the impact of edge pixel stretching on coordinate localization; automatic gain control and white balance processing are performed on the images, and a standardized color reference space is provided for subsequent extraction through brightness deviation feedback adjustment logic and color weight bias based on physiologically perceived ROI. Student target boxes are located and the coordinates of 17 skeletal key points are extracted: a multi-level convolutional backbone network (introducing a coordinate attention mechanism) is used to generate enhanced multi-scale feature maps, and the Distance Intersection over Union Non-maximum Suppression (DIoU-NMS) algorithm is used to lock the target boxes; sub-pixel level key point coordinates are output by combining heatmap regression and integral coordinate refinement techniques, and a One-Euro adaptive filter is used to perform dynamic trajectory smoothing and noise reduction. Dynamically locking physiological and behavioral dual-dimensional ROIs: Based on head keypoints, a 3D local topological structure is constructed using face alignment, and the 3D curved surface region is projected back to the 2D image plane through inverse perspective transformation. Simultaneously, a failure switching mechanism based on LBP texture analysis and PnP pose estimation is established to lock the physiologically perceived ROI. A spatial envelope of the hand-object interaction region is established based on wrist keypoints, and perspective correction is performed using a homography matrix to unify the spatial scale. Motion vector compensation and tensor encapsulation are performed: Pixel-level motion vectors are calculated using a dense optical flow algorithm based on polynomial expansion technology. Displacement noise is decoupled through a motion compensation model, and the aligned sequence is encapsulated into a four-dimensional (4D) spatiotemporal feature tensor containing time, channels, height, and width, providing a high-fidelity data source for subsequent multimodal feature extraction. Lens distortion correction is performed using a pre-calibrated camera intrinsic matrix and distortion coefficients to resample coordinates for radial and tangential distortion, eliminating the impact of edge pixel stretching on skeletal keypoint localization. To ensure the fidelity of skin tone extraction during subsequent remote photoplethysmography (rPPG) processing, automatic gain control and white balance processing are performed on each frame of the image. Automatic gain control compensates for uneven illumination by adjusting the electrical signal intensity, while white balance corrects color casts at different color temperatures, thus providing a standardized color reference space for skin tone feature extraction. The above processing can be characterized as follows: ,in The gain coefficient is dynamically updated by the following feedback adjustment logic. This is a bias term.

[0055] like Figure 4 As shown, the specific steps of dynamic locking and spatiotemporal alignment preprocessing for multi-dimensional regions of interest (ROI) are as follows: acquire the classroom video sequence, eliminate lens distortion and perform automatic gain control and white balance processing; locate the student target box and extract the spatial coordinates of the skeletal key points including the head, shoulders, elbows and wrists; lock the forehead and cheek skin areas as physiological perception ROIs, and automatically switch to the backup skin area when the front is occluded; establish motion perception ROIs with the center of the hand as the origin, and define the hand-object interaction area including the hand, pen tip and suspected mobile terminal; perform optical flow compensation on the physiological ROIs and encapsulate the behavior ROI sequence into a three-dimensional (3D) tensor.

[0056] The steps for eliminating lens distortion and performing automatic gain control and white balance processing include:

[0057] 1. Perform lens distortion correction:

[0058] (1) Obtaining calibration parameters: Using Zhang's calibration method, a pre-set checkerboard calibration board is photographed in different poses. The homography matrix between the extracted corner coordinates and the physical coordinates is used to pre-calculate and calibrate the camera's intrinsic parameter matrix (including focal length). and main points ) and distortion coefficients (including radial distortion) With tangential distortion ).

[0059] (2) Establishing a distortion mapping model: For the original video frames acquired by the infrared-enhanced high-definition camera, a nonlinear mapping model between the ideal imaging coordinate system (distortion-free) and the distorted coordinate system is established. The specific steps are: normalizing the coordinates on the ideal image plane... Mapping to distorted coordinates The mapping logic is as follows: Radial offset calculation: using the formula ,in To correct barrel or pincushion distortion caused by lens shape; tangential offset calculation: using the formula This is done to correct the eccentric distortion caused by the non-parallelism of the lens assembly mounting axes; finally, the pixel index on the original distorted image is obtained by applying the intrinsic parameter matrix to the mapped coordinates.

[0060] (3) Coordinate resampling and sub-pixel refinement: To improve resampling accuracy and adapt to key point localization in complex classroom scenarios, this invention introduces an inverse mapping iterative refinement technique based on higher-order Taylor expansion. Specifically, the established mapping model is used to deduce the floating-point coordinates of the ideal pixel in the original distorted frame. To avoid geometric stretching distortion in edge regions due to severe distortion, the system constructs a Residual Lookup Table (RUT) to dynamically compensate for errors in the resampled coordinates, ensuring that the positioning deviation of the skeletal key points in the edge region is controlled within a certain range. Within pixels;

[0061] (4) Use bilinear interpolation algorithm to complete pixel values: Bilinear interpolation finds four pixels in the original distorted image that are adjacent to the floating point coordinates and calculates their corresponding horizontal and vertical weights. The specific steps are: perform two linear interpolations in the horizontal direction, and then perform one linear interpolation in the vertical direction based on the obtained results. The gray or color component value of the point in the ideal imaging coordinate system is obtained by weighted averaging. This algorithm can correct geometric distortion while maintaining the edge smoothness of the image and eliminate the influence of edge pixel stretching on the coordinate positioning of skeletal key points.

[0062] 2. Perform Automatic Gain Control (AGC): Real-time calculation of the average brightness value of the physiologically perceived ROI region within the current video frame. and compare it with the preset target brightness reference. Comparison is performed; the analog gain coefficient of the image sensor (corresponding to the gain coefficient) is dynamically mapped using gain adjustment logic based on brightness deviation feedback. The image's electrical signal intensity is adjusted in real time using either digital magnification or other methods. The feedback adjustment logic, by introducing deviation interval weights and a gain sequence smoothing mechanism, compensates for local shadows caused by uneven lighting in the classroom while eliminating screen flicker caused by gain steps, thereby maintaining the brightness stability of the physiologically perceptual region of interest (ROI). Specifically:

[0063] Brightness deviation calculation and quantization: Real-time extraction of pixel grayscale values ​​within the physiologically perceived ROI region of the current frame, and calculation of its spatial average brightness. And calculate its brightness relative to the target reference brightness. deviation signal Gain adjustment step size calculation: using a nonlinear mapping function Calculate the gain adjustment step size; mapping function Configured as: when When the preset large error threshold is exceeded, a larger adjustment step size is used to achieve a fast illumination response; when When the error is lower than a preset small error threshold, the adjustment step size is reduced to maintain system stability and avoid high-frequency oscillation near the reference brightness; hardware gain limit constraint: constrain the calculated target gain within the physical gain limits supported by the image sensor ; when the analog gain reaches the upper limit and still cannot meet the brightness requirement, the digital gain compensation component is automatically superimposed to ensure the quality of feature extraction in low-light environments; gain sequence smoothing processing: the Exponential Weighted Moving Average (EWMA for short) algorithm is used to smooth the gain sequence between consecutive frames, and the formula is , wherein is a smoothing factor (usually 0.8 to 0.95), is the original gain value calculated currently; the visual flicker caused by gain mutation is eliminated through smoothing processing, and finally the smoothed gain is applied to the image sensor to realize the image brightness stable adjustment.

[0064] 3. Perform White Balance processing: based on the improved Gray World Hypothesis, through statistical analysis on effective pixels in a preset brightness interval within the current video frame, calculate the average component values of the red, green, and blue (RGB) three channels of the original image respectively, and derive the gain compensation coefficient of each channel accordingly; by introducing a physiological perception ROI-based color weight matrix to perform weighted correction on pixels of each channel, the color cast phenomenon under different color temperature environments is eliminated, thereby providing a standardized color reference space for subsequent skin color feature extraction in the physiological perception region; the details are as follows:

[0065] Effective pixel screening and average value statistics: to eliminate the influence of extremely bright (specular reflection) or extremely dark regions in video frames on global color temperature estimation, the system sets an effective brightness interval , and only performs RGB channel statistics on pixels within this interval; calculate the average component value of each channel in the effective pixel set ; ideal gray reference value calculation: calculate the arithmetic mean of the three-channel mean values as the ideal gray reference value , the formula is ; channel gain coefficient derivation: calculate the initial gain compensation coefficients of the red, green, and blue channels respectively , , , the calculation formula is , wherein ; introduction of physiological ROI weight bias: considering the extreme sensitivity of rPPG signals to the color fidelity of skin regions, the present invention performs secondary correction on the gain coefficients; the locked mean value of the physiological perception ROI region is used Compared with the global mean The proportional relationship is used to introduce a bias factor. (Typically a value of 0.1-0.2), final gain coefficient Defined as: Pixel-weighted mapping: This maps the final gain coefficients. , , The algorithm applies to the corresponding channels of the original image to achieve global color balance. This improved algorithm can prioritize the color consistency of physiological perception areas (forehead, cheeks) under different color temperature light sources, and reduce the behavioral-physiological feature coupling error caused by ambient light fluctuations.

[0066] The steps for locating the student target bounding box and extracting the spatial coordinates of the skeletal key points include:

[0067] 1. Student Target Detection and Boundary Locking: A lightweight target detection network based on an improved YOLO or MobileNet architecture is used to scan the video sequence. This network introduces a Feature Pyramid Network (FPN) with embedded Coordinate Attention mechanism. By embedding positional information into channel attention, the network's ability to perceive overlapping targets in the classroom is enhanced. In specific implementation, the original video frames are scaled to a uniform resolution and input into a multi-level convolutional backbone network. Multi-scale feature maps are generated through hierarchical strided convolution and pooling operations. The prediction head is used to slide and scan on the feature maps, and the center point coordinates, height, and width information of the target objects are calculated through the regression branch. Finally, the Distance Intersection over Union Non-maximum Suppression (DIoU-NMS) algorithm is used to introduce center point distance constraints while considering the overlap between boxes, thereby accurately locking the bounding rectangles of multiple individual students in the classroom and effectively solving the problem of missed detection caused by mutual occlusion between students in the front and back rows.

[0068] 2. Perform human pose estimation and key point extraction: Input the captured student target image into the human pose estimation model, which adopts a high-resolution network (HRNet) architecture. This model maintains high-resolution representation throughout the process and performs multi-scale feature parallel fusion. Specifically, a heatmap regression algorithm is used to generate response heatmaps for the corresponding channels of 17 key points. A Gaussian kernel distribution function is applied to the target pixels in the heatmap to find the local maxima with the highest heat value as the initial coordinates. Integral Regression is used to perform sub-pixel-level offset compensation for the neighborhood of the maxima, thereby outputting the coordinates of 17 standard skeletal key points, including the head center, shoulders, elbows, and wrists.

[0069] 3. Trajectory Smoothing and Denoising: A keypoint confidence scoring mechanism is introduced during the extraction process. For each frame's output coordinate point, a confidence score between 0 and 1 is assigned based on the heatmap response intensity. Abnormal detection results with scores below a preset threshold are discarded. For the remaining valid coordinate sequence, a One-Euro adaptive filter is used to perform trajectory smoothing. This mechanism utilizes low-pass filtering logic with a dynamic cutoff frequency to adjust the smoothing weights in real time based on the keypoint movement rate. This mechanism enhances the filtering strength when the student is stationary to suppress minor fluctuations caused by pixel noise, and reduces filtering latency when the student moves quickly, thereby ensuring the smoothness and real-time performance of the keypoint's spatiotemporal trajectory.

[0070] A lightweight object detection network (such as YOLO or MobileNet architecture) is used to scan the video sequence and locate the bounding rectangles of multiple individual students in the classroom. The captured student target images are then input into a human pose estimation model, and a heatmap regression algorithm is used to output a set of 17 standard skeletal key points, including the head center, both shoulders, both elbows, and both wrists. At the same time, a key point confidence scoring mechanism is introduced to eliminate abnormal coordinate jumps caused by limb overlap or environmental occlusion, ensuring the smoothness of the spatiotemporal trajectory of key points.

[0071] The specific steps for backbone extraction and multi-scale feature map generation include: Hierarchical feature extraction: Five consecutive downsampling convolution operations are performed on the input image using a backbone network (such as CSPDarknet or MobileNetV3); original feature maps with strides of 8, 16, and 32 are output from the 3rd, 4th, and 5th convolution stages of the backbone network, respectively, to capture target information under different spatial receptive fields. Cross-scale feature fusion: A top-down path aggregation is performed using a Feature Pyramid Network (FPN), upsampling the deep high-semantic feature map and then merging it with the shallow high-resolution feature map laterally; a coordinate attention module is introduced at the fusion node to retain accurate positional information by encoding features in the horizontal and vertical directions, thereby generating an enhanced set of multi-scale feature maps.

[0072] The specific steps of the Distance Intersection Over Union (DIoU-NMS) algorithm include: Candidate box sorting in descending order: Based on the classification confidence score output by the prediction head, all original candidate detection boxes of the same class are sorted in descending order; Distance Intersection Over Union (DIoU) calculation: The candidate box with the highest current score is selected iteratively. As a baseline box, calculate the remaining candidate boxes. Intersection over union with the reference frame Calculate the Euclidean distance between the center points of the two rectangles. And the diagonal distance of the smallest closure region containing these two rectangles. Judgment and Suppression: Based on the formula Calculate the distance intersection-union score; if the score exceeds a preset suppression threshold, then a candidate box is determined. Redundant boxes are identified and removed; a center point distance penalty term is introduced. This allows the system to retain the detection results of adjacent individuals based on the center point offset when two targets highly overlap, thereby reducing the false negative rate.

[0073] The specific steps of the heatmap regression algorithm include: Response heatmap construction: For the coordinates of each ground truth key point in the image... A two-dimensional Gaussian distribution heatmap is constructed on the corresponding channel of the output feature map. The Gaussian kernel distribution function is as follows: ,in, The preset standard deviation is used to control the diffusion range of the thermal point response; Model prediction and maximum extraction: The thermal map is predicted using the output of HRNet, and the location of the pixel with the largest response intensity in each channel is used as the rough spatial coordinates of the key point.

[0074] The specific steps of integral regression include: local probability normalization: selecting around the maxima of the predicted heatmap. The neighborhood window is used to normalize the response values ​​within the window using the Softmax function, transforming the thermal values ​​into a probability distribution. Expected coordinate calculation: The sub-pixel level expected coordinates of key points are calculated using integral mapping logic, as shown in the following formula: ,in, The neighborhood window is used; through this integral feedback mechanism, the quantization error caused by the limited resolution of the heat map is eliminated, and high-precision trajectory capture of students' subtle movements (such as fingertip writing and pupil micro-movements) is achieved.

[0075] For a One-Euro adaptive filter, the specific steps include: Instantaneous rate calculation: Calculating the key point positions in the current frame. Position relative to the previous frame The rate of change between them is used to derive the instantaneous motion rate of the key point. ,in The sampling time interval; rate smoothing: the instantaneous rate is smoothed using a first-order low-pass filter to obtain the smoothed rate. ,in The preset rate smoothing factor; dynamic cutoff frequency calculation: based on the smoothing rate. Real-time calculation of adaptive cutoff frequency The formula is as follows: ,in, This is the preset minimum cutoff frequency (used to suppress minor jitter when stationary). Rate sensitivity coefficient (used to adjust response speed during rapid motion); smoothing weight mapping and coordinate filtering: based on cutoff frequency. Calculate the smoothing weights for the current position Use this weight to perform filtering and updating of the keypoint coordinates: Through this mechanism, the system automatically lowers the cutoff frequency to enhance the filtering effect when students are stationary and studying, and automatically raises the cutoff frequency to reduce latency when students are turning pages or moving.

[0076] The steps for locking the physiologically sensed ROI and automatically switching to the backup skin area when the front is obscured include:

[0077] 1. Perform 3D dynamic locking of the main physiological perception ROI: Use face alignment technology to obtain the sequence of facial key points and map the key points onto a preset 3D Morphable Model (3DMM) to construct the 3D local topology of the facial region; In specific implementation, based on the aligned 3D coordinates, extract the mesh vertices of the forehead center and the bilateral cheek regions, and use inverse perspective mapping to project the 3D curved surface region back to the 2D image plane, thereby locking the main physiological perception ROI (including the forehead center sub-region and the bilateral cheek sub-regions).

[0078] 2. Real-time monitoring of occlusion weight: By calculating the confidence score of facial key points in the current frame and combining it with the Local Binary Pattern (LBP) operator to analyze the consistency of skin texture, the occlusion weight index of each sub-region is obtained. When the occlusion weight of the front area (such as the forehead or cheek) exceeds the preset threshold, it is determined that the main physiological perception ROI has been interrupted, and the system triggers an automatic failover backup mechanism.

[0079] 3. Location and Switching of Backup ROI: Utilizing a pose estimation algorithm (Perspective-n-Point, PnP) combined with the calibrated camera intrinsic parameter matrix, the system obtains the 3D rotation matrix and translation vector of the target head, and derives the 3D Euler angles (pitch, yaw, roll) accordingly. Based on the current 3D rotation matrix of the head, the system calculates the visual projection of the lateral area of ​​the head in the current camera coordinate system through spatial geometric coordinate transformation, thereby dynamically locating the skin area behind the ear or the lateral skin area of ​​the neck as a backup physiological perception ROI. Specifically, the physiological perception ROI consists of the primary physiological perception ROI and the backup physiological perception ROI. Under normal monitoring conditions, the system prioritizes extracting rPPG signals from the primary physiological perception ROI. When the primary region fails, the system seamlessly switches to the backup physiological perception ROI to maintain the continuity of signal extraction.

[0080] Based on key points of the head Using face alignment technology, a three-dimensional local topology structure is constructed in the facial region, and the unobstructed skin areas of the central forehead and both cheeks are identified as the primary physiological perception ROI (denoted as ROI). To handle head turning or hand occlusion, the system calculates occlusion weights in real time. : ,in, Here, dist represents the predicted location of the keypoint, and dist is the Euclidean distance. This is the preset standard deviation. If... If the value exceeds a preset threshold, the system automatically triggers a failover backup mechanism, using spatial geometric transformation to locate the skin area behind the ear or the skin area on the side of the neck as a backup physiological sensing ROI to ensure the continuity of physiological signal extraction.

[0081] For face alignment technology, the steps include: Key point feature regression: using a feature extraction network based on the HRNet architecture, multi-scale feature fusion is performed on the captured facial image, and the coordinates of 68 or 98 dense facial key points, including the outer edge of the eyebrows, eye sockets, bridge of the nose, and contours, are extracted using a heatmap regression algorithm; 3D model fitting: the dense facial key points are input into a 3DMM fitting model; the shape coefficient and pose parameters are iteratively optimized using the least squares method to deform the 3D average face model until the Euclidean distance residual between its projection point on the 2D plane and the detected key points reaches a preset minimum value, thereby constructing a personalized 3D local topology structure for the target object.

[0082] The steps for inverse perspective transformation include: spatial coordinate mapping to determine the corresponding forehead and cheek areas in the 3DMM model. Spatial coordinates of each grid vertex in the 3D target coordinate system Projection matrix calculation, using pre-calibrated camera intrinsic parameter matrices. and the rotation matrix obtained through pose estimation With translation vector Construct a projection transformation model from 3D space to a 2D image plane; pixel coordinate calculation: according to the formula Calculate the projected coordinates of 3D vertices onto the original 2D video frame. ,in The scale factor is used to accurately map the three-dimensional curved surface region back to the image pixel space through inverse perspective transformation, thereby achieving sub-pixel-level dynamic locking of the physiological perception ROI.

[0083] The keypoint occlusion weight calculation model specifically utilizes the Local Binary Pattern (LBP) operator to analyze skin texture consistency, including: LBP feature map generation, targeting each pixel within the physiologically perceptual ROI region. Choose its radius as Within the neighborhood Each sampling point; compare the sampling point with the center pixel. The gray value, if the gray value of the sampling point is greater than or equal to Then mark it as 1, otherwise mark it as 0; the generated The binary numbers are arranged in order and converted to decimal to serve as the LBP value for that pixel, generating a local texture feature map; Texture histogram extraction: The physiological perception ROI is divided into several sub-grids, and the distribution frequency of LBP values ​​within each grid is statistically analyzed to construct a texture statistical histogram for that region. Skin consistency verification and weight calculation: real-time extracted histogram Compared with a pre-set (or obtained during the pre-class rest period) standard skin texture baseline histogram Perform similarity comparison (e.g., using chi-square test or Bach distance); occlusion weighting. Quantification: Combining key point confidence scores The calculation logic for occlusion weight is as follows: ,in, This is a histogram similarity function; when a region is obscured by a hand or book, the texture distribution deviates from the normal skin tone pattern, leading to... The value drops significantly, thus reducing the occlusion weight. If the preset security threshold is exceeded, the subsequent backup area switching logic will be triggered.

[0084] The process for pose estimation algorithms and obtaining the 3D rotation matrix includes: establishing a 2D-3D spatial mapping: extracting the coordinates of facial 2D key points. Corresponding 3D mesh vertex coordinates in the constructed 3DMM model Perform one-to-one mapping to construct the observation sample set; establish the perspective projection equation: using the pre-calibrated camera intrinsic parameter matrix. The following perspective geometric constraint equations are established: ,in, As a scale factor, It is a translation vector. This is the 3D rotation matrix to be solved; matrix solution and refinement: using the PnP algorithm (such as the iterative Levenberg-Marquardt optimization algorithm) to minimize the reprojection error, i.e., minimizing... This allows us to calculate the optimal 3D rotation matrix for the current frame. Euler angle extraction: based on a 3D rotation matrix The target's pitch, yaw, and roll angles are derived through trigonometric function decomposition. Rotation matrix. It fully characterizes the orientation of the student's head in three-dimensional space, aiming to provide a precise coordinate transformation basis for subsequent dynamic positioning of the skin area behind the ear or on the side of the neck.

[0085] Establish motion-aware ROI and define the hand-object interaction area, including:

[0086] 1. Define the spatial envelope of the Hand-Object Interaction (HOI) region: Using the wrist and palm center key points as the three-dimensional spatial reference origin, the boundary expansion coefficient of the action observation window is dynamically adjusted by calculating the instantaneous motion velocity vector of the hand key points between consecutive frames, thus constructing an action observation window with adaptive boundary adjustment capability; Based on the relative topological relationship between the hand and the desktop reference plane, the vertical distance from the palm center to the desktop plane is calculated using a preset camera extrinsic matrix. With the palm center as the sphere center and 1.5 to 2 times this vertical distance as the dynamic radius, a local spherical or cylindrical spatial envelope model is constructed, which is defined as the Hand-Object Interaction (HOI) region; This region aims to strictly cover and monitor in real time the evolution of the student's hand posture, the writing trajectory of the pen tip, and mobile smart terminal devices that may appear on the desktop.

[0087] 2. Perform spatial scale unification processing: Semantic recognition algorithms (such as lightweight semantic segmentation networks or edge detection operators) are used to capture the highlighted writing displacement trajectory of the pen tip and the geometric edge features of device operation in real time within the HOI region. For behavior-aware ROI image sequences extracted at different shooting distances and angles, the specific steps for implementing spatial scale unification are as follows: A perspective transformation correction is performed on the local region using a homography matrix to convert the tilted viewpoint into a standard top-down viewpoint. The homography matrix is ​​obtained by selecting at least four non-collinear feature reference points within the desktop reference plane and calculating the projection mapping relationship between the reference points in the original image coordinate system and the target top-down coordinate system. A bilinear interpolation algorithm is used to resample the extracted unequal-sized image sequences, scaling them to a preset fixed pixel dimension (e.g., ...). or (pixels); mean normalization is performed to eliminate the influence of illumination and scale fluctuations on feature extraction, thereby providing high-precision spatial constraints for subsequent differentiation between periodic learning behavior and non-periodic interference behavior.

[0088] To capture features for semantic recognition algorithms, the following steps are taken: feature mask generation, which uses a pre-trained lightweight segmentation network (such as MobileNet-UNet) to perform pixel-level classification of the HOI region and generate a binary mask image of the pen tip region and the mobile terminal region; geometric feature extraction, which performs Canny edge detection and connected component analysis on the mask image to extract the centroid coordinate sequence of the pen tip and the rectangular envelope features of the mobile terminal, thereby quantifying the spatial evolution during the hand-object interaction process.

[0089] For homography matrix acquisition and perspective transformation correction, the following steps are taken: feature point pair matching, and determination of four reference points with known physical spacing within the desktop reference plane. (e.g., the four corners of a desktop), and obtain their pixel coordinates in the original video frame. Matrix solving is performed using the Direct Linear Transform (DLT) algorithm, based on... Establish a system of linear equations based on the correspondence, and solve for the results. homography matrix Pixel remapping, using the formula Each pixel from the original tilted viewpoint is mapped to a standard top-down coordinate system. This step eliminates perspective distortion caused by the camera mounting angle, ensuring that the behavioral features extracted by the subsequent 3D-ResNet model are spatially translation-invariant.

[0090] The steps of performing optical flow compensation and encapsulating the behavioral ROI sequence into a 3D tensor include:

[0091] 1. Perform motion vector calculation based on dense optical flow: Calculate the motion vectors of skin pixels within the physiologically perceived ROI between adjacent frames using dense optical flow algorithms (such as the Farneback algorithm or the TV-L1 algorithm). Specifically, first, convert consecutive video frames into grayscale images and construct a Gaussian pyramid. Multi-scale analysis is then used to capture displacements of different amplitudes. The intensity distribution of each pixel's neighborhood is approximated using the polynomial expansion technique found in the Farneback algorithm, and the horizontal displacement component of each pixel between adjacent frames is calculated. With vertical displacement components This generates a dense vector field that characterizes the minute movements on the skin surface.

[0092] 2. Correcting Pixel Offset Using Motion Compensation Model: The motion compensation model is based on global motion estimation theory. By statistically analyzing dense vector fields, it fits the overall motion trajectory of the target head using affine transformation or perspective transformation models. In practice, a transformation matrix is ​​constructed based on the calculated motion vectors, and reverse spatial mapping is performed on the physiological perception ROI of the current frame to make the image sequence accurately aligned with the reference frame in spatial dimension. This model corrects the pixel offset caused by the slight shaking of the target object, decoupling the periodic displacement of pixels in spatial coordinates from the fluctuation of pixel values ​​(brightness / color) caused by blood oxygen metabolism. This effectively separates non-pulsatile displacement noise from the original video signal and improves the signal-to-noise ratio of remote photoplethysmography (rPPG) signals.

[0093] 3. Spatiotemporal Tensor Encapsulation of Execution Behavior ROI: The processed T-frame sequence of behavior-aware ROI images is hard-aligned with millisecond-level timestamps and encapsulated into a four-dimensional (4D) spatiotemporal feature tensor of shape (T,C,H,W) using a deep learning framework (such as PyTorch or TensorFlow). Here, T is the time step (number of frames), C is the number of color channels, and H and W are the height and width of the image, respectively. While preserving the high-resolution spatial features of the local region, the tensor fully records the evolution information of the action on the time axis, aiming to adapt to the nonlinear mapping and semantic extraction of fine-grained behavioral features of students by subsequent deep learning models.

[0094] This involves polynomial augmentation techniques and displacement vector calculations, including: local intensity polynomial modeling, which uses a second-order polynomial to approximate the intensity distribution of the neighborhood of each pixel in the image, assuming the local coordinates of the pixel are... Then its brightness value Represented as: ,in, It is a symmetric matrix. For vectors, The coefficients are scalar coefficients; the coefficients are obtained by performing weighted least squares fitting on the neighboring pixels; inter-frame polynomial transform analysis, let the two adjacent frames be... and If the second frame is displaced relative to the first frame Then there is By substituting into the above polynomial model, the relationship matrix between the polynomial coefficients of the two frames is derived using identity transformations; pixel-level displacement vector calculation is performed, and based on the evolution of the coefficient matrix, a relationship matrix for the displacement vector is constructed. The system of linear equations: ,in The average value of the coefficient matrices of the two frames is used to solve the equation and obtain the dense displacement vector of each pixel, thus providing an accurate motion reference for subsequent reverse spatial mapping and pixel offset correction.

[0095] S2. Based on the physiological sensing region after optical flow compensation, the original blood volume pulsation signal is reconstructed and the heart rate variability index is calculated. Simultaneously, hand writing frequency features are extracted from the encapsulated behavior sensing region to generate behavior features.

[0096] In this optional embodiment, S2 includes:

[0097] S21. After optical flow compensation, all pixels in the physiological perception area are subjected to mean dimensionality reduction in the spatial dimension, and continuous sampling is performed along the time dimension to obtain the color time sequence signal.

[0098] S22. The color time sequence signal is reconstructed using the color method orthogonal projection technology and adaptive fusion technology to obtain the original blood volume pulsation signal. The original blood volume pulsation signal is then filtered using a preset Butterworth bandpass filter. Based on the filtered original blood volume pulsation signal, the heart rate variability index is calculated.

[0099] S23. Based on the encapsulated behavior perception region, generate a spatiotemporal feature tensor and input the spatiotemporal feature tensor into the improved three-dimensional residual network to extract the spatiotemporal features of hand movements. Combine the fast Fourier transform to extract the hand writing frequency features and obtain the frequency domain features.

[0100] In this optional embodiment, the frequency features of handwriting are extracted by combining fast Fourier transform to obtain frequency domain features, including:

[0101] Based on the hand-object interaction area, the coordinate displacement of key points at the center of the wrist and palm is located between consecutive video frames, generating a hand motion vector sequence representing the kinetic energy state of the hand. A fast Fourier transform is performed on the hand motion vector sequence to calculate the power spectral density and obtain the kinetic energy ratio of the target frequency band. The kinetic energy ratio is compared with a preset energy threshold. If the kinetic energy ratio is greater than the preset energy threshold, the hand motion is determined to be a periodic writing behavior, and the hand writing frequency features are extracted. If the kinetic energy ratio is less than or equal to the preset energy threshold, it is determined to be a non-periodic writing behavior. Combining the kinetic energy ratio and the periodic writing behavior determination results, the frequency domain features of the hand motion are generated.

[0102] S24. The spatiotemporal features and frequency domain features are concatenated and feature mapped to generate behavioral features that include attitude angle, line of sight direction and periodicity of movement indicators.

[0103] It should be noted that pixel spatial averaging is performed on the physiologically perceived ROI: Chromatic orthogonal projection technology is used to remove ambient light diffuse reflection interference in the chromaticity space and reconstruct a high-fidelity original blood volume pulsation (BVP) signal; a zero-phase Butterworth bandpass filter is used to purify the signal, eliminating baseline drift while ensuring that the peak temporal position does not shift, and the heart rate variability (HRV) index reflecting the state of the autonomic nervous system is calculated accordingly; multi-dimensional behavioral semantic features are extracted: the corrected behavioral ROI sequence is input into a three-dimensional residual network (3D-ResNet) integrated with spatiotemporally deformable convolution to capture the nonlinear deformation and temporal evolution features of hand movements in spatial morphology; simultaneously, the power spectral density (PSD) of the hand motion vector is calculated using Fast Fourier Transform (FFT) to obtain the energy distribution features of the target frequency band; behavioral feature vectors are generated: the extracted spatial motion features and frequency domain distribution features are concatenated and feature mapped to generate a unified behavioral semantic feature vector containing posture angle, gaze direction, and movement periodicity index. This provides a high-dimensional feature representation for subsequent collaborative compensation and state discrimination.

[0104] like Figure 5 As shown, the specific steps of feature extraction and action frequency domain representation based on multimodal parallel branching are as follows: spatial mean reduction is performed on the physiological perception ROI pixels to obtain the original signal sequence of red, green, and blue (RGB) channels changing over time. The original BVP signal was reconstructed in the color space using POS or CHROM algorithms through orthogonal projection. The original BVP signal was then purified, and the HRV index, reflecting the state of the autonomic nervous system, was calculated. A three-dimensional convolutional kernel was used to simultaneously capture hand grip posture and movement evolution features. These hand grip posture and movement evolution features constituted the underlying spatiotemporal features that distinguish actions such as note-taking, turning pages, and operating mobile terminals, aiming to provide high-dimensional feature support for constructing a mind-body joint feature space and performing fine-grained behavioral semantic mapping. Based on the coordinate displacement of the wrist and palm center key points between consecutive frames, a hand motion vector representing the hand's kinetic energy state was generated. An FFT was performed on the hand motion vector, and the in-band energy ratio was used to determine whether the action possessed periodic writing characteristics. The extracted dominant action frequency was used as the hand writing frequency feature, aiming to provide a key reference input frequency for subsequent removal of physiological signal motion artifacts. After global average pooling, the extracted spatiotemporal and frequency domain features were dimensionality reduced to generate a behavioral semantic feature vector. .

[0105] like Figure 9 As shown, the steps for BVP reconstruction of the original pulse wave signal are as follows: Obtain the preprocessed normalized color time-series signal C(t); C(t) is a column vector containing the red (R), green (G), and blue (B) channel observations after mean and normalization processing, which characterizes the temporal evolution of skin region color in the three-dimensional color space; calculate the chromaticity space projection vector P using a preset orthogonal projection matrix: P is a two-dimensional vector containing two mutually orthogonal components. The design logic of this matrix is ​​that its row vectors are orthogonal to the normal direction of skin reflection (i.e., the normalized unit vector [1,1,1]). Through this linear mapping, the three-dimensional color space is projected onto a two-dimensional orthogonal plane perpendicular to the skin color normal vector. This maps the in-phase noise in C(t) that is severely affected by light intensity to the normal direction of the projection plane and removes it, retaining the first projection component P1 (corresponding to the calculation result of the first row of the matrix) and the second projection component P2 (corresponding to the calculation result of the second row of the matrix) with physiological physical significance. The original pulse wave signal S is reconstructed through adaptive fusion technology. BVP The calculation formula is as follows: , among which, S BVP This refers to the pulsation signal of the original blood volume after reconstruction. The standard deviation of the corresponding component within the sliding time window is represented by the ratio of the standard deviations of the two orthogonal components. This ratio is used as a dynamic scaling factor to scale and inversely subtract the residual motion artifact components in P1 and P2. Since artifacts caused by limb movements are strongly correlated in the two orthogonal dimensions, and pulse signals caused by hemoglobin absorption have specific phase differences in different chromaticity dimensions, this differential fusion operation can effectively cancel out non-physiological displacement noise. This process maps multi-channel color vectors to a chromaticity-sensitive dimension highly correlated with hemoglobin absorption characteristics, achieving high-fidelity dynamic reconstruction of the pulse wave and providing a high signal-to-noise ratio input source for subsequent signal purification and index calculation.

[0106] For signal reconstruction using the chromatic projection technique, the following steps are taken: RGB channel normalization to obtain the original red, green, and blue three-channel time-series signals extracted from the physiologically perceived ROI. The mean value within the sliding time window is used for normalization to obtain the mean-free signal. Chromaticity feature projection uses a preset orthogonal projection matrix to map the normalized signal to two mutually orthogonal chromaticity components. and The calculation formula is as follows: , Adaptive signal fusion utilizes the standard deviation of the two chromaticity components within a sliding time window. Calculate adaptive weighting factors The original pulse wave signal is obtained through differential fusion. This step utilizes the difference in projection of skin color reflection components in orthogonal space to effectively remove non-physiological diffuse reflection interference.

[0107] The original BVP signal was filtered using a Butterworth bandpass filter in the range of 0.7–4.0 Hz to remove respiratory rate and high-frequency electronic noise. The pulse peak-to-peak interval (IBI) was detected based on the filtered signal, and the heart rate variability index RMSSD was calculated using the following formula: Among them, IBI is the time interval between two consecutive pulse beats, and RMSSD reflects parasympathetic nerve activity and is a key physiological parameter for assessing cognitive load and psychological stress state.

[0108] like Figure 8 As shown, signal purification using a Butterworth bandpass filter includes: filter parameter configuration, and constructing a... A Butterworth bandpass filter of order 4 (usually 4th or 6th order) has its passband frequency set to... This frequency band corresponds to the normal human pulse rate range (42 bpm to 240 bpm); transfer function calculation: the difference equation coefficients of the filter are solved using the bilinear transform method based on the preset cutoff frequency; zero-phase filtering: to avoid the impact of phase lag caused by the filter on the subsequent heart rate variability (HRV) peak location, the system uses forward-backward filtering logic. The signal is processed to ensure that the filtered signal maintains a high signal-to-noise ratio while the peak position does not shift in the time domain, thereby improving the accuracy of HRV index extraction.

[0109] like Figure 7 As shown, the encapsulated behavior-aware ROI spatiotemporal tensor is input into the improved 3D-ResNet network. The improvements, technical motivations, and specific implementation methods of the improved 3D-ResNet network compared to the standard 3D-ResNet include:

[0110] 1. Introduction of Spatio-Temporal Deformable Residual Block: Addressing the technical challenges of subtle, non-linear trajectories, irregular posture deformations, and minute ranges in student hand-object interaction (HOI) actions in classrooms (such as pen spinning and book turning), the fixed sampling grid of standard convolutional kernels struggles to capture these complex local deformations. Implementation: The fixed sampling convolutional kernels in standard 3D-ResNet are replaced with deformable 3D convolution layers. Specifically, a lightweight offset learning branch is set up alongside the main convolutional branch, using a 1×1×1 convolutional kernel to process the input feature map and calculate the offset set of sampling points in the spatial dimension (x,y) and temporal dimension (t) in real time. The bilinear interpolation algorithm adaptively deforms the sampling position of the convolution kernel according to the offset set, so that its sampling points can dynamically fit the irregular edges and motion envelopes of the hand, pen tip and mobile terminal, thereby enhancing the network's accuracy in capturing fine writing trajectories and device operation action features.

[0111] 2. Multi-Scale Temporal Enhancement Module: This module aims to enhance the system's sensitivity to sudden key actions (such as quickly hiding a phone or flipping through a book instantly) and suppress relatively static redundant information in the video background through differentiated modeling in the temporal dimension. Implementation: A temporal gating unit based on a self-attention mechanism is embedded in the residual connection path. This unit utilizes a set of one-dimensional dilated convolutions with different dilation rates to extract inter-frame momentum differences, capturing multi-scale temporal evolution features. A temporal weight coefficient sequence is generated using a sigmoid activation function, dynamically reweighting the feature map on the time axis. Through this mechanism, the system can adaptively increase the feature gain of frames containing key actions while weakening static background interference. A temporal gating unit based on a self-attention mechanism is embedded in the residual connection path to quantify the importance of action features within different time steps. This module utilizes one-dimensional dilated convolution to extract inter-frame momentum differences and dynamically reweights the feature map using generated temporal weight coefficients, thereby automatically suppressing static redundant information in the video background and enhancing the feature response to sudden key actions (such as turning a book or hiding a phone). The network uses the aforementioned improved mechanism to simultaneously capture the spatial morphology of hand grip posture and the evolution of the action over time. It outputs a high-dimensional behavioral feature map through multi-layer convolution and downsampling operations, and then generates a behavioral semantic feature vector through global average pooling, providing a highly discriminative behavioral descriptor for subsequent mind-body consistency verification.

[0112] Perform a Fast Fourier Transform (FFT) on the hand motion displacement vector D(t) to obtain the power spectral density. Calculate the energy percentage of the target frequency band (1.5-4.0Hz, corresponding to 90-240 actions per minute). : ,like If the value is >0.6, the action is determined to be a note-taking behavior with periodic writing characteristics; otherwise, it is classified as a non-periodic action (such as playing on a mobile phone, wiping sweat, etc.).

[0113] The extracted high-dimensional spatiotemporal feature vector is fused with the output frequency domain feature index through multimodal feature fusion to form a unified behavioral semantic feature vector. The generation method and specific steps of the behavioral semantic feature vector are as follows: Obtain the high-dimensional spatiotemporal feature vector representing the evolution of hand movements and grip posture after global average pooling. Extract the calculated energy percentage And action periodic labels generated based on threshold determination (For example, define note-taking as 1 and non-periodic actions as 0); use feature concatenation technology to... , as well as Vectors are concatenated according to a predefined dimensional order, and prior weights including pose angle and gaze direction are introduced. A fully connected layer is then used for linear mapping to generate a behavior semantic feature vector with uniform dimensions. , ,in, This represents a vector concatenation operation. This represents a linear mapping function. The generated function... It fully encompasses the spatial movement patterns, temporal evolution trends, and frequency domain periodic features of individual students, aiming to provide multi-dimensional feature inputs with strong semantic descriptive capabilities for the subsequent physical-psychological consistency determination in step four.

[0114] S3. Based on the original blood volume pulsation signal and hand writing frequency characteristics, a phase-energy dual-dimensional verification mechanism is constructed, and the verification mechanism is used to distinguish between real motion artifacts and false motion artifacts to generate physiological characteristics.

[0115] In this optional embodiment, S3 includes:

[0116] Using the original blood volume pulsation signal and hand writing frequency features as inputs, the overlap between the physiological energy spectrum and the behavioral action frequency distribution is calculated using a normalized cross-correlation formula. Simultaneously, the phase-locking value between the behavioral displacement vector and the physiological pixel fluctuation signal within the corresponding frequency band of the hand writing frequency features is calculated. Combining the overlap, phase-locking value, and kinetic energy ratio, a phase-energy dual-dimensional verification mechanism is constructed. This mechanism is used to determine motion artifacts, distinguishing between real and spurious motion artifacts. If determined to be a real motion artifact, the hand writing frequency features and the original blood volume pulsation signal are used as inputs, and an improved recursive least squares filter and harmonic suppression constraints are used to remove motion artifacts and generate physiological features. If determined to be a spurious motion artifact, the original blood volume pulsation signal is used as the physiological feature.

[0117] It should be noted that the extracted handwriting frequency features Perform time-domain correlation analysis with the original pulse wave signal; construct a... A recursive least squares (RLS) adaptive filter is used as the reference input; artifacts caused by hand movements are removed from the physiological signal as noise components to ensure the accuracy of physiological parameter extraction in dynamic teaching scenarios.

[0118] The specific steps for motion artifact compensation of physiological signals based on behavior feedback are as follows: determine whether the micro-displacement of the physiological ROI region is caused by the kinetic energy conduction of hand movements by calculating the cross-correlation coefficient; and use a recursive least squares filter to remove periodic or non-periodic artifact noise caused by limb movements from the original BVP signal.

[0119] Using handwriting frequency features extracted from behavior branches The writing action phase-energy dual-dimensional trigger determination is performed to accurately distinguish between true and false motion artifacts. The specific operation steps are as follows: Perform pre-trigger correlation calculation to obtain the reconstructed original BVP signal and the obtained hand dominant frequency. The overlap between physiological energy spectrum and behavioral frequency distribution was calculated using the normalized cross-correlation formula. Perform phase consistency verification and calculate... The phase locking value (PLV) between the behavioral displacement vector and the physiological pixel fluctuation signal within the corresponding frequency band is calculated. By determining whether the average phase difference within adjacent time windows is within a preset stable drift range, false triggering caused by occasional overlap between the writing frequency and the physiological pulse is eliminated. Kinetic energy conduction intensity is determined, and calculations are performed. Percentage of motion kinetic energy within the frequency band Only when the overlap And the phase lock value is within the preset threshold range, and When the kinetic energy conduction intensity threshold is met, it is determined that the current physiological signal is severely contaminated by the true and false shadows induced by writing, and the system automatically triggers an adaptive noise compensation mechanism.

[0120] An improved recursive least squares (RLS) filter with integrated harmonic suppression constraints is constructed, which... The composite eigenvector formed by its second and third harmonics is used as the reference input. The original noisy BVP signal is used as the desired input. Filter weight vector The update logic is as follows: Gain vector and prior error calculation: using the formula Calculate the gain and solve for the prior estimation error. Introducing harmonic suppression constraints: A harmonic regularization term is introduced into the weight update formula. This term applies a penalty function to the non-dominant frequency components of the reference signal to prevent the filter from mistakenly removing harmonic components of the physiological signal related to the writing frequency as noise; weight iterative update: the final update formula is revised to: ,in, Step size factor The mathematical expression for the constraint operator for the harmonic distribution of hand movements is as follows: ,in, The upper limit of the harmonic suppression order is preferably 3, corresponding to the fundamental frequency, second harmonic, and third harmonic; The harmonic order is... For the first Write the harmonic frequencies in order; This is a frequency domain impulse function used to locate harmonic frequencies; It is the first Weights of the first harmonic penalty term; This represents the weight vector of the filter at the previous time step. The filter precisely removes motion artifacts caused by limb writing movements from the original BVP signal by minimizing the weighted sum of squared errors including constraint terms, outputting a high-fidelity pulse wave signal. This ensures the completeness of HRV indicator extraction in highly dynamic scenarios.

[0121] With the frequency of hand writing The relevant action features are used as the reference input vector x(k), and the original noisy BVP signal is used as the desired input d(k). Adaptive noise cancellation is performed using a recursive least squares (RLS) filter. The update formula for the filter weight vector w(k) is as follows: , , Where k(k) is the gain vector, For prior estimation error, The forgetting factor (usually taken as 0.95~0.9995) is used. Through iterative updates, the filter subtracts the periodic or quasi-periodic motion noise components induced by limb movements from the original BVP signal, and outputs a high-fidelity physiological characteristic signal to ensure the accuracy of heart rate variability index extraction for students in dynamic scenarios such as writing.

[0122] S4. Integrate behavioral and physiological characteristics, perform multi-level physical and mental state assessment based on the integration results, and calibrate the assessment results in conjunction with the preset individual physiological benchmark profile to generate student physical and mental state evaluation results.

[0123] In this optional embodiment, S4 includes:

[0124] S41. Align behavioral and physiological features according to preset timestamps, and perform tensor splicing in the feature channel dimension to generate joint mind-body features;

[0125] S42. Compare the pitch angle in Euler angles with the preset angle threshold, and combine it with the periodic writing behavior judgment result in the mind-body joint feature to obtain the head-down learning branch.

[0126] S43. Based on the head-down learning branch, the physiological cognitive load score is calculated using a weighted fusion formula, and the physiological cognitive load score is compared with the preset high and low arousal thresholds and resting level thresholds. The focus state is distinguished based on the comparison results. Among them, the focus state includes deep focus state and false focus state.

[0127] S44. Compare the pitch angle with the preset rest threshold, and extract the respiratory rate feature based on the original blood volume pulsation signal for cross-validation to obtain the sleep state;

[0128] S45. Based on the student's focus and sleep status, combined with the preset individual physiological baseline profile, generate the evaluation results of the student's physical and mental state.

[0129] It should be noted that the compensated physiological feature vector and behavioral feature vector are concatenated by tensors to construct a joint mind-body feature space; consistency verification, conflict judgment, and multidimensional fatigue judgment are performed through a multi-level state decision tree; and the final evaluation result is output by calculating the percentage of parameter offset based on the individual physiological baseline profile established before class rest.

[0130] like Figure 6 As shown, the specific operational steps of multi-level state determination and individual benchmark calibration based on mind-body consistency verification are as follows: align and fuse physiological and behavioral feature vectors in the time dimension; identify deep focused learning, latent distraction, and deep sleep states based on the consistency between behavioral recognition results and physiological arousal indicators; calculate the offset of real-time physiological indicators relative to the individual resting baseline to eliminate the influence of individual differences on psychological stress evaluation.

[0131] The compensated physiological feature vector (Including metrics such as HR, RMSSD, LF / HF, etc.) and the generated behavioral semantic feature vector Tensor concatenation is performed along the time dimension to construct a joint feature vector. This joint space preserves the temporal correspondence between physical behavior and physiological arousal, providing a foundation for subsequent consistency verification. A hierarchical decision-making logic is executed based on the joint feature vector, utilizing a mind-body inconsistency verification mechanism to identify different learning states of students. The specific steps and judgment logic are as follows: Calculate the head posture angle. And perform preliminary classification: using the Perspective-n-Point (PnP) algorithm, based on the extracted facial key points such as the eyes, nose tip, and corners of the mouth, solve for the rotation matrix between the camera coordinate system and the head target coordinate system; extract the pitch angle as the head pose angle. ;like And the output If the characteristics of periodic writing are met, the student enters the head-down learning sub-branch; otherwise, the student's focus is determined by the direction of their gaze; the physiological cognitive load score is then verified. In the sub-branch of head-down learning, a weighted fusion model is used to calculate the physiological cognitive load score. The calculation formula is as follows: ,in, and For preset weighting coefficients, and To establish individual physiological baseline values; if the calculated values ​​are... Being in a high arousal state, i.e., satisfying A preset high arousal threshold (corresponding to a stress state where RMSSD is significantly lower than the baseline and LF / HF is significantly increased) is used to determine a state of deep focus; if In the low wake-up baseline state, that is, satisfying To approximate the resting sleep threshold, a mind-body inconsistency verification mechanism is triggered, correcting the identification result to a state of hidden distraction or false focus; a deep sleep state determination is then performed: if the head tilt angle... Continuous deviation exceeding a preset static threshold (e.g.) And the duration exceeds The system synchronously retrieves respiratory rate (RR) features extracted from the physiological branch. The specific implementation process is as follows: bandpass filtering is performed on the reconstructed BVP signal and the respiratory envelope is extracted. The respiratory frequency is identified using Fast Fourier Transform (FFT). If the RR shows a significant periodic slowing down and tends to stabilize (the standard deviation of RR is lower than the preset variation threshold), it is finally determined to be a deep sleep state.

[0132] In this optional embodiment, the method further includes:

[0133] Within a pre-set resting time window, multidimensional physiological parameters of the target subject are acquired. Statistical analysis is then used to extract the steady-state heart rate value and the fundamental frequency index of heart rate variability from these parameters, constructing an individual physiological baseline profile of the target subject. During the real-time monitoring phase in class, physiological characteristics and heart rate variability are used as parameters, and the individual physiological baseline profile is used as a reference standard. The absolute difference between the parameters and the reference standard is calculated, and a normalization mapping mechanism is used to convert the absolute difference into a percentage offset, generating a relative offset vector reflecting the real-time psychological stress level of the target subject.

[0134] It should be noted that, as Figure 10As shown, this step serves as the baseline construction stage for the entire process. Before executing the real-time monitoring task, it utilizes a preset resting time window to obtain the target's resting baseline profile B. This profile B is used as a global reference variable for calculating the physiological cognitive load score C, thereby achieving dynamic correction of the real-time monitoring data. The specific process includes: within the preset resting time window before the start of the teaching task, acquiring the original sequence of multidimensional physiological parameters of the target object under a state of no cognitive load through non-contact sensing methods; extracting the heart rate steady-state value and HRV fundamental frequency index from the original sequence using statistical analysis methods, and constructing the individual physiological baseline profile B of the target object based on this, to characterize the individual's basal physiological metabolic level and autonomic nervous system equilibrium state; this profile B can effectively decouple the systematic interference of basal heart rate differences during the adolescent growth and development stage on the evaluation results. During the classroom real-time monitoring stage, the high-confidence pulse wave signal after suppressing motion artifacts is used as a real-time parameter. The absolute difference between the value and the corresponding indicator in the baseline profile B is calculated using a subtractor; a normalization mapping is performed using a divider, and the absolute difference is divided by the baseline value B to calculate the percentage offset. The calculation formula is as follows: Using percentage offset as the core input variable for judging physical and mental state, this normalization mapping mechanism converts the absolute value of physiological signals into a relative offset vector reflecting the degree of psychological stress. The relative offset vector serves as a standardized input variable for multi-level physical and mental state judgment. Its function is to uniformly map the physiological characteristics of different individuals to a dimensionless feature space representing stress state, effectively eliminating the systematic bias caused by differences in baseline heart rate, heterogeneity of physical fitness, and fluctuations in metabolism among students. This enables standardized assessment of students' psychological load and learning focus across individual dimensions.

[0135] In addition, in one embodiment, the physical and mental state assessment method was simulated and its performance evaluated using a self-developed classroom behavior dataset.

[0136] 1. Experimental Environment and Data Foundation: Hardware Environment: The experiment used high-definition infrared enhanced cameras deployed at the front end of the teaching scenario, and the back-end computing platform used an edge inference terminal with high parallel computing capabilities (equipped with NVIDIA RTX series graphics cards). Data Source: A self-developed classroom behavior dataset was used to evaluate the algorithm's effectiveness; this dataset contains multi-dimensional video sequences of 21 subjects in a real classroom environment, covering typical actions such as looking up to listen to the lesson, looking down to write, operating mobile terminals, and large arm swings; the data samples were spatially aligned and encapsulated using tensors to ensure the consistency of input for physiological and behavioral feature extraction.

[0137] 2. Performance Evaluation of the Action Recognition Branch: Comparative tests were conducted on the self-developed dataset for the improved 3D-ResNet network. The experiments focused on examining the ability to capture fine-grained actions after introducing spatiotemporally deformable convolution and temporal gating mechanisms; the test results are shown in Table 1.

[0138] Table 1 - Test Results

[0139] Results analysis: Experimental data demonstrate that by introducing an offset learning branch into the residual block, the system can more accurately fit the nonlinear deformation of the hand holding the tool, effectively solving the technical problem that traditional methods are prone to confusion when distinguishing between writing and non-learning hand movements.

[0140] 3. Analysis of typical manifestations of physiological signal compensation mechanisms:

[0141] For the proposed behavior feedback-driven RLS adaptive filter, its performance was representatively verified by simulating the signal extraction process of students taking notes under dynamic interference:

[0142] Before compensation: Without behavioral branch intervention, skin micro-displacement noise caused by the subject's limb movements is directly superimposed on the physiological signal, resulting in a mean absolute error (MAE) of approximately 8.5 bpm for heart rate extraction. After compensation: By introducing the collaborative compensation mechanism of this invention, using hand movement frequency as a reference input, the system can effectively filter out periodic motion artifacts. In this embodiment, the MAE of heart rate extraction is significantly reduced to approximately 2.2 bpm. Conclusion: The experiment verifies that the technical logic of using behavioral features as a physiological denoising reference source has extremely high stability in complex dynamic environments.

[0143] 4. Case analysis of determining the state of mind-body consistency:

[0144] The source analysis of typical cases of hidden distraction in the experimental sample revealed the following:

[0145] Perceptual characteristics: The target subject maintains a fixed head-down or steady-paced note-taking posture for an extended period, which is easily misjudged as deep focus in traditional purely visual evaluation. Physiological essence: Extracted physiological indicators show the individual's physiological cognitive load score. The student remained in a low-arousal baseline state (close to resting level), and the HRV (Human Reactivity Value) index did not show stress-induced fluctuations. The system, based on multi-level decision-making logic and through conflict detection of psychosomatic characteristics, successfully corrected the identification result to false focus. This case demonstrates the unique technical advantages of this invention in revealing the deep psychological state of students.

[0146] like Figure 2As shown, according to another embodiment of the present invention, a student physical and mental state evaluation system based on behavior-physiological synergistic compensation is also provided, the system comprising:

[0147] The perception region determination module 1 is used to preprocess the acquired video sequence, determine the physiological perception region and the behavior perception region based on the preprocessed video sequence, and perform optical flow compensation and encapsulation processing on the physiological perception region and the behavior perception region respectively.

[0148] Behavioral feature generation module 2 is used to reconstruct the original blood volume pulsation signal and calculate the heart rate variability index based on the physiological sensing region after optical flow compensation, and simultaneously extract hand writing frequency features from the encapsulated behavioral sensing region to generate behavioral features.

[0149] Physiological feature generation module 3 is used to construct a phase-energy dual-dimensional verification mechanism based on the original blood volume pulsation signal and hand writing frequency characteristics, and to use the verification mechanism to distinguish between real motion artifacts and false motion artifacts to generate physiological features.

[0150] The evaluation result acquisition module 4 is used to integrate behavioral characteristics and physiological characteristics, perform multi-level physical and mental state judgment based on the integration results, and calibrate the judgment results in combination with the preset individual physiological benchmark profile to generate student physical and mental state evaluation results.

[0151] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.

Claims

1. A method for evaluating the physical and mental state of students based on behavioral-physiological co-compensation, characterized in that, The method includes: S1. Preprocess the acquired video sequence, determine the physiological perception region and the behavior perception region based on the preprocessed video sequence, and perform optical flow compensation and encapsulation processing on the physiological perception region and the behavior perception region respectively. S2. Based on the physiological sensing region after optical flow compensation, the original blood volume pulsation signal is reconstructed and the heart rate variability index is calculated. Simultaneously, hand writing frequency features are extracted from the encapsulated behavior sensing region to generate behavior features. S3. Based on the original blood volume pulsation signal and hand writing frequency characteristics, a phase-energy dual-dimensional verification mechanism is constructed, and the verification mechanism is used to distinguish between real motion artifacts and false motion artifacts to generate physiological characteristics. S4. Integrate behavioral and physiological characteristics, perform multi-level physical and mental state assessment based on the integration results, and calibrate the assessment results in conjunction with the preset individual physiological benchmark profile to generate student physical and mental state evaluation results. S3 includes: Using the original blood volume pulsation signal and hand writing frequency features as inputs, the overlap between the physiological energy spectrum and the behavioral action frequency distribution is calculated using the normalized cross-correlation formula, and the phase lock value between the behavioral displacement vector and the physiological pixel fluctuation signal within the corresponding frequency band of the hand writing frequency features is calculated simultaneously. By combining overlap, phase lock value and kinetic energy ratio, a phase-energy dual-dimensional verification mechanism is constructed, and the verification mechanism is used to judge motion artifacts to distinguish between real motion artifacts and false motion artifacts. If it is determined to be a real motion artifact, the hand writing frequency characteristics and the original blood volume pulsation signal are used as inputs, and the motion artifact is removed by using an improved recursive least squares filter and harmonic suppression constraints to generate physiological characteristics. If it is determined to be a false motion artifact, the original blood volume pulsation signal is taken as a physiological feature.

2. The student physical and mental state evaluation method based on behavior-physiological synergistic compensation according to claim 1, characterized in that, S1 includes: S11. Obtain the video sequence of the classroom scene, and perform lens distortion correction, automatic gain control and white balance processing on the video sequence in sequence to complete the preprocessing of the video sequence and generate standardized video frames. S12. Input the standardized video frames into the preset lightweight target detection network to locate the student target box and obtain the student target image. Input the student target image into the preset human pose estimation model and output the coordinates of the skeletal key points by combining the heat map regression algorithm. S13. Extract facial key point coordinates from skeletal key point coordinates, extract facial sub-regions based on facial key point coordinates, and perform occlusion judgment on facial sub-regions to obtain physiological perception regions; wherein, the physiological perception regions include main physiological perception regions and backup physiological perception regions. S14. Based on the coordinates of the key points of the skeleton, construct the hand-object interaction area, use the semantic recognition algorithm to capture the geometric edge features of the hand-object interaction area, and perform perspective transformation correction, resampling and normalization processing on the geometric edge features in sequence to generate the behavior perception area. S15. Perform optical flow compensation on the physiological perception region using the dense optical flow algorithm, and encapsulate the behavior perception region using a preset deep learning framework.

3. The student physical and mental state evaluation method based on behavior-physiological synergistic compensation according to claim 2, characterized in that, S13 includes: S131. Based on the coordinates of the skeletal key points, the facial key point coordinates are obtained using face alignment technology, and the facial key point coordinates are mapped to a preset three-dimensional deformable model to construct a three-dimensional local topological structure. S132. Based on the three-dimensional local topology, extract the grid vertex coordinates of the facial sub-regions, and project the grid vertex coordinates using inverse perspective transformation to generate the main physiological perception region; wherein, the facial sub-regions include the forehead center sub-region and the bilateral cheek sub-regions. S133. Calculate the confidence score of facial key points and use the local binary mode operator to analyze the skin texture consistency to obtain the occlusion weight of each facial sub-region. S134. Compare the occlusion weight with a preset threshold. If the occlusion weight exceeds the preset threshold, trigger the failover backup mechanism and obtain the backup physiological perception area as the physiological perception area. If the occlusion weight does not exceed the preset threshold, use the main physiological perception area as the physiological perception area.

4. The student physical and mental state evaluation method based on behavior-physiological synergistic compensation according to claim 3, characterized in that, The triggered failure switching backup mechanism, in which the backup physiological sensing region is obtained as the physiological sensing region, includes: Based on the coordinates of facial key points and mesh vertices, the three-dimensional rotation matrix of the target head is obtained using a pose estimation algorithm and a preset camera intrinsic parameter matrix. The three-dimensional rotation matrix is ​​decomposed into trigonometric functions to obtain the Euler angles of the target head in three-dimensional space; wherein the Euler angles include pitch, yaw and roll angles. Based on Euler angles, the area behind the ear or the skin area on the side of the neck is dynamically located through spatial geometric coordinate transformation to obtain a backup physiological sensing area, and the backup physiological sensing area is used as the physiological sensing area.

5. The student physical and mental state evaluation method based on behavior-physiological synergistic compensation according to claim 1, characterized in that, S2 includes: S21. After optical flow compensation, all pixels in the physiological perception area are subjected to mean dimensionality reduction in the spatial dimension, and continuous sampling is performed along the time dimension to obtain the color time sequence signal. S22. The color time sequence signal is reconstructed using the color method orthogonal projection technology and adaptive fusion technology to obtain the original blood volume pulsation signal. The original blood volume pulsation signal is then filtered using a preset Butterworth bandpass filter. Based on the filtered original blood volume pulsation signal, the heart rate variability index is calculated. S23. Based on the encapsulated behavior perception region, generate a spatiotemporal feature tensor and input the spatiotemporal feature tensor into the improved three-dimensional residual network to extract the spatiotemporal features of hand movements. Combine the fast Fourier transform to extract the hand writing frequency features and obtain the frequency domain features. S24. The spatiotemporal features and frequency domain features are concatenated and feature mapped to generate behavioral features that include attitude angle, line of sight direction and periodicity of movement indicators.

6. The student physical and mental state evaluation method based on behavior-physiological synergistic compensation according to claim 5, characterized in that, The step of extracting handwriting frequency features using Fast Fourier Transform to obtain frequency domain features includes: Based on the hand-object interaction area, the coordinate displacement of key points at the center of the wrist and palm is located between consecutive video frames, and a hand motion vector sequence representing the kinetic energy state of the hand is generated. Perform a fast Fourier transform on the hand motion vector sequence to calculate the power spectral density and obtain the kinetic energy ratio of the target frequency band; The kinetic energy ratio is compared with a preset energy threshold. If the kinetic energy ratio is greater than the preset energy threshold, the hand movement is determined to be a periodic writing behavior, and the hand writing frequency feature is extracted. If the kinetic energy ratio is less than or equal to the preset energy threshold, it is determined to be a non-periodic writing behavior. By combining the kinetic energy ratio with the results of periodic writing behavior determination, frequency domain features of hand movements are generated.

7. The student physical and mental state evaluation method based on behavior-physiological synergistic compensation according to claim 1, characterized in that, S4 includes: S41. Align behavioral and physiological features according to preset timestamps, and perform tensor splicing in the feature channel dimension to generate joint mind-body features; S42. Compare the pitch angle in Euler angles with the preset angle threshold, and combine it with the periodic writing behavior judgment result in the mind-body joint feature to obtain the head-down learning branch. S43. Based on the head-down learning branch, a weighted fusion formula is used to calculate the physiological cognitive load score, and the physiological cognitive load score is compared with the preset high and low arousal thresholds and resting level thresholds. Based on the comparison results, the focus state is distinguished; wherein, the focus state includes deep focus state and false focus state. S44. Compare the pitch angle with the preset rest threshold, and extract the respiratory rate feature based on the original blood volume pulsation signal for cross-validation to obtain the sleep state; S45. Based on the student's focus and sleep status, combined with the preset individual physiological baseline profile, generate the evaluation results of the student's physical and mental state.

8. The student physical and mental state evaluation method based on behavior-physiological synergistic compensation according to claim 1, characterized in that, The method also includes: Within the pre-set resting time window, multidimensional physiological parameters of the target subjects are obtained, and statistical analysis is used to extract heart rate steady-state value and heart rate variability fundamental frequency index from the multidimensional physiological parameters to construct an individual physiological baseline profile of the target subjects. During the real-time classroom monitoring phase, physiological characteristics and heart rate variability indicators are used as parameters, and individual physiological baseline profiles are used as reference standards. The absolute difference between the parameters and the reference standards is calculated, and the absolute difference is converted into a percentage offset through a normalization mapping mechanism to generate a relative offset vector that reflects the real-time psychological stress level of the target object.

9. A student physical and mental state evaluation system based on behavior-physiological synergistic compensation, used to implement the student physical and mental state evaluation method based on behavior-physiological synergistic compensation as described in any one of claims 1-8, characterized in that, The system includes: The perception region determination module is used to preprocess the acquired video sequence, determine the physiological perception region and the behavior perception region based on the preprocessed video sequence, and perform optical flow compensation and encapsulation processing on the physiological perception region and the behavior perception region respectively. The behavioral feature generation module is used to reconstruct the original blood volume pulsation signal and calculate the heart rate variability index based on the physiological sensing area after optical flow compensation, and simultaneously extract the hand writing frequency features from the encapsulated behavioral sensing area to generate behavioral features. The physiological feature generation module is used to construct a phase-energy dual-dimensional verification mechanism based on the original blood volume pulsation signal and hand writing frequency characteristics, and to use the verification mechanism to distinguish between real motion artifacts and false motion artifacts to generate physiological features. The evaluation result acquisition module is used to integrate behavioral and physiological characteristics, perform multi-level physical and mental state assessments based on the integrated results, and calibrate the assessment results in conjunction with preset individual physiological benchmark profiles to generate student physical and mental state evaluation results.

Citation Information

Patent Citations

  • Analysis method and device based on multi-modal information, equipment and medium

    CN120746783A

  • Pen holding posture tremor spectrum analysis method and system based on intelligent wearable device

    CN122440178A