Student mental health prediction method and system based on action recognition
By constructing a multi-scene spatiotemporal database and social relationship map, combined with graph neural networks and long-term memory networks, the limitations of identifying students' depressive behavior in the existing technology are solved, early warning of students' mental health and accurate identification of group communication risks, and the scientificity and efficiency of mental health management are improved.
Patent Information
- Application Number
- CN202510466293.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-15
- Publication Date
- 2025-07-29
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The existing technology has limitations in identifying students' depressive behaviors, and it is impossible to effectively distinguish between exam tension and persistent depression, ignore the transmission effect of depression in the group, and traditional methods such as wearable devices trigger resistance, the psychological scale is limited by students' concealment, and the misjudgment rate based on single-modal expression analysis is high.
By obtaining video data in multiple scenarios, building a spatiotemporal database, establishing a social relationship map, quantifying the physical proximity, interaction synchronization and emotional propagation delay among students, using a two-way timing modeling method, combining graph neural networks and long-term memory networks to make mental health predictions, generating a dynamic social relationship map weight matrix and feedback intervention suggestions in real time.
It has realized early warnings for students' depressive behavior, accurately identified individual abnormalities and group transmission risks, improved the objectivity and scientific nature of mental health monitoring, optimized the allocation of psychological counseling resources, and enhanced the initiative and systemicity of campus mental health management.
Smart Images

Figure CN120388684A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of student mental health prediction, and particularly to a method and system for predicting student mental health based on action recognition. Background Art
[0002] The mental health crisis among teenagers has become a global educational problem. Data from the Blue Book of Mental Health in China shows that 14.8% of teenagers are at risk of depression, and among them, the proportion of hidden depression in the classroom scenario is as high as 63%. Such psychological problems often first manifest through abnormal action behaviors: students with depressive tendencies show significant attenuation of body language (such as a 72% reduction in daily gestures), an increase in social distance (the average distance from classmates increases by 1.5 times), and a loss of action synchronization (response delay exceeds 3 seconds during group activities). However, traditional assessment methods have serious limitations - psychological scales are limited by students' subjective concealment (32% of respondents admitted to fabricating questionnaires), wearable devices cause resistance (68% of students refused to wear EEG caps for a long time in a pilot project in a middle school), and the misjudgment rate of methods based on single-modal expression analysis in group scenarios is as high as 41% (the proportion of cases misjudging fatigue-induced eye closure as low mood is 27%).
[0003] Existing technical solutions have systematic defects in depressive behavior recognition: Although heart rate variability detection (HRV) based on wearable devices can capture physiological fluctuations caused by anxiety, it cannot distinguish the differences between exam stress (instantaneous HRV drops by 28%) and persistent depression; speech emotion analysis completely fails when depressive students keep silent in class; mainstream pose estimation models (such as OpenPose) can detect limb movements, but lack the correlation modeling of persistent micro-movements (such as unilateral shoulder adduction for 20 minutes) and social avoidance behaviors (such as deliberately turning back to organize schoolbags more than 5 times per break while facing away from classmates). More critically, existing systems generally ignore the transmission effect of depressive emotions in groups. For example, an important transmission feature is that the negative action imitation rate of peers within 3 meters in the front and back rows of a depressive student increases by 39%. Summary of the Invention
[0004] The present invention overcomes the deficiencies of the prior art and provides a method and system for predicting student mental health based on action recognition.
[0005] To achieve the above object, the technical solution adopted by the present invention is: A method for predicting student mental health based on action recognition, comprising the following steps:
[0006] S1: Obtain video data in multiple school scenarios respectively, process and track the students in the video data, and construct a spatio-temporal database;
[0007] S2: Associate the spatio-temporal databases of different students, set influence weights according to the interaction frequency and physical distance, and establish a social relationship graph;
[0008] S3: Obtain the characteristics of group behavior in the spatio-temporal database, and perform interactive activity analysis according to the social relationship graph;
[0009] S4: Establish a weight matrix for the dynamic social relationship graph based on the interactive activity analysis, and input the data of each item of interactive activity into the matrix for two-way prediction between students;
[0010] S5: Construct a heat map of the class according to the two-way prediction in S4, and feedback intervention suggestions to the teacher in real time.
[0011] In a preferred embodiment of the present invention, in S1, the video data is specifically the dynamic interactive behavior and static behavior in the whole time period, and the multi-scenarios include scenarios with social behavior records such as the school cafeteria, teaching building, and library.
[0012] In a preferred embodiment of the present invention, in S1, the processing is specifically to use a high frame rate of 10 - 20 fps for the interactive behavior in the video data and a low frame rate of 1 - 10 fps for the static behavior, and associate the student identities in different scenarios with the student seats, and construct a spatio-temporal database of students through face recognition and gait recognition.
[0013] In a preferred embodiment of the present invention, in S2, the associated database is specifically to associate the spatio-temporal databases of different students in different scenarios, and use the seats in the classroom as the initial spatio-temporal database, establish two-dimensional coordinates for each student, mark different students through the two-dimensional coordinates, and make cross-scenario connections through the marks.
[0014] In a preferred embodiment of the present invention, in S2, the establishment of the social relationship graph is specifically to calculate the physical distance between students through the marks, calculate the interaction frequency through the spatio-temporal database, set the weight relationship through the interaction frequency and physical distance, and establish the social relationship graph.
[0015] In a preferred embodiment of the present invention, in S3, the group behavior characteristics are specifically to calculate the action response delay of the target student and the surrounding students or all students, and detect the emotional differences through facial recognition; the interactive activity is specifically to identify the actions of the lips, limbs, and head respectively and analyze according to these actions.
[0016] In a preferred embodiment of the present invention, in S4, the weight matrix of the dynamic social relationship graph uses the behavior characteristics of individuals and groups as input data, builds a model framework for the spread of influence and the capture of temporal behavior trends between students, and thus predicts the mental health between students.
[0017] In a preferred embodiment of the present invention, in S5, the class heat map specifically marks the health risk levels with different colors based on classroom seats, and monitors the depression emotion transmission path in real time.
[0018] A student mental health prediction system based on action recognition includes a data acquisition module, a preprocessing module, a target detection and tracking module, a social relationship modeling module, a group behavior analysis module, a two-way influence prediction module, and an early warning and intervention module.
[0019] In a preferred embodiment of the present invention, the target detection and tracking module adopts a dual-mode adaptive detection algorithm for breaks / classes; the two-way influence prediction module is provided with a graph neural network and a time series model.
[0020] The present invention solves the defects existing in the background technology and has the following beneficial effects:
[0021] (1) The present invention provides a student mental health prediction method and system based on action recognition. By quantifying multi-dimensional features such as the physical proximity, interaction synchrony, and emotion transmission delay among students, a dynamically evolving social influence weight matrix is constructed. This technology not only captures the typical social withdrawal behaviors of depressed students (such as deliberately keeping a distance and reducing physical interactions), but also can identify the ripple effects they trigger in the group (such as the decrease in the activity level of surrounding students' actions and the delay in facial expression feedback). By introducing a two-way time series modeling method, while analyzing the gradual change of individual behaviors (such as the attenuation of gesture frequency over consecutive days), the emotion conduction path at the group level is synchronously tracked (such as the three-degree propagation law of depression emotion along the social network), enabling early warning to focus on both individual abnormalities and group transmission risks.
[0022] (2) The present invention provides a student mental health prediction method and system based on action recognition. It quantifies multi-dimensional social features such as physical proximity, interaction frequency, and emotion synchrony into computable graph structure data, and continuously updates the social relationship strength through a sliding time window mechanism, making the modeling of the depression emotion transmission path conform to the laws of social dynamics and have real-time response capabilities. The design of the two-way attention mechanism breaks through the traditional one-way influence assumption. By separating the attention weights of forward propagation and reverse feedback, it accurately identifies the dominant direction and key nodes of emotion contagion, effectively distinguishing real emotion infection from accidental behavior similarity. Combining with the standardized mapping method of the SCL-90 scale, it converts unstructured behavior features into quantifiable mental health indicators, establishing an interpretive bridge from body language to mental state, making the setting of the early warning threshold conform to both clinical psychology standards and the characteristics of the educational scenario.
[0023] (3) The present invention provides a method and system for predicting students' mental health based on action recognition. Through multi-modal data fusion and an end-to-end computing architecture, the time required to convert the original video stream into actionable intervention strategies is compressed to a single break period, enabling early detection and rapid response to mental health problems. The generation technology of dynamic heatmaps converts complex social relationship and mental state data into an intuitive spatial visualization solution, helping educators quickly locate hotspots of emotional transmission. The introduction of the federated learning framework enables the mining of mental health trends across classes and grades while ensuring the privacy and security of behavioral data, providing data support for schools to formulate hierarchical intervention strategies. This technical system not only improves the objectivity and scientific nature of mental health monitoring, but also significantly enhances the initiative and systematicness of campus mental health management by promptly blocking the chain of negative emotion transmission and optimizing the allocation of psychological counseling resources. BRIEF DESCRIPTION OF THE DRAWINGS
[0024] To more clearly illustrate the technical solutions in the embodiments of the present invention or in the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings described below are only some embodiments recorded in the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0025] Figure 1 is a flowchart of a preferred embodiment of the present invention;
[0026] Figure 2 is a schematic diagram of a prediction system of a preferred embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0027] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, rather than all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the protection scope of the present invention.
[0028] Many specific details are set forth in the following description in order to fully understand the present invention, but the present invention can also be implemented in other ways different from those described herein. Therefore, the protection scope of the present invention is not limited by the specific embodiments disclosed below.
[0029] As shown in the figure, a method for predicting students' mental health based on action recognition includes the following steps:
[0030] S1: Obtain video data in multiple scenarios of the school respectively, process and track the students in the video data, and construct a spatio-temporal database;
[0031] In the present invention, in S1, the video data specifically includes dynamic interactive behaviors and static behaviors throughout the entire period. The multi-scenarios include scenarios with social behavior records such as the school cafeteria, teaching building, and library. In S1, the processing specifically involves using a high frame rate of 10 - 20 fps for the interactive behaviors in the video data and a low frame rate of 1 - 10 fps for the static behaviors, and associating the student identities in different scenarios with the student seats. A spatio-temporal database of students is constructed through face recognition and gait recognition.
[0032] It should be noted that first, video data is collected by monitoring devices deployed in multiple school scenarios, and hierarchical processing is performed according to the dynamic characteristics of different scenarios. In high-dynamic scenarios such as the cafeteria and corridor during breaks, the camera captures fast interactive behaviors such as students running and physical contact at a frame rate of 10 - 20 fps; while in low-dynamic scenarios such as the classroom and the self-study area of the library, a frame rate of 1 - 10 fps is used to record static behaviors such as writing at a desk or lowering the head for a long time. To achieve adaptive frame rate adjustment, the system calculates the motion intensity of the scene in real time through the Motion Energy Image (MEI) algorithm. When it detects that the pixel change between consecutive frames exceeds the dynamic threshold, it automatically switches to the high frame rate mode to balance computational resources and behavior capture accuracy.
[0033] After the video data is collected, the system associates the identities of students across scenarios through multi-modal biometric technology. In the real-time face recognition stage, the lightweight RetinaFace-MobileNet model quickly locates the face area and extracts key points, and then the ArcFace-ResNet50 model generates high-precision feature vectors, which are matched with the face database containing the face features of students. For side face or partially occluded scenarios, 3D face reconstruction technology is introduced to generate virtual frontal images to improve the recognition rate. At the same time, gait recognition technology is combined to make up for the limitations of face recognition: based on the YOLOv7-Pose model, human body key points are detected, the gait cycle is segmented and features are extracted, and the GaitSet model is used to convert the gait sequence into feature vectors, which are cross-validated with the face recognition results to ensure the accuracy of cross-scenario identity association. For example, when a student's face is blocked by a desk in the classroom, the system can match the gait characteristics when he enters the classroom with the face data in the cafeteria scenario to complete identity confirmation.
[0034] To establish a spatio-temporal behavior benchmark, the system maps classroom seats to a two-dimensional grid coordinate system, and each seat is assigned a unique coordinate label (such as A3). In non-classroom scenarios, a global point cloud map is constructed through visual SLAM technology, and the real-time positions of students are mapped to a unified coordinate system. The behavior data of each student is marked with a spatio-temporal label in the format of "student ID_scene code_timestamp_coordinate". For example, "S001_C1_202310011430_A3" means that student S001 is at seat A3 in the classroom at 14:30 on October 1st. These spatio-temporal labels and biometric data (face and gait feature vectors) together constitute a time-series database. Among them, high-frequency dynamic behavior data is temporarily stored in the in-memory database Redis, and low-frequency static data is persistently stored in MySQL, and the biometric privacy is protected by AES-256 encryption. The database design combines spatio-temporal indexes to support fast query of the behavior trajectories of specific students in different scenarios, providing a data basis for subsequent social relationship modeling and mental health prediction.
[0035] S2: Associate the spatio-temporal databases of different students, set influence weights according to the interaction frequency and physical distance, and establish a social relationship graph;
[0036] In the present invention, in S2, the associated database specifically associates the spatio-temporal databases of different students in different scenarios, and uses the seats in the classroom as the initial spatio-temporal database. A two-dimensional coordinate is established for each student, and different students are marked differently through the two-dimensional coordinates. Cross-scene connections are made through the marks. In S2, the establishment of the social relationship graph is specifically as follows: the physical distance between students is calculated through the marks, and the interaction frequency is calculated through the spatio-temporal database. A weight relationship is set through the interaction frequency and physical distance to establish a social relationship graph.
[0037] It should be noted that in step S2, the system first maps the position of each student to a unique identifier in the global coordinate system based on the two-dimensional coordinates of the seats in the classroom scene. Specifically, according to the classroom seat layout table provided by the school, the seats are divided into grids by rows and columns (such as 6 rows × 8 columns), and a coordinate label is assigned to each seat (such as C3 for the second column in the third row). When a student moves or changes seats in the classroom, their coordinates are updated in real time through the target detection and tracking module. For non-classroom scenarios (such as the cafeteria, library), the system constructs a three-dimensional point cloud map based on visual SLAM technology, and aligns the local coordinate systems of different scenarios to the global coordinate system through a coordinate transformation matrix to ensure the consistency of cross-scene positions. For example, the real-time position of a student in the cafeteria can be mapped to (X = 12.5, Y = 7.2) in the global coordinate system after coordinate conversion, which is unified with the coordinate (C3) in the classroom into the same spatial reference system.
[0038] After completing the cross-scenario coordinate mapping, the system extracts students' behavioral data from the spatio-temporal database to quantify the physical distance and interaction frequency. The physical distance is calculated based on the Euclidean distance formula, using the square root of the difference in students' coordinates as the real-time distance between two people. For example, if two students are located at coordinates A2 (X = 2, Y = 1) and B4 (X = 4, Y = 3) respectively, then the distance between them is √[(4 - 2) 2 +(3 - 1) 2 = √8 ≈ 2.83 meters. The interaction frequency is statistically analyzed by analyzing the overlapping behaviors in spatio-temporal data. For example, when two students show synchronized actions (such as raising their hands at the same time, turning their heads to look at each other for more than 2 seconds) or the physical distance is less than 1 meter (such as walking side by side, talking face to face) within the same time period (such as within 5 minutes), it is counted as one effective interaction. The system accumulates the interaction times on a daily basis and generates a dynamic interaction frequency value by combining a time decay factor (such as the interaction weight for the current day is 1, and for the previous day is 0.8).
[0039] Based on the above indicators, the system constructs the node and edge weights of the social relationship graph. Each student serves as a node in the graph, and the node attributes include their coordinates, interaction frequency statistical value, and behavioral feature vector; the edge weight between nodes is determined by a weighted function of physical distance and interaction frequency. The specific formula is: weight W = α × (1 / d) + (β × f), where d is the average physical distance, f is the normalized interaction frequency, and α and β are adjustable parameters (the default values are α = 0.6 and β = 0.4). For example, if the average distance between two students is 3 meters and the single-day interaction frequency is 5 times, then W = 0.6 × (1 / 3) + 0.4 × (5 / 10) = 0.2 + 0.2 = 0.4. To capture the dynamic changes in social relationships, the system updates the graph weights every hour and uses a graph attention network (GAT) to perform embedded learning on node features to identify potential influence propagation paths (such as a student's anxiety emotion spreading to neighboring seat classmates through high-frequency interactions). The finally generated social relationship graph is stored in JSON format, including node coordinates, edge weights, and timestamp metadata, providing structured input for subsequent mental health prediction.
[0040] S3: Obtain the characteristics of group behaviors in the spatio-temporal database, and conduct interaction activity analysis based on the social relationship graph;
[0041] In a preferred embodiment of the present invention, in S3, the group behavior characteristics specifically include calculating the action response delay between the target student and surrounding students or all students, and detecting emotional differences through facial recognition; the interaction activity specifically includes separately identifying the actions of the lips, limbs, and head and analyzing based on these actions.
[0042] It should be noted that in step S3, the system extracts the behavior data of the target student and their surrounding group (i.e., the students around, up, down, left, and right) from the spatio-temporal database, and quantifies the interaction activity through multi-dimensional action feature analysis and emotion difference detection. First, for the calculation of action response delay, the system extracts the 17 key-point time-series data of the human body of the target student and the surrounding students (such as the coordinates of the hands, shoulders, and hips) based on the AlphaPose skeleton key-point detection algorithm, and uses the dynamic time warping (DTW) algorithm to align the time axes of the action sequences, and calculates the trigger delay of specific actions. For example, when the teacher asks a question, the timestamp of the target student's hand-raising action is t1, and the neighboring student raises their hand at time t2. The system uses DTW to match the similarity of the two skeleton-point trajectories and obtains the response delay Δt = |t2 - t1|. If Δt < 3 seconds, it is determined as active interaction; otherwise, it is regarded as a lagged response. For spontaneous behaviors that are not clearly triggered (such as turning the head, getting up), sliding window cross-correlation analysis is used to statistically calculate the synchronization probability of the actions of the target student and the group within 5 minutes, and generate an action response heat map.
[0043] Secondly, the detection of facial emotion differences is achieved through a pre-trained ResNet-18 micro-expression recognition model. The system intercepts the face region image from the video stream. After gray-scale normalization and affine transformation alignment, it is input into the model to output 7 types of emotion scores (anger, anxiety, calm, happiness, sadness, disgust, surprise). To eliminate the interference of light and angle, the model integrates adaptive histogram equalization (CLAHE) and 3D face reconstruction technology to synthesize virtual frontal faces for side face and head-down scenarios. The calculation of emotion differences uses cosine similarity comparison: within the same time period, if the anxiety score of the target student is more than 2 standard deviations higher than the average value of the surrounding student group, it is marked as "emotionally isolated"; if the correlation coefficient between its happiness score and the student group within 3 meters is > 0.8, it is determined as "emotion contagion".
[0044] The comprehensive analysis of interaction activity focuses on the fine-grained recognition of lip, limb, and head movements. Lip movements are analyzed by the LipNet model to identify the optical flow features of the lip-reading area, recognize the opening and closing frequency and duration, and determine whether students are participating in the conversation (for example, opening and closing 3-5 times per second is normal conversation, and less than 2 times is silence); limb movements are based on YOLOv7-Pose to detect the key points of the hands and torso, and count the number and amplitude of large movements (such as waving and patting on the shoulder) within a unit time, and combine the physical distance weights in the social relationship graph (for example, the limb interaction weight within 1 meter is 0.7, and outside 2 meters is 0.3) to calculate the activity index; head movements are analyzed by the head pose estimation model (Hopenet) to resolve the changes in yaw angle and pitch angle, and identify the frequency and duration of nodding and shaking the head. For example, if nodding continuously 5 times within 10 seconds and lasting for more than 3 seconds, it is determined as a high-participation interaction. Finally, the system fuses the three types of action feature and emotion difference data to generate a 0-1 standardized interaction activity score, and corrects the score through the edge weights in the social relationship graph (such as physical distance weight + interaction frequency weight) to form a dynamically updated group behavior feature vector library, providing input for subsequent mental health prediction.
[0045] S4: Establish a dynamic social relationship graph weight matrix based on the interaction activity analysis, and input the data of each item of interaction activity into the matrix for two-way prediction between students;
[0046] In a preferred embodiment of the present invention, in S4, the dynamic social relationship graph weight matrix takes the behavioral characteristics of individuals and groups as input data, builds a model framework for the influence propagation and capture of temporal behavior trends between students, and thus performs mental health prediction between students.
[0047] It should be noted that in step S4, the system constructs a dynamic social relationship graph weight matrix based on the interaction activity data generated in step S3, and realizes the two-way prediction of the mental health status of students through a hybrid architecture that combines a graph neural network and a temporal model. First, the system extracts the interaction activity characteristics of the target student and their associated groups from the spatio-temporal database, including indicators such as action response delay, emotion difference score, lip movement frequency, limb movement amplitude, and head pose change. After standardizing them (Z-score normalization), they are sliced by time window (default 5 minutes) to form temporal feature vectors. These feature vectors are used as node attributes to input into the dynamic social relationship graph, and combined with the existing physical distance and interaction frequency edge weights in the graph to construct a multi-dimensional weight matrix. The weight of each edge in the matrix is dynamically updated through an adaptive weighting function. For example, when the emotion difference score between two students within the current time window is lower than the threshold and the limb movement synchronization rate is higher than 80%, the edge weight will be increased by 0.2, otherwise it will be decreased by 0.1, so as to reflect the change of real-time interaction intensity.
[0048] To model the influence spread and temporal behavior trends among students, the system adopts a hybrid architecture of Graph Attention Network (GAT) and Bidirectional Long Short-Term Memory Network (Bi-LSTM). In the graph neural network part, GAT performs embedding learning on the nodes (students) and edges (interaction relationships) in the dynamic weight matrix, and calculates the influence weights of adjacent nodes through the multi-head attention mechanism (default 4 heads). For example, it identifies that the intensity of a student being affected by the emotions of the students in the front row in the social circle is 3 times that of the students in the back row. In the temporal modeling part, Bi-LSTM processes the temporal feature vectors of the nodes in a sliding window manner to capture long-term behavior patterns (such as the increasing trend of lowering the head continuously for a week) and short-term fluctuations (such as sudden silence within a single class). The outputs of GAT and Bi-LSTM are integrated through a Gate Fusion Module. Among them, the spatial dependence features extracted by GAT and the temporal features extracted by Bi-LSTM are weighted by the Sigmoid function and then input into the fully connected layer to generate the mental health prediction value.
[0049] The realization of bidirectional prediction depends on the dynamic backpropagation mechanism of the social relationship graph. The system not only analyzes the path of the target student being affected by others (forward propagation), but also tracks the radiation effect of their behavior on the surrounding group through the Reverse GAT Layer (backward propagation). For example, when it is detected that student A frequently lowers their head due to anxiety, the model not only predicts the mental health risk of student A itself, but also calculates the attention weight of their behavior on student B in the adjacent seat through backward propagation. If the weight exceeds 0.5, a two-way warning is triggered. Finally, the prediction results are mapped to the symptom dimensions of the SCL-90 scale (such as depression, anxiety, hostility), and the mental health scores (0-100 points) and risk levels (low / medium / high) of each student are output. The spread path of depressive emotions in the class is marked in real time through a visualization interface. For example, the red gradient arrow from student A to student B in the heat map represents the intensity of emotional contagion.
[0050] S5: Construct a heat map of the class according to the bidirectional prediction in S4 and provide real-time feedback on intervention suggestions to the teacher.
[0051] In the present invention, in S5, the class heat map specifically marks the health risk levels with different colors based on the classroom seats, and monitors the spread path of depressive emotions in real time.
[0052] It should be noted that in step S5, based on the two-way prediction results in step S4, the system maps the mental health status of students in the class to the two-dimensional spatial coordinates of classroom seats to generate a dynamically updated mental health heat map. Specifically, the system reads the seat coordinates of each student (such as seat C3 in the second column of the third row) from the spatio-temporal database, and converts their mental health scores (0-100 points) into color gradient values through the piecewise linear interpolation algorithm: scores ≤ 30 points (high risk) are mapped to red (RGB: 255, 0, 0), 30 < scores ≤ 70 points (medium risk) are yellow (RGB: 255, 255, 0), and scores > 70 points (low risk) are green (RGB: 0, 255, 0). The heat map is generated using the Kernel Density Estimation algorithm, and spatial smoothing is performed using a circular kernel function (Epanechnikov kernel) with a radius of 2 meters centered on each seat coordinate, so that the risk levels of adjacent seats affect each other. For example, when there are two medium-risk students (yellow) within 3 seats around a high-risk student (red), the boundary area will show an orange gradient, intuitively reflecting the local aggregation effect of emotion transmission.
[0053] To trace the propagation path of depressive emotions, the system analyzes the edge weights and attention scores of the graph attention network (GAT) in the social relationship graph to identify the key links of influence propagation. For example, if the attention weight of student A (coordinate C3) to student B (coordinate D5) is 0.78 (threshold > 0.6), and the difference in their mental health scores Δ = 15 points (A score 35 → B score 20), then it is determined that there is an emotion contagion path from A to B. The system draws an arrow from A to B on the heat map through the Bresenham line algorithm. The depth of the arrow color is determined by the product of the attention weight and the score difference (such as 0.78 × 15 = 11.7, mapped to dark red RGB: 139, 0, 0), and the width of the arrow thickens as the interaction frequency increases (increasing 1 pixel for every 10 interactions). At the same time, the ant colony algorithm is used to simulate the optimal path of emotion propagation, and a semi-transparent trajectory line is superimposed on the heat map to mark the high-probability channels for the spread of depressive emotions in the class.
[0054] The generation of real-time intervention suggestions relies on the correlation analysis between a predefined expert rule base and prediction results. The system matches intervention strategies from the rule base according to the distribution characteristics of high-risk areas in the heat map (such as three consecutive red seats arranged in a straight line) and the topological structure of the emotional propagation path (such as the cross-row propagation chain length ≥ 4): If it is detected that a certain student's risk level has risen for three consecutive days and affects more than 3 classmates, a first-level warning is triggered, and the suggestion of "Immediately arrange psychological counseling and adjust the seat to the isolation area" is pushed to the teacher's end; If an emotional propagation path involves more than 20% of the students in the class, a second-level warning is generated, and the suggestion is "Carry out group psychological counseling and strengthen the design of classroom interaction". All suggestions are converted into structured text through the Natural Language Generation module (NLG), accompanied by a screenshot of the heat map and a comparison curve of historical data, and are encrypted locally and synchronized in real time to the visualization interface on the teacher's end. Teachers can view the detailed warning reasons (such as "The frequency of this student's head-down movement has increased by 200% in the past week, and the interaction frequency with 3 classmates has decreased by 90%") by clicking on specific seats in the heat map, and directly check the intervention plans recommended by the system to carry out follow-up.
[0055] A student mental health prediction system based on action recognition, including a data acquisition module, a preprocessing module, a target detection and tracking module, a social relationship modeling module, a group behavior analysis module, a two-way influence prediction module, and a warning and intervention module.
[0056] In a preferred embodiment of the present invention, the target detection and tracking module adopts a dual-mode adaptive detection algorithm for break time / class; the two-way influence prediction module is provided with a graph neural network and a time series model.
[0057] It should be noted that a complete technical closed-loop from data acquisition to intervention feedback is constructed through the collaboration of multiple modules. The system first conducts all-time video acquisition through a camera network deployed in multiple campus scenarios (such as classrooms, cafeterias, libraries). Among them, the data acquisition module adopts a dynamic frame rate control strategy, using a high frame rate of 10 - 20fps to capture fast action details for high-interaction scenarios during breaks (such as running in the corridor, queuing in the cafeteria), while switching to a low frame rate mode of 1 - 10fps for static classroom scenarios (such as self-study, listening to lectures), focusing on identifying micro-expression features such as chin resting and head-down. This frame rate adaptive mechanism reduces the amount of raw data while ensuring the accuracy of behavior recognition, significantly alleviating the computational pressure on subsequent modules.
[0058] The preprocessing module conducts scene classification and optimization processing on the original video, and implements differential enhancement according to different scene characteristics. For example, in areas with high occlusion such as the cafeteria, the ESRGAN super-resolution model is used to repair blurred faces, and 480p low-definition images are reconstructed into 1080p available data; while in the classroom scene, the video frame rate is compressed from 5fps to 2fps through key frame extraction technology (such as segments with continuous head-down actions exceeding 5 seconds), which not only preserves the behavioral timing characteristics but also reduces the storage space by 60%. The processed data is linked to the campus one-card system. Through the dual verification of face recognition (ArcFace model) and gait characteristics (GaitSet model), the behavioral data of the same student in different scenes is associated with a unified ID, constructing a student behavior database covering the entire spatio-temporal dimension.
[0059] The target detection and tracking module uses an adaptive algorithm with dual modes of between-class / in-class to achieve precise identity binding. In the between-class mode, an improved YOLOv5 model is deployed to detect fast-moving targets, and the DeepSort algorithm is combined to track high-dynamic behaviors such as chasing and fighting across cameras, with a trajectory matching accuracy of 92%; while in the in-class mode, the lightweight MobileNetV3 network is switched to continuously track static features such as the fine-tuning of students' sitting postures and changes in facial orientations with low computational overhead. To address common occlusion problems (such as the front-row students blocking the back-row), the system introduces a gait recognition compensation mechanism, and more than 90% of the occlusion scene identity confirmations are achieved by analyzing sitting postures (such as the left leg being slightly bent and the back tilt angle), ensuring the continuity of cross-class behavioral data.
[0060] The social relationship modeling module constructs a dynamic weight matrix based on the spatio-temporal database to quantify the interaction intensity between students. This module maps classroom seats to a two-dimensional grid coordinate system (such as student A's coordinates (3, 2)), calculates the physical proximity weight through Euclidean distance (such as the weight coefficient of 0.8 for front-row and back-row students, 0.5 for students in alternate rows), and at the same time counts the cross-scene interaction frequency (such as the number of times of dining at the same table in the cafeteria and studying adjacent to each other in the library), and generates a dynamic social relationship graph in combination with a time-series decay factor (the weight decreases if there is no interaction for three consecutive days). To achieve cross-scene relationship reasoning, the system uses graph embedding technology to project the multi-dimensional behavioral characteristics of students (such as action response delay, emotional synchrony) into a low-dimensional vector space, and discovers implicit social associations through vector similarity calculation (such as students who are often accompanied in the library are given a higher weight even if their seats are far apart).
[0061] The group behavior analysis module captures mental health conduction signals through multi-modal feature fusion. This module uses the optical flow method to calculate the action response delay. For example, when the probability that student B shows a similar action within 2 seconds after depressed student A bows his head suddenly increases from 20% to 65%, it is marked as an emotion contagion event. At the same time, an improved ResNet-18 model is used to analyze facial micro-expressions to quantify the group emotion difference (such as triggering an alarm when the expression synchronization of students in the same study group is lower than the threshold). For the assessment of interaction activity, the system extracts three types of indicators: the lip opening and closing frequency (detecting lip key points through OpenPose), the number of nods (counting the changes in neck joint angles), and the gesture amplitude (analyzing the movement trajectories of limb key points), and constructs a weighted scoring model (such as 0.4 for lips, 0.3 for head, and 0.3 for limbs) to achieve the mapping from body language to mental state.
[0062] The two-way influence prediction module integrates the graph neural network (GNN) and the long short-term memory network (LSTM) to construct a dynamic model of mental health propagation. The GNN branch takes the social relationship graph as the input and calculates the influence weights between nodes through the graph attention mechanism (GAT). For example, the positive attention weight of a depressed student to the classmate in front of him can reach 0.81, while the reverse weight is only 0.32, clearly distinguishing the primary and secondary relationships of emotion propagation. The LSTM branch analyzes the temporal evolution of individual behaviors, such as detecting the gradual process of a certain student's activity decreasing by 15% for 5 consecutive days. The features of the two branches are cross-modally aligned in the fusion layer, and finally, a dual prediction result including the individual depression probability (mapped based on the SCL-90 scale) and the group contagion chain (such as "Student A → Deskmate B → Front-row C") is output. The prediction accuracy reaches 89%, which is 23% higher than the traditional single-modal method.
[0063] The early warning and intervention module converts the prediction results into actionable strategies in the educational scenario. The system generates a dynamic heat map based on the classroom seat layout, using a red-orange-yellow three-color gradient to label the risk levels. When more than 3 medium- and high-risk students gather in a certain area, it is automatically marked as an emotion propagation hot spot. The teacher's interface provides multi-dimensional intervention suggestions: short-term measures such as immediately adjusting the seats (moving the depressed student from (3,2) to the edge position (5,1) to block the propagation path), and long-term solutions recommend customized psychological counseling courses (such as the deskmate mutual assistance plan). To achieve privacy protection, all visualized data is anonymized, only coordinate information is retained, and detailed behavior records can be accessed only after double permission verification. The full-process delay of the system from data collection to generating strategies is controlled within 5 minutes, supporting teachers to complete the intervention response within a single class period, and reducing the incidence of abnormal behaviors in the target class in actual applications.
[0064] Based on the enlightenment of the ideal embodiments of the present invention, through the above description, relevant personnel can completely make various changes and modifications without departing from the technical idea of this invention. The technical scope of this invention is not limited to the content in the specification, and the technical scope must be determined according to the scope of the claims.
Claims
1. A method for predicting the mental health of students based on action recognition, characterized in that, It includes the following steps: S1: Obtain the video data in multiple scenarios of the school respectively, process and track the students in the video data, and construct a spatio-temporal database; S2: Associate the spatio-temporal databases of different students, set influence weights according to the interaction frequency and physical distance, and establish a social relationship graph; S3: Obtain the characteristics of group behaviors in the spatio-temporal database, and conduct interaction activity analysis according to the social relationship graph; S4: Establish a weight matrix of the dynamic social relationship graph according to the interaction activity analysis, and input the data of each item of interaction activity into the matrix for two-way prediction between students; S5: Construct a heat map of the class according to the two-way prediction in S4, and feedback intervention suggestions to teachers in real time.
2. The method for predicting the mental health of students based on action recognition according to claim 1, wherein: In the S1, the video data specifically refers to the dynamic interaction behaviors and static behaviors in the whole period, and the multiple scenarios include the cafeteria, teaching building, library, etc. of the school with social behavior records.
3. A method for predicting students' mental health based on action recognition according to claim 1, characterized in that: In the S1, the processing specifically means using a high frame rate of 10 - 20fps for the interaction behaviors in the video data and a low frame rate of 1 - 10fps for the static behaviors, and associating the student identities in different scenarios with the student seats, and constructing the spatio-temporal database of students through face recognition and gait recognition.
4. The method for predicting the mental health of students based on action recognition according to claim 1, wherein: In the S2, the database association specifically means associating the spatio-temporal databases of different students in different scenarios, and taking the seats in the classroom as the initial spatio-temporal database, establishing two-dimensional coordinates for each student, marking different students through the two-dimensional coordinates, and making cross-scenario connections through the marks.
5. The method for predicting the mental health of students based on action recognition according to claim 4, wherein: In the S2, the establishment of the social relationship graph specifically means calculating the physical distance between students through the marks, calculating the interaction frequency through the spatio-temporal database, setting the weight relationship through the interaction frequency and physical distance, and establishing the social relationship graph.
6. The method for predicting the mental health of students based on action recognition according to claim 1, wherein: In the S3, the group behavior characteristics specifically mean calculating the action response delay between the target student and the surrounding students or all students, and detecting the emotional differences through facial recognition; the interaction activity specifically means respectively identifying the lip, limb, and head movements and analyzing according to these movements.
7. A method for predicting students' mental health based on action recognition according to claim 1, characterized in that: In the S4, the weight matrix of the dynamic social relationship graph takes the individual and group behavior characteristics as input data, builds a model framework for the influence propagation and capturing of the temporal behavior trend between students, and conducts mental health prediction between students based on this.
8. A method for predicting the mental health of students based on action recognition according to claim 1, characterized in that: In the S5, the class heat map specifically means marking the health risk levels with different colors based on the classroom seats, and monitoring the depression emotion transmission path in real time.
9. A student mental health prediction system based on action recognition, a system constructed based on the prediction method according to any one of claims 1-8, characterized in that: It includes a data acquisition module, a preprocessing module, an object detection and tracking module, a social relationship modeling module, a group behavior analysis module, a two-way influence prediction module, and a warning and intervention module.
10. The student mental health prediction system based on action recognition according to claim 9, characterized in that: The object detection and tracking module adopts a dual-mode adaptive detection algorithm for between-class / within-class; the two-way influence prediction module is equipped with a graph neural network and a temporal model.
Citation Information
Cited By
Education scene-oriented group emotion thermodynamic diagram generation and interaction feedback system
CN121146333A
Intelligent classroom interactive teaching software and teaching system
CN122115171A