Student mental health assessment method fusing multi-source biological behavior data
By collecting multi-source biological behavior data and constructing a fusion model, and dynamically adjusting the assessment weights, the problems of insufficient data and timeliness in the existing student mental health assessment are solved, thus achieving accuracy and real-time intervention in student mental health assessment.
Patent Information
- Application Number
- CN202511802822.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-03
- Publication Date
- 2026-02-13
AI Technical Summary
Existing technologies for assessing the mental health of middle school students rely on subjective scales or single-modal data, resulting in limited data dimensions, generalized age-related characteristics, and insufficient timeliness of assessments, making it difficult to form a real-time and reliable intelligent intervention system.
Collect multi-source biological behavior data, including physiological indicators, classroom behavior videos, voice data, and campus environment interaction data, perform preprocessing, quality evaluation, and spatiotemporal alignment, construct a fusion model for joint representation learning, dynamically adjust evaluation weights, output the Mental Health Index (MHI), and trigger graded early warning and intervention.
It has improved the accuracy and reliability of student mental health assessment, reduced misjudgments and omissions, shortened response time, and formed a closed-loop intelligent intervention system.
Smart Images

Figure CN121528541A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the fields of educational informatization and artificial intelligence technology, and in particular to a method for assessing student mental health by integrating multi-source biological behavioral data. Background Technology
[0002] Student mental health assessment refers to the systematic collection, analysis, and judgment of students' psychological characteristics and behavioral performance in areas such as cognition, emotion, will, personality, and social adaptation using psychological methods and techniques. This process aims to determine their mental health level, identify potential psychological problems, and provide a basis for educational intervention, counseling support, or referral for treatment. Currently, existing technologies for student mental health assessment mainly rely on subjective scales or single-modal data (such as video facial expressions), which suffer from problems such as limited data dimensions, generalization of age-related characteristics, insufficient assessment timeliness, and inefficient resource allocation. For example, reliance on scales leads to subjectivity and missed detections; semester-based assessments struggle to promptly capture sudden psychological crises; and the mismatch between the number of mental health teachers and the number of at-risk students results in low utilization. While existing technologies propose age-based data collection or classroom video recognition, they lack key capabilities such as real-time bio-behavioral analysis, data quality verification, and multimodal temporal alignment, making it difficult to form a reliable, closed-loop intelligent intervention system.
[0003] Existing technologies for student mental health assessment, such as the publicly available technology "Student Mental Health Assessment System Based on Big Data" (publication number CN118335336A), introduce relative change indicators and key performance indicators, dynamically adjusting the weights of auxiliary recognition tasks during training. This technology primarily relies on historical assessment data (questionnaires, facial expressions, text, behavioral records, etc.) and cannot capture implicit psychological stress responses.
[0004] In addition, the published technology, such as "Patent Publication No. CN120413077A, titled 'Classification and Early Warning Method for Mental Health of College Students Based on Multimodal Deep Learning'", although it integrates text and static features (such as GPA and family background), does not collect non-perceptual biological behavioral data (such as classroom videos, voice tone, and campus card behavior trajectory). The data granularity is relatively coarse, and it is easy to miss non-verbal and non-cognitive psychological abnormal signals.
[0005] Therefore, there is an urgent need to propose a logically simple, accurate and reliable method for assessing student mental health by integrating multi-source biological behavioral data. Summary of the Invention
[0006] To address the above problems, the purpose of this invention is to provide a method for assessing student mental health by integrating multi-source biological behavioral data. The technical solution adopted by this invention is as follows: A student mental health assessment method that integrates multi-source biological behavioral data includes the following steps: Multi-source biological behavioral data of students were collected; the multi-source biological behavioral data includes physiological indicators, classroom behavior video data, voice data, and campus environment interaction data. Real-time acquisition of preset core indicators from multi-source biological behavior data; The multi-source biological behavior data is preprocessed; Quality assessment and spatiotemporal alignment were performed on the preprocessed multi-source biological behavior data to obtain multimodal features after quality verification and spatiotemporal alignment; The initial assessment weights are loaded according to the student's grade level, and the initial assessment weights are dynamically adjusted based on real-time preset core indicators to obtain dynamic assessment weights. A fusion model is constructed, and the multimodal features that have been quality-verified and spatiotemporally aligned are input into the fusion model for joint representation learning, and the emotion dimension features are output. The Mental Health Index (MHI) is derived based on dynamic evaluation weights and emotional dimension features. Preset trigger conditions and trigger graded early warnings and interventions based on the value of the Mental Health Index (MHI).
[0007] Compared with the prior art, the present invention has the following beneficial effects: This invention obtains multi-source biological behavioral data of students by collecting data including physiological indicators, classroom behavior video data, voice data, campus environment interaction data, etc. It uses biological behavioral data to ensure the accuracy and reliability of the assessment.
[0008] This invention obtains multimodal features after quality verification and spatiotemporal alignment by performing quality evaluation and spatiotemporal alignment on preprocessed multi-source biological behavior data. The time compensation ensures the reliability and fusionability of the data.
[0009] This invention loads initial assessment weights based on the student's grade level, dynamically adjusts the initial assessment weights based on real-time preset core indicators to obtain dynamic assessment weights, and dynamically calibrates the weights according to grade level and abnormal events to reduce misjudgments and omissions caused by grade level generalization.
[0010] This invention constructs a fusion model in which a parallel multi-branch convolutional module extracts multi-scale features (i.e., fine-grained local features, downsampled feature maps, large-scale local features, and medium-scale features). The output feature maps are concatenated along the channel dimension to ensure comprehensive and reliable feature extraction. Furthermore, this invention employs a main capsule layer and an emotion capsule layer, and utilizes a dynamic routing mechanism to iteratively calculate the coupling coefficient between lower-level capsules and higher-level emotion capsules. Finally, the emotion capsule layer outputs emotion-dimensional features. Moreover, the multi-branch convolutional module, main capsule layer, and emotion capsule layer of this invention employ dynamic routing to improve the robustness of emotion recognition and state estimation.
[0011] This invention calculates the Mental Health Index (MHI) based on dynamic assessment weights and emotional dimension characteristics, which facilitates risk identification and graded intervention, and shortens the response time window.
[0012] In summary, this invention has the advantages of simple logic and high accuracy and reliability, and has high practical and promotional value in the fields of educational informatization and artificial intelligence. Attached Figure Description
[0013] To more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings used in the embodiments will be briefly introduced below. It should be understood that the following drawings only show some embodiments of the present invention and should not be regarded as a limitation on the scope of protection. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.
[0014] Figure 1 This is a logic flowchart of the present invention.
[0015] Figure 2 This is a schematic diagram of the multi-source data acquisition deployment of the present invention.
[0016] Figure 3 The flowchart of the dynamic weight adjustment for school age classification in this invention.
[0017] Figure 4 The multimodal feature fusion model structure diagram of the present invention.
[0018] Figure 5 The flowchart of the graded early warning and intervention response of this invention. Detailed Implementation
[0019] To make the objectives, technical solutions, and advantages of this application clearer, the present invention will be further described below with reference to the accompanying drawings and embodiments. The embodiments of the present invention include, but are not limited to, the following embodiments. All other embodiments obtained by those skilled in the art based on the embodiments in this application without inventive effort are within the scope of protection of this application.
[0020] In this embodiment, the term "and / or" is merely a description of the relationship between related objects, indicating that there can be three relationships. For example, A and / or B can represent three situations: A exists alone, A and B exist simultaneously, and B exists alone.
[0021] The terms "first" and "second," etc., used in the specification and claims of this embodiment are used to distinguish different objects, not to describe a specific order of objects. For example, "first target object" and "second target object," etc., are used to distinguish different target objects, not to describe a specific order of target objects.
[0022] In the embodiments of this application, the terms "exemplary" or "for example" are used to indicate that something is an example, illustration, or description. Any embodiment or design that is described as "exemplary" or "for example" in the embodiments of this application should not be construed as being more preferred or advantageous than other embodiments or design. Specifically, the use of the terms "exemplary" or "for example" is intended to present the relevant concepts in a specific manner.
[0023] In the description of the embodiments in this application, unless otherwise stated, "multiple" means two or more. For example, multiple processing units means two or more processing units; multiple systems means two or more systems.
[0024] like Figures 1 to 5 As shown, this embodiment provides a method for assessing student mental health by integrating multi-source biological behavioral data. In this embodiment, a middle school student S with depressive tendencies is used as an example to conduct a student mental health assessment, which includes the following steps: The first step involved collecting multi-source biological behavioral data from students. This data included physiological indicators, classroom behavior video data, audio data, and campus environment interaction data. For example, in the classroom environment, a light field camera array (64 lenses, 8K@120fps) deployed at the four corners continuously captured student S's classroom behavior. Analysis using the ST-GCN model extracted his joint movement trajectory and revealed that his head-down rate remained above 90% for three consecutive days. Simultaneously, a directional microphone array collected his classroom audio; Whisper model analysis showed a 40% decrease in his tone entropy compared to baseline, indicating emotional flattening. In the campus public environment, the system retrieved student S's campus card records from the campus data center, finding a 300% surge in his nighttime (after 10 PM) convenience store purchases within a week. Access control records also showed an increased frequency of his early return to his dormitory after evening self-study, indicating a tendency towards social avoidance. At the individual student level, a smart bracelet worn by student S, integrating flexible electronic technology, continuously collected his heart rate variability (HRV) and skin conductance response (GSR). The data showed an upward trend in his resting physiological stress level.
[0025] The second step is to preprocess the collected multi-source biological behavior data, specifically: (21) The collected physiological signals are filtered for noise reduction, outlier processing, and characterization, including: (211) Filtering and denoising: A zero-phase Butterworth bandpass filter was used, with the passband set to 0.04-0.4Hz for heart rate variability signals in physiological indicators and 0.05-2Hz for skin conductance response signals. In this way, baseline drift and high-frequency noise were removed, and its transfer function was... satisfy: Where f is the frequency, f cHere, n is the cutoff frequency, and n is the filter order.
[0026] (212) Outlier handling: An improved interquartile range method based on a sliding window is used to identify outliers; the upper quartile Q3 and lower quartile Q1 of the signal are calculated within a window of length M, and data points S(t) that do not satisfy the following formula are considered outliers:
[0027] Where Q1 represents the lower quartile, which is the value of 25% of the data in the dataset; Q3 represents the upper quartile, which is the value of 75% of the data in the dataset; and IQR represents the interquartile range, which is used to measure the dispersion of the data. This represents the scaling factor, with a value range of 1.5 to 3.0.
[0028] Outliers are repaired using spline interpolation, and the cubic spline interpolation function satisfies: in, Indicates the first each interval cubic spline function on; This represents the spline coefficients, which are solved using the continuity condition; The x-coordinate (time) of the known data points.
[0029] (213) Characterization: Extract time-domain features, frequency-domain features, and nonlinear features from the preprocessed physiological signals, and perform Z-score normalization on all features: Where F represents the extracted original feature values (i.e., time domain features, frequency domain features, and nonlinear features). This represents the mean of student S's historical data; This represents the standard deviation of student S.
[0030] (22) Perform human joint point trajectory restoration on classroom behavior video data, including the following steps: Two-dimensional or three-dimensional keypoint coordinates are extracted from classroom behavior video data using a deep learning-based human pose estimation model. This human pose estimation model is a mature existing algorithm and will not be elaborated upon here. For coordinate loss due to occlusion, linear interpolation based on kinematic constraints is used for filling in the missing coordinates. for: in, This represents the last valid coordinate before the missing segment; Indicates the first valid coordinate after the missing segment; , and For time indexing, satisfying 0 < < .
[0031] (23) Effective segment detection of speech data includes: combining energy-based speech activity detection and a convolutional neural network-based speech / non-speech classifier to accurately segment the speech segments of the target student from a continuous audio stream. The energy-based detection uses the following expression for short-time energy E and zero-crossing rate ZCR: Where E represents short-time energy, which measures signal strength; It represents the nth sampling point of the audio signal; N represents the total number of sampling points in a frame of audio; ZCR represents the zero-crossing rate, which measures the frequency components of the signal; This represents the sign function, which returns the sign of the argument. When E is greater than the energy threshold... And ZCR is less than the zero-crossing rate threshold. If so, it is determined to be a speech frame.
[0032] (24) Perform behavioral personalization baseline calibration on campus environment interaction data, including: (241) Calculate the relative deviation between the current behavioral sequence value and its personal dynamic baseline, expressed as:
[0033] in, This represents the individual's dynamic baseline corresponding to the current behavior; This represents the dynamic assessment weight of the school-age stratification for the i-th feature; L represents the length of the historical data window, i.e., the number of historical time steps used to calculate the baseline. For example, L=14 means using data from the past 14 days. This represents the behavioral data value (such as the number of consumptions) at a historical time t−i.
[0034] (242) Utilize the individual dynamic baseline corresponding to the current behavior Obtain the normalized behavioral feature values Its expression is: Where V(t) represents the original behavior value at the current time t; ϵ represents a constant set to prevent division by zero.
[0035] The third step involves quality assessment and spatiotemporal alignment of the preprocessed multi-source biological behavior data to obtain multimodal features that have undergone quality verification and spatiotemporal alignment. Here, quality assessment includes calculating the quality coefficient of the preprocessed classroom behavior video data. Its expression is:
[0036] in, This indicates the number of valid frames in the video, that is, the number of frames in the captured video that meet the quality requirements (such as clear image, moderate lighting, no serious obstruction, etc.). This indicates the total number of video frames, referring to the total number of captured video frames. The time deviation refers to the difference between the timestamp of a video frame and the ideal timestamp. It is used to evaluate the temporal consistency of video frames; the smaller the deviation, the better the synchronization. Here, the quality coefficient of classroom behavior video data... Used to quantify the credibility of video data, with a value range of [0,1]. If the value is less than 0.6, a re-acquisition or parameter adjustment will be triggered.
[0037] In addition, the spatiotemporal alignment uses a timestamp interpolation compensation method to obtain the adjusted timestamp. Its expression is:
[0038] in, The original timestamp of the preprocessed multi-source biological behavior data is the timestamp stamped by the receiving end's system clock when the data packet is captured at the receiving end. This time point already includes the time spent transmitting the data in the network. This represents the packet loss rate, which is the proportion of data packets lost during data transmission. This represents the average round-trip time, which is the average time required for a data packet to travel from the sender to the receiver and back to the sender.
[0039] The fourth step involves loading initial assessment weights based on the student's grade level, and then dynamically adjusting these initial assessment weights based on real-time preset core indicators to obtain dynamic assessment weights. For example, the weight adjustments for each grade level are shown in Table 1: Table 1: Weight Adjustments for Each Grade Level Educational Stage Key Indicators Initial weights Dynamic adjustment rules primary school students Number of conflicts, frequency of peer interaction <![CDATA[W1=0.6]]> <![CDATA[The number of conflicts ↑ 20% → W1 is increased to 0.75]]> middle school student Academic stress index, nighttime activity duration <![CDATA[W2=0.7]]> <![CDATA[Learning duration > 8h / day → W2 is increased to 0.85]]> college students Frequency of psychological counseling, consistency of self-report <![CDATA[W3=0.5]]> <![CDATA[Self-evaluation consistency ↓ 30% → W3 improved to 0.65]]> First, based on student registration information, student S was identified as a middle school student and assigned the following core indicators for this educational stage: academic stress index and nighttime activity time, with an initial weight of W2=0.7. Real-time monitoring revealed that student S had "daily effective study time > 8 hours" and "a surge in nighttime activity time." The dynamic assessment weight for student S was then increased from 0.7 to 0.85. This updated weight was immediately applied to subsequent mental health index calculations, making the fusion model more focused on the academic stress dimension, which is most sensitive for middle school students.
[0040] The fifth step is to construct a fusion model and input the multimodal features that have been quality-verified and spatiotemporally aligned into the fusion model for joint representation learning, and output the emotion dimension features.
[0041] Here, the fusion model includes a parallel multi-branch convolutional module, a main capsule layer, and an emotion capsule layer connected in sequence; the parallel multi-branch convolutional module acquires multimodal features after quality verification and spatiotemporal alignment, and extracts multi-scale features; the main capsule layer acquires multi-scale features and converts them into capsule vectors to obtain the main capsule vector set; the emotion capsule layer acquires the main capsule vector set and emotion dimension features.
[0042] The parallel multi-branch convolution module includes four branches configured in parallel: a first branch, a second branch, a third branch, and a fourth branch. The first branch employs a first 1×1 convolutional layer and performs a linear transformation along the channel dimension to extract fine-grained local features and reduce computational complexity. Its input is a multimodal feature map. Applying a 1×1 convolution kernel, the output feature map is... .in, To determine the number of output channels, the output of the first 1×1 convolutional layer is directly passed to the channel splicing layer.
[0043] In addition, the second branch consists of a second 1×1 convolutional layer and an average pooling layer; it uses the second 1×1 convolutional layer to adjust the number of channels and the average pooling layer to capture global contextual information, and its input is also a multimodal feature map. The second branch outputs a downsampled feature map. ,in, This is the pooling kernel size for average pooling. The output of the second branch is passed to the channel splicing layer.
[0044] The third branch consists of a third 1×1 convolutional layer and a 6×6 convolutional layer. It uses the third 1×1 convolutional layer for dimensionality reduction and the 6×6 convolutional layer to capture large-scale local features, outputting a feature map. Additionally, the fourth branch is a 3×3 convolutional layer, whose output feature map... The output feature maps of the four branches are concatenated along the channel dimension to obtain the fused feature map. (That is, multi-scale features).
[0045] In this embodiment, the concatenated feature map is input into the capsule network, passing sequentially through the main capsule layer and the emotion capsule layer. A dynamic routing mechanism iteratively calculates the coupling coefficient between the lower-level capsules (i.e., the main capsule layer) and the higher-level emotion capsules (i.e., the emotion capsule layer). Finally, the emotion capsule layer outputs the emotion dimension features. The main capsule layer converts the feature map into capsule vectors, where each capsule represents an entity and its pose, and its input is the fused feature map. It generates a set of 8D capsule vectors through convolution operations. and output the main capsule vector set. . The emotion capsule layer uses a dynamic routing mechanism to map low-level capsules to high-level emotion capsules (such as valence and arousal). Here, the emotion capsule layer receives low-level capsule vectors from the main capsule layer and maps them to high-level emotion capsule vectors through the dynamic routing mechanism. Here, the input of the capsule network is the set of main capsule vectors U, and the predicted vector is calculated , . Among them, is the transformation matrix.
[0046] Here, the updated coupling coefficient has the following expression: ; ; ; Among them, represents the logit of the prior probability between the low-level capsule in the main capsule layer and the high-level capsule j in the emotion capsule layer. It is initialized to 0 at the beginning of dynamic routing and is continuously updated through iteration. represents the logit of the prior probability between the low-level capsule in the main capsule layer and the high-level capsule k in the emotion capsule layer. k represents the summation index of the high-level capsules in the emotion capsule layer, indicating that when calculating the denominator, all possible high-level capsules k need to be traversed. represents the output vector of the high-level capsule j, the final output after compression processing. represents the input vector of the high-level capsule j, the weighted prediction sum of all low-level capsules for the high-level capsule j.
[0047] Step 6: Based on the dynamic evaluation weight and emotional dimension features, the mental health index MHI is obtained, and its expression is Among them, σ represents the Sigmoid function; represents the dynamic evaluation weight of the school-age stage classification of the m-th feature; represents the m-th normalized emotional dimension feature; b represents the bias Step 7: Preset the trigger conditions and trigger hierarchical early warnings and interventions according to the value of the mental health index MHI. For example: Level 1 high-risk warning: When MHI > 0.85 and there are preset high-risk keywords in the voice data, the system starts the crisis intervention of a psychologist within 10 minutes.
[0048] Level 2 medium-risk warning: When 0.7 < MHI ≤ 0.85 and the social withdrawal behavior lasts for ≥ 3 days, the system automatically pushes the VR social training course to the student terminal and notifies the head teacher.
[0049] Level 3 attention warning: When the abnormal rate of a single core evaluation dimension index exceeds 30%, the system sends suggestions for adjusting the learning plan to the head teacher.
[0050] Step 8: Calculate the current proportion of students under warning and dynamically allocate psychological counseling resources; the number of psychological counselors required for the dynamic allocation of psychological counseling resources is... Its expression is:
[0051] in, Indicates basic configuration; Indicates the current number of people under warning; This indicates the total number of registered students.
[0052] In addition, students receiving intervention can be periodically reassessed, and intervention strategies can be optimized based on the reassessment results, forming a closed loop of assessment-early warning-intervention-reassessment. For example, after two weeks of intervention, the system reassessed student S and calculated that his MHI value had dropped to 0.62, indicating a significant reduction in risk. This result was recorded and archived by the system for optimizing future models and intervention strategies.
[0053] The above embodiments are merely preferred embodiments of the present invention and are not intended to limit the scope of protection of the present invention. Any changes made based on the design principles of the present invention, or any non-creative modifications made thereon, shall fall within the scope of protection of the present invention.
Claims
1. A student mental health assessment method integrating multi-source biological behavioral data, characterized in that, Includes the following steps: Multi-source biological behavioral data of students were collected; the multi-source biological behavioral data includes physiological indicators, classroom behavior video data, voice data, and campus environment interaction data. Real-time acquisition of preset core indicators from multi-source biological behavior data; The multi-source biological behavior data is preprocessed; Quality assessment and spatiotemporal alignment were performed on the preprocessed multi-source biological behavior data to obtain multimodal features after quality verification and spatiotemporal alignment; The initial assessment weights are loaded according to the student's grade level, and the initial assessment weights are dynamically adjusted based on real-time preset core indicators to obtain dynamic assessment weights. A fusion model is constructed, and the multimodal features that have been quality-verified and spatiotemporally aligned are input into the fusion model for joint representation learning, and the emotion dimension features are output. The Mental Health Index (MHI) is derived based on dynamic evaluation weights and emotional dimension features. Preset trigger conditions and trigger graded early warnings and interventions based on the value of the Mental Health Index (MHI).
2. The student mental health assessment method integrating multi-source biological behavioral data according to claim 1, characterized in that, Preprocessing of the multi-source biological behavior data includes: The collected physiological signals are filtered for noise reduction, outlier processing, and characterization. Human joint point trajectory restoration of classroom behavior video data; Detect valid segments from the speech data; Personalized baseline calibration of interactive data in the campus environment.
3. The student mental health assessment method integrating multi-source biological behavioral data according to claim 1, characterized in that, The process of filtering and denoising, handling outliers, and characterizing the physiological signals from the collected indicators includes: A zero-phase Butterworth bandpass filter was used to filter and denoise the physiological signals. An improved interquartile range method based on a sliding window was used to identify outliers in filtered and denoised physiological signals. Interpolation repair was performed using a spline-based interpolation method to obtain the repaired physiological signal. Extract time-domain features, frequency-domain features, and nonlinear features from the repaired physiological signals; The process of repairing human joint trajectory in classroom behavior video data includes: extracting two-dimensional or three-dimensional joint coordinates from the classroom behavior video data using a human posture estimation model, and filling them using a linear interpolation method based on kinematic constraints. The personalized baseline calibration of interactive data in the campus environment includes: calculating the relative deviation between the current behavioral sequence value and its personal dynamic baseline, expressed as: in, This represents the individual's dynamic baseline corresponding to the current behavior; represents the dynamic evaluation weight of the school age classification for the i-th feature; L represents the length of the historical data window; This represents the behavioral data value at historical time t−i; Utilize the personal dynamic baseline corresponding to the current behavior Obtain the normalized behavioral feature values Its expression is: Where V(t) represents the original behavior value at the current time t; ϵ represents a constant set to prevent division by zero.
4. The student mental health assessment method integrating multi-source biological behavioral data according to claim 1, 2, or 3, characterized in that, Quality assessment of preprocessed multi-source biological behavior data; the quality assessment includes calculating the quality coefficient of the preprocessed classroom behavior video data. Its expression is: in, Indicates the number of valid frames in the video; Indicates the total number of frames in the video; Indicates time deviation.
5. The student mental health assessment method integrating multi-source biological behavioral data according to claim 4, characterized in that, The preprocessed multi-source biological behavior data undergoes spatiotemporal alignment; the spatiotemporal alignment uses a timestamp interpolation compensation method to obtain adjusted timestamps. Its expression is: in, This represents the original timestamp of the preprocessed multi-source biological behavior data; Indicates packet loss rate; This indicates the average round-trip time.
6. The student mental health assessment method integrating multi-source biological behavioral data according to claim 1, 2, or 3, characterized in that, The fusion model includes a parallel multi-branch convolutional module, a main capsule layer, and an emotion capsule layer connected in sequence; the parallel multi-branch convolutional module acquires multimodal features after quality verification and spatiotemporal alignment, and extracts multi-scale features; The main capsule layer acquires multi-scale features and converts them into capsule vectors to obtain the main capsule vector set; the emotion capsule layer acquires the main capsule vector set and emotion dimension features.
7. The student mental health assessment method integrating multi-source biological behavioral data according to claim 6, characterized in that, The parallel multi-branch convolution module includes a first branch, a second branch, a third branch, and a fourth branch configured in parallel; the outputs of the first branch, the second branch, the third branch, and the fourth branch are concatenated using channels to obtain multi-scale features. The first branch employs a first 1×1 convolutional layer and performs a linear transformation along the channel dimension to extract fine-grained local features and output a feature map. The second branch consists of a second 1×1 convolutional layer and an average pooling layer; the second 1×1 convolutional layer adjusts the number of channels and uses the average pooling layer to capture global contextual information, outputting a downsampled feature map. The third branch consists of a third 1×1 convolutional layer and a 6×6 convolutional layer; the third 1×1 convolutional layer performs dimensionality reduction, and the 6×6 convolutional layer captures a large range of local features, outputting a feature map. The fourth branch is a 3×3 convolutional layer, and its output feature map .
8. The student mental health assessment method integrating multi-source biological behavioral data according to claim 7, characterized in that, The main capsule layer and the emotion capsule layer employ a dynamic routing mechanism to iteratively calculate and update the coupling coefficient between them. The emotion capsule layer receives low-level capsule vectors from the main capsule layer and maps them to high-level emotion capsule vectors through the dynamic routing mechanism. The updated coupling coefficient... The expression is: in, The lower capsules representing the main capsule layer The log-prior probability between the higher-level capsule j of the emotion capsule layer; The lower capsules representing the main capsule layer The log-prior probability between the emotion capsule layer and the higher-level capsule k, where k represents the summation index of the higher-level capsule in the emotion capsule layer.
9. The student mental health assessment method integrating multi-source biological behavioral data according to claim 1, 2, or 3, characterized in that, Based on dynamic evaluation weights and emotional dimension features, the Mental Health Index (MHI) is calculated, and its expression is: Where σ represents the Sigmoi function; This represents the dynamic evaluation weight of the school age classification for the m-th feature; denoted as the m-th normalized sentiment dimension feature; b represents the bias.
10. The student mental health assessment method integrating multi-source biological behavioral data according to claim 1, 2, or 3, characterized in that, Also includes: Calculate the current proportion of students under warning and dynamically allocate psychological counseling resources; The number of psychological counselors required for the dynamic allocation of psychological counseling resources is: Its expression is: in, Indicates basic configuration; Indicates the current number of people under warning; This indicates the total number of registered students.
Citation Information
Patent Citations
Student mental health assessment system based on big data
CN118335336A
College student mental health classification and early warning method based on multi-modal deep learning
CN120413077A
Cited By
Campus environment-oriented student psychological state monitoring and intervention decision-making method and system
CN121880276A