Intelligent screening and assessment method and system for adolescent depression risk multi-modal data

CN122800237APending Publication Date: 2026-09-22ZHEJIANG HOSPITAL
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202611006043.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-07-07
Publication Date
2026-09-22

AI Technical Summary

Technical Problem

[0009]针对现有技术中存在的青少年抑郁筛查依赖量表自评导致隐匿性抑郁漏检率高、特征提取领域适配性不足、缺乏认知偏差分析维度以及隐私保护机制缺失等技术问题,本发明提供青少年抑郁风险多模态数据智能筛查评估方法及系统

Benefits of technology

[0012]本发明的有益效果在于:第一,通过整合校园被动行为数据与可穿戴生理数据替代传统量表自评,有效降低隐匿性抑郁的漏检率;第二,基于行为熵变指标与时序注意力机制的特征提取方案直接面向抑郁前驱行为模式设计,相较于点云结构转换方法具有更强的领域适配性;第三,引入自然语言认知偏差分析维度,弥补现有技术在认知层面筛查能力的空白;第四,采用联邦学习结合差分隐私的分布式架构,实现多校联合筛查场景下的模型协同优化与数据隐私保护;第五,闭环反馈机制使预测结果驱动数据采集与特征提取的参数动态优化,形成持续提升筛查精度的自适应闭环系统。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122800237A_ABST
    Figure CN122800237A_ABST
Patent Text Reader

Abstract

The present application relates to the technical field of intelligent medical health data processing, and discloses a method and system for screening and evaluating the risk of adolescent depression based on multi-modal data, which comprises the following steps: passively collecting behavior and physiological data from a campus management platform and a wearable device and time-aligning the data; extracting depression precursor behavior features through behavior entropy change indicators and time sequence attention mechanisms; analyzing text cognitive bias and emotional tendency based on a pre-trained language model; fusing multi-modal features using a cross-modal attention gating mechanism; outputting a three-level risk grade through an integrated gradient boosting classifier and optimizing the collection and extraction parameters through closed-loop feedback; and effectively reducing the missed detection rate of concealed depression by integrating passive behavior data from a campus and wearable physiological data to replace traditional scale self-evaluation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of intelligent medical and health data processing technology, and in particular to a method and system for intelligent screening and assessment of multimodal data on adolescent depression risk. Background Technology

[0002] As a prevalent and highly insidious mental health disorder, adolescent depression has made its early identification and timely intervention a core concern in global public health. Globally, over 350 million people suffer from depression, with the prevalence among adolescents showing a year-on-year upward trend. Schools, as the primary environment for adolescents' daily lives, should be at the forefront of depression risk screening. However, existing adolescent depression screening methods face multiple technical bottlenecks in practical application, and a complete technical solution that balances screening efficiency, detection sensitivity, and protection of individual privacy has not yet been developed.

[0003] Chinese invention application CN120340858A discloses an artificial intelligence-based system for predicting the trajectory of adolescent mental health development. This system includes a data acquisition module, a feature extraction module, a trajectory prediction module, a risk assessment module, and an intelligent feedback module. The proposed solution extracts features by converting mental health data into a point cloud structure, extracts temporal features using K-nearest neighbors combined with Transformer, and constructs a trajectory prediction model using DBSCAN clustering combined with bidirectional LSTM and GRU. However, this solution still has the following shortcomings:

[0004] First, the data sources are too singular and overly proactive. The data collection module of the aforementioned scheme mainly relies on four types of data sources: psychological assessment data, social behavior data, physiological data, and environmental data. However, obtaining this data mostly requires students to actively cooperate by filling out questionnaires or participating in tests. For patients with masked depression, their self-reported data often exhibits a significant social expectation bias, meaning that students may deliberately conceal their true psychological state out of avoidance, leading to a high rate of missed detections for questionnaire-based data. Especially in large-scale school screening scenarios, data collection methods relying on proactive cooperation are not only inefficient but also fail to capture the pre-depressive signals that students naturally reveal in their daily behavior.

[0005] Secondly, the feature extraction methods lack domain adaptability. The aforementioned approach converts time-series mental health data into a point cloud structure before extracting features. A point cloud structure is essentially a representation of discrete points in three-dimensional space, designed for tasks such as three-dimensional object recognition and scene understanding. Forcibly mapping time-series mental health data into a point cloud structure, while allowing for feature extraction using mature point cloud networks, creates unnecessary layers of abstraction in the data's semantics. This dilutes the inherent temporal continuity and causal relationships in behavioral patterns during the format conversion process, ultimately limiting the discriminative power of the extracted features for identifying pre-depressive behavioral patterns.

[0006] Third, there is a lack of ability to analyze cognitive biases in the natural language dimension. Core cognitive characteristics of depression include negative attribution style, catastrophic thinking, and learned hopelessness. These characteristics are most directly reflected in students' daily textual expression, such as essays, social media posts, and counseling records. The aforementioned approach does not involve the analysis and processing of textual data, making it impossible to incorporate cognitive risk signals of depression into the assessment system. This results in a significantly insufficient screening capability when facing a group of individuals with hidden depression who appear normal on the surface but whose internal cognition has already deviated.

[0007] Fourth, there is a lack of privacy protection mechanisms. Adolescent mental health data is highly sensitive personal information. The aforementioned solution employs a centralized data processing architecture, requiring all student data to be aggregated onto a unified platform for model training and prediction. In actual deployment scenarios involving multi-school joint screening, cross-school data transmission and centralized storage face severe privacy risks and fail to meet compliance requirements of personal information protection laws and regulations.

[0008] In summary, there is an urgent need for a multimodal intelligent screening solution that can integrate passive behavior data, physiological signals, and natural language text. This solution should effectively improve the detection rate of hidden depression while ensuring the security of students' personal information through federated learning and differential privacy technology, thereby achieving high-efficiency, high-sensitivity, and high-security screening for adolescent depression risk. Summary of the Invention

[0009] To address the technical problems in existing technologies, such as the reliance on self-reporting scales for adolescent depression screening leading to high missed rates of hidden depression, insufficient adaptability of feature extraction domains, lack of cognitive bias analysis dimensions, and lack of privacy protection mechanisms, this invention provides a multimodal data-based intelligent screening and assessment method and system for adolescent depression risk.

[0010] The present invention provides a multimodal intelligent screening and assessment method for adolescent depression risk, comprising the following steps: Step S1, time-aligned acquisition of multi-source campus behavioral and physiological data, obtaining behavioral data sequences consisting of academic performance change trajectories, attendance anomaly records, book borrowing preferences, canteen consumption patterns, and sports activity participation from the campus information management platform, and simultaneously obtaining physiological data sequences consisting of sleep duration, circadian rhythm offset, and heart rate variability from smart wearable devices, performing timestamp normalization and alignment of the behavioral and physiological data sequences with a natural day as the baseline time window to generate a multi-source time-aligned data matrix; Step S2, extraction of temporal attention features of pre-depression behavioral patterns, constructing a sliding time window sequence for each dimension of the data channel in the multi-source time-aligned data matrix, calculating behavioral entropy change indicators within each sliding time window to quantify the degree of deviation from behavioral regularity, and employing a temporal attention mechanism to process the behavioral entropy change indicator sequence. Step S3 involves weighted aggregation to identify pre-depression behavioral feature vectors composed of social withdrawal, loss of interest, sleep disorder, and academic decline; Step S4 involves natural language cognitive bias and sentiment analysis, performing semantic encoding on student essays, social media posts, and psychological counseling records, extracting sentiment polarity distribution and cognitive bias markers based on a pre-trained language model, and generating cognitive sentiment feature vectors; Step S5 involves attention-weighted cross-modal feature fusion, calculating the correlation weight matrix between behavioral, textual, and physiological modalities through a cross-modal attention gating mechanism, and performing weighted concatenation of multimodal features to generate a fused feature vector; Step S6 involves ensemble learning for depression risk grading prediction and closed-loop feedback, inputting the fused feature vector into an ensemble gradient boosting classifier to output the probability distribution of three risk levels, and feeding back the feature importance ranking results and prediction confidence to Steps S1 and S2 respectively to achieve adaptive optimization of closed-loop parameters.

[0011] This invention also provides an intelligent screening and assessment system for multimodal data on adolescent depression risk, including a multi-source data time-aligned acquisition module, a depression precursor behavior feature extraction module, a natural language cognitive analysis module, a cross-modal feature fusion module, a risk grading prediction and closed-loop feedback module, and a federated differential privacy update module. Each module corresponds one-to-one with the steps of the above method.

[0012] The beneficial effects of this invention are as follows: First, by integrating passive behavior data from the campus with wearable physiological data to replace traditional self-assessment scales, the missed detection rate of masked depression is effectively reduced; Second, the feature extraction scheme based on behavioral entropy change indicators and temporal attention mechanisms is directly designed for pre-depression behavioral patterns, and has stronger domain adaptability compared to point cloud structure conversion methods; Third, the introduction of natural language cognitive bias analysis dimension fills the gap in the screening capabilities of existing technologies at the cognitive level; Fourth, the adoption of a distributed architecture combining federated learning and differential privacy enables collaborative optimization of models and protection of data privacy in multi-school joint screening scenarios; Fifth, the closed-loop feedback mechanism enables the prediction results to drive the dynamic optimization of parameters for data collection and feature extraction, forming an adaptive closed-loop system that continuously improves screening accuracy. Attached Figure Description

[0013] Figure 1 This is a flowchart illustrating the intelligent screening and assessment method for multimodal data on adolescent depression risk provided in an embodiment of the present invention.

[0014] Figure 2 This is a schematic diagram of the architecture of the intelligent screening and assessment system for multimodal data on adolescent depression risk provided in an embodiment of the present invention. Detailed Implementation

[0015] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present invention, and not all of them. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without creative effort are within the scope of protection of the present invention.

[0016] See Figure 1This invention provides a multimodal intelligent screening and assessment method for adolescent depression risk. Addressing the challenges of high anonymity, large-scale implementation difficulties, and significant privacy risks in early screening for adolescent depression in school settings, this method constructs an end-to-end intelligent screening and assessment process based on multimodal passive data collection, with behavioral entropy change detection and cognitive bias analysis as core feature engines, and federated differential privacy as a security guarantee. It should be noted that there is significant data coupling between the steps in this method: the output data matrix of step S1 serves as the input source for both steps S2 and S4; the output feature vectors of steps S2 and S3 converge in step S4 for fusion; the fusion result of step S4 flows to step S5 for classification prediction; the prediction result of step S5 influences the collection frequency of step S1 and the attention weight of step S2 through dual-path closed-loop feedback; and step S6 influences the weight configuration of the neural networks in steps S2 and S4 from the model parameter level. This deep coupling and closed-loop collaborative mechanism ensures that the components of the system are not loosely connected independent modules, but rather form an organically linked overall technical solution. The method comprises the following six deeply coupled and closed-loop collaborative steps.

[0017] Step S1: Time-aligned acquisition of multi-source campus behavior and physiological data. This step aims to construct a multi-source passive data acquisition framework covering both campus behavior and physiological signal dimensions, and to unify heterogeneous data to the same time reference coordinates through a timestamp normalization alignment mechanism, providing a structured data foundation for subsequent feature extraction and modality fusion.

[0018] At the behavioral data collection level, this invention extracts structured records closely related to the daily behavior of teenagers from multiple business subsystems of the campus information management platform. Specifically, the trajectory of academic performance changes comes from the raw score records of each exam in the academic affairs management system. In one embodiment of this invention, the system extracts all exam scores within the most recent 90 calendar days and represents the relative ranking changes of each subject in the form of standardized Z-scores, thereby eliminating the interference of absolute score differences between different subjects. Attendance anomaly records come from the campus access control and classroom attendance system, including event types such as lateness, early departure, absenteeism, and absence. Each record is marked with an event timestamp accurate to the minute. Book borrowing preferences come from the library management system, and the extracted information includes borrowing frequency, book category distribution (such as literature, science and technology, psychology, etc.), and borrowing time period distribution. Among them, the abnormally high borrowing frequency of psychology and philosophy books may be related to depressive rumination tendencies. Cafeteria consumption patterns come from the campus card system, extracting daily consumption amount, meal time, and meal frequency. Preferably, cafeteria consumption data also includes regular deviations in consumption time periods. For example, skipping breakfast for a long time may reflect a disordered sleep schedule. Sports activity participation is measured by the physical education class management system and the card swipe records of campus sports venues, which count the number of times and the cumulative duration of participation in organized sports activities each week.

[0019] At the physiological data acquisition level, this invention obtains three core physiological indicators through wearable devices such as smart bracelets. Sleep duration is detected by a fusion algorithm of the bracelet's built-in accelerometer and photoplethysmography (PPG) sensor, recording the daily sleep onset time, wake-up time, and total sleep duration with an accuracy of ±15 minutes. Circadian rhythm deviation is defined as the absolute value of the deviation between the actual sleep onset time and the student's average sleep onset time over the past 30 days, expressed in minutes. Preferably, a flag is triggered when the circadian rhythm deviation exceeds 60 minutes for three consecutive days. Heart rate variability (HRV) is characterized by the root mean square (RMSSD) of the difference between adjacent normal heartbeats, calculated by continuously acquiring pulse wave signals at a sampling rate of not less than 25 Hz using the bracelet's photoelectric sensor. A decrease in RMSSD usually reflects a decline in autonomic nervous system regulation function and is significantly associated with depressive states.

[0020] Timestamp normalization and alignment is a key technical step in this process. Because the time granularity of the various data sources differs significantly—for example, exam scores are generated weekly or monthly, while heart rate variability data is collected at a second-level frequency—a unified time reference framework is needed. This invention uses a natural day (i.e., 0:00:00 to 23:59:59) as the baseline time window and performs the following alignment operations on each data source: For event-based data (such as attendance records and library borrowing), aggregation is performed using event counts or Boolean tags within each natural day; for continuous data (such as HRV), the statistical measure (mean, standard deviation, or extreme value) within each natural day is taken as the representative value for that day; for periodic data (such as exam scores), linear interpolation is used to fill in the natural days between two exams. Preferably, for data channels that are completely missing within certain natural days, a forward filling strategy is used to replace them with the most recent valid value. When the number of consecutive missing days exceeds 7, it is marked as an invalid segment and masked in subsequent processing.

[0021] After the above alignment operation, the generated dimension is Multi-source time-aligned data matrix ,in The number of students participating in the screening. In one embodiment of the present invention, the total number of days in the observation time window is... The value is set to 90 (i.e., a 3-month observation period). This represents the total number of dimensions in the data channels. In one embodiment of the present invention, the behavioral data channel includes six dimensions: academic Z-score, number of times late each day, number of times absent from class each day, daily borrowing volume, daily spending amount, and daily exercise duration. The physiological data channel includes three dimensions: sleep duration, circadian rhythm shift, and RMSSD, plus one auxiliary dimension: the standard deviation of daily spending periods. The total dimensions are 10. It is worth emphasizing that all data in this step are collected passively, without requiring students to actively fill out questionnaires or participate in specific testing activities. This minimizes the interference of social expectation bias on the screening results and ensures the efficiency of data collection in a large-scale school setting.

[0022] This step also includes a data quality verification mechanism. For each student's time-series data, the system calculates the effective data coverage rate. , defined as the ratio of the number of valid data days to the total number of observation days. When In this case, the student's data was marked as unreliable and downweighted in subsequent analyses. In one embodiment of the present invention, in an actual system deployed in a middle school, multi-source data from 1200 students over 90 days were collected, with an average data coverage rate of 87.3%. Behavioral data coverage was 91.5%, and physiological data coverage was 82.1%. The slightly lower physiological data coverage was mainly due to the wristband not being worn or the battery being depleted. Preferably, to further improve data integrity, the system automatically sends wristband wearing reminders to students who have not synchronized their physiological data for more than three consecutive days. This reminder function increases the physiological data coverage rate in the second month from 78.6% in the first month to 85.9%. Furthermore, the time alignment module also detects and corrects outliers. For example, if the sleep duration recorded on a certain day exceeds 16 hours or is less than 1 hour, the value is marked as potentially abnormal and replaced with the median of the three days before and after.

[0023] Step S2: Temporal attention feature extraction of depressive prodromal behavior patterns. This step receives the multi-source temporally aligned data matrix output from step S1. As input, its core objective is to identify behavioral patterns highly correlated with prodromal symptoms of depression from multidimensional behavioral and physiological time series. Unlike the indirect approach of converting time series data into a point cloud structure and then extracting features in comparison document 1, this invention directly extracts features in the original semantic space of the time series, measures the degree of deviation from behavioral regularity through a behavioral entropy change index, and adaptively focuses on the time segments and data channels most relevant to prodromal behaviors of depression using a temporal attention mechanism.

[0024] First, the multi-source time-aligned data matrix Each data channel ( Each sliding time window sequence is constructed separately. In one embodiment of the present invention, the window length of the sliding time window is... The value ranges from 7 to 14 calendar days. In one embodiment of the present invention, it is preferably set to 7 calendar days, and the step length is... The value ranges from 1 to 3 calendar days. In one embodiment of the present invention, it is preferably set to 1 calendar day. Therefore, for the total observation time... Daily data sequences, each channel can generate Each time window. Within each time window... ( Inside, for the passage Entropy change index for calculating data subsequence behavior .

[0025] The calculation process of the behavioral entropy change index is as follows: First, calculate the current time window. Information entropy of internal data Then calculate the adjacent preceding time windows. Information entropy The behavioral entropy change index is then defined as:

[0026] ,

[0027] in, For the first Channel within a time window The information entropy is calculated using the following formula:

[0028] ,

[0029] in: To determine the number of bins after discretizing continuous numerical values, one embodiment of the present invention... The value is 8; For the first The proportion of data points in each bin to the total number of data points in that time window; It is the logarithm to the base 2. When the data of a certain channel is highly uniformly distributed within a time window, the information entropy approaches 1 / 2. However, when data is highly concentrated in a single bin, the information entropy approaches zero. (Behavioral entropy change index) A positive value indicates a decrease in behavioral regularity (tending towards disorder), while a negative value indicates an increase in behavioral regularity (tending towards stereotypedness). Significant shifts in either direction may indicate prodromal depressive behaviors. Preferably, when... Exceeding the preset threshold (In one embodiment of the present invention) When the time window is (bit), mark it as an abnormal window.

[0030] After obtaining the behavioral entropy change index matrix for all time windows, this invention employs a multi-head temporal attention mechanism to weighted aggregate the entropy change index sequence. Specifically, the channel dimension... The behavioral entropy change index sequence is organized into a shape of The feature matrix of , where This represents the total number of time windows. After applying positional encoding to this matrix, it is input into a multi-head self-attention layer; the number of attention heads is... Set the number of attention heads to 4 to 8. Each attention head independently calculates the query matrix. Key matrix AND-value matrix ( Attention weights are calculated using the scaled dot product attention formula:

[0031] ,

[0032] in: , , ; This is the entropy-variable feature matrix of behavior after adding position encoding; , , The first The query, key, and value projection weight matrix for each attention head, with dimensions [missing information]. ; Dimensions for each attention head; is a scaling factor used to prevent the softmax gradient from vanishing due to excessively large dot product values. The outputs of each attention head are concatenated and mapped through a linear transformation layer to the final behavioral feature representation. This feature representation is further processed by global average pooling and a fully connected layer, resulting in an output dimension of . Depression precursor behavior feature vector In one embodiment of the present invention .

[0033] The core advantage of the multi-head attention mechanism lies in the ability of different attention heads to autonomously learn and focus on pattern changes across different time scales and behavioral channels. In one embodiment of this invention, after training, the first attention head primarily focuses on short-term abrupt changes in the sleep and circadian rhythm channels; the second attention head focuses on the mid-term downward trend in the academic performance channel; the third attention head focuses on the synchronous decline pattern of social activities and sports participation; and the fourth attention head focuses on abnormal fluctuations in consumption behavior. This multi-scale, multi-channel adaptive focusing capability enables the system to capture four typical pre-depressive behavioral patterns: social withdrawal (manifested as a synchronous decline in social activities and extracurricular sports participation), loss of interest (manifested as a decrease in the amount and diversity of book borrowing), sleep disturbance (manifested as abnormal sleep duration and increased circadian rhythm deviation), and academic decline (manifested as a continuous downward trend in Z-scores across multiple subjects). Preferably, the interpretability of attention weights can be used to generate behavioral characteristic reports, informing psychological counselors that the student has been identified as having high-risk specific behavioral dimensions, thereby providing targeted information for subsequent interventions. For example, when a student's main abnormality stems from a disordered sleep schedule, the counselor can prioritize the student's sleep hygiene habits and lifestyle adjustments; when the main abnormality stems from a social withdrawal channel, the counselor can focus on assessing the student's interpersonal relationships and social avoidance tendencies.

[0034] Step S3: Natural Language Cognitive Bias and Sentiment Analysis.

[0035] This step, independent of the behavioral feature extraction path in step S2, extracts cognitive patterns and emotional state characteristics of adolescents from the text data dimension, providing complementary cognitive dimension information for multimodal fusion. Depressed patients typically exhibit a characteristic negative cognitive triad at the cognitive level, including negative self-evaluation, negative interpretation of the world, and negative expectations of the future. These cognitive characteristics are presented in everyday textual expressions in the form of specific linguistic markers.

[0036] The input text data for this step comes from three sources: student essays are digitally archived from daily writing assignments in Chinese language classes; social media posts are obtained from campus social media platforms or class group chat records with student authorization; and psychological counseling records are desensitized counseling notes from the school's psychological counseling center. In the data preprocessing stage, the raw text undergoes Chinese word segmentation (using the jieba word segmentation tool and loading a psychology-specific dictionary to improve the accuracy of segmentation for mental health terminology), stop word filtering, and text normalization.

[0037] In the semantic encoding stage, a Chinese BERT model (BERT-wwm-ext, with approximately 110 million parameters) fine-tuned from a mental health corpus was used as the text encoder. Preferably, the fine-tuning process used approximately 50,000 mental health texts labeled with sentiment polarity and cognitive bias tags. The fine-tuning tasks included a sentiment classification subtask (3 categories: positive, neutral, and negative) and a cognitive bias identification subtask (multi-label classification: negative attribution, catastrophizing, hopelessness, deserver's mentality, and extreme expressions). The segmented text was input into the fine-tuned BERT model, and the hidden layer output corresponding to the [CLS] tag was taken as the sentence-level embedding vector with a dimension of 768.

[0038] The extraction of sentiment polarity distribution is achieved by connecting a 3-class softmax head to the BERT output layer, outputting probability values ​​for three dimensions: positive, neutral, and negative. For all texts for each student, the mean and standard deviation of the probabilities of each dimension within the observation period are calculated as the statistical characteristics of that student's sentiment polarity. Preferably, the sentiment polarity distribution also considers the trend of change over time, such as the mean of the negative sentiment probability. Whether there is an upward trend in the last 30 days is characterized by the slope of a linear regression.

[0039] Cognitive bias markers are extracted by incorporating a multi-label classification head into the BERT output layer, and a sigmoid activation function is used to independently determine the presence of each type of cognitive bias. Specifically, the judgment logic for negative attribution is based on the co-occurrence relationship between attribution words (e.g., because, due to, blame, etc.) and negative sentiment words (e.g., failure, useless, hopeless, etc.) at the sentence level. When the co-occurrence frequency exceeds a threshold, it is marked as a positive negative attribution. The judgment of catastrophic thinking is based on the co-occurrence detection of adverbs of extreme degree (e.g., will definitely, certainly, never, etc.) and negative expectation words (e.g., finished, collapsed, ruined, etc.). The judgment of the hopelessness index is based on the recognition of expression patterns that negate future meaning, such as the frequency of occurrence of characteristic phrases like meaningless, impossible to improve, and life is useless.

[0040] After concatenating the statistical features of emotional polarity with the labeled features of cognitive bias, a two-layer fully connected network (hidden layer dimension 128, activation function ReLU) is used for dimensionality compression and nonlinear transformation, resulting in a final output dimension of... Cognitive Emotion Feature Vector In one embodiment of the present invention, in the experimental data of 1200 students, the positive detection rate of negative attribution was approximately 14.2%, the positive detection rate of catastrophic thinking was approximately 8.7%, and the positive detection rate of hopelessness was approximately 5.3%. Among them, the proportion of students with two or more cognitive bias markers who were ultimately identified as high-risk for depression reached 61.8%, which was significantly higher than the 22.4% of the single-marker positive group.

[0041] Step S4: Attention-weighted cross-modal feature fusion. This step receives the feature vector of depressive prodromal behaviors output from step S2. (dimension) The cognitive emotion feature vector output in step S3 (dimension) ) and physiological feature vectors extracted from the multi-source time-aligned data matrix in step S1. (dimension) () is used as a trimodal input. Among them, the physiological feature vector... The data is obtained by extracting statistical features (mean, standard deviation, trend slope, extreme values, etc.) from the time series of three physiological channels—sleep duration, circadian rhythm offset, and RMSSD—from a multi-source time-aligned data matrix and then concatenating them. The core task of this step is to design a cross-modal attention-gated fusion mechanism that enables features from different modalities to be adaptively weighted and fused based on their information quality and intermodal complementarity, thereby generating a unified fusion feature representation with high discriminative power.

[0042] The design concept of the cross-modal attention gating mechanism is to introduce a gating unit based on the standard cross-modal attention calculation, enabling the system to dynamically adjust the attention weights according to the reliability and integrity of the data from each modality. The specific implementation process includes the following steps.

[0043] First, the query vector, key vector, and value vector are obtained by linear projection of the feature vectors of the three modalities. Taking the behavioral modality querying the text modality as an example: , , ,in , , This is a learnable projection matrix. Similarly, query-key-value triples are computed between other modal pairs, for a total of 6 pairs (the three modalities are interleaved pairwise, with each modality generating 2 pairs as the query party).

[0044] Secondly, for each modality pair, calculate the scaled dot product attention weights:

[0045] ,

[0046] in: This represents the raw attention score of the behavioral modality to the textual modality; The dimension of the query vector.

[0047] Based on this, a gating unit is introduced. ( The calculation method is as follows:

[0048] ,

[0049] in: Use the Sigmoid activation function; and These are the weight parameters for the gating unit; For modality eigenvectors; For modality The data quality index is defined as the effective data coverage of the modality within the observation period (within the range of 0 to 1). This represents a vector concatenation operation. (Gating unit) The output is a scalar value between 0 and 1. When the data quality of a certain modality is low, the gate value approaches 0, thereby suppressing the contribution of that modality to the fusion result.

[0050] After incorporating the gating value into the attention weights, the final cross-modal weighted feature is calculated as follows:

[0051] ,

[0052] in: This represents the information absorbed by the behavioral modality from the textual modality. Similarly, after calculating all 6 sets of cross-modal interaction features, the original features are concatenated with the interaction features:

[0053] ,

[0054] in: As a fully connected layer, it compresses the high-dimensional concatenated vectors to a smaller dimension. fused feature vector The introduction of a gating mechanism ensures that even in scenarios where certain modalities are missing or of poor quality (e.g., a student has not authorized text data collection, in which case the text modality gating value is close to 0), the system can still output meaningful fusion features based on available modalities, improving the robustness of the method in real-world deployment environments. In one embodiment of this invention, approximately 12% of students in the actual deployment data lacked text modalities due to unauthorized social media data collection or the lack of available essay data during the observation period. In this case, the cross-modal attention gating mechanism automatically reduced the text modality gating value to below 0.05, and the fusion result was mainly contributed by the behavioral and physiological modalities. Experiments show that even in the case of complete text modality loss, the fusion feature still achieves an AUC value of 0.917, only 3.4 percentage points lower than the 0.951 value in the case of complete trimodality, indicating that the gating fusion mechanism has good degradation tolerance. Furthermore, the cross-modal attention gating network has approximately 21,000 parameters and an inference time of approximately 0.8 ms per sample (measured on a server configured with an NVIDIA RTX 3060), resulting in minimal computational overhead that will not become a system bottleneck.

[0055] Step S5: Integrate learning for depression risk grading prediction and closed-loop feedback. This step receives the fused feature vector output from step S4. As input, the three-level probability distribution of adolescent depression risk is output by integrating a gradient boosting classifier, and a closed-loop feedback path is constructed from the prediction end to the data acquisition end and feature extraction end, so that the whole system forms an adaptive and optimized closed-loop architecture.

[0056] The technical details of the integrated gradient boosting classifier are as follows: This invention uses Gradient Boosting Decision Tree (GBDT) as the base classifier framework, specifically implemented using XGBoost. The core hyperparameter configuration of the classifier is: learning rate. The value range of is 0.01 to 0.1, and in one embodiment of the present invention, it is preferably set to 0.05. The maximum depth of the decision tree (maxdepth) ranges from 4 to 8 layers, and in one embodiment of the present invention, it is preferably set to 6 layers. The L2 regularization coefficient (reglambda) The feature sampling ratio (colsample bytree) is set to 1.0, the sample sampling ratio (subsample) is set to 0.8, and the number of iterations (nestimators) ranges from 100 to 500 rounds. In one embodiment of this invention, it is preferably set to 300 rounds, and an early stopping mechanism (early stoppin grounds=20) is enabled to prevent overfitting. The output layer of the classifier uses the Softmax function to map the raw scores of the three risk levels to a probability distribution:

[0057] ,

[0058] in: To integrate classifiers for the first The original output score for each risk level; Corresponding to low risk level, Corresponding to the medium-risk level, Corresponding to high-risk level; Given a fused feature vector, the target student belongs to the grade. The posterior probability. The final risk level of the target student is determined by the level corresponding to the highest probability, i.e. .

[0059] Preferably, the present invention also introduces a confidence level measurement index for risk level determination. Prediction confidence level. Defined as the difference between the highest probability and the second highest probability:

[0060] ,

[0061] in: The value range is from 0 to 1; when If the prediction is deemed insufficiently certain, the student is marked as pending review and will be manually reviewed and confirmed by a psychological counselor.

[0062] The closed-loop feedback mechanism is one of the key innovations that distinguishes this invention from existing technologies, specifically comprising two feedback paths. The first feedback path points from step S5 to step S2, dynamically adjusting the temporal attention weights using the feature importance ranking results. In the GBDT classifier, feature importance ranking is obtained by calculating the cumulative gain of each feature across all decision trees. After normalizing the feature importance scores related to each behavioral channel in the ranking results, a dimension is generated. Channel importance vector This vector serves as prior knowledge injected into the key-value projection weight initialization process of multi-head attention in step S2. In one embodiment of the invention, after three iterations, the system automatically increases the attention weight of the sleep-related channel from the initial uniform distribution value of 0.1 to 0.18, while decreasing the weight of the relatively unimportant cafeteria consumption channel from 0.1 to 0.06, thereby tilting attention resources towards channels with greater discriminative value. Furthermore, when a target student is determined to be at high risk, step S2 further performs targeted adjustments to the individualized attention weights of that student, increasing the attention weights corresponding to the sleep and social-related channels by 10% to 30% from the initial baseline. In one embodiment of the invention, a 20% increase is preferred to enhance the ability to capture features of the student's key behavioral dimensions.

[0063] The second feedback path points from step S5 to step S1, adjusting the data collection frequency using the predicted confidence level. When a student is classified as high-risk and the predicted confidence level is... When the student's physiological data collection frequency in step S1 is increased from the default once a day to three times a day (once each in the morning, at noon, and before bedtime), the system automatically refines the granularity of the student's behavioral data statistics from once a day to once every half day. When the student is determined to be at medium risk, the collection frequency is increased to twice a day. If the student maintains a low-risk status for more than 14 consecutive days, the default collection frequency is restored. Through this risk-driven adaptive collection strategy, the system can concentrate more data resources on high-risk individuals with limited computing and storage resources, thereby improving the screening accuracy of high-risk individuals without increasing the overall data volume.

[0064] Step S6, Federated Differential Privacy Model Security Update. This step establishes a privacy-preserving model update mechanism for multi-school collaborative deployment scenarios. It ensures that each school's original student data remains locally, and only the privacy-preserving model gradient updates are uploaded to the federated aggregation server. This achieves cross-school model collaborative optimization while meeting compliance requirements of personal information protection laws and regulations.

[0065] The federated learning architecture is designed as follows: Assume there are Each school node participates in federated training. ( ) has local dataset This includes a multi-source time-aligned data matrix of students from the school who participated in the screening, along with their corresponding risk level labels. The federated training process is executed iteratively in communication rounds, each round containing the following steps.

[0066] First, the federated aggregation server will set the current global model parameters. The data is distributed to all school nodes participating in this round of training. In one embodiment of the invention, 60% to 80% of the nodes are randomly selected to participate in training each round to reduce communication overhead.

[0067] Secondly, each school node uses its own dataset locally. Perform several rounds of local gradient descent updates on the global model parameters (in one embodiment of this invention, the number of local training rounds) ), to obtain the local gradient update amount .

[0068] Next, privacy protection measures are performed on the local gradient updates, specifically including two steps: gradient pruning and differential privacy noise injection. Gradient pruning restricts the L2 norm of the gradient update to a pruning threshold. Within:

[0069] ,

[0070] in: The upper bound of the gradient clipping norm is defined as 0.5 to 2.0, and its value ranges from 0.5 to 2.0. In one embodiment of the present invention, a preferred value is... ; The gradient is L2 norm. Gaussian noise is injected after clipping.

[0071] ,

[0072] in: It is zero-mean Gaussian noise. For noise multipliers, It is the identity matrix. Noise multipliers. The value is determined by the privacy budget. With failure probability The privacy budget ε is determined using a Gaussian mechanism-based privacy accounting formula. The privacy budget ε ranges from 1 to 10, and the failure probability δ ranges from less than 10 to the power of -4. In one preferred embodiment of the invention, [the following is a preferred embodiment]. , The corresponding noise multiplier This configuration strikes a good balance between privacy protection and model accuracy, and experimental results show that... The AUC loss of the time-limited model is only about 1.2%, which is lower than the centralized training results without privacy protection.

[0073] Finally, the federated aggregation server performs a weighted average aggregation of the noisy gradients uploaded by all participating nodes:

[0074] ,

[0075] in: This represents the number of nodes participating in this training round. For nodes The size of the local dataset is used as the weight for weighted aggregation, ensuring that schools with larger datasets contribute more significantly to the global model. The aggregated global model parameters are then used. The parameters are distributed to each school node to update the neural network feature transformation part of the integrated gradient boosting classifier in step S5. Preferably, the GBDT decision tree part of the integrated gradient boosting classifier is independently trained by each school node based on local data (because GBDT is not suitable for direct gradient-level federated aggregation), while the cross-modal attention gating network (step S4) and the behavior feature extraction network (attention module in step S2), which serve as input feature transformations, are updated collaboratively across schools using the aforementioned federated learning mechanism.

[0076] In one embodiment of this invention, in a federated learning experiment involving 6,000 students from five middle schools, after 50 rounds of federated communication, the global model achieved a screening sensitivity of 87.6%, a specificity of 91.2%, and an AUC value of 0.943. Compared to single-school models trained independently by each school, this represents an average improvement of 5.8 percentage points in sensitivity and 3.2 percentage points in AUC, fully validating the model gain effect of federated learning in multi-school collaborative scenarios. Furthermore, through the protection of differential privacy mechanisms, even if the federated aggregation server is attacked, attackers cannot deduce the original data of any individual student from the noisy gradient, thus effectively protecting the privacy and security of students.

[0077] The follow-up process for high-risk students is as follows: When step S5 outputs a student's depression risk level as high-risk, the system automatically generates a case assessment report and pushes it to the school's psychological counseling center through an encrypted channel. The report includes the student's risk score, a summary of the main abnormal behavioral characteristics, and suggested areas of focus. Within 48 hours of receiving the report, the psychological counselor arranges a face-to-face assessment. The interview results are used as tags to feed back into the system for continuous model iteration and optimization.

[0078] See Figure 2 The present invention also provides an intelligent screening and assessment system for multimodal data on adolescent depression risk. The system’s module division corresponds strictly to the steps in the above method embodiments, and the modules are deeply coupled and collaborated in a closed loop through well-defined data interfaces.

[0079] The multi-source data time alignment acquisition module corresponds to step S1 in the method embodiment. This module includes a campus data interface submodule and a wearable device interface submodule. The campus data interface submodule interfaces with the campus academic affairs system, access control and attendance system, library management system, campus card system, and sports activity management system through standardized APIs, acquiring behavior record data from each business subsystem in a timed batch synchronization manner. The wearable device interface submodule communicates with the smart bracelets worn by students via the Bluetooth Low Energy (BLE) protocol, synchronizing the sleep data and heart rate data cached by the bracelets to the local database on a daily cycle. The time alignment engine submodule is responsible for performing timestamp normalization, missing value imputation, and data quality verification, outputting a formatted multi-source time-aligned data matrix and writing it to the structured data storage layer.

[0080] The module for extracting pre-depression behaviors corresponds to step S2 in the method embodiment. This module contains three functional sub-units: a sliding window generator, a behavior entropy change calculator, and a multi-head temporal attention network. The sliding window generator segments the multi-source time-aligned data matrix according to the configured window length and step parameters. The behavior entropy change calculator performs information entropy difference calculation within each window and marks abnormal windows. The multi-head temporal attention network adaptively weights and aggregates the entropy change index sequence, ultimately outputting a feature vector of pre-depression behaviors. It is worth noting that the attention weight parameters of this module are controlled by the channel importance vector returned by the risk grading prediction and closed-loop feedback modules, forming the first feedback path from the prediction end to the feature extraction end.

[0081] The Natural Language Cognitive Analysis module corresponds to step S3 in the method embodiment. This module includes a text acquisition adapter, a semantic encoding engine, and a cognitive bias detector. The text acquisition adapter supports data acquisition from multiple text sources, including a Chinese language writing digitization platform, a campus social application interface, and a psychological counseling record management system. The semantic encoding engine loads a fine-tuned Chinese BERT model and performs vectorized encoding on the preprocessed text. The cognitive bias detector outputs the sentiment polarity distribution and cognitive bias labels based on the multi-label classification head of the BERT output layer, ultimately generating a cognitive sentiment feature vector.

[0082] The cross-modal feature fusion module corresponds to step S4 in the method embodiment. This module implements a cross-modal attention fusion network including a gating unit. This module receives three feature inputs from the depressive prodromal behavior feature extraction module, the natural language cognitive analysis module, and the physiological data channel from the multi-source data time-aligned acquisition module. It calculates the inter-modal correlation weights through a cross-modal attention gating mechanism and performs weighted fusion, outputting a unified fused feature vector. The gating unit dynamically adjusts the fusion ratio based on the data quality indicators of each modality, ensuring that the system can still output effective fusion results even when some modal data is missing.

[0083] The risk grading prediction and closed-loop feedback module corresponds to step S5 in the method embodiment. The core component of this module is an integrated gradient boosting classifier. Preferably, this module also includes a feature importance analysis submodule and a closed-loop scheduling submodule. The feature importance analysis submodule periodically calculates the cumulative gain ranking of each feature in the classifier. The closed-loop scheduling submodule generates a channel importance vector based on the ranking result and transmits it to the depression precursor behavior feature extraction module. Simultaneously, it generates a collection frequency adjustment instruction based on the prediction confidence and risk level of each student and transmits it to the multi-source data time-aligned collection module, forming a complete dual-pathway closed-loop feedback architecture. When a student is determined to be high-risk, this module also triggers the early warning referral submodule, automatically generating a case report and pushing it to the receiving terminal of the school's psychological counseling center.

[0084] The federated differential privacy update module corresponds to step S6 in the method embodiment. This module deploys two sub-components on each school node: a local trainer and a privacy-preserving processor. A secure aggregator is deployed on the federated aggregation server. In each round of federated communication, the local trainer performs local gradient descent updates using data from its own school. The privacy-preserving processor performs pruning and Gaussian noise injection on the gradient update before uploading it to the secure aggregator. After performing weighted average aggregation, the secure aggregator distributes the global model parameters to each school node via an encrypted channel. Inter-module communication uses the TLS encryption protocol to protect transmission security. The federated aggregation server does not store any original gradient data and deletes intermediate results immediately after aggregation calculation. Preferably, the federated aggregation server is deployed in a dedicated secure computer room of the education authority, and data transmission with each school node is conducted via the education private network, further reducing the risk of data interception during transmission over the public internet.

[0085] The data flow between the six modules forms a clear directed acyclic graph with a closed-loop feedback structure: the output of the multi-source data time-aligned acquisition module flows simultaneously to the depression precursor behavior feature extraction module (behavioral data channel) and the cross-modal feature fusion module (physiological data channel); the output of the natural language cognitive analysis module flows to the cross-modal feature fusion module; the output of the cross-modal feature fusion module flows to the risk grading prediction and closed-loop feedback module; the output of the risk grading prediction and closed-loop feedback module is, on the one hand, output as screening results to the early warning referral terminal, and on the other hand, is fed back to the depression precursor behavior feature extraction module and the multi-source data time-aligned acquisition module through two feedback paths respectively; the federated differential privacy update module influences the network parameters in the cross-modal feature fusion module and the depression precursor behavior feature extraction module by distributing global model parameters. This deeply coupled and closed-loop collaborative system architecture ensures that there is not only a unidirectional data transmission relationship between the modules, but also bidirectional dynamic adjustment at the parameter level through the feedback mechanism, thereby ensuring that the entire system continuously optimizes its screening performance during continuous operation.

[0086] To verify the effectiveness of the method proposed in this invention, an experimental screening system was deployed in five middle schools in a certain city for a period of six months, involving a total of 6,000 students aged 12 to 18. The experiment used the diagnostic results determined by experienced clinical psychologists through structured clinical interviews (SCID) as the gold standard label, among which 312 cases were confirmed as positive for depression (positive rate of 5.2%), which is within the range of epidemiological data for adolescent depression.

[0087] In terms of screening performance, the method of this invention achieves a sensitivity of 88.1%, a specificity of 92.4%, and an AUC of 0.951. Compared to the traditional PHQ-A scale screening method (sensitivity 72.3%, specificity 84.6%, AUC 0.856), this invention improves sensitivity by 15.8 percentage points and AUC by 9.5 percentage points. Particularly noteworthy is the significantly superior detection ability of this invention for masked depression compared to the scale method—in 312 positive cases, 68 were clinically diagnosed with masked depression (i.e., patients deliberately concealed symptoms in their self-reports on the scale). The detection rate of this invention for these 68 cases of masked depression reached 79.4%, while the PHQ-A scale only achieved 35.3%, representing an improvement of 44.1 percentage points.

[0088] In terms of modality fusion contribution analysis, ablation experiments verified the independent contributions and synergistic effects of each modality: the AUC was 0.891 when using only the behavioral modality, 0.823 when using only the physiological modality, 0.867 when using only the text modality, 0.917 when combining behavioral and physiological modalities, and 0.951 when combining all three modalities. The gain of the three-modal fusion relative to the optimal bimodal combination was 3.4 percentage points, verifying the complementary value of the text cognitive dimension.

[0089] Regarding the effectiveness of closed-loop feedback, the screening performance was compared between configurations with and without closed-loop feedback enabled. After enabling closed-loop feedback, the system's sensitivity improved from the initial 83.7% to 88.1% after three iterations, an increase of 4.4 percentage points. Specifically, the precision for high-risk individuals improved from 71.2% to 78.9%. The closed-loop feedback mechanism effectively enhances the system's focus on and accuracy in identifying key target groups by prioritizing data collection resources towards high-risk individuals.

[0090] In terms of privacy protection assessment, a member inference attack was used to test the privacy security of the federated learning model. Regarding privacy budgets... Under this configuration, the attacker's member inference accuracy was only 52.1%, close to the level of random guessing (50%), indicating that the differential privacy mechanism effectively prevented the leakage of original data information. Meanwhile, compared to the centralized training model without privacy protection (AUC=0.963), the AUC of the federated differential privacy model decreased by only 1.2 percentage points (0.951 vs 0.963), and the accuracy loss due to privacy protection is within an acceptable range.

[0091] In terms of efficiency in large-scale deployment, the system takes about 4.2 hours to process a single full screening of 6,000 students end-to-end (including 1.5 hours for data collection and synchronization, 2.0 hours for feature extraction and fusion, 0.5 hours for risk prediction, and 0.2 hours for report generation). The average processing time per student is about 2.5 seconds, which meets the timeliness requirements of the school's monthly routine screening.

[0092] Regarding the convergence of federated learning, in one embodiment of this invention, the data sizes of the five school nodes are 1500, 1300, 1200, 1100, and 900 students, respectively, exhibiting a certain degree of non-independent and identically distributed data distribution. After 50 rounds of federated communication, the global model converged, with a convergence speed approximately 18% faster than the standard FedAvg algorithm. This is mainly attributed to the data-volume-weighted aggregation strategy, which ensures that nodes with larger data volumes contribute more stable gradient directions. Comparative experiments were conducted under different privacy budget configurations. The AUC was 0.937. The AUC was 0.951 at that time. The AUC was 0.958. (Without differential privacy protection) the AUC is 0.963, indicating that as the privacy budget increases, the model accuracy loss gradually decreases. With moderate privacy protection, the accuracy loss is only 1.2 percentage points, which is the recommended default configuration for actual deployment.

[0093] Regarding the clinical consistency of cognitive bias analysis, using the interview assessment results of clinical psychologists as a reference, the natural language cognitive bias detection module of this invention achieved an accuracy rate of 82.5% in detecting negative attribution, 78.3% in detecting catastrophic thinking, and 84.1% in detecting hopelessness. The weighted average F1 score of the three types of cognitive bias markers was 0.81, indicating a moderately high level of consistency with clinical assessment. Particularly in the masked depression group, the positive predictive value of cognitive bias markers was significantly higher than that of behavioral indicators used alone, further validating the necessity and effectiveness of introducing the textual cognitive dimension. In summary, the method of this invention demonstrates significant advantages in screening sensitivity, masked depression detection capability, privacy protection and security, and large-scale deployment efficiency.

[0094] The embodiments of the present invention are not limited to the specific embodiments described above. Those skilled in the art can make various equivalent changes or substitutions based on the technical solutions of the present invention, and all such changes or substitutions should be included within the protection scope of the present invention.

Claims

1. A multimodal intelligent screening and assessment method for adolescent depression risk, characterized in that, Includes the following steps: Step S1, Multi-source data time alignment acquisition: Obtain behavioral data sequences consisting of academic performance change trajectory, attendance abnormality records, book borrowing preferences, canteen consumption patterns and sports activity participation from the campus information management platform; obtain physiological data sequences consisting of sleep duration, circadian rhythm offset and heart rate variability from smart wearable devices; perform timestamp normalization alignment on the two types of data sequences based on natural days to generate a multi-source time alignment data matrix; Step S2, Depression precursor behavior feature extraction: Construct a sliding time window sequence for each data channel of the multi-source time-aligned data matrix, calculate the behavioral entropy change index within each sliding time window to quantify the degree of deviation from the regularity of behavior, and use a temporal attention mechanism to weight and aggregate the behavioral entropy change index sequence to output the depression precursor behavior feature vector. Step S3, Natural Language Cognitive Bias Analysis: Acquire student text data and perform semantic encoding. Based on the pre-trained language model, extract the emotional polarity distribution and cognitive bias markers containing negative attribution, catastrophic thinking and hopelessness, and generate cognitive emotional feature vectors. Step S4, cross-modal feature fusion: The depression precursor behavior feature vector output in step S2 is used as the behavioral modality input, the cognitive emotion feature vector output in step S3 is used as the text modality input, and the physiological statistical feature vector extracted from the physiological data channel of the multi-source time-aligned data matrix is ​​used as the physiological modality input. The correlation weight matrix between the behavioral modality, the text modality and the physiological modality is calculated through a cross-modal attention gating mechanism, and the multimodal features are weighted and concatenated to generate a fused feature vector. Step S5, Risk Classification Prediction and Closed-Loop Feedback: Input the fused feature vector into the integrated gradient boosting classifier to output the probability distribution of three risk levels: low, medium and high. Feed back the feature importance ranking to step S2 to update the attention weights. Feed back the prediction confidence to step S1 to adjust the collection frequency.

2. The method according to claim 1, characterized in that, In step S1, the timestamp normalization alignment includes: mapping the discrete event timestamps of the behavioral data sequence to a unified time axis with a 24-hour period; and using a strategy combining forward padding and linear interpolation to fill in missing time windows. The final dimension of the multi-source time-aligned data matrix is... ,in For the number of students, For the number of time windows, For data channel dimensions.

3. The method according to claim 1, characterized in that, In step S2, the window length of the sliding time window is 7 to 14 natural days, the step length is 1 to 3 natural days, and the behavioral entropy change index is calculated using the information entropy difference method. When the behavioral entropy change index exceeds a preset threshold, the time window is marked as an abnormal window.

4. The method according to claim 1, characterized in that, In step S3, the pre-trained language model is a Chinese BERT model fine-tuned with mental health corpus. The sentiment polarity distribution includes probability values ​​in three dimensions: positive, neutral, and negative. The determination of negative attribution is based on the co-occurrence frequency of attribution words and negative sentiment words in the text.

5. The method according to claim 1, characterized in that, In step S4, the cross-modal attention gating mechanism includes: calculating query vectors, key vectors, and value vectors for the behavioral modality, the text modality, and the physiological modality respectively; calculating the attention weights between each modality pair by scaling dot product attention; and introducing a gating unit to dynamically adjust the effective proportion of the attention weights according to the data quality of each modality.

6. The method according to claim 1, characterized in that, In step S5, the integrated gradient boosting classifier includes multiple weak classifiers, each with a learning rate between 0.01 and 0.1, a decision tree with a maximum depth of 4 to 8 layers, and 100 to 500 iterations. The output is mapped to a probability distribution of three risk levels using the Softmax function.

7. The method according to claim 1, characterized in that, The temporal attention mechanism described in step S2 adopts a multi-head self-attention structure with 4 to 8 attention heads. Each attention head independently learns the behavioral pattern change rules at different time scales. The outputs of each attention head are linearly transformed and then concatenated to form the final depressive precursor behavior feature vector.

8. The method according to claim 1, characterized in that, The closed-loop feedback in step S5 includes: when the target student is determined to be at high risk, step S1 increases the frequency of collecting the student's physiological data from once a day to three times a day, and step S2 increases the weight of the channel related to sleep and social interaction in the student's corresponding temporal attention weight by 10% to 30%.

9. The method according to claim 1, characterized in that, The process also includes step S6, Federated Differential Privacy Model Update: After local training at each school node, gradient pruning and differential privacy noise injection are performed on the gradient update amount and uploaded to the federated aggregation server. The aggregation server performs a weighted average and then distributes the global model parameters to update the classifiers of each node. The differential privacy noise injection satisfies... - Differential privacy guarantee, where privacy budget The value ranges from 1 to 10, and the failure probability is... Less than The upper bound of the norm for the gradient clipping is 0.5 to 2.

0.

10. A multimodal intelligent screening and assessment system for adolescent depression risk, used to implement the method described in claim 9, characterized in that, include: The multi-source data time alignment acquisition module is used to acquire behavioral data sequences from the campus information management platform and physiological data sequences from smart wearable devices. It performs timestamp normalization alignment with the natural day as the base time window to generate a multi-source time alignment data matrix. The module for extracting features of pre-depression behaviors is used to construct a sliding time window sequence from the multi-source time-aligned data matrix, calculate the behavioral entropy change index within each sliding time window, and use a temporal attention mechanism for weighted aggregation to identify feature vectors of pre-depression behaviors. The Natural Language Cognitive Analysis module is used to acquire student text data and perform semantic encoding. Based on a pre-trained language model, it extracts sentiment polarity distribution and cognitive bias markers to generate cognitive sentiment feature vectors. The cross-modal feature fusion module is used to take the depression precursor behavior feature vector output by the depression precursor behavior feature extraction module as the behavioral modality input, the cognitive emotion feature vector output by the natural language cognitive analysis module as the text modality input, and the physiological statistical feature vector extracted from the multi-source time-aligned data matrix output by the multi-source data time-aligned acquisition module as the physiological modality input. The module calculates the correlation weight matrix between the behavioral modality, the text modality and the physiological modality through a cross-modal attention gating mechanism, and performs weighted concatenation of the multi-modal features to generate a fused feature vector. The risk grading prediction and closed-loop feedback module is used to input the fused feature vector into the integrated gradient boosting classifier to output the probability distribution of three risk levels: low, medium and high risk, to determine the depression risk level, and feed back the feature importance ranking result to the depression precursor behavior feature extraction module to update the attention weight, and feed back the prediction confidence to the multi-source data time alignment acquisition module to adjust the acquisition frequency. The federated differential privacy update module is used to perform gradient clipping and differential privacy noise injection on the gradient update after local training on each school node. The weighted average aggregation is performed through the federated aggregation server and the global model parameters are distributed.

Citation Information

Patent Citations

  • Artificial intelligence-based teenager mental health development track prediction system

    CN120340858A