Teenager depression auxiliary diagnosis model training method and system based on multi-modal data

By collecting and processing EEG, heart rate variability and eye movement data of adolescents in VR intelligent interactive scenarios, multimodal depression recognition model is trained, and objective diagnosis of adolescent depression is solved, and more efficient identification and diagnosis of depression is achieved.

CN120432128AActive Publication Date: 2025-08-05SHANGHAI PUDONG NEW AREA MENTAL HEALTH CENT (SHANGHAI PUDONG NEW AREA PSYCHOLOGICAL COUNSELING CENT)

Patent Information

Application Number
CN202510507996.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-22
Publication Date
2025-08-05
Estimated Expiration
2045-04-22

AI Technical Summary

Technical Problem

The prior art is difficult to effectively identify adolescent depression, and diagnosis depends on subjective evaluation, and there are missed and misdiagnosed, and there is a lack of objective biological diagnostic standards.

Method used

The VR intelligent interactive scenario was constructed, and the EEG, heart rate variability and eye movement data of adolescent depression and normal control subjects were collected simultaneously. Through data preprocessing and feature extraction, single-modal and multimodal depression recognition models were trained, and shared representations were generated by automatic encoder for classifier training, combining modal attention and contrast enhancement models to improve recognition accuracy.

Benefits of technology

It realizes objective and accurate identification of adolescent depression in a virtual reality environment, improves the scientificity and efficiency of diagnosis, reduces missed diagnosis and misdiagnosis, and provides a more effective depression identification model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120432128A_ABST
    Figure CN120432128A_ABST
Patent Text Reader

Abstract

The invention provides a teenager depression auxiliary diagnosis model training method and system based on multi-modal data, and relates to the technical field of mental health. According to the method, firstly, a VR intelligent interaction scene is constructed, and physiological data of juvenile depression and physiological data of a normal contrast subject in the VR intelligent interaction scene are synchronously collected; data preprocessing and feature extraction are conducted on the electroencephalogram data, the heart rate variability data and the eye movement data, a single-mode depression recognition model is obtained through classifier training, and then a multi-mode depression recognition model and a cross-mode depression recognition model are obtained. The teenager depression auxiliary diagnosis model obtained through machine learning is helpful for early discovery and timely treatment of teenager depression.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of mental health, and in particular relates to a training method and system for an auxiliary diagnosis model of adolescent depression. Background Art

[0002] Depression refers to a mental disorder caused by a variety of factors, with clinical features primarily manifesting as low mood, pessimism, and suicidal attempts. According to data from the "Healthy China Action Plan (2019-2030)," the prevalence of depression in my country is as high as 2.1%, and the age of onset is getting younger and younger. Globally, over 300 million people suffer from depression. It is estimated that by 2030, depression will become the world's leading disease burden. In recent years, the global prevalence of depression has surged due to the impact of the COVID-19 pandemic, leading to the phenomenon of "COVID-19 depression." Recent survey data shows that over the past two years of the pandemic, the incidence of depression in adolescents has more than doubled compared to pre-pandemic levels. Globally, one in four adolescents has experienced depressive symptoms, and the rate of increase in adolescents has surpassed that of adults. In my country, according to the "China National Mental Health Development Report (2019-2020)" released in March 2021, the detection rate of depression among adolescents in 2020 was 24.6%, and the detection rate of severe depression reached 7.4%. Adolescent depression is characterized by high morbidity, a high incidence of adverse events, and a slow and insidious onset. However, the current diagnosis and assessment of adolescent depression is primarily based on symptomatology and scale assessments, which often fail to effectively identify adolescents with depression. The diagnosis of depression lacks a biological diagnostic "gold standard," and diagnosis is largely uncertain, relying primarily on subjective assessments by psychiatrists. Statistics show that general practitioners can only correctly identify 47.3% of patients with depression, resulting in a significant proportion of missed diagnoses and misdiagnoses, which can lead to inadequate or overtreatment. Therefore, there is an urgent need to develop an objective auxiliary diagnostic model for adolescent depression that does not rely on self-assessment or observation by others, acting as a "lie detector for depression" to identify adolescents with depression at the source. Summary of the Invention

[0003] To this end, the technical problem to be solved by the present invention is to provide a training method and system for an auxiliary diagnosis model of adolescent depression based on multimodal data, which can construct an effective adolescent depression recognition model, thereby objectively and accurately evaluating adolescent depression.

[0004] In a first aspect, the present invention provides a method for training an auxiliary diagnosis model for adolescent depression based on multimodal data, comprising:

[0005] Step S1, constructing a VR intelligent interaction scene, and synchronously collecting physiological data of adolescent depression subjects and normal control subjects in the VR intelligent interaction scene, wherein the physiological data includes EEG data, heart rate variability data, and eye movement data;

[0006] Step S2, performing data preprocessing and feature extraction on the EEG data, heart rate variability data, and eye movement data to obtain EEG feature data, heart rate variability feature data, and eye movement feature data;

[0007] Step S3, performing classifier training based on the EEG feature data, the heart rate variability feature data, and the eye movement feature data, respectively, to obtain a unimodal depression recognition model; the unimodal depression recognition model includes a depression recognition model based on the EEG data, a depression recognition model based on the heart rate variability data, and a depression recognition model based on the eye movement data;

[0008] Step S4, performing feature layer fusion on the EEG feature data, the heart rate variability feature data, and the eye movement feature data, and then performing classifier training to obtain a multimodal depression recognition model;

[0009] Step S5: Using the EEG feature data, heart rate variability feature data, and eye movement feature data as inputs to the autoencoder to generate their respective shared representations; only providing one of the feature data when training the classifier, and testing the other two feature data on the trained classifier, thereby obtaining a cross-modal depression recognition model.

[0010] Furthermore, in step S2,

[0011] EEG data preprocessing includes: applying a 0.5-100 Hz bandpass filter to the raw EEG signal to remove low-frequency drift and high-frequency noise; applying a 50 Hz notch filter to eliminate power supply interference; performing signal denoising through discrete wavelet transform; calculating the SNR and RMSE to evaluate signal quality; and finally separating the processed signal into five frequency bands: Delta, Theta, Alpha, Beta, and Gamma for power spectral density analysis.

[0012] Heart rate variability data preprocessing includes: applying 0.04-5 Hz bandpass filtering to the raw photoplethysmography signal to remove baseline drift and high-frequency noise; performing signal denoising through discrete wavelet transform; identifying heartbeat intervals through peak detection algorithm; and extracting time domain and frequency domain indicators;

[0013] The eye movement data preprocessing includes: if the missing value of a data tuple in the eye movement data exceeds a set threshold, the data tuple is discarded; if the missing value of the data tuple does not exceed the set threshold, the data tuple is filled with the average value or the median.

[0014] Furthermore, in step S2, the feature extraction adopts a correlation-based feature selection method combined with a best-first search strategy.

[0015] Furthermore, before step S3, Gaussian filtering formula is used for denoising, Gaussian noise is added for data enhancement, weighted cross entropy loss function is used for sample rebalancing, and support vector distance discriminant function formula is used for edge sample identification and resampling; the significance P value of the evaluation feature is defined, and the difference between the evaluation feature and the target label is tested to eliminate features with insignificant differences, and finally K-fold cross validation is performed to determine the optimal feature subset based on the comprehensive performance on the training set.

[0016] Furthermore, in step S3, any one of the EEG feature data, heart rate variability feature data and eye movement feature data is used as the input layer of the autoencoder, noise is added to the original input, and a shared representation is generated through encoding as input data for classifier training, and then the original input is reconstructed through decoding.

[0017] Furthermore, in step S4, the feature layer fusion step is: directly connecting the EEG feature data, heart rate variability feature data and eye movement feature data together, inputting them into the autoencoder to generate a shared representation; using the unsupervised backpropagation algorithm to fine-tune the weights and biases of the autoencoder to generate the final shared representation for actual training of the classifier.

[0018] Furthermore, the fused multimodal feature vector is recorded as:

[0019] z fusion ∈R d

[0020] Among them, z fusion is the feature vector after multimodal fusion;

[0021] Define four parallel sub-models f1, f2, f3, and f4 and output the predicted probability:

[0022] pi=fi(·),i=1,2,3,4

[0023] Among them, f i is the i-th sub-model, which is used to output depression predictions from different angles;

[0024] p i is the predicted probability output by the i-th sub-model;

[0025] Fusion predicts the probability of depression:

[0026]

[0027] in,

[0028] Among them, p(y=1|z fusion ): Indicates that at a given z fusion Under the condition of , the probability that the sample belongs to category 1, category 1 is depression;

[0029] λ i is the weight coefficient of the i-th sub-model, satisfying

[0030]

[0031] p i is the output probability of the i-th sub-model, specifically:

[0032] p1 is the output of the standard MLP sub-model;

[0033] p2 is the output of the modal attention mechanism model;

[0034] p3 is the output of the contrast enhancement model;

[0035] p4 is the output of the auxiliary task model;

[0036] Sub-model f1 is a standard MLP model, which is used to perform nonlinear modeling on the fusion features:

[0037] p1=f1(z fusion )=Softmax(W1·z fusion +b1)

[0038] W1 is the weight matrix;

[0039] b1 is the bias vector;

[0040] Softmax(·) is used to transform the output into probability;

[0041] Sub-model f2 is the modal attention modeling MLP, which is used for modal weight adjustment:

[0042]

[0043] p2=f2(z fusion )=Softmax(W2·z fusion +b2)

[0044] in:

[0045] z i is the representation of the i-th mode;

[0046] w is the trainable vector parameter in the attention mechanism;

[0047] w is a trainable weight vector used to calculate the importance of different modalities;

[0048] Tanh(·) is the hyperbolic tangent function, which is used as an activation function to introduce nonlinear transformation to the input features;

[0049] exp(·) is an exponential function used to amplify the difference and form a softmax format;

[0050] α i Represents the attention weight of the i-th modality. The larger the value, the more important the modality is to the final classification task.

[0051] M is the total number of modes;

[0052] Sub-model f3 is a contrast enhancement model, which uses the difference between positive and negative stimulus responses in emotional tasks as the enhancement feature modeling:

[0053]

[0054] z f (positive) It is the fusion feature representation extracted by the model under positive emotional stimulation;

[0055] z f (negative) It is the fusion feature representation extracted by the model under negative emotional stimulation;

[0056] Δz is the difference vector between the two states, which is used to capture the intensity of the subject's response to different emotional stimuli and is an enhancement feature;

[0057] [z fusion ,Δz] represents the concatenation of the original fusion features and the emotion difference features, which is used to enhance the input expression ability of the model;

[0058] w3 is the weight matrix of sub-model f3;

[0059] b3 is the bias term;

[0060] f3(·) is the contrast enhancement modeling module;

[0061] p3 is the output probability of sub-model f3;

[0062] Sub-model f4 is an auxiliary task modeling module, which is used to predict individual behavioral variables in parallel with the execution of the main task:

[0063] f4=f4 (main) (z fusion )=Softmax(W4·z fusion +b4)

[0064] Among them, W4 and b4 are learnable parameters;

[0065] The output is probability p4, which indicates the predicted probability that the sample is depressed;

[0066]

[0067] r^ is the predicted reaction time;

[0068] s^ is the predicted behavioral score;

[0069] w r ,w s and b r ,b s are trainable parameters for auxiliary tasks;

[0070]

[0071] Calculate the mean square error MSE for the two behavioral variables separately;

[0072] λ r ,λ s represents the weighting coefficient of auxiliary loss;

[0073] The four sub-model outputs are combined in a weighted manner:

[0074]

[0075] Add Beta distribution to model uncertainty:

[0076] α=exp(W α ·z fusion ), β=exp(W β ·z fusion )

[0077]

[0078] Beta distribution parameters α and β are used to describe the confidence of the model in the prediction results;

[0079] Expectations for predicted outcomes;

[0080] Var(y) is the variance of the prediction results, which is used to measure uncertainty. The larger the value, the more uncertain the model is.

[0081] Furthermore, curriculum learning and adversarial training mechanisms are used to optimize the model.

[0082] Furthermore, the VR intelligent interaction scene is implemented by building a Web VR environment using the A-Frame framework, and a mind-to-mind dialogue space is constructed through the Claude API.

[0083] On the other hand, the present invention also provides a training system for an auxiliary diagnosis model for adolescent depression based on multimodal data. The technical solution of this system is as follows:

[0084] A virtual reality subsystem is used to construct a VR intelligent interactive scene and synchronously collect physiological data from adolescent depression subjects and normal controls in the VR intelligent interactive scene. The physiological data includes EEG data, heart rate variability data, and eye movement data.

[0085] a data processing subsystem, performing data preprocessing and feature extraction on the EEG data, heart rate variability data, and eye movement data to obtain EEG feature data, heart rate variability feature data, and eye movement feature data;

[0086] a unimodal depression recognition model training subsystem for performing classifier training based on the EEG feature data, the heart rate variability feature data, and the eye movement feature data to obtain a unimodal depression recognition model; the unimodal depression recognition model includes a depression recognition model based on the EEG data, a depression recognition model based on the heart rate variability data, and a depression recognition model based on the eye movement data;

[0087] A multimodal depression recognition model training subsystem performs feature-layer fusion on the EEG feature data, the heart rate variability feature data, and the eye movement feature data, and then performs classifier training to obtain a multimodal depression recognition model;

[0088] The cross-modal depression recognition model training subsystem uses the EEG feature data, heart rate variability feature data, and eye movement feature data as inputs to an autoencoder to generate their respective shared representations; only one of the feature data is provided when training a classifier, and the other two feature data are tested on the trained classifier, thereby obtaining a cross-modal depression recognition model.

[0089] Furthermore, the virtual reality subsystem includes a Biopac MP160 physiological multichannel instrument and an aSee A8 portable telemetry eye tracker to obtain the EEG data, eye movement data and heart rate variability data.

[0090] Beneficial effects:

[0091] The multimodal data-based adolescent depression diagnostic model training method proposed in this paper uses virtual reality technology to create a more ecological scenario with enhanced control over irrelevant variables. This method accurately collects EEG, heart rate variability, and eye movement data from adolescents with depression and normal controls, thereby constructing a more effective adolescent depression identification model. This model enables rapid and scientific identification of adolescent depression, providing guidance for treatment and saving valuable time.

[0092] The present invention is based on a unimodal depression recognition model based on EEG, eye movement and heart rate variability features. The unimodal EEG / eye movement / heart rate variability features are input into a denoising autoencoder to generate their respective shared representations for training a classifier, thereby obtaining a unimodal depression recognition result.

[0093] By utilizing the complementarity of different modalities, a multimodal depression model is trained based on EEG, eye movement, and heart rate variability features. Two modal fusion strategies (feature fusion and hidden layer fusion) are used to perform feature layer fusion of EEG, eye movement, and heart rate variability data, which improves the classification accuracy and builds a more accurate depression recognition model.

[0094] The cross-modal adolescent depression model training conducted by the present invention uses the shared representation learned by the autoencoder to determine whether it has a strong correlation with adolescent depression and a weak correlation with the signal form of EEG / eye movement or heart rate variability.

[0095] In addition, the adolescent depression auxiliary diagnosis model training system based on multimodal data provided by the present invention has the same technical effect as the above-mentioned training method. BRIEF DESCRIPTION OF THE DRAWINGS

[0096] In order to make the contents of the present invention more clearly understood, the present invention is further described in detail below based on specific embodiments of the present invention in conjunction with the accompanying drawings.

[0097] Figure 1 This is a flow chart of the training method for the adolescent depression auxiliary diagnosis model according to Example 1 of the present invention;

[0098] Figure 2 This is a flow chart of the support vector machine (SVM) model used in the adolescent depression auxiliary diagnosis model training method according to Example 1 of the present invention;

[0099] Figure 3 This is a graph showing the calculation results of normality test of collected physiological indicators in Example 1 of the present invention;

[0100] Figure 4 This is the ROC curve diagram of Example 1 of the present invention. DETAILED DESCRIPTION

[0101] The present invention will be described in detail below with reference to the accompanying drawings and in conjunction with embodiments. The principles and features of the present invention will be described below in conjunction with the accompanying drawings. It should be noted that the embodiments and features of the embodiments in this application may be combined with each other unless there is a conflict. The embodiments are provided only to illustrate the present invention and are not intended to limit the scope of the invention.

[0102] With the development of wearable technology, the real-time, non-invasive collection and analysis of physiological signals has become possible. This has prompted increasing research attention to the feasibility of using physiological signals from the body itself to monitor and assess stress, aiming to achieve objective measurement of psychological stress. Existing studies have explored the specific manifestations of depression through physiological and behavioral signals, exploring objective indicators for depression assessment. Domestic and international studies have shown that non-invasive, objective physiological signals such as EEG, heart rate variability, and eye movement can reflect a person's psychological state to varying degrees. The electroencephalogram (EEG) is an objective and reliable method for assessing brain function, offering advantages such as high sensitivity, low cost, and portability. Some researchers have explored its application in depression identification. Previous studies have demonstrated that using support vector machines to classify the three EEG bands (alpha, beta, and theta) and the full-band power spectrum can achieve classification accuracies of 71.7% and 88.6%, respectively, achieving good results in depression identification. Numerous studies have demonstrated that using EEG to examine alpha wave activity (8-13 Hz) in the left and right frontal lobes can analyze resting frontal EEG asymmetry. This metric can serve as an important neural marker for depression, with good predictive value. However, the interpretation of frontal lateralization remains relatively crude, with few studies addressing and accurately analyzing the specific meaning of lateralization scores. Therefore, the current use of frontal EEG lateralization scores to differentiate between depressed and non-depressed individuals remains difficult to implement in clinical practice. Recent research has focused on phase-amplitude coupling (PAC), which is implicated in cognitive processes and mental health. Studies have found that cross-coupling between neural oscillations of different frequencies reflects synchronization between local and global brain networks, and abnormal PAC patterns are associated with depression in adults. Decreased resting-state theta-gamma PAC may be a biomarker of poor mental health, particularly depression. Furthermore, sustained changes in PAC may underlie TMS treatment, with reductions in depressive symptoms after TMS associated with increases in PAC. Therefore, research on frontal lobe lateralization and PAC needs further development, and establishing a more accurate depression prediction model based on EEG is an urgent problem to be solved.

[0103] Heart rate variability (HRV) is a commonly used non-invasive biomarker that is easy to measure and can be collected non-invasively using wearable devices. HRV not only reflects the impact of the internal and external environment on the cardiovascular system but also reflects the corresponding adjustments made by the cardiovascular system under autonomic and humoral regulation. Because changes in emotional state directly affect the autonomic nervous system and humoral regulation, individual mood changes can be reflected in HRV, meaning that HRV can be used for emotion recognition. HRV refers to the subtle differences in the duration of successive heartbeats, which is regulated by both the sympathetic and parasympathetic nervous systems and reflects the balance of autonomic nervous system function. When sympathetic nerve activity decreases or vagal nerve activity increases, HRV increases, and vice versa. The neurohumoral factors reflected in HRV may be valuable indicators for predicting sudden cardiac death and arrhythmic events: reduced parasympathetic activity and increased sympathetic activity lead to lower HRV levels, a lower threshold for ventricular fibrillation (VF), and a higher risk of malignant arrhythmias. Existing studies generally find that depression is closely associated with decreased HRV, with patients with severe depression having lower HRV than those with mild depression. Successful treatment of depression is associated with increased HRV. However, some studies have also found that decreased HRV in depressed patients may be related to other heart diseases, not simply due to depression alone, and that HRV in depressed patients fluctuates significantly with age. Numerous studies have found that abnormal HRV is associated with a bidirectional association between depression and cardiovascular disease, with depressed patients being more susceptible to cardiovascular disease than healthy individuals, and vice versa. Therefore, there is no unified and effective standard for predicting depression using HRV, and whether HRV can be reliably associated with specific life phenomena in other areas remains unknown. This paper further explores the reliability and sensitivity of HRV as a physiological indicator for depression identification. Eye movement (EM) is considered another effective physiological signal for depression identification. Numerous studies using eye tracking technology have shown that affective disorders are characterized by attentional biases toward emotional stimuli. Eye tracking technology allows for relatively direct and continuous measurement of visual attention, is non-invasive, economical, and simple, and offers high temporal and spatial resolution. Eye movement abnormalities in patients with depression have been extensively studied. Compared with healthy controls, depressed patients exhibit abnormal horizontal pursuit eye movements, weaker correlation between pursuit and saccadic movements, higher blink rates, and abnormal saccades, suggesting impairments in their ocular motor systems. Depressed patients exhibit atypical visual scanning patterns in free viewing tests, characterized by longer fixations, more fixations, fewer saccades, and an attentional preference for negative emotional faces. However, studies based on eye movement tests, such as pursuit tests, have shown inconsistent findings. Therefore, many eye movement-specific features of depression require further exploration and in-depth research.

[0104] In recent years, virtual reality (VR) has demonstrated significant potential in the assessment and diagnosis of mental illness. Virtual environments have been shown to demonstrate significant differences in outcome measures between patients with mental illness and healthy controls. VR technology provides a multisensory, three-dimensional environment that fully immerses users in a simulated world. Users perceive three-dimensional images through a head-mounted display (HMD), and spatial location within the visual environment is determined by motion-tracking sensors within the helmet. Simultaneously, users hear sounds through headphones and interact with virtual objects using input devices such as joysticks, wands, and data gloves. Compared to more passive media such as radio and television, VR offers a higher level of cognitive, social, and physical interaction, thereby enhancing the impact of the virtual environment on users. VR environments can stimulate and simultaneously measure symptoms of mental illness, contributing to improved reliability in assessing mental illness. Incorporating other measures into VR measurements, such as physical activity, eye movements, distance tracking, and physiological measurements, can further enhance the objectivity of VR measurements. These measures can often be obtained without distracting participants, thus maximizing immersion in the VR environment. Furthermore, the introduction of VR offers the advantages of creating standardized environments to ensure consistency and replacing standard laboratory environments with immersive technology to motivate participants. VR technology has been shown to complement subjective assessments and neurophysiological measurements in psychological assessments. Numerous studies have found significant correlations between VR-based outcome measures and traditional diagnostic measures for patients with schizophrenia, ADHD, obsessive-compulsive disorder, and other mental disorders. However, research on adolescent depression remains limited, and the manifestation of depression symptoms in virtual environments remains unclear. Predictive models for depression based on VR technology remain underdeveloped.

[0105] In summary, this paper aims to develop an objective, accurate, convenient, and practical method for assessing adolescent depression. Based on the variability of EEG, heart rate, and eye movement in depressed and healthy controls, combined with VR measurement technology, it is hoped that a more effective depression identification model will be constructed.

[0106] This invention represents an innovation and breakthrough in the noninvasive collection and analysis of physiological signals of depression. By using virtual reality (VR) technology to create a more ecological scenario with enhanced control over irrelevant variables, it accurately collects EEG, heart rate variability, and eye movement data from both depressed and normal controls, thereby constructing a more effective depression identification model.

[0107] The search for noninvasive EEG biomarkers of depression is crucial, as they could potentially facilitate a more objective diagnosis of the disorder, which is often diagnosed using questionnaires that rely on both professional and patient subjectivity. However, this is challenging, as depression manifests with diverse symptoms and a high prevalence, particularly associated with anxiety disorders, leading to inconsistent biomarker findings across the literature. Given the complexity of depression, understanding the latest discoveries in its biomarkers is crucial, and multimodal / cross-modal depression identification based on EEG, heart rate variability, and eye movement patterns can provide key insights.

[0108] Furthermore, most existing research on multimodal / cross-modal depression identification takes place in the laboratory. While this effectively controls for irrelevant variables, it also significantly weakens the required ecological context. This invention utilizes virtual reality (VR) technology to create a more ecological experimental environment while controlling for irrelevant variables, making the research results more relevant and effectively applicable in the real world.

[0109] Example 1

[0110] This embodiment provides a method for training an auxiliary diagnosis model for adolescent depression based on multimodal data, comprising the following steps:

[0111] 1. The key technology that this method first addresses is VR technology: combining human-computer interaction and user experience methods with generative artificial intelligence technology to activate adolescents' multi-sensory channels such as vision and hearing, promote their immersion in a friendly virtual environment, and then create a more ecological scene with enhanced control over irrelevant variables, so as to achieve the purpose of accurately collecting subjects' EEG, heart rate variability, and eye movement data.

[0112] Using the A-Frame framework to build a Web VR environment and the Claude API to implement features like "Expert Voice," this scenario creates a heart-to-heart dialogue space to help identify students experiencing depression. This scenario integrates the immersive experience of VR technology with the natural conversation capabilities of a large-scale language model. Through natural environments, soothing sounds, and interactive elements, it creates a safe and private virtual space, encouraging students to express recent pain and hope, and providing crucial information for identifying depression.

[0113] While the subject is immersed in the VR scene, a Biopac MP160 physiological multichannel instrument (EEG and PPG acquisition system) and an aSee A8 portable telemetry eye tracker will be connected to the subject simultaneously. These instruments collect EEG, pulse, and other data to monitor psychological processes and dynamic changes, track the activation and relaxation of adolescent brain regions, understand adolescents' attentional preferences for VR intelligent interactive scenarios, and study changes in their autonomic nervous system. This embodiment uses a synchronized EEG, eye movement, and heart rate variability acquisition network to ensure that the EEG and heart rate variability data recorded simultaneously while the subject freely browses emotional faces are synchronized with millisecond precision. This is the basis for meaningful synchronized analysis of EEG, eye movement, and heart rate variability.

[0114] As a preferred embodiment, the present application corrects the physiological data collected in the VR intelligent interaction scene. Although VR has great potential in the assessment and diagnosis of mental illness, and can provide a multi-sensory three-dimensional environment that allows people to be completely immersed in a simulated world, we found in our study that there are some differences between adolescent depression and normal control subjects in the VR intelligent interaction scene and the real environment. When using the VR intelligent interaction scene, there will be a delay in adolescents entering the virtual environment at the beginning of the experiment, and there will be situations where they are free from the original environment during the experiment. This causes a difference between the psychological reaction and the real environment, affecting the accuracy of the collected test data.

[0115] This application fully considers this influence and modifies the physiological data by adding VR influence factors to reduce the difference in psychological impact between depression and normal control subjects in real environments and VR intelligent interaction scenarios.

[0116] The VR impact factor is obtained through the following steps:

[0117] 1) Construct a typical real environment for comparison and establish a typical VR intelligent interaction scene that is the same as the typical real environment;

[0118] 2) Subjects with depression and normal controls were tested in typical real-world environments and typical VR intelligent interaction scenarios, and their physiological data were collected.

[0119] 3) Analyze and compare the physiological data of depression subjects and normal control subjects in typical real environments and typical VR intelligent interaction scenarios to obtain VR influencing factors.

[0120] For example, a typical real-world environment for an easily implemented comparative experiment on heart rate variability data can be constructed. Select a group of subjects with depression and a control group with normal heart rate variability to collect heart rate variability data. Analyze the collected heart rate variability data. If a linear relationship is found between the two, a simple ratio of the heart rate variability data in the two cases can be used as a VR impact factor. In the actual experiment, this VR impact factor can be multiplied by the heart rate variability data to obtain heart rate variability data that is closer to the real environment.

[0121] If the relationship between the two is nonlinear, data fitting can be used to establish a data relationship expression or corresponding relationship table between the two as the VR influencing factor. During the formal experiment, the heart rate variability data can be added to the VR influencing factor for correction.

[0122] 2. After data collection, pre-processing operations are performed, including:

[0123] 1) EEG signal preprocessing

[0124] Because EEG signals are relatively weak, low-frequency, and poorly resistant to interference, they can be interfered with by other physiological signals, such as electrooculography and electromyography, during data collection, affecting their signal quality. Therefore, before analyzing and processing the data, it is necessary to remove noise from the raw EEG signals.

[0125] EEG signal preprocessing utilizes a multi-step optimization process, including linear interpolation to handle NaN values, 0.5-100Hz bandpass filtering to remove baseline drift and high-frequency myoelectric noise, and 50Hz notch filtering to eliminate power supply interference when the sampling rate is sufficient. Artifact detection and repair are performed on the filtered signal to improve detection accuracy. A discrete wavelet transform (DWT) algorithm is used for deep denoising, which is particularly effective in removing complex artifacts such as electrooculogram (EOG). The processed signal is decomposed into five standard frequency bands: Delta (0.5-4Hz), Theta (4-8Hz), Alpha (8-13Hz), Beta (13-30Hz), and Gamma (30-50Hz). The power of each frequency band is calculated using optimized power spectral density analysis. The quality of preprocessing is assessed using signal-to-noise ratio and root mean square error to ensure the reliability of subsequent feature extraction and analysis.

[0126] 2) Heart rate variability preprocessing

[0127] When photoplethysmography (PPG) signals are used to calculate heart rate variability, they undergo a series of processing steps to achieve accurate results. PPG signals typically have a frequency range of 0.5-10 Hz. During acquisition, they are subject to various noise artifacts, including motion artifacts, ambient light interference, and baseline drift. These artifacts are removed through wavelet transforms, bandpass filtering, and median filtering. The processed signals undergo peak detection to obtain the pulse interval sequence, which is then used to calculate various heart rate variability metrics, including time-domain and frequency-domain analysis.

[0128] 3) Eye movement signal preprocessing

[0129] This embodiment uses three classic preprocessing methods in data mining: filling missing values, removing outliers, and data normalization. During the experiment, if the subject is not focused on the scene or blinks too frequently or closes his eyes for too long, the eye tracking device may lose the capture of the subject's pupil, resulting in missing values in some recorded data tuples. Generally, there are two different strategies to deal with these data with missing values. If a data tuple contains a large number of missing values (such as more than 30% of the attributes), the tuple will be discarded. If a data tuple contains a small number of missing values (such as less than 30% of the attributes), it will be filled with the average or median of the attribute of the subject.

[0130] In this embodiment, the Biopac MP160 physiological multichannel instrument and the aSee A8 portable telemetry eye tracker are used to collect EEG signals, eye movement signals, and ECG signals, and then data preprocessing and feature extraction are performed on them respectively.

[0131] The extracted EEG and eye movement features were selected using the CFS-based Bestfirst feature selection method, and the selected features were used in subsequent single-modal / multi-modal / cross-modal depression recognition research.

[0132] The experimental process is as follows:

[0133] Fifty-one adolescents with depression and 64 healthy controls were recruited (all participants were right-handed, with normal / corrected-to-normal vision, no astigmatism, and no brain or cardiovascular disease. The depressed group had no current or lifetime diagnosis of bipolar disorder, schizophrenia, autism spectrum disorder, attention deficit hyperactivity disorder, or other serious mental disorders, while the healthy control group had no current or lifetime diagnosis of a mental illness). These participants were then subjected to three types of data collection: EEG, eye movement, and heart rate variability, in a VR environment. Data analysis revealed specific differences in these three markers of depression, which were then used for modeling.

[0134] After informing the subjects of their rights and interests and completing the informed consent form, the following procedures were performed:

[0135] (1) Baseline testing;

[0136] (2) VR scene experience: By immersing the subjects in a VR scene and completing a conversation task, the participants’ status is monitored in real time while EEG, eye movement, and heart rate variability data are monitored.

[0137] VR is a web-based VR interactive scene with a beautiful environment.

[0138] Under the guidance of the person in charge, the subjects entered a VR immersive environment, in which they could freely explore the VR space; have conversations with the elves in the VR world, expressing their inner feelings, thoughts and confusions; and engage in interactive activities in the scene to stay away from pain.

[0139] (3) Post-experimental testing: After the depressed and control subjects complete the experiment, the collected EEG, eye movement, and heart rate variability data will be pre-processed and feature extracted for subsequent modeling.

[0140] 3. Unimodal Depression Identification

[0141] In research on unimodal depression recognition based on EEG, EM, and HRV features, the unimodal EEG / EM / HRV features are used as the input layer of a denoising autoencoder. Noise is added to the original input, and a shared representation (hidden layer) of the EEG / EM / HRV is generated through the encoding process. The original input is then reconstructed through the decoding process. The generated shared representation is the high-level depression-related features learned by the autoencoder and serves as the input data for classifier training.

[0142] Exemplarily, during the classifier training process, this embodiment adopts a 5-fold cross-validation validation strategy repeated three times. Specifically, the data of 115 subjects (64 in the control group and 51 in the depression group) were randomly divided into five non-overlapping subsets. Each time, four subsets (92 data items) were selected as training sets to train the SVM classifier, and the remaining subset (23 data items) was used as the test set for evaluation to obtain the classification accuracy. This process was repeated three times to ensure the stability of the results, and the average classification accuracy of the three experiments was finally taken as the evaluation indicator.

[0143] To comprehensively evaluate the model's performance, we conducted a final test using the trained SVM classifier on all 115 data points and calculated the overall classification accuracy. We also recorded the classification accuracy of the control and depression groups to assess the differences in model performance across different categories. Throughout all experiments, we maintained strict between-group sample balance to ensure the reliability of the model evaluation.

[0144] 4. Multimodal Depression Identification

[0145] This embodiment uses a multimodal denoising autoencoder (MDAE) to perform feature layer fusion on EEG, eye movement, and heart rate feature data. Specifically, the present invention uses two modality fusion strategies (feature fusion and hidden layer fusion).

[0146] a) Structural diagram of the feature fusion strategy: EEG features, EM features, and HRV features are directly concatenated and fed into an autoencoder to generate a shared representation. An unsupervised backpropagation algorithm is then used to fine-tune the weights and biases of the autoencoder. Finally, the generated shared representation is used for the actual training of the classifier.

[0147] b) Structural diagram of the hidden layer fusion strategy: EEG features, EM features, and HRV features are input into the autoencoder to generate their respective shared representations. The unsupervised backpropagation algorithm is still used to fine-tune the weights and biases of the autoencoder. The shared representation of EEG is directly connected to the shared representation of EM to synthesize a new shared representation. Finally, this synthesized shared representation is used as the input data for six classifiers to train the classifiers.

[0148] The classifier training process was the same as that used for unimodal depression recognition, employing a 5-fold cross-validation strategy repeated three times. The entire subject data set was randomly divided into 5 folds, with 4 folds used for training the SVM classifier each time and the remaining 1 fold used for testing. This was repeated three times to ensure stability. Finally, the trained model was evaluated on the entire dataset, and the average classification accuracy was calculated as the final result for multimodal depression recognition.

[0149] 5. Cross-modal Depression Identification

[0150] High-level features (shared representations) are learned using EEG, eye movement, and heart rate variability features. EEG (delta, theta, alpha, beta, gamma, and full-band) features and EM features are used as input to autoencoders to generate their respective shared representations. During classifier training, only data from a single modality is provided, while data from other modalities are used to test the trained classifier.

[0151] In this embodiment, the process is mainly divided into three parts:

[0152] EM training: The shared representation generated by EM is used as training data, and the shared representation generated by EEG / HRV is used as test data.

[0153] EEG training: The shared representation generated by EEG is used as training data, and the shared representation generated by EM / HRV is used as test data.

[0154] HRV training: The shared representation generated by HRV is used as training data, and the shared representation generated by EEG / EM is used as test data.

[0155] The classifier training process was the same as for unimodal and multimodal depression recognition, using a 5-fold cross-validation strategy repeated three times. Similarly, to verify the stability of the model, the SVM training and evaluation process was repeated three times, and the classification accuracy was obtained three times. The average accuracy and standard deviation were calculated as the final classification result for cross-modal depression recognition.

[0156] The classifier algorithm of this embodiment adopts the support vector machine SVM model. Figure 2 .

[0157] The 55 collected physiological indicators were tested for normality, with a significance level of p < 0.05. Independent sample t-tests were performed on indicators that conformed to the normal distribution, and Mann-Whitney U tests were performed on indicators that did not conform to the normal distribution. A p-value less than 0.05 indicated that there was a significant difference in the indicator between the two groups.

[0158] Specific formula:

[0159] t-value formula:

[0160]

[0161] in:

[0162] M1 and M2 are the means of the two samples.

[0163] o is the joint variance (pooled variance):

[0164]

[0165] o and is the variance of the two samples, n1 and m2 are the sample sizes.

[0166] οDegrees of freedom df=n1+n2-2.

[0167] The calculation results are shown in Figure 3 .

[0168] Note: (*p<0.05, **p<0.01, ***p<0.001)

[0169] A binary classification model capable of effectively distinguishing between "normal" and "depressed" individuals was constructed using a support vector machine (SVM). The model's hyperparameters were optimized through 5-fold cross-validation, ultimately selecting the radial basis kernel (RBF) function with parameters set to C = 100 and gamma = 0.1. Physiological indicators with significant differences at the 0.01 level (p < 0.01) were selected for model construction. Overall, with a class weight of 1, the model achieved an accuracy of 81.74% and an AUC (Area under the ROC curve) of 0.921, demonstrating good classification capabilities. Figure 4 .

[0170] Get the best result for the feature:

[0171] avg_saccade_distance(px), avg_saccade_duration(s), theta_beta_ratio, lfhf ratio

[0172] In a preferred embodiment, the present invention adopts the following specific steps to perform model training:

[0173] (1) Data processing

[0174] In addition to conventional denoising and normalization, this method combines data augmentation and sample rebalancing techniques to effectively improve model performance under conditions of limited data or class imbalance. Data augmentation expands the training sample space by adding noise and slicing, simulating a wider range of physiological states and improving the model's generalization ability. Sample rebalancing adjusts class weights in the loss function, allowing the model to focus more on minority class samples during training and alleviating bias caused by class imbalance. Furthermore, a strategy for identifying and resampling edge samples is introduced, leveraging the support vector information of the SVM to identify samples with blurred judgment boundaries. Local synthesis and interpolation are then used to increase the proportion of such samples in the training set.

[0175] For feature selection, we employ multi-level statistical methods. We use P-values to assess feature significance and ANOVA F-tests to assess the correlation between features and target labels. We eliminate redundant features with insufficient information and select features with significant discriminatory power. Finally, we use K-fold cross-validation to comprehensively analyze performance on the training set to determine the optimal feature subset, reducing noise interference and improving model stability and interpretability.

[0176] (1.1) Denoising

[0177] Denoising uses the Gaussian filter formula to clean invalid noise in real data, improve data quality, and enable the model to learn meaningful features more accurately:

[0178]

[0179] (1.2) Data enhancement

[0180] Data augmentation uses a Gaussian noise addition formula to artificially increase disturbances, so that the model can still correctly classify under slight changes, thereby improving robustness and generalization ability:

[0181]

[0182] (1.3) Sample rebalancing

[0183] Sample rebalancing uses a weighted cross entropy loss function to adjust the model's focus on different categories and alleviate the bias problem caused by category imbalance:

[0184]

[0185] (1.4) Marginal sample identification and resampling strategy

[0186] The edge sample identification and resampling strategy uses the support vector distance discriminant function formula to improve the model's ability to identify fuzzy areas of classification boundaries and reduce misjudgments:

[0187]

[0188] (1.5) Multi-level feature selection

[0189] F-value calculation formula:

[0190]

[0191] K-fold cross validation accuracy calculation:

[0192]

[0193] (2) Model construction

[0194] In multimodal emotion recognition tasks, traditional models often use only a single architecture (such as MLP or LSTM) to model fused features. While these models can achieve basic performance, they often exhibit instability or lack generalization when faced with heterogeneous data distributions, high-noise inputs, and clinically borderline individuals. Therefore, this paper proposes a heterogeneous multi-model ensemble classifier (HEC) architecture to comprehensively assess depression risk from multiple dimensions, improving the system's overall discriminability, robustness, and clinical applicability.

[0195] The classifier consists of four sub-models with distinct structures and functional focuses. The discriminant logic is designed from four perspectives: global modeling, modal attention, state-differentiation enhancement, and multi-task auxiliary supervision. The four sub-models operate in parallel, and their outputs are fused using a trainable weighted approach to produce the final prediction probability. This makes the system highly adaptable to different sample types (e.g., typical depression, subclinical states, mild mood disorders, etc.).

[0196] (2.1) Overview of model structure

[0197] The fused multimodal feature vector is recorded as:

[0198] z fusion ∈R d

[0199] Among them, z fusion It is the feature vector after multimodal fusion, integrating electroencephalogram (EEG), eye movement (EM) and heart rate variability (HRV) features.

[0200] Define four parallel sub-models f1, f2, f3, f4 and output the predicted probability:

[0201] pi=fi(·),i=1,2,3,4

[0202] f i : The i-th sub-model is used to output depression predictions from different angles (i = 1, 2, 3, 4)

[0203] p i : The predicted probability output by the i-th sub-model (obtained through Softmax).

[0204] The final fusion is the depression prediction probability:

[0205] in

[0206] p(y=1|z fusion ): Indicates that given the fusion feature vector z fusion Under the condition of , the probability that the sample belongs to category 1 (i.e., "depression") is the final output depression prediction probability.

[0207] λ i : The weight coefficient of the i-th sub-model, which indicates the importance of the model in the final prediction, satisfying

[0208] (i.e. weighted average)

[0209] p i : The output probability of the i-th sub-model,

[0210] Specifically:

[0211] p1: output of the standard MLP submodel

[0212] p1: Output of the modal attention mechanism model

[0213] p3: Output of the contrast enhancement model (using the difference between positive and negative emotional stimuli)

[0214] p4: Output of the auxiliary task model

[0215] (2.2) Sub-model f1: Standard MLP model

[0216] As the most basic branch, f1 directly performs nonlinear modeling on the fusion features:

[0217] p1=f1(z fusion )=Softmax(W1·z fusion +b1)

[0218] W1: weight matrix;

[0219] b1: bias vector;

[0220] Softmax(·): used to transform the output into probability;

[0221] f1(·): The first sub-model, which is a standard multi-layer perceptron (MLP) model.

[0222] p1: The predicted probability output by the model (e.g., the probability of depression).

[0223] (2.3) Sub-model f2: Modal Attention Modeling MLP

[0224] In order to cope with the dynamic changes in the importance of different modalities, the attention mechanism is introduced to adjust the modal weights:

[0225]

[0226] p2=f2(z fusion )=Softmax(W2·z fusion +b2)

[0227] This formula is used to calculate the attention weight α of the i-th modality i It measures the different modal features z through a scoring function with nonlinear activation (tanh) i the importance of.

[0228] in:

[0229] z i : Representation of the i-th mode;

[0230] w: trainable vector parameter in the attention mechanism;

[0231] The numerator is the score of the current modality, and the denominator is the normalization of the scores of all modalities, forming a Softmax structure overall, so that all α i The total is 1.

[0232] z i : Embedding representation or feature vector representing the iii-th modality (e.g., EEG, HRV, or EM).

[0233] w: A trainable weight vector used to calculate the importance (attention score) of different modalities.

[0234] tanh(·): Hyperbolic tangent function, used as an activation function to introduce nonlinear transformation to the input features.

[0235] exp(·): exponential function used to amplify the difference and form a softmax format.

[0236] α i : Represents the attention weight of the i-th modality. The larger the value, the more important the modality is to the final classification task.

[0237] M: The total number of modalities, which is 3 in this task (EEG, eye movement, and heart rate variability).

[0238] (2.4) Sub-model f3: contrast enhancement model

[0239] This module uses the difference between "positive stimulus" and "negative stimulus" responses in emotional tasks as enhanced feature modeling:

[0240]

[0241]

[0242] zf(positive): The fusion feature representation extracted by the model under positive emotional stimulation.

[0243] zf(negative): The fused feature representation extracted by the model under negative emotional stimulation.

[0244] Δz: The difference vector between the two states (negative feature minus positive feature), which is used to capture the intensity of the subject's response to different emotional stimuli and is an enhancement feature.

[0245] [z fusion ,Δz]: represents the concatenation of the original fusion features and the emotion difference features, which is used to enhance the input expression ability of the model.

[0246] w3: weight matrix of sub-model f3.

[0247] b3: bias term.

[0248] f3(·): Contrast enhancement modeling module.

[0249] p3: The output probability of sub-model f3, which comprehensively considers the basic features and the response to emotional sensitivity.

[0250] (2.5) Sub-model f4: Auxiliary task model

[0251] Sub-model f4 is an auxiliary task modeling module, which aims to predict individual behavioral variables (such as reaction time and score) in parallel with the execution of the main task (depression classification) to achieve multi-task joint learning, thereby enhancing the generalization and discrimination capabilities of the main task.

[0252] The model predicts behavioral variables (such as reaction time r and score s) in parallel and combines the main task:

[0253] p4=f4 (main) (z fusion )=Softmax(W4·z fusion +b4)

[0254]

[0255] This formula indicates that the sub-model f4 performs the main task classification (predicting the probability of depression) on the fusion feature zfusion\mathbf{z}_{fusion}zfusion.

[0256] W4, b4 are learnable parameters;

[0257] The output is probability p4, which indicates the predicted probability that the sample is "depressed".

[0258] r^: predicted reaction time;

[0259] s^: predicted behavioral score;

[0260] w r ,w s and b r ,b s are trainable parameters for the auxiliary task.

[0261] The mean square error (MSE) was calculated for each of the two behavioral variables;

[0262] λ r ,λ s : Represents the weighting coefficient of auxiliary loss.

[0263] (2.6) Output Fusion and Uncertainty Modeling

[0264] The four sub-model outputs are combined in a weighted manner:

[0265]

[0266] To improve the confidence of diagnosis, Beta distribution is added to model uncertainty:

[0267] α=exp(W α ·z fusion ),β=exp(W β ·z fusion )

[0268]

[0269] Beta distribution parameters α and β are used to describe the confidence of the model in the prediction results.

[0270] The expectation of the predicted outcome (i.e., the average predicted probability);

[0271] Var(y): Variance of the prediction results, used to measure uncertainty. A larger value indicates a more uncertain model.

[0272] (3) Model optimization

[0273] (3.1) Course learning mechanism

[0274] This method introduces a curriculum learning mechanism. This strategy constructs sample difficulty levels and controls the training sequence, allowing the model to gradually learn from easy to difficult. Specifically, in the initial training set, "simple samples" with significant discrimination are prioritized to help the model quickly establish initial classification capabilities. As training progresses, "complex samples" with less discriminative power or blurred boundaries are gradually introduced to enhance the model's ability to judge critical samples. This method can effectively improve the model's convergence speed and generalization ability, avoiding being trapped in local optima.

[0275] Sample difficulty assessment:

[0276]

[0277] Training step-by-step construction process:

[0278]

[0279] (3.2) Adversarial training mechanism

[0280] Furthermore, to enhance the model's robustness to noise and atypical data, this method introduces an adversarial training mechanism. By applying small perturbations to the original samples to generate adversarial examples, the model simultaneously learns from both real samples and their "opponent" examples during training, thereby improving its ability to recognize small perturbations. Furthermore, a generative adversarial network (GAN) is used to synthesize realistic pseudo-physiological data, expanding the training sample pool, effectively alleviating sample scarcity issues and enhancing the model's generalization performance under various conditions.

[0281] Adversarial perturbation generation:

[0282]

[0283] Generate adversarial network loss function:

[0284]

[0285] The content implemented in this embodiment includes:

[0286] (1) Based on the dynamic interaction mechanism between adolescents and the environment, effective VR intelligent interaction scenarios and measurement indicators, and in accordance with the characteristics of adolescent physical and mental development, a three-dimensional interactive VR scenario was designed. This involved VR scenario script design, multi-sensory channel activation (visual and auditory), interaction between adolescents and key figures, and interaction between adolescents and the environment. A multi-factor mixed experimental design was used to manipulate the core influencing variables, providing a stable and controllable experimental scenario.

[0287] (2) Using a synchronous acquisition network for EEG, eye movement, and heart rate variability, we ensure that the EEG and heart rate variability data recorded simultaneously during the subjects' free browsing of emotional faces are synchronized with millisecond accuracy, which is the basis for meaningful synchronous analysis of EEG, eye movement, and heart rate variability.

[0288] (3) Use signal processing methods to process EEG data, mainly targeting the three bands (alpha, beta and theta) that can better identify depression.

[0289] (4) After completing the study on unimodal depression recognition based on EEG, eye movement, and heart rate variability features, the unimodal EEG / eye movement / heart rate variability features are input into the denoising autoencoder to generate their respective shared representations for training the classifier to obtain the unimodal depression recognition results.

[0290] (5) Taking advantage of the complementarity of different modalities, a multimodal depression recognition study based on EEG, eye movement, and heart rate variability features was completed. Two modal fusion strategies (feature fusion and hidden layer fusion) were used to perform feature layer fusion on EEG, eye movement, and heart rate variability data to improve classification accuracy and build a more accurate depression recognition model.

[0291] (6) Complete a cross-modal depression recognition study based on EEG, eye movement, and heart rate variability features to explore whether the shared representation learned by the autoencoder is strongly correlated with adolescent depression and weakly correlated with the signal form of EEG / eye movement or heart rate variability.

[0292] Example 2

[0293] This embodiment is a system for training an auxiliary diagnosis model for adolescent depression based on multimodal data, including:

[0294] The virtual reality subsystem is used to construct VR intelligent interaction scenarios and simultaneously collect physiological data from adolescent depression and normal control subjects in VR intelligent interaction scenarios. The physiological data includes EEG data, heart rate variability data, and eye movement data;

[0295] The data processing subsystem performs data preprocessing and feature extraction on the EEG data, heart rate variability data, and eye movement data to obtain EEG feature data, heart rate variability feature data, and eye movement feature data;

[0296] The unimodal depression recognition model training subsystem trains classifiers based on EEG feature data, heart rate variability feature data, and eye movement feature data to produce a unimodal depression recognition model. The unimodal depression recognition model includes a depression recognition model based on EEG data, a depression recognition model based on heart rate variability data, and a depression recognition model based on eye movement data.

[0297] The multimodal depression recognition model training subsystem performs feature-layer fusion on EEG feature data, heart rate variability feature data, and eye movement feature data, and then performs classifier training to obtain a multimodal depression recognition model.

[0298] The cross-modal depression recognition model training subsystem uses EEG feature data, heart rate variability feature data, and eye movement feature data as inputs to the autoencoder to generate their respective shared representations; only one of the feature data is provided when training the classifier, and the other two feature data are tested on the trained classifier, thereby obtaining a cross-modal depression recognition model.

[0299] The virtual reality subsystem includes the Biopac MP160 physiological multichannel instrument and the aSee A8 portable telemetry eye tracker to obtain EEG data, eye movement data, and heart rate variability data.

[0300] Obviously, the above embodiments are merely examples for clarity of explanation and are not intended to limit the implementation methods. Those skilled in the art will readily appreciate that other variations or modifications based on the above descriptions are possible. It is not necessary and impossible to enumerate all implementation methods here. Obvious variations or modifications arising therefrom remain within the scope of protection of the present invention.

Claims

1. A method for training an auxiliary diagnosis model for adolescent depression based on multimodal data, characterized in that: include: Step S1, constructing a VR intelligent interaction scene, and synchronously collecting physiological data of adolescent depression subjects and normal control subjects in the VR intelligent interaction scene, wherein the physiological data includes EEG data, heart rate variability data, and eye movement data; Step S2, performing data preprocessing and feature extraction on the EEG data, heart rate variability data, and eye movement data to obtain EEG feature data, heart rate variability feature data, and eye movement feature data; Step S3, performing classifier training based on the EEG feature data, the heart rate variability feature data, and the eye movement feature data, respectively, to obtain a unimodal depression recognition model; the unimodal depression recognition model includes a depression recognition model based on the EEG data, a depression recognition model based on the heart rate variability data, and a depression recognition model based on the eye movement data; Step S4, performing feature layer fusion on the EEG feature data, the heart rate variability feature data, and the eye movement feature data, and then performing classifier training to obtain a multimodal depression recognition model; Step S5: Using the EEG feature data, heart rate variability feature data, and eye movement feature data as inputs to the autoencoder to generate their respective shared representations; only providing one of the feature data when training the classifier, and testing the other two feature data on the trained classifier, thereby obtaining a cross-modal depression recognition model.

2. The method according to claim 1, characterized in that In the step S2, EEG data preprocessing includes: applying a 0.5-100 Hz bandpass filter to the raw EEG signal to remove low-frequency drift and high-frequency noise; applying a 50 Hz notch filter to eliminate power supply interference; performing signal denoising through discrete wavelet transform; calculating the SNR and RMSE to evaluate signal quality; and finally separating the processed signal into five frequency bands: Delta, Theta, Alpha, Beta, and Gamma for power spectral density analysis. Heart rate variability data preprocessing includes: applying 0.04-5 Hz bandpass filtering to the raw photoplethysmography signal to remove baseline drift and high-frequency noise; performing signal denoising through discrete wavelet transform; identifying heartbeat intervals through peak detection algorithm; and extracting time domain and frequency domain indicators; The eye movement data preprocessing includes: if the missing value of a data tuple in the eye movement data exceeds a set threshold, the data tuple is discarded; if the missing value of the data tuple does not exceed the set threshold, the data tuple is filled with the average value or the median.

3. The method according to claim 2, characterized in that In step S2, the feature extraction adopts a correlation-based feature selection method combined with a best-first search strategy.

4. The method according to claim 1, wherein Before step S3, a Gaussian filter formula is used for denoising, Gaussian noise is added for data enhancement, a weighted cross entropy loss function is used for sample rebalancing, and a support vector distance discriminant function formula is used for edge sample identification and resampling; a significance P value of the evaluation feature is defined, and the difference between the evaluation feature and the target label is tested to eliminate features with insignificant differences, and finally a K-fold cross validation is performed to determine the optimal feature subset based on the comprehensive performance on the training set.

5. The method according to claim 1, wherein In step S3, any one of the EEG feature data, heart rate variability feature data and eye movement feature data is used as the input layer of the autoencoder, noise is added to the original input, and a shared representation is generated through encoding as the input data for classifier training, and then the original input is reconstructed through decoding.

6. The method according to claim 1, characterized in that In step S4, the feature layer fusion step is: directly connecting the EEG feature data, heart rate variability feature data and eye movement feature data together, inputting them into the autoencoder to generate a shared representation; using the unsupervised backpropagation algorithm to fine-tune the weights and biases of the autoencoder to generate the final shared representation for actual training of the classifier.

7. The method according to claim 6, characterized in that The fused multimodal feature vector is recorded as: With fusion ∈R d Among them, z fusion is the feature vector after multimodal fusion; Define four parallel sub-models f1, f2, f3, and f4 and output the predicted probability: pi=fi(·),i=1,2,3,4 Among them, f i is the i-th sub-model, which is used to output depression predictions from different angles; p i is the predicted probability output by the i-th sub-model; Fusion predicts the probability of depression: in, Among them, p(y=1|z fusion ): Indicates that at a given z fusion Under the condition of , the probability that the sample belongs to category 1, category 1 is depression; λ i is the weight coefficient of the i-th sub-model, satisfying p i is the output probability of the i-th sub-model, specifically: p1 is the output of the standard MLP sub-model; p2 is the output of the modal attention mechanism model; p3 is the output of the contrast enhancement model; p4 is the output of the auxiliary task model; Sub-model f1 is a standard MLP model, which is used to perform nonlinear modeling on the fusion features: ·p1=f1(z fusion )=Softmax(W1 z fusion +b1) W1 is the weight matrix; b1 is the bias vector; Softmax(·) is used to transform the output into probability; Sub-model f2 is the modal attention modeling MLP, which is used for modal weight adjustment: <h2 style=";text-align:left;direction:ltr">p2=f2(z<h2 style=";text-align:left;direction:ltr"> fusion <h2 style=";text-align:left;direction:ltr"> )=Softmax(W2·z<h2 style=";text-align:left;direction:ltr"> fusion <h2 style=";text-align:left;direction:ltr"> +b2) in: z i is the representation of the i-th mode; w is the trainable vector parameter in the attention mechanism; w is a trainable weight vector used to calculate the importance of different modalities; Tanh(·) is the hyperbolic tangent function, which is used as an activation function to introduce nonlinear transformation to the input features; exp(·) is an exponential function used to amplify the difference and form a softmax format; α i Represents the attention weight of the i-th modality. The larger the value, the more important the modality is to the final classification task. M is the total number of modes; Sub-model f3 is a contrast enhancement model, which uses the difference between positive and negative stimulus responses in emotional tasks as the enhancement feature modeling: in, z f (positive) is the fusion feature representation extracted by the model under positive emotional stimulation; z f (negative) is the fusion feature representation extracted by the model under negative emotional stimulation; Δz is the difference vector between the two states, which is used to capture the intensity of the subject's response to different emotional stimuli and is an enhancement feature; [z fusion ,Δz] represents the concatenation of the original fusion features and the emotion difference features, which is used to enhance the input expression ability of the model; w3 is the weight matrix of sub-model f3; b3 is the bias term; f3(·) is the contrast enhancement modeling module; p3 is the output probability of sub-model f3; Sub-model f4 is an auxiliary task modeling module, which is used to predict individual behavioral variables in parallel with the execution of the main task: f4=f4 (main) (z fusion )=Softmax(W4·z fusion +b4) Among them, W4 and b4 are learnable parameters; The output is probability p4, which indicates the predicted probability that the sample is depressed; r^ is the predicted reaction time; s^ is the predicted behavioral score; w r ,w s and b r ,b s are trainable parameters for auxiliary tasks; Calculate the mean square error MSE for the two behavioral variables separately; λ r ,λ s represents the weighting coefficient of auxiliary loss; The four sub-model outputs are combined in a weighted manner: Add Beta distribution to model uncertainty: α=exp(W α ·With fusion ),β=exp(W β ·With fusion ) Beta distribution parameters α and β are used to describe the confidence of the model in the prediction results; Expectations for predicted outcomes; Var(y) is the variance of the prediction results, which is used to measure uncertainty. The larger the value, the more uncertain the model is.

8. The method according to claim 7, characterized in that The model is optimized using curriculum learning and adversarial training mechanisms.

9. The method according to claim 1, characterized in that The VR intelligent interaction scene is implemented by building a Web VR environment using the A-Frame framework, and a mind dialogue space is constructed through the Claude API.

10. A training system for the auxiliary diagnosis model of adolescent depression based on multimodal data, characterized by: include: A virtual reality subsystem is used to construct a VR intelligent interactive scene and synchronously collect physiological data from adolescent depression subjects and normal controls in the VR intelligent interactive scene. The physiological data includes EEG data, heart rate variability data, and eye movement data. a data processing subsystem, performing data preprocessing and feature extraction on the EEG data, heart rate variability data, and eye movement data to obtain EEG feature data, heart rate variability feature data, and eye movement feature data; a unimodal depression recognition model training subsystem for performing classifier training based on the EEG feature data, the heart rate variability feature data, and the eye movement feature data to obtain a unimodal depression recognition model; the unimodal depression recognition model includes a depression recognition model based on the EEG data, a depression recognition model based on the heart rate variability data, and a depression recognition model based on the eye movement data; A multimodal depression recognition model training subsystem performs feature-layer fusion on the EEG feature data, the heart rate variability feature data, and the eye movement feature data, and then performs classifier training to obtain a multimodal depression recognition model; The cross-modal depression recognition model training subsystem uses the EEG feature data, heart rate variability feature data, and eye movement feature data as inputs to an autoencoder to generate their respective shared representations; only one of the feature data is provided when training a classifier, and the other two feature data are tested on the trained classifier, thereby obtaining a cross-modal depression recognition model.

Citation Information

Patent Citations

  • Depression identification method and system based on multi-modal data fusion model

    CN116010901A

  • Screening device for computer-assisted depressive symptom diagnosis

    CN116110578A

  • Multi-mode fusion depression identification auxiliary decision-making system based on EEG and voice signals

    CN117116468A

  • Multi-modal emotion recognition method and system based on confidence fusion

    CN117591967A

  • Audio-visual multi-modal emotion recognition method and system based on cross-modal attention mechanism

    CN117909885A

Cited By

  • Teenager depression prediction system

    CN121587724A

  • Adolescent depression prediction system

    CN121587724B