A depression assessment and intervention system based on multi-modal data and large models

CN122822318APending Publication Date: 2026-09-25BEIJING INST OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610835653.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-06-10
Publication Date
2026-09-25

AI Technical Summary

Technical Problem

尽管这些工具在临床实践中被广泛使用,但它们存在固有的局限性:首先,患者的回答极易受到主观因素的影响,情绪波动、认知偏差和社会期望可能会阻碍个体准确传达其真实的心理状态,从而增加了误诊和漏诊的风险

Benefits of technology

1.突破单模态生理特征识别瓶颈,显著提升抑郁状态判别精度。现有的脑电信号具有较高的噪声敏感性,且缺乏跨模态的统一融合框架 。本发明提出基于跨模态交互的脑电表示增强机制。创新性地利用大语言模型提取的语音韵律与文本模态结构化语义信息来深层强化客观的脑电特征,有效克服了单一脑电信号易受干扰的缺点,显著提高了模型对抑郁症相关神经模式的提取效率与判别力。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122822318A_ABST
    Figure CN122822318A_ABST
Patent Text Reader

Abstract

The application discloses a depression auxiliary evaluation and intervention system based on multi-modal data and a large model. The system takes electroencephalogram signals, voice signals and their transcription text signals as multi-modal inputs, enhances electroencephalogram representation through cross-modal interaction, improves the generalization ability of the model under different subjects and different interaction paradigms through cross-domain knowledge transfer, and further combines a large language model, a knowledge graph and retrieval enhancement generation technology to realize longitudinal tracking of the user's psychological state, risk stratification and personalized closed-loop intervention.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of auxiliary diagnosis and treatment of depression, mental health assessment and intelligent intervention technology, specifically involving a depression auxiliary assessment and intervention system based on multimodal data and large models. Background Technology

[0002] Depression is a highly prevalent mental health disorder characterized by persistent low mood or loss of interest in daily activities. Currently, traditional diagnostic methods for depression primarily rely on subjective self-report questionnaires and clinical interviews. While these tools are widely used in clinical practice, they have inherent limitations: First, patients' responses are highly susceptible to subjective factors; mood swings, cognitive biases, and social expectations can hinder individuals from accurately conveying their true psychological state, thus increasing the risk of misdiagnosis and missed diagnosis. Second, differences in how individuals understand and interpret questionnaire items further compromise the validity and consistency of diagnostic results.

[0003] In terms of disease intervention, existing intervention methods generally lack personalization and real-time adaptability. While drug therapy is effective for some patients, it is often accompanied by significant side effects, and individual treatment responses vary greatly. Psychotherapy, although providing support, relies heavily on the expertise and experience of psychiatrists, and treatment plans often lack the individualization required to meet the diverse needs of different patients. More importantly, traditional interventions are usually sporadic and discontinuous, making it difficult to continuously monitor the patient's mental state or adaptively adjust treatment strategies in real-time based on individual fluctuations in mental status. In recent years, the emergence of multimodal recognition technology has opened new avenues for the accurate identification and personalized intervention of depression. Electroencephalography (EEG) signals can capture neural activity in real time, providing rich physiological and psychological information and serving as an important objective basis for assessing depression.

[0004] However, the current field of EEG and multimodal assisted diagnosis and treatment of depression still faces three key technical challenges that urgently need to be addressed: First, the lack of precise cross-modal physiological feature enhancement strategies. Although EEG signals are objectively real, they are highly sensitive to noise. Existing multimodal indicators have not yet formed a unified cross-modal fusion framework, making it difficult to effectively utilize the structured semantic information of modalities such as speech and text to enrich and enhance EEG representations, resulting in models struggling to extract highly discriminative depression-related neural patterns. Second, severe cross-domain data bias and individual heterogeneity limit the generalization ability of models. EEG signals exhibit significant individual differences (such as differences in data distribution caused by factors like skull thickness and scalp impedance), and in real-world deployments, they face domain bias issues from controlled experimental paradigms (such as resting-state stimulation) to natural scenarios (such as conversational states). This makes it difficult for models to achieve high-precision identification on new real-world subjects with limited data. Third, the lack of an adaptive closed-loop intervention and modulation framework. Although the concept of data-driven precision psychiatry is emerging, current intervention systems still heavily rely on predefined procedures or static assessments based on questionnaires. Existing systems lack a truly personalized interaction paradigm based on multi-source dynamic data, making it impossible to accurately track users' dynamic states at the dialogue level, and consequently, unable to provide objective, physiologically driven, and empathetic real-time closed-loop intervention and feedback.

[0005] Currently, the clinical diagnosis of depression mainly relies on subjective self-report questionnaires and clinical interviews. However, patients' answers are easily influenced by subjective factors such as emotional fluctuations, cognitive biases, and social expectations. This not only increases the risk of misdiagnosis and missed diagnosis but also further impairs the consistency of diagnostic results due to differences in individual understanding of questionnaire items. To introduce objective assessment criteria, electroencephalography (EEG) signals have been widely studied because they can capture neural activity in real time. However, EEG signals themselves are highly sensitive to noise, and existing multimodal recognition technologies have not formed a unified cross-modal fusion framework, making it difficult to effectively utilize the structured semantic information of modalities such as speech and text to deeply enhance EEG features. This results in a serious deficiency in the model's efficiency and discriminative power in extracting neural patterns related to depression. When applying physiological features such as EEG to actual auxiliary diagnosis and treatment, the model also faces a severe bottleneck in generalization due to individual heterogeneity and cross-paradigm domain shift. Due to physical differences such as skull thickness and electrode impedance, traditional models often tend to learn the identity-specific characteristics of the subjects rather than the essential pathological characteristics of depression, resulting in extremely poor recognition accuracy across subjects. More seriously, most existing models are trained based on controlled resting-state or stimulus-state paradigms. When deployed in natural interaction or dialogue scenarios, electromyographic artifacts caused by speech and changes in cognitive load can lead to severe data distribution shifts, resulting in a sharp decline in recognition performance and making it difficult to meet the dynamic monitoring needs of real clinical environments. At the disease intervention level, existing treatment methods generally lack deep personalization and real-time adaptability. Traditional drug treatment and psychological counseling are usually sporadic and discontinuous, making it difficult not only to achieve continuous, longitudinal monitoring of the patient's psychological state, but also to adaptively adjust intervention strategies in real-time interactions based on the individual's fluctuating mental state over time. Existing digital intervention systems also mostly rely on static assessment data or preset fixed procedures, lacking personalized interaction mechanisms driven by multi-source dynamic physiological data, and are unable to track user status at the dialogue level and provide closed-loop dynamic feedback with professional empathy. Based on the numerous shortcomings of existing technologies, the technical problem this invention aims to solve is: how to overcome the bottlenecks of existing technologies, such as strong diagnostic subjectivity, weak discrimination of single-modal physiological features, and poor generalization ability across subjects and interactive scenarios, and thus construct a precise closed-loop intervention system that moves from static recognition to dynamic interaction. Specifically, this invention aims to propose an adaptive framework that integrates multimodal data and a large language model. It seeks to enhance the physiological and semantic representation of EEG features through cross-modal interaction, eliminate signal interference caused by individual physical differences and paradigm shifts through cross-domain knowledge transfer, and ultimately achieve dialogue-level longitudinal state tracking and highly personalized, clinically evidence-based real-time emotion intervention based on a large language model and professional knowledge graph. Summary of the Invention

[0006] In view of this, the purpose of this invention is to provide a depression-assisted assessment and intervention system based on multimodal data and a large model. This system uses electroencephalogram (EEG) signals, speech signals, and their transcribed text signals as multimodal inputs. It enhances EEG representation through cross-modal interaction, improves the model's generalization ability across different subjects and interaction paradigms through cross-domain knowledge transfer, and further combines large language models, knowledge graphs, and retrieval-enhanced generation techniques to achieve longitudinal tracking of users' psychological states, risk stratification, and personalized closed-loop intervention.

[0007] 1. A depression auxiliary assessment and intervention system based on multimodal data and a large model, characterized in that it includes:

[0008] The multimodal data acquisition and preprocessing module is used to acquire the subject's electroencephalogram (EEG) signals, speech signals, and text signals transcribed from the speech signals, and to preprocess the data of different modalities. The base feature extraction module extracts deep representations of EEG, speech, and text. The EEG representation enhancement module based on cross-modal interaction is used to receive deep EEG representations, deep speech representations, and deep text representations, and generate enhanced EEG representations through cross-modal semantic alignment and dynamic fusion mechanisms. The cross-domain knowledge transfer and personalized modeling module is used to perform subject invariant feature learning, cross-paradigm domain alignment, target domain pseudo-label self-training, and adaptive anchor feature calibration based on the enhanced EEG representation, so as to improve the system's generalization ability under different subjects and different interaction scenarios. The large language model closed-loop state tracking and intervention module transforms the multimodal recognition results, which include depression risk probability and physiological characteristics, into a structured context that the large language model can understand, and combines knowledge graphs and retrieval enhancement generation technology to generate personalized intervention responses. The risk stratification and safety constraint module is used to classify the user's current state into different risk levels based on the probability of depression identification, historical state trends, structured risk expression, and multimodal physiological indicators, and to trigger corresponding safety intervention processes.

[0009] Furthermore, it also includes an intervention feedback update module, which is used to continue collecting multimodal data for the next round after each round of intervention, and to use the changes in the state before and after the intervention as feedback signals to update the user's historical state and the next round of intervention strategy.

[0010] Furthermore, it also includes a personalized long-term baseline management module, which is used to maintain a long-term individual baseline for each user and compare the current status with the individual baseline to assist in assessment and intervention decisions.

[0011] Preferably, the EEG representation enhancement module based on cross-modal interaction includes: Feature projection units are used to map deep EEG representations, deep speech representations, and deep text representations to a unified semantic space; The similarity calculation unit is used to calculate the semantic relevance between EEG features and projected speech features and projected text features; A dynamic weight allocation unit is used to generate dynamic fusion weights for EEG, speech, and text based on the semantic relevance. A cross-modal fusion unit is used to perform a weighted summation of EEG, speech, and text features according to the dynamic fusion weights to obtain an enhanced EEG representation; and The multimodal contrast constraint unit is used to force the alignment of EEG, speech and text modalities in a unified semantic space through the multimodal contrast loss function.

[0012] Preferably, the dynamic weight allocation unit uses a Softmax function with a temperature coefficient to generate dynamic fusion weights; when the quality of a certain modality is lower than a preset threshold or the modality is missing, the cross-modal fusion unit renormalizes the fusion weights according to the modality quality score and uses the remaining effective modalities to complete the enhanced representation calculation; when the EEG modality is missing or the quality score is lower than a set threshold, the system pauses the cross-modal similarity calculation and directly uses the deep features of the remaining effective modalities for pooling or attention fusion.

[0013] Preferably, the cross-domain knowledge transfer and personalized modeling module includes: 1) Construct a domain adversarial neural network including a feature extractor, a depression classifier, and a subject identity discriminator. After training, obtain a feature extractor trained with subject identity invariance constraints. and its corresponding depression classifier ; 2) The source domain data of the controlled experiment paradigm and the target domain data of the natural interaction paradigm are compared with the feature extractor in 1). and its corresponding depression classifier The first fine-tuning was performed to enable it to extract depression-related features consistent across paradigms, thus obtaining the cross-paradigm aligned model. 3) Use the cross-paradigm aligned model to make preliminary predictions on the target domain data of the natural interaction paradigm to obtain the initial prediction probability of the target domain sample; when the prediction probability is higher than the set threshold, the sample is used as a reliable pseudo-label sample of the target domain. Continue using labeled samples from the source domain and reliable pseudo-labeled samples from the target domain to extract features. Depression classifier The model was fine-tuned a second time; 4) Calculate the Pearson correlation coefficient between each feature dimension and the depression label in the source domain, and select the feature dimension with a high correlation with the depression label as the anchor feature; In the target domain, the feature extractor with the second fine-tuning is used. Deep features are extracted from the target domain samples and input into the depression classifier after a second fine-tuning. The output initial probabilities are used to perform soft clustering on the target domain samples, and weighted voting is performed based on the difference in centroids of anchor features to determine the calibration labels of the target domain samples.

[0014] Preferably, in the cross-domain knowledge transfer and personalized modeling module, the threshold is set as a dynamically adjustable threshold, defined as:

[0015] in, δ represents the baseline confidence threshold, and δ represents the threshold relaxation factor. Indicates the source domain prediction confidence. This indicates the high confidence threshold. This represents the model's predicted label results for subjects in the source domain data.

[0016] The weighted voting method, which combines the differences in centroids of anchor points to determine the calibration labels of target domain samples, includes: Soft clustering yielded two candidate clusters. and ; The centroid difference between the two clusters on each anchor feature dimension is:

[0017] in, and Representing candidate clusters and In the Centroid of each anchor point feature dimension;

[0018] Where d represents the feature dimension. Represents the set of anchor point features. This represents the Pearson correlation coefficient. The centroid difference between the two clusters at each anchor feature dimension; If V is greater than 0, then determine It belongs to the depression-related positive cluster; otherwise, it is judged as follows: Let the positive clusters related to depression be denoted by . The centroid of the negative cluster is The overall feature center of the target domain samples of the subjects is Then the calibration label for the target domain sample is: 。

[0019] Preferably, the large language model closed-loop state tracking and intervention module includes: Calculate the cosine similarity between the features of the current round of dialogue and the features of previous rounds. The Softmax function is used to generate historical round weights. Where N represents the total number of historical rounds; τ is the temperature coefficient; and the historical comprehensive state vector is:

[0020] Current state characteristics With historical aggregation characteristics Converted into natural language description, a summary of historical emotional evolution is obtained; The numerical features of the multimodal recognition results are transformed into structured natural language descriptions to form objective multimodal indicator text. ; Using large language models or medical text encoders to extract data from user statements, historical summaries, and multimodal indicator texts. Extracting structured psychological state information ; System commands for virtual mental health assistants or virtual psychiatrists Summary of the Evolution of Historical Emotions Objective multimodal indicator text Structured psychological state information and the current user statement Concatenate them into a unified multimodal query context; Retrieve relevant entities, relationships, and unstructured evidence from mental health knowledge graphs and authoritative corpora based on multimodal query context; and Based on the obtained multimodal query context, related entities, relationships, and unstructured evidence, a large language model is invoked to generate personalized intervention responses under the joint constraints of professionalism, empathy, security, personalization, and boundary constraints.

[0021] Preferably, the At least including relative power changes in EEG frequency bands, characteristics of EEG asymmetry between the left and right hemispheres, probability of depression risk, risk level, changes in speech rate and pauses, abnormalities in tone and rhythm, expression of negative emotions in text, expression of self-denial, expression of despair, and trends in state changes compared to the user's historical baseline. The extracted structured psychological state information fields include emotional state, depressive symptoms, triggering events, negative automatic thoughts, cognitive distortion type, sleep and appetite-related descriptions, social support status, time information, risk expression, and content related to self-harm, suicide, or harming others. The entity types in the knowledge graph include diseases, symptoms, emotional states, cognitive distortions, treatment methods, intervention techniques, risk factors, life events, behavioral activation methods, cognitive behavioral therapy techniques, relaxation training methods, and crisis intervention measures; the relationship types in the knowledge graph include manifestation, possible trigger, related, available, belonging, mitigation methods, treatment methods, based on, and included. Unstructured evidence includes explanations of depressive symptoms, recommendations for cognitive behavioral therapy, behavioral activation methods, sleep management recommendations, emotion regulation techniques, and crisis intervention guidelines; The professionalism constraint requires responses to be based on retrieved mental health knowledge and intervention principles; the empathy constraint requires responses to use supportive, non-judgmental, gentle, and respectful language; and the safety constraint requires responses to identify crisis signals and trigger crisis intervention procedures.

[0022] Preferably, in the multimodal data acquisition and preprocessing module, the preprocessing of the EEG signal includes a first type of preprocessing and a second type of preprocessing: the first type of preprocessing is used for cross-modal representation enhancement, including baseline drift removal, power frequency filtering, bandpass filtering, resampling, standardization, and segmentation, and organizing it into a time patch tensor adapted to the input of the EEG pre-training model; the second type of preprocessing is used for closed-loop intervention and dialogue state tracking, using a short time window to extract power spectral density and relative power features.

[0023] The present invention has the following beneficial effects: 1. Overcoming the bottleneck of single-modal physiological feature recognition and significantly improving the accuracy of depressive state discrimination. Existing EEG signals are highly sensitive to noise and lack a unified cross-modal fusion framework. This invention proposes an EEG representation enhancement mechanism based on cross-modal interaction. It innovatively utilizes phonological prosody and textual modal structured semantic information extracted from a large language model to deeply enhance objective EEG features, effectively overcoming the susceptibility of single EEG signals to interference, and significantly improving the model's extraction efficiency and discriminative power for depression-related neural patterns.

[0024] 2. Overcoming individual differences and cross-domain bias to achieve high generalization in real-world interactive scenarios. Addressing significant individual physical differences in EEG signals (such as skull thickness and impedance) and severe cross-domain data bias caused by the shift from controlled experimental paradigms to natural dialogue interaction paradigms, this invention designs a cross-domain knowledge transfer and personalized modeling mechanism. Through subject-identity-invariant feature learning, cross-paradigm domain alignment, dynamic pseudo-label self-training, and adaptive anchor point calibration mechanisms, the model effectively eliminates signal interference caused by physical differences and paradigm bias, achieving high-precision zero-shot generalization across scenarios with limited data and new real-world subjects.

[0025] 3. Breaking the limitations of static assessment, achieving "dialogue-level" dynamic closed-loop interaction and intervention. This invention breaks through the limitations of traditional intervention systems that heavily rely on static assessments or preset fixed procedures, and constructs a large language model closed-loop state tracking and intervention module. The system transforms objective physiological indicators of multimodal recognition into explicit structured text constraints that the large model can understand. Combined with mental health knowledge graphs and Retrieval Enhancement Generation (RAG) technology, it can not only accurately track the dynamic state evolution of users in vertical multi-turn dialogues, but also output highly empathetic and adaptive emotional intervention responses with professional clinical medical basis, forming a closed-loop feedback mechanism of "listening—assessment—personalized intervention—listening again".

[0026] 4. A robust safety net mechanism and personalized long-term baseline management. This invention introduces a risk stratification and safety constraint module, which can perform multi-dimensional comprehensive risk assessment based on multimodal physiological indicators, historical trends, and structured risk expressions. When crisis signals such as self-harm, suicide, or severe despair are detected, a mandatory safety intervention process is triggered first, ensuring the system's safety in clinical application. Simultaneously, the personalized long-term baseline management module maintains a unique long-term physiological and cognitive baseline for each user. By comparing the deviation between the current state and the individual baseline, it keenly detects potential increases in risk, truly achieving proactive, personalized, dynamic, intelligent, and closed-loop intelligent assessment and intervention for depression. Attached Figure Description

[0027] Figure 1 This is a schematic diagram of the overall framework of the MM-LLM of the present invention; Figure 2 This is a structural diagram of the EEG representation enhancement module of the present invention; Figure 3 This is the LLM response generation graph based on knowledge enhancement and empathy guidance of the present invention. Detailed Implementation

[0028] The present invention will now be described in detail with reference to the accompanying drawings and embodiments.

[0029] The system comprises: a multimodal data acquisition and preprocessing module, a base feature extraction module, an EEG representation enhancement module based on cross-modal interaction, a cross-domain knowledge transfer and personalized modeling module, a large language model closed-loop state tracking and intervention module, a risk stratification and safety constraint module, an intervention feedback update module, and an individualized long-term baseline management module. These modules are sequentially connected according to data flow, collectively forming a closed-loop system from multimodal perception and auxiliary assessment of depressive states to personalized intervention feedback.

[0030] 1. Multimodal data acquisition and preprocessing module The multimodal data acquisition and preprocessing module is used to acquire EEG signals, speech signals, and text signals transcribed from speech signals of subjects under controlled experimental paradigms and natural interaction paradigms, and to perform standardized preprocessing on data from different modalities.

[0031] The EEG signals can originate from controlled experimental paradigms such as resting state, stimulus state, and task state, or from natural interaction paradigms such as speaking state, question-and-answer state, and natural dialogue state; the speech signals are the user's verbal expressions during question-and-answer or dialogue processes; and the text signals are transcribed from the speech signals by an automatic speech recognition system. By simultaneously collecting EEG, speech, and text data, the system can comprehensively characterize the subject's depression-related state from three levels: neurophysiological activity, speech prosody, and language semantic content.

[0032] 1) For EEG signals, this module includes two types of preprocessing procedures. The first type is the EEG preprocessing procedure for cross-modal representation enhancement. The system first removes the signal mean and low-frequency baseline drift, uses a notch filter to remove power frequency interference, uses a bandpass filter to retain the effective EEG frequency band, and resamples and normalizes the EEG signal. Subsequently, the continuous EEG signal is segmented into fixed-length segments and further organized into time patch tensors adapted to the input of the EEG pre-training model. In one specific implementation, the EEG signal is processed by a 1-40Hz bandpass filter, resampled to 200Hz, and segmented into non-overlapping 30-second samples. Each sample is further reshaped into a normalized input matrix or tensor for input to the EEG pre-training encoder.

[0033] Let the preprocessed EEG input be represented as:

[0034] in, For batch size, The length of the time series. The feature dimensions for each time segment are defined. To adapt to the input format of the EEG pre-trained model, these dimensions are expanded to obtain:

[0035] The newly added dimension is used to standardize the model input format.

[0036] The second type of processing targets EEG signals used for closed-loop intervention and dialogue state tracking. The system can extract interpretable physiological indicators using a shorter time window. Specifically, firstly, the signal mean and baseline drift are removed; a 50Hz notch filter is used to remove power line interference; and a 0.5-45Hz second-order Butterworth bandpass filter is used to remove low-frequency drift and high-frequency noise. Then, the Welch method is used to calculate the power spectral density over a 2-second sliding window, extracting power characteristics from five standard frequency bands (δ, θ, α, β, γ), and calculating relative power to reduce the influence of individual physical differences such as skull thickness, scalp impedance, and electrode contact status on EEG amplitude. Relative power can be expressed as:

[0037] Where PSD represents power spectral density and band represents the corresponding EEG frequency band. If three EEG channels and five frequency bands are used, 15-dimensional relative power features can be obtained in each 2-second window, which can be used for subsequent dialogue-level state tracking and intervention feedback updates.

[0038] 2) Preprocessing the speech signal. The acquired speech signal is first resampled to meet the input requirements of the speech pre-training model. In one specific implementation, the speech signal is uniformly resampled to 16kHz and segmented into fixed-length speech segments. Let the speech input be:

[0039] in, This indicates the length of the speech sequence. Speech signals contain information such as speech rate, pauses, pitch, loudness, rhythm, and prosody. This information can reflect characteristics that may appear in a depressive state, such as a lowered voice, slower speech rate, increased pauses, and reduced emotional expression.

[0040] 3) Preprocessing the text signal. The text signal is obtained by transcribing the user's speech using an automatic speech recognition model. In one specific implementation, the system uses the Whisper automatic speech recognition model to transcribe the user's speech, and then performs word segmentation, lexicalization, truncation, padding, and encoding on the transcribed text to obtain an input sequence that can be processed by the text encoder. Let the text input be:

[0041] in, This indicates the length of the text lexical sequence. The text signal is used to extract semantic information from user expressions, such as negative emotion words, self-denial, helplessness, loss of interest, sleep and appetite-related descriptions, risk events, and expressions of potential self-harm or suicide risks.

[0042] 2. Base Feature Extraction Module The base feature extraction module is connected to the multimodal data acquisition and preprocessing module, and is used to call the pre-trained base models corresponding to different modalities to perform deep representation extraction on EEG, speech and text data.

[0043] 1) For EEG signals, in one specific implementation, this module inputs the preprocessed EEG tensor into a pre-trained CBraMod model, which can also be replaced by other pre-trained EEG encoders. The EEG input is first encoded by the PatchEmbedding module, a process that includes convolutional projection, frequency domain coding, and positional coding, which can be represented as:

[0044] in, This indicates an input projection operation. This represents the frequency domain coding of the Fast Fourier Transform. This represents the location encoding. Subsequently, the encoded EEG patch representation is input into the Transformer encoder for deep spatiotemporal interaction:

[0045] Remove the final task projection layer of the pre-trained model and retain the Transformer encoder output as a deep EEG representation:

[0046] in:

[0047] C represents the channel or embedding dimension, T represents the time dimension, and D represents the feature dimension. Since subsequent integration with speech and text features requires a unified spatial fusion, the system further performs pooling, flattening, or linear projection on the deep EEG representation to obtain a unified-dimensional EEG vector representation:

[0048] in, The EEG representation transformation network outputs the following:

[0049] in This refers to the "feature dimension" of the brainwave signal after it has passed through a transformation network.

[0050] 2) For speech signals, the Wav2Vec2.0 model is used to extract contextual hidden states containing speech rate, pitch, pauses, rhythm, and emotional prosody. Alternatively, HuBERT, WhisperEncoder, or other pre-trained speech models can be used. The speech encoder output is:

[0051] in:

[0052] Indicates batch size, Representing the unified dimensions of EEG features Represents the number of time steps of the speech signal This represents the dimension of the hidden layer features output by the speech encoder.

[0053] To obtain a fixed-length speech-level representation, average pooling is performed on the time dimension, ultimately yielding speech features:

[0054]

[0055] in, This represents the speech hiding state at the nth time step: 3) For text signals, the RoBERTa model is used to extract semantic features. Alternatively, BERT, DeBERTa, Qwen encoders, or other pre-trained language models can be used. The text encoder outputs the hidden states:

[0056] The system extracts special classification tags or pooling layer outputs as text semantic representations:

[0057] in:

[0058] This represents the hidden layer feature dimension of the speech pre-trained model output; this text semantic representation can encode information such as emotional tendencies, cognitive patterns, life events, negative automatic thoughts, self-denial, helplessness, despair, and potential risk expressions in user expressions.

[0059] 3. EEG representation enhancement module based on cross-modal interaction The EEG representation enhancement module based on cross-modal interaction is connected to the base feature extraction module to receive deep EEG representations, deep speech representations, and deep text representations, and to generate enhanced EEG representations through cross-modal semantic alignment and dynamic fusion mechanisms.

[0060] Specifically, as shown in the appendix Figure 2As shown, the EEG representation enhancement module consists of three input branches, a feature projection unit, a similarity calculation unit, a dynamic weight allocation unit, a cross-modal fusion unit, and a multimodal contrast constraint unit. The three input branches receive EEG representations from CBraMod, speech representations from Wav2Vec2.0, and text representations from RoBERTa, respectively. The EEG representation serves as the primary modality, used to preserve neurophysiological activity features; the speech and text representations serve as auxiliary modalities, used to provide prosodic and semantic cues related to depression.

[0061] First, the feature projection unit in this module maps EEG, speech, and text features to a unified semantic space. Since the original feature dimensions of EEG, speech, and text differ, direct similarity calculation and weighted fusion cannot be performed. Therefore, the system first maps the three modalities to a unified semantic space. This is because, in the preceding module, the deep EEG representation has already been transformed into a unified-dimensional EEG vector through a transformation network. Therefore, this unit only needs to map the text features and speech features to the EEG semantic space through a linear projection network: Text features are mapped to the EEG semantic space via a linear projection network:

[0062] Speech features are mapped to the EEG semantic space via a linear projection network:

[0063] in, and For learnable projection matrices, and This is the bias term. The three modal features after projection satisfy:

[0064] This ensures that EEG, speech, and text features can be compared and fused within the same semantic space.

[0065] Secondly, similarity calculation and dynamic fusion are performed. The system uses cosine similarity to calculate the semantic relevance between EEG features and projected cross-modal features. For any two feature vectors a and b, the cosine similarity is:

[0066] Each feature association score used to generate dynamic weights is calculated separately, specifically including EEG self-modal preservation scores and cross-modal similarity scores between EEG and text, and between EEG and speech:

[0067]

[0068]

[0069] in, This represents an EEG self-modal gating network used to calculate the self-modal preservation score of EEG features. In other implementations, EEG self-modal scores can also be generated from self-similarity, learnable parameters, or modal quality scores.

[0070] Subsequently, the system uses the Softmax function with a temperature coefficient to generate dynamic fusion weights:

[0071] Where τ is a temperature coefficient used to adjust the smoothness of the weight distribution. The fusion weights for EEG, text, and speech are defined, and all three satisfy the following conditions:

[0072] The final enhanced EEG representation was obtained through weighted summation:

[0073] in, This is an enhanced EEG representation that integrates speech prosody information and text semantic information. This enhanced EEG representation not only retains the objective physiological characteristics of EEG signals but also incorporates structured semantic information related to depression from speech and text, thereby improving the ability of EEG features to discriminate depressive states. The enhanced EEG representation e and its corresponding depression risk prediction results are then passed to the cross-domain knowledge transfer and personalized modeling module.

[0074] 4. Cross-domain knowledge transfer and personalized modeling module The cross-domain knowledge transfer and personalized modeling module receives the enhanced EEG representation output by the EEG representation enhancement module based on cross-modal interaction. The processing results are then transmitted to the downstream large language model closed-loop state tracking and intervention module and the risk stratification and safety constraint module. This is used for subject-invariant feature learning, cross-paradigm domain alignment, target domain pseudo-label self-training, and adaptive anchor feature calibration based on enhanced EEG representation, to improve the system's generalization ability under different subjects and different interaction scenarios.

[0075] 1) First, this module performs subject-invariant feature learning.

[0076] Because of significant individual differences in EEG signals, traditional models tend to learn subject-specific characteristics rather than the essential pathological features related to depression. To reduce the influence of individual differences caused by factors such as skull thickness, scalp impedance, baseline EEG rhythm, and electrode contact status, this invention introduces a domain adversarial neural network. This network includes a feature extractor, a depression classifier, and a subject identity discriminator. Let the dataset be:

[0077] in, This represents the input feature of the i-th sample. Indicates depression label, This represents the subject's identity label. The feature extractor output is:

[0078] in, This is a deep feature representation used for subsequent depression identification and adversarial training.

[0079] The depression classifier outputs a feature extractor. As input, its output is:

[0080] The subject identification discriminator outputs the feature extractor. As input, its output is:

[0081] Here, GRL stands for Gradient Reversal Layer. The Gradient Reversal Layer maintains the identity mapping during forward propagation, and multiplies the gradient from the subject identity discriminator by a negative coefficient during backward propagation. This makes the subject identity discriminator try to distinguish the subjects' identities as much as possible, while the feature extractor tries to confuse the subjects' identities, thereby learning features that are unrelated to the subjects' identities but are related to the depression label.

[0082] The classification loss is defined as:

[0083] Subject identity adversarial loss is defined as:

[0084] The overall training objective is:

[0085] in, To counteract the loss of weight, The weights are for regularization. Due to the presence of the gradient reversal layer, although the overall loss function is optimized in the form of a weighted sum, for the feature extractor, the adversarial loss term is equivalent to maximizing the subject identity discrimination error, thereby suppressing subject identity-related interference and improving the model's generalization ability on unseen subjects.

[0086] After this step of training is completed, the output is a feature extractor trained with subject identity invariance constraints. and its corresponding depression classifier Among them, feature extractor As part of the shared backbone network, the depression classifier continues to propagate to subsequent cross-paradigm domain alignment steps. It will then be used as the classification head for subsequent source domain supervised training and target domain pseudo-label generation.

[0087] 2) Next, perform cross-paradigm domain alignment.

[0088] Although feature extractor While individual physical differences have been eliminated, in real-world deployments, models typically need to migrate from controlled experimental paradigms such as resting or stimulus states to natural interaction paradigms such as speaking, question-and-answer, or natural dialogue states. Speaking actions in natural dialogue states introduce electromyography artifacts, and the cognitive load from proactive communication differs from that in controlled experimental paradigms, resulting in a significant distributional shift between the source and target domains.

[0089] To solve this problem, this step uses the feature extractor output in step 1). Depression classifier The training continues on this basis, and a paradigm domain discriminator is added to determine whether a sample belongs to the source domain or the target domain. It should be noted that the paradigm domain discriminator here has a different function from the subject identity discriminator in step 1: the subject identity discriminator in step 1) is used to determine which subject the sample comes from, with the aim of eliminating individual differences; the paradigm domain discriminator in this step is used to determine whether the sample comes from a controlled experimental paradigm or a natural interaction paradigm, with the aim of eliminating domain shifts caused by different interaction scenarios.

[0090] Let the source domain be controlled experimental paradigm data:

[0091] in, Indicates the first in the source domain Input features of each sample This indicates the depression status label corresponding to this sample. This represents the total number of labeled samples in the source domain.

[0092] The target domain is data based on the natural interaction paradigm.

[0093] in, Indicates the first in the target domain Input features of each sample This represents the total number of unlabeled samples in the target domain.

[0094] The source domain is labeled, while the target domain may be unlabeled. The system inherits the feature extractor trained in the previous step. As the backbone, domain adversarial training is employed to align the overall distribution of the source and target domains, and CORAL loss is used to explicitly align the second-order statistical features of the source and target domains. The overall loss is:

[0095] in, , Let represent the weighting coefficients of the domain discriminative adversarial loss and the covariance alignment loss, respectively. Represents classification loss, For domain discrimination adversarial loss between the source and target domains, The covariance alignment loss is defined as:

[0096] in, and Let d represent the covariance matrices of the source and target domain features, respectively, and d represent the feature dimension. Let Frobenius norm be represented. The covariance matrix can be expressed as:

[0097] Where Z represents the feature matrix within a batch, and n represents the batch size. This represents the mean of the characteristics of this batch.

[0098] Through the above training, the feature extractor On the one hand, it maintains the ability to classify depression on labeled data in the source domain; on the other hand, it reduces the distributional differences between the controlled experimental paradigm and the natural interaction paradigm through paradigm domain adversarial loss and CORAL loss. In other words, this step does not reconstruct a new depression recognition model, but rather uses the feature extractor obtained in step 1). Depression classifier Based on this, further fine-tuning is performed to enable the extraction of depression-related features consistent across paradigms. The output of this step consists of two parts: first, the feature extractor after cross-paradigm global distribution alignment. Depression classifier Second, the aligned model is used to make preliminary predictions on natural interaction samples in the target domain, obtaining the initial prediction probabilities of the target domain samples: The initial predicted probability is not directly output as the final diagnostic result, but is passed to step 3 as prior information after cross-paradigm alignment for screening and self-training of pseudo-label samples in the target domain.

[0099] 3) Next, perform dynamic pseudo-label self-training in the target domain.

[0100] After completing the global distribution alignment of the source and target domains based on domain adversarial and CORAL loss, the system further utilizes unlabeled samples from the target domain for pseudo-label self-training to achieve instance-level fine-grained adaptation.

[0101] Specifically, this step continues to use the same feature extractor output from step 2). Depression classifier For target domain samples Output predicted probability:

[0102] in, Used to represent the probability that a target domain sample belongs to different categories of depressive state.

[0103] When the predicted probability is higher than a set threshold, the sample is added to the training set as a reliable pseudo-label sample.

[0104] To reduce the risk of missed detections during cross-paradigm transitions, this invention further proposes a dynamic confidence threshold strategy based on source domain priors (i.e., dynamically adjusting the aforementioned screening thresholds). If a subject has demonstrated a high risk of depression in the source domain resting-state or stimulus-state data, the system appropriately lowers the pseudo-label screening threshold in the target domain dialogue state to improve the sensitivity of identifying potential depressive states. The dynamic threshold is defined as:

[0105] in, δ represents the baseline confidence threshold, and δ represents the threshold relaxation factor. Indicates the source domain prediction confidence. This indicates the high confidence threshold. This represents the model's predicted label result for the subject in the source domain (such as stimulus state) data (a value of 1 indicates that the model predicts that the subject has a risk of depression).

[0106] Subsequently, the system uses source domain labeled samples and target domain pseudo-labeled samples (i.e., the reliable sample set selected using the threshold above). (Selected reliable samples) jointly fine-tuned including feature extractor Depression classifier Overall network model:

[0107] in, Represents the cross-entropy loss function. This represents the target domain pseudo-label loss weight. This represents the supervised cross-entropy loss on labeled samples in the source domain. This represents the cross-entropy loss on pseudo-labeled samples in the target domain.

[0108] This strategy can utilize the distribution information of the target domain data itself to enable the feature extractor Depression classifier This step further adapts the model to real-world interaction scenarios such as natural dialogue, question-and-answer, or speaking modes, thereby improving its recognition sensitivity in the target domain. The output of this step is the feature extractor, fine-tuned through self-training with pseudo-labels in the target domain. Depression classifier To avoid problems such as probability compression, probability shift, or confidence distortion in unlabeled target domains, the system does not directly use the predicted probability output by the model as the final judgment result. Instead, it further extracts the high-dimensional deep feature vector output by the last layer of the model or the layer before the classifier as the input for the subsequent step 4) adaptive anchor feature calibration.

[0109] 4) Next, perform adaptive anchor point feature calibration.

[0110] Since the predicted probability of the target domain may shift or be compressed after cross-paradigm transfer, and the fixed probability threshold may no longer be reliable, this invention proposes an unsupervised calibration strategy based on adaptive anchor features. First, the Pearson correlation coefficient between each feature dimension and the depression label is calculated in the source domain:

[0111] Where d represents the feature dimension. This indicates that all samples in the source domain feature matrix are at the th... The set of values ​​for each feature dimension This represents the depression label of the source domain sample. Anchor feature sets are composed of feature dimensions that are highly correlated with the depression label.

[0112] in, This represents the relevance threshold. This anchor feature set is used to represent physiological or semantic features in the source domain that are stably associated with depressive states.

[0113] In the target domain, the system uses the feature extractor fine-tuned in step 3). Deep features of the target domain samples are extracted and combined with a depression classifier. The output initial probabilities are used to perform soft clustering on the target domain samples, resulting in two candidate clusters. and It should be noted that the clustering results here are only used to determine the relative distribution structure of the target domain samples; the two candidate clusters themselves cannot directly determine which cluster corresponds to the depression-related positive class.

[0114] Then, the centroid difference between the two clusters at each anchor feature dimension is calculated:

[0115] in, and Representing candidate clusters and In the Centroid (mean) of each anchor point feature dimension.

[0116] Weighted voting is applied to the soft clustering results of the target domain samples based on the correlation direction of the source domain:

[0117] If V is greater than 0, then determine It belongs to the depression-related positive cluster; otherwise, it is judged as follows: This is a positive cluster related to depression. Let the centroid of the positive cluster be... The centroid of the negative cluster is The overall feature center of the target domain samples of the subjects is Then the calibration label for the target domain sample is:

[0118] Through this mechanism, the model's target domain determination no longer relies solely on the probability output that may lead to deviation, but further incorporates stable depression-related physiological anchor features in the source domain, thereby improving the reliability of cross-paradigm adaptation.

[0119] At this point, the cross-domain knowledge transfer and personalized modeling module is complete. The final output of this module includes at least: First, robust deep feature representation after subject identity invariance constraint, cross-paradigm domain alignment, target domain pseudo-label self-training and anchor point calibration; Second, the depression state result after adaptive anchor feature calibration, the depression state result may include binary classification label, depression risk probability, risk level or a combination thereof; Third, the state change information and model confidence information corresponding to the natural interaction data of the target domain.

[0120] The above output will directly flow into the multimodal state input unit of the large language model closed-loop state tracking and intervention module, and will be converted into objective multimodal indicator text. This objective multimodal indicator text will serve as an important constraint for the large language model to generate professional intervention scripts, risk stratification, and closed-loop feedback strategies.

[0121]

[0122] At this point, module 4) is complete. Its final output, robust features after eliminating multiple noise and anchor-calibrated depressive state results, will directly flow into the multimodal state input unit of the large language model closed-loop state tracking and intervention module, transforming it into objective multimodal indicator text. This becomes the core constraint for generating professional intervention rhetoric in large-scale models.

[0123] 5. Large Language Model Closed-Loop State Tracking and Intervention Module The large language model closed-loop state tracking and intervention module is connected to the cross-domain knowledge transfer and personalized modeling module. This module transforms the multimodal recognition results, including depression risk probability and physiological characteristics, output by the cross-domain knowledge transfer and personalized modeling module after adaptive calibration into a structured context understandable by the large language model. It then combines this with knowledge graphs and retrieval enhancement generation techniques to generate personalized intervention responses. This module corresponds to the attached... Figure 3 LLM response generation graph based on knowledge enhancement and empathy guidance.

[0124] As attached Figure 3 As shown, this module includes a user input unit, a multimodal state input unit, a historical state summary unit, a structured psychological state extraction unit, a knowledge graph retrieval unit, a RAG retrieval unit, a constraint prompt construction unit, and a large language model response generation unit. The user input unit receives the user's current statement; the multimodal state input unit receives EEG, speech, and text features, as well as depression risk indicators, output by the aforementioned model; the historical state summary unit provides information on the emotional evolution of the user during multiple rounds of dialogue; the knowledge graph retrieval unit retrieves structured mental health knowledge; the RAG retrieval unit retrieves authoritative mental health corpora and clinical intervention evidence; the constraint prompt construction unit adds professionalism, empathy, safety, personalization, and boundary constraints to the prompt words; and finally, the large language model generates an intervention response with knowledge basis and empathetic expression.

[0125] First, this module performs longitudinal historical state tracking. In multi-turn dialogue scenarios, a single turn can only reflect the user's current state, making it difficult to fully depict the evolution trend of the user's psychological state. Therefore, this invention constructs a longitudinal historical state tracking mechanism. Assume the current dialogue has K turns, and the first K... The set of historical characteristics of a single round of dialogue is as follows:

[0126] in, It can be composed of enhanced EEG features, phonological prosody features, textual semantic features, depression risk probability, and structured psychological state indicators. The system calculates the historical features and current features for each round. Cosine similarity between them:

[0127] Then, the Softmax function is used to generate historical round weights:

[0128] Where τ is the temperature coefficient. The historical integrated state vector is:

[0129] in, This represents the aggregated results of historical states that are most relevant to the current state.

[0130] The system will display the current state characteristics. With historical aggregation characteristics Convert the data into natural language description and input it into a summary generation model to obtain a summary of the user's historical emotional evolution:

[0131] Historical Summary This can include the user's recent emotional trends, recurring negative events, typical cognitive patterns, changes in symptoms, changes in risk, and feedback after the previous round of intervention.

[0132] Secondly, multimodal indicators are textualized. To enable the large language model to understand the objective indicators output by the EEG, speech, and text models, this invention transforms the numerical features output by the multimodal recognition module into structured natural language descriptions, forming objective multimodal indicator text. . This includes at least changes in the relative power of EEG frequency bands, characteristics of EEG hemisphere asymmetry, probability of depression risk, risk level, changes in speech rate and pauses, abnormalities in tone and rhythm, negative emotional expressions in the text, expressions of self-denial, expressions of despair, and trends in state changes compared to the user's historical baseline. For example, the system can generate a structured description in the form of: "In the current round of dialogue, the relative power of the EEG alpha band is lower than the individual's historical baseline, the speech rate is slower and pauses are more frequent, negative self-evaluation expressions appear in the text, and the overall risk of depression is at a moderate level." In this way, complex physiological and behavioral signals are converted into explicit textual constraints that can be understood by a large language model.

[0133] Next, structured psychological state information is extracted. The system utilizes a large language model or medical text encoder to extract structured psychological state information from current user statements, historical summaries, and multimodal indicator text. Extracted fields include emotional state, depressive symptoms, triggering events, negative automatic thoughts, cognitive distortion types, sleep and appetite-related descriptions, social support status, time information, risk expression, and content related to self-harm, suicide, or harming others. The extraction results can be represented as follows:

[0134] Here, Symptom represents symptoms, Emotion represents emotional state, Event represents triggering event, Cognition represents cognitive pattern, Risk represents risk level, Time represents time information, and Intensity represents severity. To ensure the consistency of the extraction results, the system performs post-processing, including merging semantically repetitive descriptions, mapping colloquial expressions to standard medical or psychological terms, resolving contradictory descriptions based on contextual confidence, updating state records based on time relevance, and prioritizing the labeling of high-risk expressions.

[0135] Next, construct a multimodal query context. The system will then incorporate system instructions from the virtual mental health assistant or virtual psychiatrist. Summary of the Evolution of Historical Emotions Objective multimodal indicator text Structured psychological state information and the current user statement Concatenate them into a unified multimodal query context:

[0136] in, This is a static preset text used to define model roles, response boundaries, tone style, and safety rules; Used to provide information on historical state changes; Used to provide objective physiological and behavioral indicators; Used to provide structured records of mental states; This is used to provide the user's current round of natural language input. Through this query context, the large language model can simultaneously understand the user's current expression, historical trends, objective physiological state, and professional system constraints.

[0137] Next, knowledge graph retrieval and retrieval enhancement are performed. The system is based on multimodal query context. Retrieve relevant entities and relationships within a knowledge graph related to mental health and depression intervention. Entity types in the knowledge graph include disease, symptom, emotional state, cognitive distortion, treatment methods, intervention techniques, risk factors, life events, behavioral activation methods, cognitive behavioral therapy techniques, relaxation training methods, and crisis intervention measures. Relationship types include manifestation, potential trigger, related, available, belonging to, relief methods, treatment methods, based on, and inclusion. The knowledge graph retrieval results are represented as follows:

[0138] Simultaneously, the system retrieves relevant unstructured text evidence from authoritative mental health corpora, clinical guidelines, psychological intervention manuals, or reviewed medical documents, resulting in enhanced search output:

[0139] in, This can include explanations of depressive symptoms, recommendations for cognitive behavioral therapy, behavioral activation methods, sleep management suggestions, emotion regulation techniques, and crisis intervention guidelines. Knowledge graph retrieval and RAG retrieval can be performed in parallel to obtain both structured knowledge and unstructured evidence simultaneously.

[0140] Next, a constrained intervention response is generated. After obtaining the multimodal query context, knowledge graph results, and RAG retrieval results, the system calls a large language model to generate a personalized intervention response. In one specific implementation, the large language model can be Qwen2.5-VL; in other implementations, it can be replaced with GPT, Claude, Llama, GLM, or other large language models with natural language understanding, reasoning, and generation capabilities. The system's response generation is subject to professional constraints. Empathy constraints Security constraints Personalized constraints and boundary constraints The system faces several common constraints. These include: a professional constraint requiring responses to be based on retrieved mental health knowledge and intervention principles, avoiding the fabrication of medical facts; an empathy constraint requiring responses to use supportive, non-judgmental, gentle, and respectful language; a safety constraint requiring the identification of crisis signals such as self-harm, suicide, harming others, and severe despair, and triggering crisis intervention procedures; a personalization constraint requiring the generation of personalized suggestions based on the user's historical status, current risk level, triggering events, and preferences; and a boundary constraint requiring the system to state that it only provides auxiliary mental health support and cannot replace professional medical diagnosis and treatment. The generation objective of the large language model can be expressed as:

[0141] Where Y represents a candidate response. This represents the final generated intervention response. The formula is essentially a conditional probability maximization process, meaning that, given a multimodal query context, retrieved structured and unstructured evidence, and set multi-dimensional security and professional constraints, the system evaluates all possible candidate responses. The sequence with the highest generation probability (i.e., the degree to which it conforms to the logic and constraints) is selected as the final intervention response. The generated responses can include empathetic responses to the user's current emotions, brief restatements and confirmations of the user's expression, cognitive restructuring questions based on cognitive behavioral therapy, behavioral activation suggestions, guidance on breathing relaxation or mindfulness exercises, sleep and routine suggestions, social support suggestions, encouragement to seek professional help, crisis intervention prompts when risks increase, and questions to guide the next round of dialogue.

[0142] 6. Risk Stratification and Security Constraint Module The risk stratification and safety constraint module is connected to the large language model closed-loop state tracking and intervention module. It is used to classify the user's current state into low risk, medium risk, high risk and crisis risk levels based on the probability of depression identification, historical state trend, structured risk expression and multimodal physiological indicators.

[0143] Let the probability of depression be... Historical trend score is Crisis expression score: The multimodal anomaly score is The comprehensive risk score can then be expressed as:

[0144] in, and The system uses weighted parameters. For low-risk users, it provides emotional support, active listening, and daily mental health advice; for medium-risk users, it provides cognitive restructuring, behavioral activation, and emotion regulation exercises; for high-risk users, it recommends contacting a professional psychologist or psychiatrist and increases the frequency of status tracking; for users at crisis risk, the system triggers a crisis intervention process, prompting the user to immediately contact family, friends, a professional doctor, the local emergency hotline, or a crisis intervention hotline. When the system detects suicidal ideation, self-harm tendencies, risk of harming others, or expressions of severe acute crisis, it does not output general reassurance advice but prioritizes safety intervention content, and reminds users that the system cannot replace professional medical diagnosis and treatment.

[0145] 7. Intervention Feedback Update Module The intervention feedback update module is connected to the large language model closed-loop state tracking and intervention module and the risk stratification and safety constraint module. It is used to continue to collect the user's EEG, speech and text data for the next round after each round of intervention, and use the state changes before and after the intervention as feedback signals to update the user's historical state and the next round of intervention strategy.

[0146] Let the state before the Nth round of intervention be... After the Nth round of intervention, the state in the next round is The state change is:

[0147] The system calculates the intervention feedback score based on the change in state, user text feedback, and changes in multimodal risk:

[0148] Where g represents the feedback evaluation function. and These represent the user's textual expressions before and after the intervention. If the user's risk decreases, emotional expression improves, speech rhythm stabilizes, or abnormal EEG indicators weaken after the intervention, the system reduces the intensity of subsequent interventions; if the risk continues to rise or crisis signals appear, the system increases the risk level and triggers a stronger intervention strategy. Thus, this invention forms a closed-loop process of "listening—assessment—knowledge retrieval—personalized intervention—listening again—feedback update".

[0149] 8. Personalized Long-Term Baseline Management Module The individualized long-term baseline management module is used to maintain a long-term individual baseline for each user to improve the personalization of the system. The individual baseline includes the EEG frequency band power baseline, EEG asymmetry baseline, speech rate and tone baseline, text emotion expression baseline, historical mean risk of depression, historical state fluctuation range, common triggering events, common negative cognitive patterns, and intervention preferences and effective intervention methods.

[0150] Let the set of user historical states be:

[0151] in, This indicates the total number of times the user's historical status has been recorded.

[0152] An individual's long-term baseline can be represented as:

[0153] The offset between the current state and the individual baseline is:

[0154] The system compares the current state not only with the group model but also with the user's own long-term baseline. When the current state deteriorates significantly relative to the individual baseline, even if the group risk score has not yet reached the high-risk threshold, the system can keenly detect the increase in potential risk at the individual level at the "auxiliary assessment" level, thereby updating the assessment conclusion; and triggering a matching personalized early intervention strategy at the "closed-loop intervention" level. By introducing this long-term baseline mechanism, this system truly achieves a leap from passive, uniform, static diagnosis to proactive, personalized, dynamic depression assessment and intelligent closed-loop intervention.

[0155] Through the above technical solutions, this invention enables a shift from single, static depression assistance identification to continuous, dynamic, closed-loop intervention. On one hand, this invention enhances the ability of EEG features to discriminate depressive states through cross-modal interaction enhancement of EEG, speech, and text. On the other hand, this invention reduces the impact of individual subject differences and natural interaction scenario domain shifts on model performance through domain adversarial training, cross-paradigm domain alignment, dynamic pseudo-labels, and adaptive anchor point calibration. Simultaneously, this invention transforms multimodal identification results into personalized intervention responses with professional basis, empathic expression, and safety constraints through large language models, knowledge graphs, and retrieval enhancement generation techniques, thus providing a continuously updated closed-loop intelligent method for depression assistance assessment and mental health support.

[0156] In summary, the above are merely preferred embodiments of the present invention and are not intended to limit the scope of protection of the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.

Claims

1. A depression auxiliary assessment and intervention system based on multimodal data and a large model, characterized in that, include: The multimodal data acquisition and preprocessing module is used to acquire the subject's electroencephalogram (EEG) signals, speech signals, and text signals transcribed from the speech signals, and to preprocess the data of different modalities. The base feature extraction module extracts deep representations of EEG, speech, and text. The EEG representation enhancement module based on cross-modal interaction is used to receive deep EEG representations, deep speech representations, and deep text representations, and generate enhanced EEG representations through cross-modal semantic alignment and dynamic fusion mechanisms. The cross-domain knowledge transfer and personalized modeling module is used to perform subject invariant feature learning, cross-paradigm domain alignment, target domain pseudo-label self-training, and adaptive anchor feature calibration based on the enhanced EEG representation, so as to improve the system's generalization ability under different subjects and different interaction scenarios. The large language model closed-loop state tracking and intervention module transforms the multimodal recognition results, which include depression risk probability and physiological characteristics, into a structured context that the large language model can understand, and combines knowledge graphs and retrieval enhancement generation technology to generate personalized intervention responses. The risk stratification and safety constraint module is used to classify the user's current state into different risk levels based on the probability of depression identification, historical state trends, structured risk expression, and multimodal physiological indicators, and to trigger corresponding safety intervention processes.

2. The system according to claim 1, characterized in that, It also includes an intervention feedback update module, which is used to continue collecting multimodal data for the next round after each round of intervention, and uses the changes in status before and after the intervention as feedback signals to update the user's historical status and the next round of intervention strategy.

3. The system according to claim 1, characterized in that, It also includes a personalized long-term baseline management module, which is used to maintain a long-term individual baseline for each user and compare the current status with the individual baseline to assist in assessment and intervention decisions.

4. The system according to claim 1, 2 or 3, characterized in that, The EEG representation enhancement module based on cross-modal interaction includes: Feature projection units are used to map deep EEG representations, deep speech representations, and deep text representations to a unified semantic space; The similarity calculation unit is used to calculate the semantic relevance between EEG features and projected speech features and projected text features; A dynamic weight allocation unit is used to generate dynamic fusion weights for EEG, speech, and text based on the semantic relevance. A cross-modal fusion unit is used to perform a weighted summation of EEG, speech, and text features according to the dynamic fusion weights to obtain an enhanced EEG representation; and The multimodal contrast constraint unit is used to force the alignment of EEG, speech and text modalities in a unified semantic space through the multimodal contrast loss function.

5. The system according to claim 4, characterized in that, The dynamic weight allocation unit uses a Softmax function with a temperature coefficient to generate dynamic fusion weights. When the quality of a certain modality is lower than a preset threshold or the modality is missing, the cross-modal fusion unit renormalizes the fusion weights according to the modality quality score and uses the remaining effective modalities to complete the enhanced representation calculation. When the EEG modality is missing or the quality score is lower than a set threshold, the system pauses the cross-modal similarity calculation and directly uses the deep features of the remaining effective modalities for pooling or attention fusion.

6. The system according to claim 1, 2 or 3, characterized in that, The cross-domain knowledge transfer and personalized modeling module includes: 1) Construct a domain adversarial neural network including a feature extractor, a depression classifier, and a subject identity discriminator. After training, obtain a feature extractor trained with subject identity invariance constraints. and its corresponding depression classifier ; 2) The source domain data of the controlled experiment paradigm and the target domain data of the natural interaction paradigm are compared with the feature extractor in 1). and its corresponding depression classifier The first fine-tuning was performed to enable it to extract depression-related features consistent across paradigms, thus obtaining the cross-paradigm aligned model. 3) Use the cross-paradigm aligned model to make preliminary predictions on the target domain data of the natural interaction paradigm to obtain the initial prediction probability of the target domain sample; when the prediction probability is higher than the set threshold, the sample is used as a reliable pseudo-label sample of the target domain. Continue using labeled samples from the source domain and reliable pseudo-labeled samples from the target domain to extract features. Depression classifier The model was fine-tuned a second time; 4) Calculate the Pearson correlation coefficient between each feature dimension and the depression label in the source domain, and select the feature dimension with a high correlation with the depression label as the anchor feature; In the target domain, the feature extractor with the second fine-tuning is used. Deep features are extracted from the target domain samples and input into the depression classifier after a second fine-tuning. The output initial probabilities are used to perform soft clustering on the target domain samples, and weighted voting is performed based on the difference in centroids of anchor features to determine the calibration labels of the target domain samples.

7. The system according to claim 6, characterized in that, In the cross-domain knowledge transfer and personalized modeling module, the threshold is set as a dynamically adjustable threshold, defined as: in, δ represents the baseline confidence threshold, and δ represents the threshold relaxation factor. Indicates the source domain prediction confidence. This indicates the high confidence threshold. This represents the model's predicted label results for subjects in the source domain data. The weighted voting method, which combines the differences in centroids of anchor points to determine the calibration labels of target domain samples, includes: Soft clustering yielded two candidate clusters. and ; The centroid difference between the two clusters on each anchor feature dimension is: in, and Representing candidate clusters and In the Centroid of each anchor point feature dimension; Where d represents the feature dimension. Represents the set of anchor point features. This represents the Pearson correlation coefficient. The centroid difference between the two clusters at each anchor feature dimension; If V is greater than 0, then determine It belongs to the depression-related positive cluster; otherwise, it is judged as follows: Let the positive clusters related to depression be denoted by . The centroid of the negative cluster is The overall feature center of the target domain samples of the subjects is Then the calibration label for the target domain sample is: 。 8. The system according to claim 1, 2 or 3, characterized in that, The large language model closed-loop state tracking and intervention module includes: Calculate the cosine similarity between the features of the current round of dialogue and the features of previous rounds. The Softmax function is used to generate historical round weights. Where N represents the total number of historical rounds; τ is the temperature coefficient; and the historical comprehensive state vector is: Current state characteristics With historical aggregation characteristics Converted into natural language description, a summary of historical emotional evolution is obtained; The numerical features of the multimodal recognition results are transformed into structured natural language descriptions to form objective multimodal indicator text. ; Using large language models or medical text encoders to extract data from user statements, historical summaries, and multimodal indicator texts. Extracting structured psychological state information ; System commands for virtual mental health assistants or virtual psychiatrists Summary of the Evolution of Historical Emotions Objective multimodal indicator text Structured psychological state information and the current user statement Concatenate them into a unified multimodal query context; Retrieve relevant entities, relationships, and unstructured evidence from mental health knowledge graphs and authoritative corpora based on multimodal query context; and Based on the obtained multimodal query context, related entities, relationships, and unstructured evidence, a large language model is invoked to generate personalized intervention responses under the joint constraints of professionalism, empathy, security, personalization, and boundary constraints.

9. The system according to claim 8, characterized in that, The At least including relative power changes in EEG frequency bands, characteristics of EEG asymmetry between the left and right hemispheres, probability of depression risk, risk level, changes in speech rate and pauses, abnormalities in tone and rhythm, expression of negative emotions in text, expression of self-denial, expression of despair, and trends in state changes compared to the user's historical baseline. The extracted structured psychological state information fields include emotional state, depressive symptoms, triggering events, negative automatic thoughts, cognitive distortion type, sleep and appetite-related descriptions, social support status, time information, risk expression, and content related to self-harm, suicide, or harming others. The entity types in the knowledge graph include diseases, symptoms, emotional states, cognitive distortions, treatment methods, intervention techniques, risk factors, life events, behavioral activation methods, cognitive behavioral therapy techniques, relaxation training methods, and crisis intervention measures; the relationship types in the knowledge graph include manifestation, possible trigger, related, available, belonging, mitigation methods, treatment methods, based on, and included. Unstructured evidence includes explanations of depressive symptoms, recommendations for cognitive behavioral therapy, behavioral activation methods, sleep management recommendations, emotion regulation techniques, and crisis intervention guidelines; The professionalism constraint requires responses to be based on retrieved mental health knowledge and intervention principles; the empathy constraint requires responses to use supportive, non-judgmental, gentle, and respectful language; and the safety constraint requires responses to identify crisis signals and trigger crisis intervention procedures.

10. The system according to claim 1, 2 or 3, characterized in that, In the multimodal data acquisition and preprocessing module, the preprocessing of EEG signals includes a first type of preprocessing and a second type of preprocessing: the first type of preprocessing is used for cross-modal representation enhancement, including baseline drift removal, power frequency filtering, bandpass filtering, resampling, standardization and segmentation, and organizing it into a time patch tensor that adapts to the input of the EEG pre-training model; The second type of preprocessing is used for closed-loop intervention and dialogue state tracking, employing a short time window to extract power spectral density and relative power features.