Multimodal large model oriented to mental disorders identification and intelligent intervention method and system
By employing multimodal data acquisition and fusion technologies, combined with large language models and cross-modal attention mechanisms, the heterogeneity and computational resource challenges in the diagnosis and intervention of mental disorders have been addressed, enabling precise identification and personalized intervention, thereby improving the accuracy of mental disorder diagnosis and treatment outcomes.
Patent Information
- Application Number
- CN202510473405.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-16
- Publication Date
- 2025-11-21
- Estimated Expiration
- 2045-04-16
AI Technical Summary
Existing technologies for the diagnosis of mental disorders have limitations due to their reliance on traditional methods, leading to a high risk of misdiagnosis and missed diagnosis. Multimodal data fusion also faces challenges such as heterogeneity, insufficient temporal modeling capabilities, and high computational resource requirements, making it difficult to achieve accurate identification and personalized intervention.
Employing a multimodal data acquisition, preprocessing, feature extraction, cross-modal feature fusion, and intelligent intervention system, combined with a large language model and cross-modal attention mechanism, this system utilizes EEG, ECG, electrodermal conductance, facial expression, and text data to achieve accurate identification and personalized intervention for mental disorders.
It improves the accuracy of mental disorder identification and the effectiveness of personalized intervention. Through cross-modal spatiotemporal attention mechanism and dynamic time window adaptive selection, it optimizes identification accuracy and generalization ability, constructs a closed-loop mechanism for personalized non-drug intervention, and improves treatment effect and patient compliance.
Smart Images

Figure CN120015351B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of mental disorder intelligent identification and intervention, in particular to a multi-modal large model identification and intelligent intervention method and system for mental disorders. BACKGROUND
[0002] Mental disorders are a wide range of serious mental illnesses, including depression, anxiety, bipolar disorder, schizophrenia and other types. At present, the diagnosis and identification of mental disorders mainly rely on traditional clinical methods, including patient's self-reported symptoms, standardized psychological scales and clinical doctors' interviews and observations. Although these methods can provide effective diagnostic basis in some cases, there are also significant limitations. Patients' self-reports are usually subjective and can be influenced by illness, cognitive ability and social and cultural background, which may lead to inaccurate information. In addition, the evaluation of clinical doctors usually depends on personal experience and skills, which makes the accuracy and consistency of diagnosis biased, especially in complex or atypical cases, the risk of misdiagnosis and missed diagnosis is higher.
[0003] In recent years, with the rapid development of artificial intelligence and deep learning technology, research has gradually shifted to intelligent identification methods based on multi-modal data. Multi-modal data integrates information from different physiological and behavioral signal sources, such as electroencephalogram, galvanic skin response, facial expression, speech and text data, etc. These signals can reflect the individual's physiological and psychological state from multiple angles. By integrating these multi-modal data, not only can the individual's mental state be captured more comprehensively, but also the accuracy of mental disorder identification can be improved, providing more accurate basis for early screening and intervention.
[0004] However, the application of multimodal data fusion technology in the identification of mental disorders and non-pharmacological interventions faces numerous challenges, mainly in terms of data heterogeneity, insufficient temporal modeling capabilities, data redundancy and noise interference, and excessive computational resource requirements. First, different modalities of data, such as electroencephalograms (EEGs), electrodermal responses, and facial expressions, exhibit significant heterogeneity, including differences in signal scale, noise characteristics, and acquisition methods, making data fusion complex and difficult to achieve effectively. Second, existing temporal modeling methods are still insufficient in processing multimodal data. The identification of mental disorders typically requires accurate capture of the temporal features of the brain and autonomic nervous system, but the temporal features of different modalities are often difficult to align and model uniformly. Third, due to the potential presence of noise and redundant information in the data, extracting effective features while ensuring data quality, and ensuring the accuracy and interpretability of the fusion results, has become a major technical challenge. Finally, although multimodal data fusion has great potential, existing methods have high computational resource requirements and strict real-time requirements, making it difficult to balance efficiency and effectiveness in practical applications. The multimodal large-scale model identification and intelligent intervention method and system for mental disorders proposed in this invention not only breaks through the limitations of existing technologies, but also provides innovative solutions for early screening, accurate diagnosis and personalized intervention of mental disorders. Summary of the Invention
[0005] The purpose of this invention is to provide a multimodal large-scale model identification and intelligent intervention method and system for mental disorders. It aims to achieve accurate identification of mental disorders and dynamically optimize personalized intervention strategies by integrating multimodal signals from EEG, ECG, electrodermal conductance, facial expression analysis and text data, combined with a large language model and cross-modal attention mechanism.
[0006] To achieve the above objectives, this invention provides a multimodal large-scale model identification and intelligent intervention system for mental disorders, comprising:
[0007] Multimodal data acquisition module: Simultaneously acquires multimodal physiological data such as electroencephalogram (EEG), electrocardiogram (ECG), electrodermal conductance (EDA), and facial expressions through multiple sensors;
[0008] Data preprocessing module: performs cleaning and standardization processing on the collected multimodal data;
[0009] Feature extraction module: Extracts key spatiotemporal features from the data of each modality to prepare for subsequent analysis;
[0010] Cross-modal feature fusion module: Utilizes a cross-modal spatiotemporal attention mechanism to fuse data features from different modalities and integrate multi-source information;
[0011] A mental disorder recognition module: using a large language model-based recognition framework and deep learning algorithms, the integrated multi-modal data is intelligently processed and classified into mental disorders;
[0012] A personalized non-pharmacological intervention module: based on the mental disorder recognition results, combined with cognitive behavioral therapy theory, a dialogue-feedback-adjustment closed-loop mechanism is constructed for personalized non-pharmacological intervention.
[0013] The application also provides a multi-modal large model recognition and intelligent intervention method for mental disorders, comprising the following steps:
[0014] S1, a multi-modal data acquisition stage, in which the system synchronously acquires electroencephalogram, electrocardiogram, electrodermal and expression multi-modal physiological data through multiple sensors, and each modality of signal represents the physiological and psychological state of the individual in different dimensions, ensuring that the state change of the subject can be fully reflected;
[0015] S2, a data preprocessing stage, in which the acquired multi-modal data is cleaned and standardized to ensure the reliability and accuracy of the data and provide effective data support for subsequent analysis;
[0016] S3, a feature extraction stage, in which key spatio-temporal features are extracted from the data of each modality, and through these features, the physiological and psychological state of the individual can be accurately reflected;
[0017] S4, a cross-modal feature fusion stage, in which a cross-modal spatio-temporal attention mechanism is used to fuse the data features of different modalities, a deep learning method is used to optimize the information collaboration between modalities, a time series modeling capability is combined to accurately capture the semantic relationship between different modalities, and the recognition accuracy and generalization ability are improved;
[0018] S5, a mental disorder recognition stage, in which a large language model (LLM)-based recognition framework is used to intelligently process the integrated multi-modal data, a deep learning algorithm is used for mental disorder classification, similarity constraints and collaborative loss functions in the model are optimized to improve recognition accuracy, and the stability and reliability of the recognition results are ensured;
[0019] S6, a personalized non-pharmacological intervention stage, in which a dialogue-feedback-adjustment closed-loop mechanism is constructed based on the mental disorder recognition results and combined with cognitive behavioral therapy (CBT) theory; the intervention strategy is adjusted in real time according to the individual's physiological feedback signals (such as electrodermal response and electrocardiogram change); based on a reward model, the effect of the current intervention strategy is evaluated according to the feedback signal, and through the combination of proximal policy optimization (PPO) algorithm, the intervention process is dynamically optimized to achieve precise personalized intervention, improve the patient's compliance and treatment effect.
[0020] Preferably, the multi-modal data acquisition stage of step S1 comprises:
[0021] S11, the system designs an intelligent interaction experiment based on a large language model, the subjects participate in the experiment by text dialogue with the large language model, the system can select single modal text data acquisition or simultaneously collect electroencephalogram, electrocardiogram, skin electricity and facial expression physiological signals, comprehensively reflect the physiological and psychological changes of the subjects in the interaction process, in the experiment process, the system designs different situations and topics, stimulates the emotional fluctuation of the subjects, so that the data is more representative and diverse;
[0022] S12, the activation degree of the frontal lobe region of the emotional disorder patient group is different from that of the normal group, and the signal-to-noise ratio is higher and the motion artifact is less. The electroencephalogram signal is collected by a multi-channel electroencephalogram device, and a reference electrode is provided. The electrocardiogram signal is collected by an electrocardiogram device simultaneously, the heart electrical activity and heart rate variability are recorded, the skin electricity signal is monitored by a skin electricity device in real time, the facial expression data is collected by a facial expression recognition system, the facial expression is converted into a quantitative index of emotional state, the emotional feedback and real-time evaluation of emotional state are provided, the system synchronously collects the text data generated by the subjects in the interaction process with the large language model, and the text content reflects the language emotion, cognitive reaction and intention of the subjects; Through the synchronous collection of these multi-modal signals, the system can comprehensively and accurately capture the multi-dimensional changes of the subjects in emotion and cognition, and provide sufficient basis for subsequent data processing, feature extraction and intervention strategy formulation.
[0023] Preferably, the data preprocessing stage of step S2 includes preprocessing of electroencephalogram signal, electrocardiogram signal, skin electricity signal, expression data and text data;
[0024] The preprocessing process of electroencephalogram signal includes band-pass filtering and artifact removal. The band-pass filtering range is 1Hz-40Hz to remove direct current drift and high frequency noise. In the artifact removal process, first, the quality of the original electroencephalogram signal is checked, and the artifact signals caused by body movement, blinking and eye movement, etc. are removed, and then the independent component analysis (ICA) method is used to further remove the residual artifacts, to ensure the purity and reliability of the signal;
[0025] The preprocessing process of electrocardiogram signal, first, high-pass filtering is used to remove baseline drift, then low-pass filtering is used to remove high-frequency noise, R-wave detection algorithm is used for heart beat period segmentation, and abnormal waveforms are removed to ensure signal quality;
[0026] The preprocessing of skin electricity signal includes denoising, outlier removal and normalization. First, low-pass filtering is used to remove high-frequency noise, then the coefficient of variation (CV) of skin electricity signal is calculated, abnormal channels and abnormal trial data are identified and removed, and finally data normalization is performed to reduce the influence of physiological differences between individuals;
[0027] The preprocessing process of expression data includes data screening and normalization. In the data screening process, the facial key point tracking algorithm is used to detect abnormal frames and remove them, and the expression data is standardized to remove individual differences to ensure the comparability of expression data of different individuals.
[0028] The preprocessing of text data includes tokenization, stop word removal, text cleaning, and vectorization. First, appropriate tokenization algorithms are used for different languages (such as Jieba for Chinese and NLTK Tokenizer for English) to segment the text into words or sub-word units. Then, stop words, punctuation marks, and low-frequency words are removed to reduce noise interference. During the text cleaning process, the case is unified, special characters are deleted, and spelling checking is performed. Finally, word vectors (Word2Vec, GloVe) or deep language models (BERT) are used for text vectorization to obtain high-dimensional semantic representations and provide structured input for subsequent analysis.
[0029] Preferably, the feature extraction stage of step S3 is divided into electroencephalogram signal, electrocardiogram signal, electrodermal signal, expression data, and text data feature extraction.
[0030] The electroencephalogram signal feature extraction includes the following steps:
[0031] Linear features: extract the root mean square and the margin factor, the specific calculation method is as follows:
[0032] ;
[0033] wherein, represents the margin factor, represents the root mean square value, represents the maximum value of the absolute value in a set of data { } ;
[0034] Nonlinear features: extract C0 complexity and Lempel-Ziv complexity (LZC); C0 complexity is used to measure the irregularity of time series, and the calculation process involves power spectrum analysis and inverse Fourier transform. LZC complexity is used to evaluate the randomness of time series, and the higher the value, the closer the signal is to a random process, and the more diverse the frequency components are.
[0035] The electrocardiogram signal feature extraction includes heart rate variability (HRV) features:
[0036] Calculate the time domain features to measure the trend of heart rate changes:
[0037] ;
[0038] wherein, represents the standard deviation of all normal sinus intervals, N represents the recordedRR the total number of intervals, representing the first interval, RR representing the average of all intervals, RR intervals;
[0039] The frequency domain features are calculated to evaluate the activities of the sympathetic and parasympathetic nervous systems.
[0040] The skin electrical signal feature extraction includes extracting the baseline conductance level (SCL), skin conductance response amplitude (SCR amplitude), skin conductance response frequency (SCR frequency), etc. of the skin electrical signal, reflecting the activation degree of the individual's autonomic nervous system.
[0041] The expression feature extraction includes extracting the key motion information of facial expression changes using the optical flow method, including facial muscle displacement vector, mouth corner motion amplitude, eyelid opening degree, etc., representing the individual's emotional expression pattern.
[0042] The text feature extraction includes extraction of multiple aspects such as lexical features, syntactic features, semantic features, and emotional features.
[0043] Preferably, the cross-modal feature fusion stage of step S4 adopts a cross-modal spatio-temporal attention mechanism (CSTAM) and combines a deep learning method to optimize the information collaboration between modalities, specifically including:
[0044] S41, dynamic time window selection and time sequence coding: first, an adaptive dynamic time window selection method based on local window is used to dynamically adjust the window size according to the sampling rate and change rate of each modality data, aligning the modality data to similar time scales, then using a time sequence coding model based on Transformer, using position encoding and time sequence information to learn the global time sequence dependence of multi-modal data, generating a unified feature representation for subsequent modality fusion;
[0045] S42, cross-modal attention alignment: a cross-modal attention mechanism is introduced to construct the alignment relationship between modalities by calculating the correlation matrix between different modalities, assuming that the attention weights between modality and modality are:
[0046]
[0047] wherein, and are the query and key representations of modality and , For feature dimension, Used to calculate weighted values;
[0048] S43. Multimodal Feature Collaborative Learning: This method performs collaborative learning of multimodal features through similarity constraints in the deep feature space. To ensure consistency of information between modalities, a collaborative loss function is used. Achieving collaborative learning of multimodal data:
[0049] ;
[0050] in, This indicates that the two modalities come from the same sample. For similarity threshold, Represents the distance metric between modalities. Indicates the first The coordinate loss value of each sample.
[0051] Preferably, step S5 performs identification and analysis based on the fused multimodal features, specifically including:
[0052] S51. The K-fold cross-validation method is used to train and optimize the parameters of the mental disorder recognition model based on aligned multimodal features. Specifically, the fused multimodal feature data and corresponding labels are randomly divided into K equal parts. K-1 parts are used as the training set in turn, and the remaining part is used as the test set. The model performance is trained and verified in turn. During the model construction process, a fully connected neural network is used for recognition. The network processes the multimodal features through the fully connected layer and makes the final prediction.
[0053] S52. Use the optimized multimodal mental disorder recognition model to predict the alignment feature data in the test set.
[0054] Preferably, step S6, based on the results of mental disorder identification and combined with cognitive behavioral therapy theory, constructs a dialogue-feedback-adjustment mechanism. This mechanism dynamically adjusts the intervention strategy by collecting real-time physiological feedback signals from individuals and considering their emotional fluctuations and cognitive states. Specifically:
[0055] Real-time monitoring of multimodal data is used to obtain feedback signals for adjusting intervention strategies. A reward model is employed to evaluate the effectiveness of the current intervention strategy. The model's reward function is as follows:
[0056] ;
[0057] in, , , and They represent at time 10:00 and 11:00 respectively Feedback values from skin conductance, electrocardiogram, electroencephalogram, and facial expressions. , , and These are preset weighting coefficients;
[0058] The proximal strategy optimization algorithm is used to dynamically optimize the adjustment of the intervention strategy, and the strategy is set. The policy parameters are adjusted by maximizing the following objective function:
[0059] ;
[0060] in, Indicates time step Expectations; It is the probability ratio of the strategy. It is an advantage estimate. It is a hyperparameter used to limit the magnitude of strategy updates. Through this optimization process, the intervention strategy is dynamically adjusted based on real-time physiological feedback, thereby achieving precise and personalized intervention.
[0061] Therefore, the present invention employs the above-mentioned multimodal large-scale model identification and intelligent intervention method and system for mental disorders, which has the following beneficial effects:
[0062] (1) A cross-modal spatiotemporal attention mechanism and dynamic time window adaptive selection technology are proposed, which can efficiently extract core features from multimodal data such as EEG, EKG, ECG, and facial expressions, and accurately model the temporal features of individuals, thereby optimizing the recognition accuracy and generalization ability of mental disorders;
[0063] (2) By introducing similarity constraints in the deep feature space, the collaborative loss of multimodal data is optimized, further improving the fusion effect of different modal information, making the recognition process more accurate and stable;
[0064] (3) In terms of non-pharmacological intervention, by constructing a closed-loop mechanism of "dialogue-feedback-adjustment" and combining it with the theory of cognitive behavioral therapy (CBT), the intervention strategy is dynamically adjusted through individual physiological feedback signals to achieve precise and personalized intervention; through the proximal strategy optimization (PPO) algorithm, the intervention process can be adaptively adjusted to improve the flexibility and efficiency of the intervention method and further enhance patient compliance and treatment effect.
[0065] (4) The reward model based on physiological feedback optimizes the intervention strategy of the large language model, making the intervention process more intelligent and personalized.
[0066] The technical solution of the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. Attached Figure Description
[0067] Figure 1The flowchart shows the multimodal large-scale model identification and intelligent intervention method for mental disorders according to the present invention.
[0068] Figure 2 This is a schematic diagram of cross-modal feature fusion according to an embodiment of the present invention;
[0069] Figure 3 This is a flowchart illustrating the identification process for mental disorders according to an embodiment of the present invention.
[0070] Figure 4 This is a flowchart illustrating personalized non-drug interventions in an embodiment of the present invention. Detailed Implementation
[0071] The following detailed description of embodiments of the invention provided in the accompanying drawings is not intended to limit the scope of the claimed invention, but merely to illustrate selected embodiments of the invention. All other embodiments obtained by those skilled in the art based on the embodiments of the invention without inventive effort are within the scope of protection of the invention.
[0072] Example
[0073] like Figure 1 As shown, a multimodal large-scale model (LLM) approach for the identification and intelligent intervention of mental disorders is presented. First, an intelligent interactive experiment based on a large language model (LLM) is designed. Subjects interact with a computer in real time, inputting text using a keyboard to engage in dialogue with the LLM, simulating everyday communication scenarios. This approach combines text-based dialogue with multimodal physiological signals to comprehensively assess the subjects' physiological and psychological responses during mental disorder identification and intelligent intervention. During this process, the system simultaneously collects multimodal data such as EEG signals, skin conductance, ECG signals, and facial expressions, and preprocesses them to remove noise and standardize signals. Next, using a cross-modal spatiotemporal attention mechanism and dynamic time window adaptive selection technology, temporal features are extracted from each modality's data, and a deep learning model is used for accurate identification of mental disorders. In the decision fusion stage, a personalized intervention strategy is generated using the LLM, and adjustments are made in real time using physiological feedback signals. To optimize the intervention effect, a proximal policy optimization (PPO) algorithm is used to adaptively adjust the intervention process, making the system more intelligent and personalized. By applying this system to clinical testing, an ideal accuracy rate for mental disorder identification was achieved, and patient treatment compliance and intervention effects were significantly improved. Specifically, it includes:
[0074] S1. In the multimodal data acquisition phase, the system simultaneously acquires multimodal physiological data such as electroencephalogram (EEG), electrocardiogram (ECG), electrical conductance of skin (EDA), and facial expressions through multiple sensors. Each modality's signal represents an individual's physiological and psychological state in different dimensions, ensuring a comprehensive reflection of the subject's state changes; specifically:
[0075] The system automatically generates various scenarios and topics according to the experimental purpose, designs different emotional stimulus scenarios, and sets challenging questions or discussions for each scenario, aiming to trigger the emotional fluctuations of the subjects and observe their physiological and psychological reactions. To ensure the diversity and representativeness of the data, the system regularly changes the scenario settings, including emotional changes (such as anxiety, depression, anger, etc.), cognitive responses (such as the difficulty and complexity of questions), and the emotional orientation of the conversation content. The emotional fluctuations of the subjects are fully reflected through their physiological reactions and language expressions.
[0076] During the experiment, the system synchronously collects various physiological signals and expression, text data, comprehensively covering the emotional and cognitive states of the subjects. Specifically, the system collects multi-modal physiological signals (including 3-channel electroencephalogram signals, 1-channel electrocardiogram signals, and 1-channel skin conductance signals) and facial expression data on more than 30 subjects. The electroencephalogram signals are collected through a multi-channel electroencephalogram device, focusing on the electroencephalogram activity in the frontal lobe region of the subjects, in order to capture the changes in brain waves related to emotional fluctuations, such as alpha waves, beta waves, and gamma waves. These data provide detailed reflections of the brain activity of the subjects during emotional changes. The electrocardiogram signals are synchronously collected through an electrocardiogram device, recording the electrical activity of the subject's heart and heart rate variability (HRV), which reveals the influence of emotional fluctuations on the autonomic nervous system. The skin conductance signals are monitored through a skin conductance device to reflect the changes in skin conductance, reflecting the physiological responses of the subjects under emotional fluctuations, with high sensitivity and effectively capturing the small physiological changes caused by emotions. In addition, facial expression data are collected through a high-precision facial expression recognition system, which can analyze the emotional performance of the subjects' faces in real time and convert them into quantitative emotional state indicators such as happiness, sadness, and anger. These synchronously collected multi-modal data comprehensively capture the emotional, cognitive responses, and physiological state changes of the subjects, providing a solid foundation for subsequent feature extraction, model training, and the development of individualized intervention strategies.
[0077] S2, in the data preprocessing stage, the collected multi-modal data are cleaned and standardized, including pseudo-trace removal and band-pass filtering for electroencephalogram signals, quality inspection and noise removal for electrocardiogram, skin conductance, and expression data, to ensure the reliability and accuracy of the data and provide effective data support for subsequent analysis; specifically:
[0078] The preprocessing process of electroencephalogram signals includes band-pass filtering and artifact removal. Band-pass filtering uses a frequency range of 1 Hz to 40 Hz to remove low-frequency DC drift and high-frequency noise. The artifact removal process eliminates artifact signals generated by body movement, blinking, and eye movement through quality checks, and uses independent component analysis (ICA) method to further eliminate residual artifacts to ensure the purity and reliability of the signals.
[0079] In the preprocessing process of electrocardiogram signals, first, high-pass filtering (cutoff frequency 0.5 Hz) is used to remove baseline drift, and then low-pass filtering (cutoff frequency 40 Hz) is used to remove high-frequency noise. The R-wave detection algorithm is used for heart beat period segmentation, and abnormal waveforms are removed to ensure the accuracy of the data and the quality of the signals.
[0080] The preprocessing of electrodermal signals includes denoising, normalization and outlier removal. First, low-pass filtering with a cutoff frequency of 5 Hz is used to remove high-frequency noise, then the coefficient of variation (CV) of electrodermal signals is calculated to identify and remove abnormal channels and abnormal trial data. Finally, data normalization is performed to reduce the influence of individual physiological differences.
[0081] The preprocessing process of expression data includes data screening and normalization, using facial key point tracking algorithm to detect abnormal frames and remove them, and standardizing the expression data to eliminate individual differences, so that the expression data of different individuals have high comparability.
[0082] The preprocessing of text data first performs word segmentation, using Jieba for Chinese and NLTKTokenizer for English, to divide the text into words or sub-word units. Then, stop words, punctuation marks and low-frequency words are removed to reduce irrelevant noise interference. Text cleaning steps include unifying case, deleting special characters and spelling checking to ensure text standardization. Finally, word vector methods (such as Word2Vec, GloVe) or deep language models (such as BERT) are used to vectorize the text to obtain high-dimensional semantic representation, providing structured input for subsequent analysis. After the above detailed preprocessing process, the quality of each modality data is significantly improved, ensuring the stability of the data and laying a solid foundation for multi-modal data fusion and subsequent analysis.
[0083] S3, feature extraction stage, extract key spatio-temporal features from data of each modality, for example, extract linear and nonlinear features from electroencephalogram signals, extract heart rate variability (HRV) features from electrocardiogram signals, extract skin conductance level (SCL) features from electrodermal signals, and extract optical flow features of facial expression changes from expression data. Through these features, the physiological and psychological state of the individual can be accurately reflected. The energy, complexity and frequency domain features are extracted from electroencephalogram and electrocardiogram signals, the change amplitude and dynamic features are extracted from skin conductance and expression data, and the key information is obtained by combining vocabulary, syntax, semantics and sentiment analysis of text data. Through multi-modal feature extraction, high-quality time sequence feature input is provided for mental disorder recognition; specifically including:
[0084] 1. Feature extraction of electroencephalogram signals
[0085] (1) Linear features
[0086] Extract the root mean square (RMS) and the coefficient of fluctuation (CF). The root mean square reflects the energy size of the signal, and the coefficient of fluctuation is calculated by the ratio of the peak value of the signal to the root mean square, which measures the fluctuation degree of the signal:
[0087] ;
[0088] Wherein, CF represents the coefficient of fluctuation, RMS represents the root mean square value, max{ |x|} represents the maximum value of the absolute value in a set of data ;
[0089] (2) Nonlinear features
[0090] Extract C0 complexity and Lempel-Ziv complexity (LZC). C0 complexity is used to measure the irregularity of time series, and the calculation process involves power spectrum analysis and inverse Fourier transform. LZC complexity is used to evaluate the randomness of time series. The higher the value, the closer the signal is to a random process, and the more diverse the frequency components are.
[0091] 2. Feature extraction of electrocardiogram signals
[0092] (1) Heart rate variability (HRV) features
[0093] Calculate time domain features (mean RR interval, , and root mean square of interbeat interval RMSSD ) to measure the trend of heart rate change:
[0094] ;
[0095] Wherein, SDNN represents the standard deviation of all normal sinus interbeat intervals,N Total number of intervals, RR Total number of intervals, Total number of intervals, Total number of intervals, RR Total number of intervals, Total number of intervals, RR Total number of intervals,
[0096] Calculate frequency domain features (low frequency power LF, high frequency power HF and their ratio LF / HF) to assess the activity of sympathetic and parasympathetic nervous system.
[0097] 3. Feature extraction of electrodermal signals
[0098] Extract baseline skin conductance level (SCL), skin conductance response amplitude (SCR amplitude), skin conductance response frequency (SCR frequency) and other features of electrodermal signals to reflect the activation degree of individual autonomic nervous system.
[0099] 4. Facial expression features
[0100] Extract key motion information of facial expression changes using optical flow method, including facial muscle displacement vector, mouth corner movement amplitude, eyelid opening degree, etc., to characterize individual emotional expression patterns.
[0101] 5. Feature extraction of text data
[0102] (1) Lexical features
[0103] Calculate word frequency (TF), inverse document frequency (IDF), TF-IDF and other indicators to measure the importance of different words.
[0104] Extract lexical richness (such as Type-Token Ratio, TTR) to evaluate the language complexity of the text.
[0105] (2) Syntactic features
[0106] Statistical sentence length, average word length, syntactic tree depth and other structural information to characterize the language style of the text.
[0107] Calculate the complexity of subject-verb-object structure to analyze the collocation pattern of sentence components.
[0108] (3) Semantic features
[0109] Use word embedding models (Word2Vec, GloVe, BERT) to map text to high-dimensional semantic space and capture the context relationship between words.
[0110] Calculate text similarity (such as cosine similarity, Jaccard similarity) to measure the relevance of different texts.
[0111] (4) Sentiment features
[0112] The sentiment polarity of the text is analyzed using an emotional dictionary (such as NRC, HOWNET, SentiWordNet), and the emotional intensity score is extracted.
[0113] Through the above feature extraction method, the physiological and psychological state of the individual in the multi-modal data can be accurately captured, providing reliable feature input for subsequent pattern recognition and emotional computing.
[0114] S4, cross-modal feature fusion stage, through cross-modal spatio-temporal attention mechanism, the data features of different modalities are fused, this stage uses deep learning method to optimize the information collaboration between modalities, combines the time series modeling ability, accurately captures the semantic relationship between different modalities, improves the recognition accuracy and generalization ability, such as Figure 2 As shown in the figure, specifically comprising:
[0115] 1, dynamic time window selection and time series encoding
[0116] In view of the inconsistency of different modal data (such as electroencephalogram, electrocardiogram, electrodermal, expression and text) in time granularity, this stage first adopts an adaptive dynamic time window selection method based on local window, dynamically adjusts the window size according to the sampling rate and change rate of each modality data, and aligns each modality data to similar time scale. Through this method, different modal data can obtain consistent granularity on the time axis. Then, a time series encoding model based on Transformer is used to learn the global time series dependence of multi-modal data using position encoding and time series information, and generate a unified feature representation for subsequent modal fusion.
[0117] 2, cross-modal attention alignment
[0118] Since different modal data may express different psychological semantic information at the same time, this stage introduces a cross-modal attention mechanism (Cross-Attention Mechanism) to build the alignment relationship between modalities by calculating the correlation matrix between different modal data. Let the attention weight between modality and modality be:
[0119]
[0120] Where, and are the query and key representations of modality and , is the feature dimension, The weighted values are used to calculate the alignment features, which are used for information transmission and fusion between modalities.
[0121] 3. Multi-modal feature collaborative learning
[0122] In this stage, multi-modal feature collaborative learning is performed through similarity constraints in the deep feature space. To ensure the consistency of information between modalities, a collaborative loss function is used The collaborative learning of multi-modal data is achieved:
[0123] ;
[0124] wherein, indicates that the two modalities come from the same sample, is a similarity threshold, denotes the distance metric between modalities. By optimizing this collaborative loss function, the features of different modalities are finally aligned in the feature space, ensuring the complementarity and consistency of information between modalities. The optimized features are fused by concatenation operation as the input for subsequent mental disorder recognition.
[0125] S5, mental disorder recognition stage, as shown in Figure 3 is shown, first, based on the fused multi-modal features (such as electroencephalogram, electrocardiogram, electrodermal, expression and text, etc.) for recognition analysis, these features are preprocessed by adaptive time alignment and time series modeling method, to ensure unified representation on similar time scale. The system uses a fully connected neural network to classify and predict each sample to reflect the individual's mental state. In the model training stage, K-fold cross-validation method is used to train and optimize the parameters of the recognition model based on the aligned multi-modal features. Specifically, the data is randomly divided into K equal parts, each time K-1 parts are used as the training set and 1 part is used as the test set, and the model is trained and verified in turn. In each fold training, a fully connected neural network is used to process the multi-modal features, and a ReLU activation function and a softmax output classification probability. By evaluating the accuracy, sensitivity and specificity of the model on the test set, the model parameters are optimized and the best performing model is selected. In the model testing stage, the optimized recognition model is used to predict the aligned feature data in the test set, calculate and generate performance indicators such as accuracy, sensitivity and specificity, and comprehensively evaluate the actual recognition effect of the model to ensure its effectiveness and accuracy in mental disorder recognition.
[0126] S6, personalized non-pharmacological intervention stage, based on the mental disorder recognition results, combined with cognitive behavioral therapy (CBT) theory, a "dialogue-feedback-adjustment" closed-loop mechanism is constructed, as shown in Figure 4The system dynamically adjusts the intervention strategy based on real-time physiological feedback signals, emotional fluctuations, and cognitive states of the individual, achieving personalized and precise intervention. The system monitors multiple modalities of data in real-time to obtain feedback signals for adjusting the intervention strategy. Skin conductance reflects emotional fluctuations, and electrocardiogram signals are used to assess emotional state and stress levels. The feedback signals at each moment are input into the system through a reward model, forming the basis for real-time intervention adjustment. The reward function is defined as:
[0127] ;
[0128] where, , , and represent the feedback values of skin conductance, electrocardiogram, electroencephalogram, and facial expression at time , , , and are preset weight coefficients reflecting the contribution of each physiological signal to the intervention effect.
[0129] To dynamically optimize the intervention strategy, the system uses the Proximal Policy Optimization (PPO) algorithm, which adjusts the strategy parameters adaptively based on the current strategy performance to maximize the objective function of the intervention effect:
[0130] ;
[0131] where, represents the expectation for time step , is the probability ratio of the strategy, is the advantage estimation, is a hyperparameter that limits the amplitude of strategy updates. Through this optimization process, the intervention strategy is dynamically adjusted based on real-time physiological feedback, achieving more precise personalized intervention.
[0132] During the intervention process, the patient engages in text dialogue with the large language model. The model generates personalized responses based on the patient's emotional and psychological changes, providing emotional support and cognitive reconstruction to help the patient gradually restore psychological balance. When the system detects an increase in skin conductance (indicating an increase in emotional fluctuations) or an increase in heart rate (indicating an increase in anxiety), the intervention strategy will automatically adjust, possibly by reducing intervention intensity, adjusting dialogue content, or changing interaction frequency to help the patient regain emotional stability. Real-time feedback mechanism is crucial, the system needs to constantly evaluate the individual's response to ensure the precision and adaptability of the intervention.
[0133] Therefore, the application adopts the above-mentioned multi-modal large model for mental disorders recognition and intelligent intervention method and system, through real-time collection of multi-modal data such as skin electricity, electrocardiogram, electroencephalogram and facial expression, combined with cognitive behavioral therapy (CBT) theory, a“dialogue-feedback-adjustment”closed loop mechanism is constructed. The system dynamically adjusts the intervention strategy according to the individual's physiological feedback through the reward model and the proximal policy optimization (PPO) algorithm, ensuring the accuracy and individualization of the intervention process. In addition, the interaction between the patient and the large language model further enhances the adaptability and effectiveness of the intervention strategy, helping the patient gradually restore psychological balance. This method not only provides efficient and individualized intervention solutions for patients with mental disorders, but also has strong real-time and operability, providing an innovative solution for future mental health intervention technology.
[0134] Finally, it should be noted that: the above examples are only used to illustrate the technical solutions of the present application and not to limit it, although the present application has been described in detail with reference to the preferred embodiments, those skilled in the art should understand: it can still modify or equivalently replace the technical solutions of the present application, and these modifications or equivalent replacements cannot make the modified technical solutions deviate from the spirit and scope of the technical solutions of the present application.
Claims
1. A multimodal large-scale model identification and intelligent intervention method for mental disorders, characterized by: Multimodal data acquisition module: Simultaneously acquires multimodal physiological data such as electroencephalogram (EEG), electrocardiogram (ECG), electrodermal conductance, and facial expressions through multiple sensors; Data preprocessing module: Cleans and standardizes the collected multimodal data; Feature extraction module: Extracts key spatiotemporal features from the data of each modality; Cross-modal feature fusion module: Utilizes a cross-modal spatiotemporal attention mechanism to fuse data features from different modalities; Mental Disorder Identification Module: Utilizes a recognition framework based on a large language model and deep learning algorithms to intelligently process and classify mental disorders from fused multimodal data; Personalized non-pharmacological intervention module: Based on the results of mental disorder identification and combined with cognitive behavioral therapy theory, a dialogue-feedback-adjustment closed-loop mechanism is constructed to carry out personalized non-pharmacological intervention. Includes the following steps: S1. In the multimodal data acquisition phase, the system simultaneously acquires multimodal physiological data such as electroencephalogram (EEG), electrocardiogram (ECG), electrodermal conductance, and facial expressions through multiple sensors. S2. In the data preprocessing stage, the collected multimodal data is cleaned and standardized. S3, Feature extraction stage: Extract key spatiotemporal features from the data of each modality; S4, Cross-modal feature fusion stage: Through cross-modal spatiotemporal attention mechanism, data features from different modalities are fused. S5, the mental disorder identification stage, adopts an identification framework based on a large language model, intelligently processes the fused multimodal data, and uses deep learning algorithms to classify mental disorders; S6. In the personalized non-pharmacological intervention stage, based on the results of mental disorder identification and combined with cognitive behavioral therapy theory, a dialogue-feedback-adjustment closed-loop mechanism is constructed. Step S6, based on the identification results of mental disorders and combined with cognitive behavioral therapy theory, constructs a dialogue-feedback-adjustment mechanism. This mechanism dynamically adjusts intervention strategies by collecting real-time physiological feedback signals from individuals and considering their emotional fluctuations and cognitive states. Specifically: Real-time monitoring of multimodal data is used to obtain feedback signals for adjusting intervention strategies. A reward model is employed to evaluate the effectiveness of the current intervention strategy. The model's reward function is as follows: ; in, , , and They represent the time intervals. Feedback values from skin conductance, electrocardiogram, electroencephalogram, and facial expressions. , , and These are preset weighting coefficients; The proximal strategy optimization algorithm is used to dynamically optimize the adjustment of the intervention strategy, and the strategy is set. The policy parameters are adjusted by maximizing the following objective function: ; in, Indicates time step Expectations; It is the probability ratio of the strategy. It is an advantage estimate. It is a hyperparameter used to limit the magnitude of strategy updates. Through this optimization process, the intervention strategy is dynamically adjusted based on real-time physiological feedback, thereby achieving precise and personalized intervention.
2. The method for multimodal large-scale model identification and intelligent intervention for mental disorders according to claim 1, characterized in that, Step S1, the multimodal data acquisition stage, includes: S11. Design an intelligent interaction experiment based on a large language model, in which subjects participate in the experiment by engaging in text-based dialogue with the large language model; S12. Collect the subject's EEG signals, ECG signals, ESC signals and facial expression data through a multi-lead EEG device, ECG device, ESC device and facial expression recognition system; and simultaneously collect the text data generated during the subject's interaction with the large language model.
3. The method for multimodal large-scale model identification and intelligent intervention for mental disorders according to claim 1, characterized in that, Step S2, the data preprocessing stage, includes the preprocessing of EEG signals, ECG signals, electrodermal signals, facial expression data, and text data. Specifically, it includes: The preprocessing of EEG signals includes bandpass filtering and artifact removal. Bandpass filtering removes low-frequency DC drift and high-frequency noise. The artifact removal process removes artifact signals through quality checks and uses independent component analysis to further remove residual artifacts. The ECG signal preprocessing process is as follows: First, the ECG signal is subjected to high-pass filtering to remove baseline drift, then low-pass filtering is used to remove high-frequency noise, the R-wave detection algorithm is used to segment the heart cycle, and abnormal waveforms are removed. Preprocessing of the electrodermal signal includes denoising, outlier removal, and normalization. First, low-pass filtering is used to remove high-frequency noise, then the coefficient of variation of the electrodermal signal is calculated, and finally the data is normalized. The preprocessing of facial expression data includes data screening and normalization; Preprocessing of text data includes word segmentation, stop word removal, text cleaning, and vectorization.
4. The method for multimodal large-scale model identification and intelligent intervention for mental disorders according to claim 1, characterized in that, Step S3, the feature extraction stage, is divided into feature extraction of electroencephalogram (EEG) signals, electrocardiogram (ECG) signals, electrodermal conductance (EDC) signals, facial expression data, and text data, specifically as follows: Electroencephalogram (EEG) signal feature extraction includes the following steps: Linear features: Extracting root mean square and gap factor: ; in, Represents the gap factor. Represents the root mean square value. Represents a set of data { The maximum absolute value in}; Nonlinear features: Extraction complexity of C0 and Lempel-Ziv; ECG signal feature extraction includes heart rate variability features: Calculate time-domain characteristics to measure the trend of heart rate changes: ; in, This represents the standard deviation of all normal sinus intervals. N Indicates a record RR The total number of intervals Indicates the first indivual RR Interval, Indicates all RR The average value of the interval; Calculate frequency domain characteristics to assess the activity of the sympathetic and parasympathetic nervous systems; Electrodermal signal feature extraction includes: extracting the baseline conductivity level, amplitude, and frequency of the electrodermal signal; Facial expression feature extraction includes: extracting key motion information of facial expression changes using optical flow; Text feature extraction includes the extraction of lexical features, syntactic features, semantic features, and sentiment features.
5. The method for multimodal large-scale model identification and intelligent intervention for mental disorders according to claim 1, characterized in that, Step S4, the cross-modal feature fusion stage, employs a cross-modal spatiotemporal attention mechanism and combines it with deep learning methods to optimize information collaboration between modalities. Specifically, it includes: S41. An adaptive dynamic time window selection method based on local windows is adopted. The window size is dynamically adjusted according to the sampling rate and change rate of each modality data to align each modality data to a similar time scale. Then, a time-series coding-based model is used to jointly model the multimodal data to generate a unified feature representation. S42. Introduce a cross-modal attention mechanism, which constructs the alignment relationship between modalities by calculating the correlation matrix between data from different modalities. Let the modalities be... and modality The attention weights between them are: ; in, and They are modal and The query and key representation, For feature dimension, Used to calculate weighted values; S43. Multimodal Feature Collaborative Learning: Using a Collaborative Loss Function Achieving collaborative learning of multimodal data: ; in, This indicates that the two modalities come from the same sample. Time indicates that the two modalities come from different samples. For similarity threshold, Represents the distance metric between modalities. Indicates the first The coordinate loss value of each sample.
6. The method for multimodal large-scale model identification and intelligent intervention for mental disorders according to claim 1, characterized in that, Step S5 performs identification and analysis based on the fused multimodal features, specifically including: S51. The K-fold cross-validation method is used to train and optimize the parameters of the mental disorder recognition model based on aligned multimodal features. Specifically, the fused multimodal feature data and corresponding labels are randomly divided into K equal parts. K-1 parts are used as the training set in turn, and the remaining part is used as the test set. The model performance is trained and verified in turn. During the model construction process, a fully connected neural network is used for recognition. The network processes the multimodal features through the fully connected layer and makes the final prediction. S52. Use the optimized multimodal mental disorder recognition model to predict the alignment feature data in the test set.
Citation Information
Patent Citations
Photoelectric multi-mode decision fusion method and system for mental disorder recognition
CN119128803A