Mental disorder-oriented multi-mode large model identification and intelligent intervention method and system
Through multimodal large-modal model identification and intelligent intervention methods and systems for mental disorders, problems such as data heterogeneity and insufficient timing modeling capabilities for mental disorder diagnosis and identification in the existing technology are solved, and accurate identification and personalized intervention are achieved, and treatment effect is improved.
Patent Information
- Application Number
- CN202510473405.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-16
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2045-04-16
AI Technical Summary
The prior art has problems such as data heterogeneity, insufficient timing modeling capabilities, data noise and redundant interference, and high computing resource requirements in the diagnosis and identification of mental disorders, resulting in deviations in identification accuracy and consistency, and a high risk of misdiagnosis and misdiagnosis.
Using multimodal large-modal model identification and intelligent intervention methods and systems for mental disorders, we combine multimodal signals of EEG, ECG, skin, facial expression analysis and text data, combined with large language models and cross-modal attention mechanisms, to achieve accurate identification of mental disorders, and dynamically optimize personalized intervention strategies.
It improves the accuracy and generalization ability of mental disorders, enhances the personalization and accuracy of non-pharmaceutical interventions, reduces the risk of misdiagnosis and missed diagnosis, and improves patient compliance and treatment effect.
Smart Images

Figure CN120015351A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of intelligent identification and intervention of mental disorders, and in particular to a multimodal large model identification and intelligent intervention method and system for mental disorders. Background Art
[0002] Mental disorders are a broad and serious category of psychological illnesses, covering depression, anxiety, bipolar disorder, schizophrenia and many other types. At present, the diagnosis and identification of mental disorders mainly rely on traditional clinical methods, including patients' self-reported symptoms, standardized psychological scales, and clinician interviews and observations. Although these methods can provide effective diagnostic evidence in some cases, they also have significant limitations. Patient self-reports are usually highly subjective and easily affected by factors such as disease condition, cognitive ability, and sociocultural background, which may lead to inaccurate information. In addition, clinicians' assessments usually rely on personal experience and skills, which leads to deviations in the accuracy and consistency of diagnosis, especially when the symptoms are complex or atypical, and the risk of misdiagnosis and missed diagnosis is high.
[0003] In recent years, with the rapid development of artificial intelligence and deep learning technology, research has gradually shifted to intelligent recognition methods based on multimodal data. Multimodal data integrates information from different physiological and behavioral signal sources, such as electroencephalograms, galvanic skin responses, facial expressions, voice, and text data. These signals can reflect the physiological and psychological state of individuals from multiple angles. By integrating these multimodal data, we can not only capture the mental state of individuals more comprehensively, but also improve the accuracy of mental disorder identification, providing a more accurate basis for early screening and intervention.
[0004] However, the application of multimodal data fusion technology in the identification of mental disorders and non-drug intervention faces many challenges, mainly reflected in the heterogeneity of data, insufficient time series modeling capabilities, data redundancy and noise interference, and excessive demand for computing resources. First, different modal data such as electroencephalogram, galvanic skin response and facial expression have significant heterogeneity, including differences in signal scale, noise characteristics and acquisition methods, which makes the fusion of these data complicated and difficult to achieve effectively. Secondly, the existing time series modeling methods are still insufficient in processing multimodal data. The identification of mental disorders usually requires accurate capture of the time series characteristics of the brain and autonomic nervous system, and the time series characteristics of different modal data are often difficult to align and uniformly model. Furthermore, since there may be noise and redundant information in the data, how to extract effective features while ensuring data quality and ensure the accuracy and interpretability of the fusion results has become a major technical problem. Finally, although multimodal data fusion has great potential, the existing methods have high demand for computing resources and strict real-time requirements, which makes it difficult to balance its efficiency and effect in practical applications. The multimodal large model recognition and intelligent intervention method and system for mental disorders proposed in the present invention not only breaks through the limitations of existing technologies, but also provides innovative solutions for early screening, accurate diagnosis and personalized intervention of mental disorders. Summary of the invention
[0005] The purpose of the present invention is to provide a multimodal large-model recognition and intelligent intervention method and system for mental disorders, aiming to achieve accurate identification of mental disorders and dynamically optimize personalized intervention strategies by integrating multimodal signals of EEG, ECG, skin electricity, facial expression analysis and text data, combining large language models and cross-modal attention mechanisms.
[0006] To achieve the above objectives, the present invention provides a multimodal large model recognition and intelligent intervention system for mental disorders, comprising: Multimodal data acquisition module: synchronously collects multimodal physiological data such as EEG, ECG, EDA and facial expressions through multiple sensors; Data preprocessing module: cleaning and standardizing the collected multimodal data; Feature extraction module: extract key spatiotemporal features from the data of each modality to prepare for subsequent analysis; Cross-modal feature fusion module: uses the cross-modal spatiotemporal attention mechanism to fuse data features of different modalities and integrate multi-source information; Mental disorder recognition module: uses a recognition framework based on a large language model and a deep learning algorithm to intelligently process the fused multimodal data and classify mental disorders; Personalized non-drug intervention module: Based on the results of mental disorder identification and combined with cognitive behavioral therapy theory, a dialogue-feedback-adjustment closed-loop mechanism is constructed to carry out personalized non-drug intervention.
[0007] The present invention also provides a multimodal large model recognition and intelligent intervention method for mental disorders, comprising the following steps: S1, multimodal data collection stage, the system synchronously collects EEG, ECG, skin electricity and facial expression multimodal physiological data through multiple sensors. The signal of each modality represents the physiological and psychological state of the individual in different dimensions, ensuring that the state changes of the subject can be fully reflected; S2, data preprocessing stage, the collected multimodal data are cleaned and standardized to ensure the reliability and accuracy of the data and provide effective data support for subsequent analysis; S3, feature extraction stage, extracting key spatiotemporal features from the data of each modality, through which the physiological and psychological states of individuals can be accurately reflected; S4, cross-modal feature fusion stage, through the cross-modal spatiotemporal attention mechanism, the data features of different modalities are fused. In this stage, deep learning methods are used to optimize the information coordination between modalities, combined with the time series modeling capabilities, to accurately capture the semantic relationship between different modalities, and improve the recognition accuracy and generalization ability; S5, mental disorder recognition stage, adopts a recognition framework based on large language model (LLM) to intelligently process the fused multimodal data, uses deep learning algorithms to classify mental disorders, and improves recognition accuracy by optimizing similarity constraints and collaborative loss functions in the model to ensure the stability and reliability of recognition results; S6, personalized non-drug intervention stage, based on the results of mental disorder identification and combined with cognitive behavioral therapy (CBT) theory, a dialogue-feedback-adjustment closed-loop mechanism is constructed; the intervention strategy is adjusted in real time according to the individual's physiological feedback signals (such as skin electrical response, ECG changes, etc.), based on the reward model, the effectiveness of the current intervention strategy is evaluated according to the feedback signal, and the intervention process is dynamically optimized by combining the proximal strategy optimization (PPO) algorithm to achieve precise personalized intervention and improve patient compliance and treatment effects.
[0008] Preferably, step S1 multimodal data collection stage includes: S11. The system designs an intelligent interaction experiment based on a large language model. The subjects participate in the experiment by having a text conversation with the large language model. The system can choose to obtain single-modal text data or synchronously collect multiple physiological signals such as EEG, ECG, skin electricity and facial expressions to fully reflect the physiological and psychological changes of the subjects during the interaction. During the experiment, the system designs different scenarios and topics to stimulate the emotional fluctuations of the subjects, making the data more representative and diverse. S12. There is a difference in the activation level of the prefrontal region between patients with emotional disorders and normal people, and the signal-to-noise ratio is higher and there are fewer motion artifacts. EEG signals are collected through a multi-lead EEG device equipped with a reference electrode. ECG signals are collected synchronously through an ECG device to record cardiac electrical activity and heart rate variability. Skin electrical signals are monitored in real time through skin electrical equipment. Facial expression data is collected through a facial expression recognition system, which converts facial expressions into quantitative indicators of emotional state, providing emotional feedback and real-time evaluation of emotional state. The system synchronously collects text data generated during the interaction between the subject and the large language model. The text content reflects the subject's language emotion, cognitive response and intention. Through the synchronous collection of these multimodal signals, the system can comprehensively and accurately capture the multidimensional changes in the subject's emotions and cognition, providing a sufficient basis for subsequent data processing, feature extraction and formulation of intervention strategies.
[0009] Preferably, the data preprocessing stage of step S2 includes preprocessing of EEG signals, ECG signals, galvanic skin signals, expression data and text data; The preprocessing of EEG signals includes bandpass filtering and artifact removal. The bandpass filtering range is 1Hz-40Hz to remove DC drift and high-frequency noise. In the artifact removal process, the original EEG signal is firstly checked for quality to remove artifact signals caused by body movement, blinking and eye movement, and then the independent component analysis (ICA) method is used to further remove residual artifacts to ensure the purity and reliability of the signal. In the ECG signal preprocessing process, high-pass filtering is first used to remove baseline drift, and then low-pass filtering is used to remove high-frequency noise. The R-wave detection algorithm is used to segment the cardiac cycle and remove abnormal waveforms to ensure signal quality. The preprocessing of skin electrical signals includes denoising, outlier removal and normalization. First, a low-pass filter is used to remove high-frequency noise, then the coefficient of variation (CV) of the skin electrical signal is calculated, abnormal channels and abnormal trial data are identified and removed, and finally data normalization is performed to reduce the impact of physiological differences between individuals. The preprocessing process of expression data includes data screening and normalization. During the data screening process, the facial key point tracking algorithm is used to detect abnormal frames and remove them, and the expression data is standardized to remove individual differences to ensure that the expression data of different individuals are comparable. The preprocessing of text data includes word segmentation, stop word removal, text cleaning and vectorization. First, appropriate word segmentation algorithms are used for different languages (such as Jieba word segmentation for Chinese and NLTK Tokenizer for English) to segment the text into words or subword units. Subsequently, stop words, punctuation marks and low-frequency words are removed to reduce noise interference. During the text cleaning process, the case is unified, special characters are deleted, and spelling is checked. Finally, word vectors (Word2Vec, GloVe) or deep language models (BERT) are used to vectorize the text to obtain high-dimensional semantic representations and provide structured input for subsequent analysis.
[0010] Preferably, the feature extraction stage in step S3 is divided into feature extraction of EEG signals, ECG signals, galvanic skin signals, expression data and text data; EEG signal feature extraction includes the following steps: Linear features: Extract the root mean square and clearance factor. The specific calculation method is as follows: ; in, represents the clearance factor, represents the root mean square value, Represents a set of data { } the maximum absolute value; Nonlinear features: Extract C0 complexity and Lempel-Ziv complexity (LZC); C0 complexity is used to measure the irregularity of the time series, and the calculation process involves power spectrum analysis and inverse Fourier transform. LZC complexity is used to evaluate the randomness of the time series. The higher the value, the closer the signal is to the random process and the more diverse the frequency components.
[0011] ECG signal feature extraction includes heart rate variability (HRV) features: Calculate time domain features and measure the trend of heart rate changes: ; in, represents the standard deviation of all normal sinus intervals, N Indicates the record RR The total number of intervals, Indicates indivual RR Interval, Indicates all RR The mean of the intervals; Frequency domain features were calculated to assess the activity of the sympathetic and parasympathetic nervous systems.
[0012] The skin electrical signal feature extraction includes: extracting the baseline conductance level (SCL), skin electrical response amplitude (SCR amplitude), skin electrical response frequency (SCR frequency), etc. of the skin electrical signal, which reflects the degree of activation of the individual's autonomic nervous system.
[0013] Expression feature extraction includes: using optical flow method to extract key motion information of facial expression changes, including facial muscle displacement vector, mouth corner movement amplitude, eyelid opening and closing degree, etc., to characterize the individual's emotional expression pattern.
[0014] Text feature extraction includes extraction of lexical features, syntactic features, semantic features, sentiment features and other aspects.
[0015] Preferably, the cross-modal feature fusion stage in step S4 adopts a cross-modal spatio-temporal attention mechanism (CSTAM) and combines a deep learning method to optimize information collaboration between modalities, specifically including: S41, Dynamic time window selection and time series coding: First, an adaptive dynamic time window selection method based on local windows is used to dynamically adjust the window size according to the sampling rate and change rate of each modal data, and align the modal data to a similar time scale. Then, a Transformer-based time series coding model is used to learn the global temporal dependency of multimodal data using position coding and time series information, and generate a unified feature representation for subsequent modal fusion. S42, Cross-modal Attention Alignment: The cross-modal attention mechanism is introduced to construct the alignment relationship between modalities by calculating the correlation matrix between different modal data. and modal The attention weight between is:
[0016] in, and The modal and The query and key representation of is the feature dimension, Used to calculate weighted values; S43. Collaborative learning of multimodal features: Collaborative learning of multimodal features is performed through similarity constraints in the deep feature space. In order to ensure the consistency of information between modalities, a collaborative loss function is used. Enable collaborative learning of multimodal data: ; in, When , it means that the two modalities come from the same sample. is the similarity threshold, represents the distance measure between modes, Indicates The coordinate loss value of each sample.
[0017] Preferably, step S5 performs recognition analysis based on the fused multimodal features, specifically including: S51. The K-fold cross-validation method is used to train and optimize the parameters of the mental disorder recognition model based on aligned multimodal features. Specifically, the fused multimodal feature data and the corresponding labels are randomly divided into K equal parts, and K-1 parts are used as training sets and the remaining 1 part is used as a test set in turn. The model performance is trained and verified in turn. During the model construction process, a fully connected neural network is used for recognition. The network processes the multimodal features through the fully connected layer and makes the final prediction. S52. Use the optimized multimodal mental disorder recognition model to predict the aligned feature data in the test set.
[0018] Preferably, step S6 is based on the results of mental disorder identification, combined with cognitive behavioral therapy theory, to construct a dialogue-feedback-adjustment mechanism, by real-time collection of individual physiological feedback signals, combined with the individual's emotional fluctuations and cognitive state, to dynamically adjust the intervention strategy, specifically: Monitor multimodal data in real time, obtain feedback signals to adjust intervention strategies, and use reward models to evaluate the effectiveness of current intervention strategies. The model reward function is as follows: ; in, , , and Respectively represent the time Feedback values of skin electricity, electrocardiogram, electroencephalogram and facial expression, , , and is the preset weight coefficient; Use the proximal strategy optimization algorithm to dynamically optimize the adjustment of intervention strategies and set strategies , the policy parameters are adjusted by maximizing the following objective function: ; in, Represents the time step expectations; is the probability ratio of the strategy, is the advantage estimate, is a hyperparameter used to limit the amplitude of strategy updates. Through this optimization process, the intervention strategy is dynamically adjusted according to real-time physiological feedback, thereby achieving refined personalized intervention.
[0019] Therefore, the present invention adopts the above-mentioned multimodal large model recognition and intelligent intervention method and system for mental disorders, which has the following beneficial effects: (1) A cross-modal spatiotemporal attention mechanism and dynamic time window adaptive selection technology are proposed, which can efficiently extract the core features of multimodal data such as EEG, skin electricity, ECG, and facial expressions, and accurately model the temporal characteristics of individuals, thereby optimizing the recognition accuracy and generalization ability of mental disorders; (2) The similarity constraint of the deep feature space is introduced to optimize the collaborative loss of multimodal data, further improving the fusion effect of different modal information and making the recognition process more accurate and stable. (3) In terms of non-drug intervention, by building a closed-loop mechanism of "dialogue-feedback-adjustment" and combining cognitive behavioral therapy (CBT) theory, the intervention strategy can be dynamically adjusted through individual physiological feedback signals to achieve precise personalized intervention; through the proximal strategy optimization (PPO) algorithm, the intervention process can be adaptively adjusted to improve the flexibility and efficiency of the intervention method, further improving patient compliance and treatment effect; (4) The reward model based on physiological feedback optimizes the intervention strategy of the large language model, making the intervention process more intelligent and personalized.
[0020] The technical solution of the present invention is further described in detail below through the accompanying drawings and embodiments. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] Figure 1 A flowchart of the multimodal large model identification and intelligent intervention method for mental disorders of the present invention; Figure 2 Schematic diagram of cross-modal feature fusion according to an embodiment of the present invention; Figure 3 A flowchart of mental disorder identification according to an embodiment of the present invention; Figure 4 This is a flow chart of personalized non-drug intervention according to an embodiment of the present invention. DETAILED DESCRIPTION
[0022] The following detailed description of the embodiments of the present invention provided in the accompanying drawings is not intended to limit the scope of the invention claimed for protection, but merely represents selected embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0023] Example like Figure 1 As shown in the figure, a multimodal large model identification and intelligent intervention method for mental disorders is first designed. First, by designing an intelligent interactive experiment based on a large language model, the subjects interact with the computer in real time, input text using the keyboard, and have a conversation with the large language model to simulate daily communication scenarios. The method combines text conversation with multimodal physiological signals to comprehensively evaluate the physiological and psychological reactions of the subjects in the process of mental disorder identification and intelligent intervention. In this process, the system synchronously collects multimodal data such as EEG signals, skin electricity, ECG signals and facial expressions, and preprocesses them to remove noise and standardize signals. Then, the cross-modal spatiotemporal attention mechanism and dynamic time window adaptive selection technology are used to extract time series features from each modal data, and the deep learning model is used to accurately identify mental disorders. In the decision fusion stage, a personalized intervention strategy is generated in combination with a large language model (LLM), and physiological feedback signals are used for real-time adjustment. In order to optimize the intervention effect, the proximal policy optimization (PPO) algorithm is used to adaptively adjust the intervention process, making the system more intelligent and personalized. By applying the system to clinical tests, an ideal accuracy rate for mental disorder identification is obtained, and the treatment compliance and intervention effect of patients are significantly improved. Specifically include: S1. In the multimodal data collection phase, the system uses multiple sensors to synchronously collect multimodal physiological data such as EEG, ECG, EDA, and facial expressions. The signal of each modality represents the physiological and psychological state of the individual in different dimensions, ensuring that the state changes of the subject can be fully reflected; specifically: The system automatically generates a variety of scenarios and topics based on the purpose of the experiment, designs different emotional stimulation scenarios, and sets challenging questions or discussions for each scenario, aiming to trigger the subjects' emotional fluctuations and observe their physiological and psychological reactions. To ensure the diversity and representativeness of the data, the system will regularly change the scenario settings, including emotional changes (such as anxiety, depression, anger, etc.), cognitive reactions (such as the difficulty and complexity of the questions, etc.), and the emotional orientation of the conversation content. The emotional fluctuations of the subjects are fully reflected through their physiological reactions and language expressions.
[0024] During the experiment, the system simultaneously collected a variety of physiological signals, expressions, and text data, fully covering the emotional and cognitive states of the subjects. Specifically, the system collected multimodal physiological signals (including 3-lead EEG signals, 1-lead ECG signal, and 1-lead skin electrical signal) and facial expression data from more than 30 subjects. EEG signals were collected through multi-lead EEG equipment, with special attention paid to the EEG activity in the prefrontal lobe area of the subjects, in order to capture EEG changes related to emotional fluctuations, such as Wave, wave and Waves. These data provide a detailed reflection of the brain activity of the subjects during emotional changes. ECG signals are synchronously collected through ECG equipment, recording the subjects' cardiac electrical activity and heart rate variability (HRV), which are used to reveal the impact of emotional fluctuations on the autonomic nervous system. Skin electrical signals monitor changes in skin conductance through skin electrical equipment, reflecting the physiological reactions of the subjects under emotional fluctuations. They have high sensitivity and can effectively capture subtle physiological changes caused by emotions. In addition, facial expression data are collected through a high-precision facial expression recognition system. The system can analyze the emotional expressions on the subjects' faces in real time and convert them into quantitative emotional state indicators, such as happiness, sadness, anger, etc. These synchronously collected multimodal data comprehensively capture the subjects' emotions, cognitive reactions and physiological state changes, providing a solid foundation for subsequent feature extraction, model training and the formulation of personalized intervention strategies.
[0025] S2. In the data preprocessing stage, the collected multimodal data are cleaned and standardized, including artifact removal and bandpass filtering of EEG signals, quality inspection and noise removal of ECG, skin electricity and facial expression data, to ensure the reliability and accuracy of the data and provide effective data support for subsequent analysis; specifically: The preprocessing of EEG signals includes bandpass filtering and artifact removal. Bandpass filtering uses a frequency range of 1Hz to 40Hz to remove low-frequency DC drift and high-frequency noise. The artifact removal process removes artifact signals caused by body movement, blinking, and eye movement through quality inspection, and further removes residual artifacts using the independent component analysis (ICA) method to ensure the purity and reliability of the signal.
[0026] During the preprocessing of ECG signals, high-pass filtering (cutoff frequency 0.5Hz) is first used to remove baseline drift, and low-pass filtering (cutoff frequency 40Hz) is then used to remove high-frequency noise. The cardiac cycle segmentation uses the R wave detection algorithm, and abnormal waveforms are removed to ensure data accuracy and signal quality.
[0027] The preprocessing of skin electrical signals includes denoising, normalization, and outlier removal. First, a low-pass filter with a cutoff frequency of 5 Hz is used to remove high-frequency noise. Then, the coefficient of variation (CV) of the skin electrical signal is calculated to identify and remove abnormal channels and abnormal trial data. Finally, data normalization is performed to reduce the impact of physiological differences between individuals.
[0028] The preprocessing process of expression data includes data screening and normalization. The facial key point tracking algorithm is used to detect abnormal frames and eliminate them. The expression data is standardized to eliminate individual differences, making the expression data of different individuals highly comparable.
[0029] The preprocessing of text data first involves word segmentation. Jieba word segmentation is used for Chinese, and NLTKTokenizer is used for English to divide the text into words or subword units. Next, stop words, punctuation marks, and low-frequency words are removed to reduce irrelevant noise interference. The text cleaning steps include unifying uppercase and lowercase letters, deleting special characters, and performing spelling checks to ensure the standardization of the text. Finally, the text is vectorized using word vector methods (such as Word2Vec, GloVe) or deep language models (such as BERT) to obtain high-dimensional semantic representations and provide structured input for subsequent analysis. After the above detailed preprocessing process, the quality of each modality data has been significantly improved, ensuring the stability of the data and laying a solid foundation for multimodal data fusion and subsequent analysis.
[0030] S3, feature extraction stage, extract key spatiotemporal features from each modality of data, for example, extract linear and nonlinear features from EEG signals, extract heart rate variability (HRV) features from ECG signals, extract skin electrodermal activity (SCL) features from skin electrodermal signals, and extract optical flow features of facial expression changes from expression data. These features can accurately reflect the physiological and psychological state of an individual. EEG and ECG signals extract energy, complexity, and frequency domain features, skin electrodermal and expression data extract change amplitude and dynamic features, and text data combines vocabulary, syntax, semantics, and sentiment analysis to obtain key information. Through multimodal feature extraction, high-quality temporal feature input is provided for mental disorder identification; specifically, it includes: 1. Feature extraction of EEG signals (1) Linear characteristics Extract the root mean square (RMS) and clearance factor (CF). The root mean square reflects the energy of the signal, and the clearance factor is calculated by the ratio of the signal peak to the root mean square, which measures the degree of signal fluctuation: ; in, represents the clearance factor, represents the root mean square value, Represents a set of data { } the maximum absolute value; (2) Nonlinear characteristics Extract C0 complexity and Lempel-Ziv complexity (LZC). C0 complexity is used to measure the irregularity of the time series, and the calculation process involves power spectrum analysis and inverse Fourier transform. LZC complexity is used to evaluate the randomness of the time series. The higher the value, the closer the signal is to the random process and the more diverse the frequency components are.
[0031] 2. Feature extraction of ECG signals (1) Heart rate variability (HRV) characteristics Calculate time domain features (mean RR Interval, , RMS of heart rate interval RMSSD ), measure the trend of heart rate changes: ; in, represents the standard deviation of all normal sinus intervals, N Indicates the record RR The total number of intervals, Indicates indivual RR Interval, Indicates all RR The mean of the intervals; Frequency domain features (low frequency power LF, high frequency power HF and their ratio LF / HF) were calculated to evaluate the activities of the sympathetic and parasympathetic nervous systems.
[0032] 3. Feature extraction of skin electrical signals The baseline conductance level (SCL), skin galvanic response amplitude (SCR amplitude), skin galvanic response frequency (SCR frequency) and other data of the skin electrical signal are extracted to reflect the degree of activation of the individual's autonomic nervous system.
[0033] 4. Expression characteristics The optical flow method is used to extract key motion information of facial expression changes, including facial muscle displacement vector, mouth corner movement amplitude, eyelid opening and closing degree, etc., to characterize the individual's emotional expression pattern.
[0034] 5. Feature extraction of text data (1) Lexical features Calculate indicators such as term frequency (TF), inverse document frequency (IDF), TF-IDF, etc. to measure the importance of different words.
[0035] Extract vocabulary richness (such as Type-Token Ratio, TTR) to evaluate the language complexity of the text.
[0036] (2) Syntactic features Structural information such as sentence length, average word length, and syntax tree depth are collected to characterize the language style of the text.
[0037] Calculate the complexity of the subject-verb-object structure and analyze the collocation patterns of sentence components.
[0038] (3) Semantic features Word embedding models (Word2Vec, GloVe, BERT) are used to map text into a high-dimensional semantic space to capture the contextual relationship between words.
[0039] Calculate text similarity (such as cosine similarity, Jaccard similarity) to measure the relevance of different texts.
[0040] (4) Emotional characteristics Sentiment lexicons (such as NRC, HOWNET, and SentiWordNet) are used to analyze the sentiment polarity of the text and extract the sentiment intensity score.
[0041] Through the above feature extraction method, the physiological and psychological states of individuals in multimodal data can be accurately captured, providing reliable feature input for subsequent pattern recognition and emotion calculation.
[0042] S4, cross-modal feature fusion stage, through the cross-modal spatiotemporal attention mechanism, the data features of different modalities are fused. In this stage, deep learning methods are used to optimize the information coordination between modalities, combined with the time series modeling ability, to accurately capture the semantic relationship between different modalities, and improve the recognition accuracy and generalization ability, such as Figure 2 As shown, specifically including: 1. Dynamic time window selection and timing encoding In order to solve the problem of inconsistent temporal granularity of different modal data (such as EEG, ECG, skin electricity, facial expressions and text), this stage first adopts an adaptive dynamic time window selection method based on local windows, dynamically adjusts the window size according to the sampling rate and change rate of each modal data, and aligns each modal data to a similar time scale. Through this method, different modal data can obtain consistent granularity on the time axis. Then, a Transformer-based temporal coding model is used to learn the global temporal dependency of multimodal data using position coding and temporal information, and generate a unified feature representation for subsequent modal fusion.
[0043] 2. Cross-modal attention alignment Since different modal data may express different psychological semantic information at the same time, the cross-attention mechanism is introduced in this stage to construct the alignment relationship between modalities by calculating the correlation matrix between different modal data. and modal The attention weight between is:
[0044] in, and The modal and The query and key representation of is the feature dimension, Used to calculate weighted values; the aligned features obtained by cross-modal attention weighted calculation will be used for information transfer and fusion between modalities.
[0045] 3. Collaborative learning of multimodal features In this stage, the multimodal features are collaboratively learned through similarity constraints in the deep feature space. In order to ensure the consistency of information between modalities, the collaborative loss function is used. Enable collaborative learning of multimodal data: ; in, When , it means that the two modalities come from the same sample. is the similarity threshold, Represents the distance metric between modalities. By optimizing this collaborative loss function, the features of different modalities are finally aligned in the feature space to ensure the complementarity and consistency of information between modalities. The optimized features are fused through a splicing operation and used as the input for subsequent mental disorder recognition.
[0046] S5, mental disorder identification stage, such as Figure 3 As shown in the figure, firstly, recognition analysis is performed based on the fused multimodal features (such as EEG, ECG, skin electricity, expression and text, etc.). These features are preprocessed by adaptive time alignment and time series modeling methods to ensure uniform representation on a similar time scale. The system uses a fully connected neural network to classify and predict each sample to reflect the individual's mental state. In the model training stage, the K-fold cross-validation method is used to train and optimize the parameters of the recognition model based on aligned multimodal features. Specifically, the data is randomly divided into K equal parts, and K-1 parts are used as training sets and 1 part is used as test sets each time, and the model performance is trained and verified in turn. In each fold training, the multimodal features are processed using a fully connected neural network, and the classification probability is output through the ReLU activation function and softmax. By evaluating the accuracy, sensitivity and specificity of the model on the test set, the model parameters are optimized and the best performing model is selected. During the model testing phase, the optimized recognition model is used to predict the aligned feature data in the test set, and performance indicators such as accuracy, sensitivity, and specificity are calculated and generated to comprehensively evaluate the actual recognition effect of the model to ensure its effectiveness and accuracy in identifying mental disorders.
[0047] S6, personalized non-drug intervention stage, based on the results of mental disorder identification, combined with cognitive behavioral therapy (CBT) theory, to build a "dialogue-feedback-adjustment" closed-loop mechanism, such as Figure 4As shown. In this stage, the individual's physiological feedback signals are collected in real time, and the intervention strategy is dynamically adjusted in combination with the individual's emotional fluctuations and cognitive state to achieve personalized and precise intervention. The system monitors multimodal data in real time and obtains feedback signals to adjust the intervention strategy. The electrical skin response reflects emotional fluctuations, and the ECG signal is used to assess emotional state and stress level. The feedback signal at each moment is input into the system through the reward model to form a real-time basis for intervention adjustment. The reward function is defined as: ; in, , , and Respectively represent the time Feedback values of skin electricity, electrocardiogram, electroencephalogram and facial expression, , , and It is the preset weight coefficient, which reflects the contribution of each physiological signal to the intervention effect; To dynamically optimize the intervention strategy, the system uses the Proximal Policy Optimization (PPO) algorithm, which adaptively adjusts the strategy parameters according to the performance of the current strategy to maximize the objective function of the intervention effect: ; in, Represents the time step expectations; is the probability ratio of the strategy, is the advantage estimate, is a hyperparameter used to limit the amplitude of strategy updates. Through this optimization process, the intervention strategy is dynamically adjusted according to real-time physiological feedback, thereby achieving more refined personalized intervention.
[0048] During the intervention process, the patient has a text conversation with the large language model, which generates personalized responses based on the patient's emotions and psychological changes, provides emotional support and cognitive reconstruction, and helps the patient gradually restore psychological balance. When the system detects an increase in skin electrical response (indicating increased emotional fluctuations) or an increase in heart rate (indicating increased anxiety), the intervention strategy will automatically adjust, which may help the patient restore emotional stability by reducing the intensity of the intervention, adjusting the content of the conversation, or changing the frequency of interaction. The real-time feedback mechanism is crucial, and the system needs to continuously evaluate individual responses to ensure the accuracy and adaptability of the intervention.
[0049] Therefore, the present invention adopts the above-mentioned multimodal large model recognition and intelligent intervention method and system for mental disorders, and constructs a "dialogue-feedback-adjustment" closed-loop mechanism by real-time collection of multimodal data such as skin electricity, electrocardiogram, electroencephalogram, facial expressions, etc., combined with cognitive behavioral therapy (CBT) theory. The system dynamically adjusts the intervention strategy according to the individual's physiological feedback through the reward model and the proximal policy optimization (PPO) algorithm to ensure the accuracy and personalization of the intervention process. In addition, the interaction between the patient and the large language model further enhances the adaptability and effectiveness of the intervention strategy, helping the patient to gradually restore psychological balance. This method can not only provide patients with mental disorders with efficient and personalized intervention plans, but also has strong real-time and operability, and provides innovative solutions for intervention technologies in the field of mental health in the future.
[0050] Finally, it should be noted that the above embodiments are only used to illustrate the technical solution of the present invention rather than to limit it. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that they can still modify or replace the technical solution of the present invention with equivalents, and these modifications or equivalent replacements cannot cause the modified technical solution to deviate from the spirit and scope of the technical solution of the present invention.
Claims
1. A multimodal large model recognition and intelligent intervention system for mental disorders, characterized by: include: Multimodal data acquisition module: synchronously collects multimodal physiological data such as EEG, ECG, skin electricity and facial expressions through multiple sensors; Data preprocessing module: cleans and standardizes the collected multimodal data; Feature extraction module: extract key spatiotemporal features from the data of each modality; Cross-modal feature fusion module: uses the cross-modal spatiotemporal attention mechanism to fuse data features of different modalities; Mental disorder recognition module: uses a recognition framework based on a large language model and a deep learning algorithm to intelligently process the fused multimodal data and classify mental disorders; Personalized non-drug intervention module: Based on the results of mental disorder identification and combined with cognitive behavioral therapy theory, a dialogue-feedback-adjustment closed-loop mechanism is constructed to carry out personalized non-drug intervention.
2. A multimodal large model recognition and intelligent intervention method for mental disorders, applied to the multimodal large model recognition and intelligent intervention system for mental disorders described in claim 1, characterized in that: The following steps are involved: S1, multimodal data collection stage, the system synchronously collects multimodal physiological data of EEG, ECG, skin electricity and facial expression through multiple sensors; S2, data preprocessing stage, cleaning and standardizing the collected multimodal data; S3, feature extraction stage, extracting key spatiotemporal features from the data of each modality; S4, cross-modal feature fusion stage, through the cross-modal spatiotemporal attention mechanism, the data features of different modalities are fused; S5, mental disorder identification stage, using a recognition framework based on a large language model to intelligently process the fused multimodal data and use deep learning algorithms to classify mental disorders; S6, personalized non-drug intervention stage, based on the results of mental disorder identification and combined with cognitive behavioral therapy theory, a dialogue-feedback-adjustment closed-loop mechanism is constructed.
3. The multimodal large model recognition and intelligent intervention method for mental disorders according to claim 2, characterized in that: Step S1: multimodal data collection phase includes: S11. Design an intelligent interaction experiment based on a large language model, where the subjects participate in the experiment by having a text conversation with the large language model; S12. The facial expression recognition system uses a multi-lead EEG device, an ECG device, and an galvanic skin device to collect the subject's EEG signal, ECG signal, galvanic skin signal, and facial expression data respectively; and simultaneously collects the text data generated during the interaction between the subject and the large language model.
4. The multimodal large model recognition and intelligent intervention method for mental disorders according to claim 2, characterized in that: Step S2: Data preprocessing includes the preprocessing of EEG signals, ECG signals, galvanic skin signals, expression data and text data, specifically including: The preprocessing process of EEG signals includes bandpass filtering and artifact removal. Bandpass filtering removes low-frequency DC drift and high-frequency noise. The artifact removal process removes artifact signals through quality inspection and uses independent component analysis to further remove residual artifacts. The ECG signal preprocessing process is as follows: first, the ECG signal is high-pass filtered to remove baseline drift, then low-pass filtered to remove high-frequency noise, the R-wave detection algorithm is used to segment the cardiac cycle, and abnormal waveforms are eliminated; The preprocessing of skin electrical signals includes denoising, outlier removal and normalization. First, a low-pass filter is used to remove high-frequency noise, then the coefficient of variation of the skin electrical signal is calculated, and finally the data is normalized. The preprocessing of expression data includes data screening and normalization; The preprocessing of text data includes word segmentation, removal of stop words, text cleaning and vectorization.
5. The multimodal large model recognition and intelligent intervention method for mental disorders according to claim 2, characterized in that: Step S3: The feature extraction stage is divided into feature extraction of EEG signals, ECG signals, skin electrical signals, expression data and text data, specifically: EEG signal feature extraction includes the following steps: Linear Features: Extract RMS and Clearance Factors: Where CF is the clearance factor, x rms represents the root mean square value, max(|x i |) represents a set of data {x i } the maximum absolute value; Nonlinear features: Extract C0 complexity and Lempel-Ziv complexity; ECG signal feature extraction includes heart rate variability features: Calculate time domain features and measure the trend of heart rate changes: Among them, SDNN represents the standard deviation of all normal sinus intervals, N represents the total number of recorded RR intervals, and RR i represents the i-th RR interval, represents the average value of all RR intervals; Calculate frequency domain features to assess the activity of the sympathetic and parasympathetic nervous systems; Skin electrical signal feature extraction includes: extracting the baseline conductance level, skin electrical response amplitude and skin electrical response frequency of the skin electrical signal; Expression feature extraction includes: using optical flow method to extract key motion information of facial expression changes; Text feature extraction includes: extraction of lexical features, syntactic features, semantic features and sentiment features.
6. The multimodal large model recognition and intelligent intervention method for mental disorders according to claim 2, characterized in that: Step S4, the cross-modal feature fusion stage, adopts a cross-modal spatiotemporal attention mechanism and combines deep learning methods to optimize information coordination between modalities, including: S41. Adopting an adaptive dynamic time window selection method based on local windows, dynamically adjusting the window size according to the sampling rate and change rate of each modal data, aligning each modal data to a similar time scale, and then using a model based on temporal coding to jointly model the multimodal data and generate a unified feature representation; S42, introduce the cross-modal attention mechanism, and construct the alignment relationship between modalities by calculating the correlation matrix between different modal data. Suppose modality m i and mode m j The attention weight between is: Among them, Q i and K j They are mode m i and m j The query and key representation of dim is the feature dimension, and softmax is used to calculate the weighted value; S43. Collaborative learning of multimodal features: using collaborative loss function L coord Enable collaborative learning of multimodal data: Among them, y = 1 means that the two modalities come from the same sample, y = 0 means that the two modalities come from different samples, θ is the similarity threshold, d is the distance measure between the modalities, Represents the coordinate loss value of the i-th sample.
7. The multimodal large model recognition and intelligent intervention method for mental disorders according to claim 2, characterized in that: Step S5 performs recognition analysis based on the fused multimodal features, specifically including: S51. The K-fold cross-validation method is used to train and optimize the parameters of the mental disorder recognition model based on aligned multimodal features. Specifically, the fused multimodal feature data and the corresponding labels are randomly divided into K equal parts, and K-1 parts are used as training sets and the remaining 1 part is used as a test set in turn. The model performance is trained and verified in turn. During the model construction process, a fully connected neural network is used for recognition. The network processes the multimodal features through the fully connected layer and makes the final prediction. S52. Use the optimized multimodal mental disorder recognition model to predict the aligned feature data in the test set.
8. The multimodal large model recognition and intelligent intervention method for mental disorders according to claim 2, characterized in that: Step S6 is based on the results of mental disorder identification and combined with cognitive behavioral therapy theory to build a dialogue-feedback-adjustment mechanism. By collecting individual physiological feedback signals in real time and combining the individual's emotional fluctuations and cognitive state, the intervention strategy is dynamically adjusted. Specifically: Monitor multimodal data in real time, obtain feedback signals to adjust intervention strategies, and use reward models to evaluate the effectiveness of current intervention strategies. The model reward function is as follows: R t =α·EDA t +β·ECG t +γ·EEG t +d·Facial Expression t ; Among them, EDA t , ECG t , EEG t and Facial Expression t Respectively represent the feedback values of skin electricity, electrocardiogram, brain electricity and facial expression at time t, α, β, γ and δ are the preset weight coefficients; The proximal policy optimization algorithm is used to dynamically optimize the adjustment of the intervention strategy. The strategy π is set and the policy parameters are adjusted by maximizing the following objective function: in, represents the expectation at time step t; r t (θ) is the probability ratio of the strategy, is the advantage estimate, ∈ is a hyperparameter used to limit the amplitude of strategy update. Through this optimization process, the intervention strategy is dynamically adjusted according to real-time physiological feedback, thereby achieving refined personalized intervention.
Citation Information
Patent Citations
Photoelectric multi-mode decision fusion method and system for mental disorder recognition
CN119128803A
Cited By
Nerve development disorder behavior intervention system based on large language model
CN120183677A
Early warning method and system for physiological stress of onboard personnel
CN120360521A
A method and system for early warning of physiological stress among onboard personnel
CN120360521B
Self-injury behavior early warning and intervention method and system based on multi-modal physiological data and AI
CN120565057A
Medical data processing method and system for reversible adhesion of hydrogel brain electrode
CN121237449A