Multi-modal mental health detection method based on thinking chain prompt and feature enhancement

By employing a multimodal mental health detection method based on thought chain cues and feature enhancement, this study addresses the problem of insufficient multimodal information fusion in social media data. It achieves deep capture of the deep semantic relationships between images and text, thereby improving the accuracy and robustness of mental health detection.

CN120878080APending Publication Date: 2025-10-31DALIAN UNIV OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510878977.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-27
Publication Date
2025-10-31

AI Technical Summary

Technical Problem

In existing technologies, mental illness detection models based on social media data suffer from insufficient multimodal information fusion and difficulty in recognizing implicit expressions, making it difficult to deeply capture the deep semantic relationship between images and text, resulting in decreased detection accuracy.

Method used

This paper adopts a multimodal mental health detection method based on thought chain cues and feature enhancement. It extracts semantic information through a multimodal large model, processes image and text data using a visual encoder and a text encoder, and combines thought chain cues to deeply fuse deep semantic information from different modalities to accurately judge mental health status.

Benefits of technology

It significantly improves the model's detection accuracy and robustness in complex multimodal scenarios, enabling more accurate identification of users' mental health status and providing technical support for early psychological intervention and public mental health monitoring.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120878080A_ABST
    Figure CN120878080A_ABST
Patent Text Reader

Abstract

According to the multi-modal mental health detection method based on thinking chain prompt and feature enhancement, deep semantic information including emotional tendency, symptom expression, hidden metaphor and the like contained in images and texts is deeply mined through thinking chain prompt by using a large model prompt module, so that the detection accuracy is improved; therefore, the analysis capability of the model on the implicit content in the user expression is improved. Besides, information obtained through a large model prompt module is fused by dominant text and image features, so that single-mode key information can be ensured not to be diluted, implicit association between the text and the image can be captured, and information of different modes can be effectively integrated. The method is superior to a current multi-modal large model in identification precision and the like, has high accuracy and robustness, is suitable for various scenes such as early psychological intervention and clinical medical assistance, can effectively monitor the psychological health state of the user, provides technical support for public psychological health detection, and has a good application prospect.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of natural language processing, and in particular to a multimodal mental health detection method based on thought chain cues and feature enhancement. Background Technology

[0002] Mental illness refers to a persistent functional disorder in an individual's cognition, emotion, will, or behavior, manifested as abnormal states such as depressed mood, anxiety, and paranoid thinking. It not only seriously affects an individual's social adaptability and quality of life, but may also induce extreme behaviors such as self-harm and suicide, posing a potential threat to both the individual and society.

[0003] In recent years, with the popularization of social media, the content posted by users on these platforms has become a unique window for researchers to explore users' psychological states. Unlike traditional clinical diagnosis, due to the anonymity of social media platforms, the content posted by users is closer to their true feelings, and analyzing this data can help researchers more effectively understand users' inner world.

[0004] Currently, the detection of mental illnesses based on social media data has attracted the attention of researchers, and many researchers have proposed methods such as text-based sentiment analysis and semantic mining. However, traditional single-modal detection techniques have limitations: relying solely on text analysis makes it difficult to fully capture the overall emotional tendency expressed by users; and existing multimodal detection models, when processing multimodal data such as images and text, often ignore the deep semantic connections between different modalities. Furthermore, simply splicing and fusing features from different modalities for detection is prone to semantic drift due to differences in information dimensionality, leading to the loss or misinterpretation of key psychological signals. This shallow fusion approach also makes it difficult for the model to delve into deeper semantics, resulting in decreased accuracy in model predictions.

[0005] Therefore, there is an urgent need to build a new model that can deeply capture the semantic relationships between multimodal data such as text and images. Through cross-modal semantic alignment and deep fusion technology, the deep semantic information hidden in them can be analyzed, thereby accurately identifying signals of mental illness and providing more effective technical support for early psychological intervention and public mental health monitoring. Summary of the Invention

[0006] To overcome the shortcomings of existing technologies, this invention provides a multimodal mental health detection method based on thought chain cues and feature enhancement. This method aims to make accurate judgments on the user's current mental health status by making full use of the deep semantic information contained in both image and text modalities.

[0007] The technical solution of this invention is as follows: A multimodal mental health detection method based on thought chain cues and feature enhancement includes the following steps: S1, Obtain a sample dataset for mental health testing, wherein each sample in the dataset includes text data and image data corresponding to the text; S2 invokes a multimodal large model and uses the mind chain prompting engineering to assist the multimodal large model in extracting semantic information, thereby enhancing the model's ability to extract deep semantic information contained in different modalities. Specifically, this includes: Using the first prompt statement, extract the emotional information from the text data; Using prompt statement two, obtain symptom information from the PHQ-9 scale that matches the text data; Using the third prompt statement, based on the emotional polarity information, the symptom information, the text data, and the image data corresponding to the text, metaphorical information is obtained; S3, use a visual encoder to process the image data to obtain an image embedding vector; The text data, sentiment information, symptom information, and metaphorical information are processed using a text encoder to obtain a text embedding vector; S4, map the image embedding vector to the same feature space as the text modality to obtain the mapped image encoding; The metaphorical part of the text embedding vector is fused with the mapped image encoding to obtain an image-dominated feature representation; The sentiment and symptom portions of the text embedding vector are fused with the text portion encoding of the text embedding vector to obtain a text-dominant feature representation; The encoding of the text embedding vector is fused with the encoding of the mapped image to obtain a hybrid feature representation; S5, classify the image-dominated feature representation, text-dominated feature representation and mixed feature representation to obtain the corresponding psychological state classification results, and finally make a final judgment on the user's psychological health status based on the voting of the three classification results.

[0008] Furthermore, the multimodal large model described in S2 is LLaVA-v1.5-7B.

[0009] Furthermore, the prompt statement one expresses the message "Please determine the sentiment polarity of the overall content according to the text content, and select one from positive, neutral, and negative." Specifically, the prompt statement sent is: "Please determine the sentiment polarity of the overall content according to the content in the text, and select one from positive, neutral, and negative.", thus generating a judgment on the overall sentiment tendency.

[0010] Furthermore, the second prompt statement expresses the following content: "Candidate symptoms: [1. Lack of interest, 2. Low mood, 3. Sleep disorder, 4. Lack of energy, 5. Eating disorder, 6. Low self-esteem, 7. Attention problems, 8. Hyperactivity / reduction, 9. Self-harm, 10. None], Please select the appropriate symptom description from the candidate symptoms according to the content in the text." Specifically, the prompt statement sent is: "Candidate symptoms: [1. Lack of interest, 2. Feeling down, 3. Sleeping disorder, 4. Lack of energy, 5. Eating disorder, 6. Low self-esteem, 7. Concentration problem, 8. Hyper / Lower activity, 9. Self-harm, 10. None], Please select the appropriate symptom description from the candidate symptoms according to the content in the text.", prompting the candidate to select a symptom that matches the current text description.

[0011] Furthermore, the prompt statement 30 expresses the following: "Emotional information: {Response1}, Symptom information: {Response2}. Based on this information and in conjunction with the text and image content released by the user, please determine whether there is a metaphorical expression. If there is, please provide the target domain and the source domain; if not, please answer 'No'." Specifically, the prompt statement is sent as follows: "Please, based on this information and combined with the content of the text and the picture released by the user, determine whether there is a metaphorical expression. If there is, please provide the target domain and the source domain. If not, please answer 'No'", so that it can determine whether the current text and image content contains a metaphorical expression while considering the emotional and symptom information.

[0012] Furthermore, the visual encoder employs the VIT model to obtain a global embedding representation of the image.

[0013] Furthermore, the text encoder employs the XLM-RoBERTa model to encode the text and extract text features.

[0014] Furthermore, S4 maps the image embedding vector to the same feature space as the text modality, specifically by using a linear layer with a GeLU activation function to map the image embedding vector to the same feature space as the text modality.

[0015] Furthermore, the classification described in S5 specifically involves using a linear layer and a softmax classifier for classification.

[0016] Furthermore, the loss function used in the method is... for: In the formula, in, This indicates the number of samples in the dataset. Represents the cross-entropy loss function. Represents the actual value. This represents the psychological state classification result obtained by classifying the dominant feature representation of the image. This represents the psychological state classification result obtained by classifying the text-dominant feature representation. This represents the psychological state classification result obtained by classifying the hybrid feature representation.

[0017] The beneficial effects of this invention are that it provides a multimodal mental health detection method based on thought chain cues and feature enhancement, solving the problems of insufficient multimodal information fusion and difficulty in recognizing implicit expressions in existing technologies. Experimental results show that this method significantly improves the model's performance in complex multimodal scenarios and has broad application potential. Specifically: (1) Enhance deep semantic understanding ability: Through pre-designed thought chain prompts, the large model can imitate the human thinking process and deeply explore the deep semantic information such as emotional tendencies, symptom manifestations and hidden metaphors contained in images and texts, thereby improving the model's ability to analyze the implicit content in the user's expression.

[0018] (2) Deep multimodal fusion: The information obtained by the large model prompting module is fused by text and image features respectively. This ensures that the key information of the single modality is not diluted, and captures the implicit relationship between text and image, thereby effectively integrating information from different modalities.

[0019] (3) Significantly improve recognition accuracy: Experimental results show that the method of the present invention outperforms multimodal baseline models such as PromptHate and UnifiedTMSC on the Twitter-Depression dataset, and is also superior to current large multimodal models such as Qwen2.5-VL-7B, with high accuracy and robustness.

[0020] (4) Wide range of applications: This invention is applicable to various scenarios such as early psychological intervention and clinical medical assistance. It can effectively monitor the mental health status of users and provide technical support for public mental health testing. Attached Figure Description

[0021] Figure 1 This is a technical roadmap in the embodiments.

[0022] Figure 2 It is a multimodal mental health detection model based on thought chain cues and feature enhancement.

[0023] Figure 3 It is a specific thought chain prompt template.

[0024] Figure 4 These are examples of depressed and non-depressed individuals from the dataset. Detailed Implementation

[0025] The implementation steps of the present invention are described in detail below with reference to the accompanying drawings.

[0026] A multimodal mental health detection method based on thought chain cues and feature enhancement includes the following steps: S1. Obtain a sample dataset for mental health testing. Each sample in the dataset includes text data and corresponding image data. This embodiment uses the publicly available Twitter-Depression dataset, which is primarily used for detecting depression in mental illnesses. Other preprocessed multimodal datasets can also be used. This dataset identifies and collects users who posted on Twitter claiming to have depression or having previously had depression, along with their posts, through text matching. It includes posts from 1402 depressed users and 5160 non-depressed users. Gui et al. further expanded the dataset into a multimodal dataset by obtaining the corresponding images based on the ID information of the images in these posts through the API provided by Twitter. This embodiment uses the expanded multimodal dataset.

[0027] Table 1 Dataset For portions of the dataset containing only text content and lacking corresponding images, a pre-trained stable-diffusion-3.5-large model is used to generate images that accurately reflect the original meaning of the text based on a prompt: "Please generate pictures based on the following text. The pictures should conform to the original meaning of the text and adopt a realistic style as much as possible." Finally, the dataset was obtained. in, Indicates the first Text data of one sample, This represents the corresponding image data. Labels indicating whether someone is not depressed or depressed.

[0028] S2 invokes the multimodal large model (LLaVA-v1.5-7B) and uses the mind chain prompting engineering to assist the multimodal large model in extracting semantic information, thereby enhancing the model's ability to extract deep semantic information contained in different modalities. Specifically, this includes: For input text and images, First, pre-designed prompts are used to extract emotional information from the text. The prompt is: "Please determine the sentiment polarity of the overall content according to the content in the text, and select one from positive, neutral, and negative." This step is described as follows: Then, the model selects symptom information that matches the text based on the PHQ-9 scale, using the corresponding prompt statements. The prompt is: "Candidate symptoms: [1. Lack of Interest, 2. Feeling Down, 3. Sleeping Disorder, 4. Lack of Energy, 5. Eating Disorder, 6. Low Self-Esteem, 7. Concentration Problem, 8. Hyper / Lower Activity, 9. Self-Harm, 10. None], Please select the appropriate symptom description from the candidate symptoms according to the content in the text. (Candidate symptoms: [1. Lack of Interest, 2. Low Mood, 3. Sleep Disorder, 4. Lack of Energy, 5. Eating Disorder, 6. Low Self-Esteem, 7. Attention Problem, 8. Hyperactivity / Reduction of Activity, 9. Self-Harm, 10. None], Please select the appropriate symptom description from the candidate symptoms according to the text content.)" This step is described as follows: Finally, the model is combined with the judgment results from the previous two steps. and To extract metaphorical information from text and images, the prompts used are... The prompt is: "Sentiment information: {Response1}, Symptom information: {Response2}. Please, based on this information and combined with the content of the text and the picture released by the user, determine whether there is a metaphorical expression. If there is, please provide the target domain and the source domain. If not, please answer 'No'." This step is described as follows: S3, use a visual encoder to process the image data to obtain an image embedding vector; Visual encoder: refers to the encoding process of input images using a pre-trained image model to generate an embedding vector representation of the image and map it into the same feature space as the text embedding, which facilitates the alignment and fusion of cross-modal features.

[0029] In the visual encoder part, this embodiment uses the VIT (VisionTransformer) model. VIT is an open-source vision model proposed by the Google team, specifically designed for image classification and visual feature extraction tasks. This model segments images into a sequence of regular patches, allowing it to share sequence processing logic with the Transformer architecture in natural language processing. This means that the image patch embeddings generated by VIT can be directly adapted to the Transformer's self-attention mechanism for global feature modeling. This unified sequence processing framework promotes the model's superior performance in image semantic understanding and long-distance dependency capture. After pre-training on large-scale image datasets, VIT outperforms traditional Convolutional Neural Networks (CNNs) on multiple vision tasks, making it a milestone model that successfully introduced the Transformer architecture into the field of vision. The calculation formula is: in, This represents the image embedding vector obtained by processing the input image using the VIT model.

[0030] The text data, sentiment information, symptom information, and metaphorical information are processed using a text encoder to obtain a text embedding vector; A text encoder refers to using a pre-trained text model to encode input text, generate an embedding vector representation of the text, and obtain the overall semantic representation of the text through pooling operations.

[0031] In the text encoder section, this embodiment uses the XLM-RoBERTa model. XLMR is an open-source cross-lingual pre-trained model proposed by MetaAI (formerly Facebook AI), designed specifically for multilingual natural language understanding tasks. Based on the BERT architecture, this model is trained on large-scale unsupervised text in over 100 languages, allowing different languages ​​to share the same set of word units and deep semantic space. This means that the language embeddings generated by XLMR can achieve cross-lingual alignment and interaction within a unified feature space. The shared architecture and semantic space contribute to the model's superior performance in tasks such as cross-lingual text classification, natural language reasoning, and named entity recognition. XLMR significantly outperforms earlier models in cross-lingual transfer capabilities and is one of the representative models in the current multilingual NLP field. The calculation formula is: in, This indicates that the input text is processed through the XLMR model and extracted via the thought chain. The resulting text embedding vector is then obtained.

[0032] S4, a linear layer with GeLU activation is used to map the image embedding vector into the same feature space as the text modality, resulting in the mapped image encoding. GeLU (Gaussian Error Linear Unit) is a common neural network activation function that enhances the model's expressive power by performing a smooth, non-linear transformation on the input based on a Gaussian distribution. GeLU is frequently used in the hidden layers of modern deep learning models such as Transformer and BERT to help improve model performance. This process is represented as follows: Then, the metaphorical portion of the text embedding vector is encoded. Image encoding after mapping The fusion yields image-dominated feature representations. The process formula is as follows: Encode the sentiment and symptom portions of the text embedding vector. Encoding of the text portion with the text embedding vector The fusion yields text-dominant feature representations. The process formula is as follows: The encoding of the text embedding vector and the encoding of the mapped image are fused to obtain a hybrid feature representation. The process formula is as follows: S5, the image-dominated feature representation, text-dominated feature representation, and mixed feature representation are classified using a linear layer and a softmax classifier to obtain the corresponding psychological state classification results. Finally, a final judgment on the user's mental health status is made based on a vote among the three classification results. Softmax is an activation function commonly used in classification tasks. It transforms each element of a vector into a probability value between 0 and 1, with the sum of all elements being 1. Softmax is typically used in the output layer of multi-class classification problems. The process formula is as follows: The method employs a loss function for: In the formula, in, This indicates the number of samples in the dataset. Represents the cross-entropy loss function. Represents the actual value. This represents the psychological state classification result obtained by classifying the dominant feature representation of the image. This represents the psychological state classification result obtained by classifying the text-dominant feature representation. This represents the psychological state classification result obtained by classifying the hybrid feature representation.

[0033] Experimental parameters: The experimental parameters and environment settings in this example are as follows: Two NVIDIA RTX 4090 24G GPUs were used for model training. Specific parameter settings are as follows: Imageembeddingsize(768): Sets the dimension of the image embedding to 768. This setting determines the representation capability of each image block in the embedding space, affecting the expression and fusion effect of image information.

[0034] Textmaxlength(512): Sets the maximum length of the input text to 512, ensuring that the length of each input sequence is consistent and avoiding difficulties in model processing caused by differences in text length.

[0035] Batchsize(8): The number of samples processed during each training session is set to 8. This parameter affects training speed and memory usage.

[0036] Learningrate (5e-5~1e-5): The learning rate is set to a range of 5e-5 to 1e-5. This setting ensures that the model can update parameters smoothly during training, avoiding instability or slow convergence caused by excessively large or small learning rates. The experimental parameter settings are shown in Table 2.

[0037] Table 2 Experimental parameter settings Comparative example: In the experiment, the performance of three baseline models in mental health testing tasks was compared: pure text model, pure image model and multimodal model. Their performance was compared with the method proposed in this embodiment to evaluate the performance of the method in this embodiment.

[0038] (1) Pure text model: This type of model mainly processes text data and is used for the detection of mental illnesses.

[0039] RoBERTa: This is an improved version of the BERT model that enhances model performance by increasing the scale of training data, improving the quality of training data annotation, and extending training time.

[0040] PHQ9-PLUS: This is a symptom classifier built on the PHQ-9 questionnaire. The model shows good generalization and interpretability on depression detection tasks on different datasets.

[0041] HAN-BERT: This model proposes a psychiatric scale-guided risk post screening method that can capture risk posts related to dimensions defined by the Clinical Depression Scale, enhancing the model's ability to detect early risks of depression. (2) Pure image model: This type of model focuses on image data processing and analyzes mental illnesses through images.

[0042] VGG: This is a classic deep convolutional neural network model that enhances the network's depth and feature representation capabilities by stacking multiple 3x3 convolutional and pooling layers.

[0043] ResNet-50: This is a 50-layer convolutional neural network with residual connections for image classification tasks. It solves the vanishing and exploding gradient problems during the training of deep neural networks by using residual block connections.

[0044] (3) Multimodal model: Since there are few multimodal detection models for mental health tasks, this paper also introduces models of other similar tasks as a baseline.

[0045] SenseMood: This model consists of a CNN-based classifier and BERT, which extracts deep features from user-posted images and text respectively, and combines visual and textual features to reflect the user's emotional expression, thereby performing depression detection and classification. PromptHate: This model uses implicit knowledge from a pre-trained RoBERTa language model to classify hate memes by constructing prompts and providing examples, taking into account text and image captioning information.

[0046] UnifiedTMSC: This model converts task descriptions into seed cues and uses paraphrasing rules to obtain different paraphrasing cues, while introducing a multimodal sentiment analysis framework with image prefix cues.

[0047] LLaVA: LLaVA is a large-scale multimodal model (LMM) developed by researchers, built by connecting the open-source visual encoder CLIP and the language decoder LLaMA. In the comparison, LLaVA-7b was chosen as the baseline model. The model is guided to output a judgment on the user's mental health status by providing specific cues, either as a single input image or a combination of image and text.

[0048] Qwen2.5-VL-7B-Instruct: Qwen2.5-VL is the latest multimodal large-scale model released by Tongyi Qianwen, available in several different sizes. Qwen2.5-VL enhances the model's ability to recognize real-world scales and spatial relationships by using coordinate values ​​corresponding to the actual size of the input image to represent bounding boxes and points. In the comparison, the Qwen2.5-VL-7B-Instruct model was used, with both an image and its corresponding title as input during the experiment.

[0049] Evaluation metrics: This example and the comparative example are analyzed using four metrics: accuracy, precision, recall, and F1 score.

[0050] Accuracy: Accuracy represents the proportion of samples correctly predicted by the model out of the total number of samples. It is one of the most commonly used metrics for evaluating the performance of classification models.

[0051] The calculation formula is: in: TP (TruePositive): The number of samples that the model correctly predicted as positive. TN (TrueNegative): The number of samples that the model correctly predicted as negative examples. FP (False Positive): The number of samples that the model incorrectly predicts as positive. FN (False Negative): The number of false negatives that the model incorrectly predicted as negatives. Precision: Precision represents the proportion of samples that the model predicts to be positive but are actually positive. It is used to measure the accuracy of the model when predicting positive.

[0052] The calculation formula is: Recall: Recall represents the proportion of samples that a model can correctly predict as positive out of all samples that are actually positive. It is used to measure the model's ability to identify positive samples.

[0053] The calculation formula is: F1 score: The F1 score is the harmonic mean of precision and recall, which takes into account both the precision and recall of the model and is suitable for imbalanced class situations.

[0054] The calculation formula is: For a given predicted label, precision and recall can be quickly calculated using a confusion matrix. The confusion matrix is ​​used to evaluate the performance of a model for binary classification problems. It contains four basic prediction outcomes: true positives (TP), true negatives (TN), false positives (FP), and false negatives (FN). In the confusion matrix, columns represent the true labels, and rows represent the predicted labels. The intersections of rows and columns correspond to these four basic outcomes, helping us to more intuitively analyze the model's classification performance. Table 3 shows the structure of the confusion matrix.

[0055] Table 3. Character Meanings in the Confusion Matrix Table 4 shows the experimental results on the Twitter-Depression dataset. Table 4 presents the experimental results of different models on the mental health detection task in the Twitter-Depression dataset. The results reveal the advantages of the proposed solution in the mental health detection task. Multimodal models, by combining information from text and images, can provide more comprehensive data analysis, thereby improving the accuracy and reliability of detection. The mental health detection model incorporating thought chain cues proposed in this embodiment demonstrates significant advantages. Comparative experiments show that this method outperforms traditional methods relying on single-modal analysis in capturing overall sentiment, mining symptom information, and decoding metaphorical information, and is also superior to early fusion models that simply splice multimodal features. The results indicate that the thought chain cue extraction mechanism plays a core role in deep semantic parsing, systematically improving the model's ability to extract complex information related to mental health. By analyzing the semantics of text and images layer by layer, the model can identify implicit emotional contradictions in language, associate content related to psychopathological features in descriptions, and use this information to decode metaphorical expressions, analyze the psychological mapping behind these symbolic expressions, and thus make more accurate judgments.

[0056] Finally, it should be noted that the above embodiments are intended to illustrate the technical solutions of the present invention and do not constitute any limitation on the present invention. Those skilled in the art should fully understand that modifications to the technical solutions described in the foregoing embodiments or equivalent substitutions for any part or all of the technical features are entirely feasible. Such modifications or substitutions, as long as they do not depart from the scope of protection defined by the claims of the present invention, should be considered reasonable extensions of the present invention.

Claims

1. A multimodal mental health detection method based on thought chain cues and feature enhancement, characterized in that, Includes the following steps: S1, Obtain a sample dataset for mental health testing, wherein each sample in the dataset includes text data and image data corresponding to the text; S2 invokes the multimodal large model, using the mind chain prompting engineering to assist the multimodal large model in extracting semantic information, specifically including: Using the first prompt statement, extract the emotional information from the text data; Using prompt statement two, obtain symptom information from the PHQ-9 scale that matches the text data; Using the third prompt statement, based on the emotional polarity information, the symptom information, the text data, and the image data corresponding to the text, metaphorical information is obtained; S3, use a visual encoder to process the image data to obtain an image embedding vector; The text data, sentiment information, symptom information, and metaphorical information are processed using a text encoder to obtain a text embedding vector; S4, map the image embedding vector to the same feature space as the text modality to obtain the mapped image encoding; The metaphorical part of the text embedding vector is fused with the mapped image encoding to obtain an image-dominated feature representation; The sentiment and symptom portions of the text embedding vector are fused with the text portion encoding of the text embedding vector to obtain a text-dominant feature representation; The encoding of the text embedding vector is fused with the encoding of the mapped image to obtain a hybrid feature representation; S5, classify the image-dominated feature representation, text-dominated feature representation and mixed feature representation to obtain the corresponding psychological state classification results, and finally make a final judgment on the user's psychological health status based on the voting of the three classification results.

2. The method according to claim 1, characterized in that, The multimodal large model described in S2 is LLaVA-v1.5-7B.

3. The method according to claim 1, characterized in that, The prompt statement reads, "Please determine the overall emotional polarity of the text and choose one from positive, neutral, or negative." 4. The method according to claim 1, characterized in that, The second prompt statement reads: "Candidate symptoms: [1. Loss of interest, 2. Depressed mood, 3. Sleep disturbance, 4. Lack of energy, 5. Eating disorder, 6. Low self-esteem, 7. Attention problems, 8. Hyperactivity / reduction, 9. Self-harm, 10. None]. Please select the appropriate symptom description from the candidate symptoms based on the text." 5. The method according to claim 1, characterized in that, The prompt statement three expresses the following information: "Emotional information: {Response1}, Symptom information: {Response2}. Based on this information and the text and image content posted by the user, please determine whether there is a metaphorical expression. If so, please provide the target domain and source domain; If it does not exist, please answer 'No'.

6. The method according to claim 1, characterized in that, The visual encoder uses the VIT model.

7. The method according to claim 1, characterized in that, The text encoder uses the XLM-RoBERTa model.

8. The method according to claim 1, characterized in that, S4 maps the image embedding vector to the same feature space as the text modality, specifically by using a linear layer with a GeLU activation function to map the image embedding vector to the same feature space as the text modality.

9. The method according to claim 1, characterized in that, The classification described in S5 specifically involves using a linear layer and a softmax classifier for classification.

10. The method according to claim 1, characterized in that, The method employs a loss function for: In the formula, in, This indicates the number of samples in the dataset. Represents the cross-entropy loss function. Represents the actual value. This represents the psychological state classification result obtained by classifying the dominant feature representation of the image. This represents the psychological state classification result obtained by classifying the text-dominant feature representation. This represents the psychological state classification result obtained by classifying the hybrid feature representation.