Mental health assessment method and system based on large language model and multi-modal data
By combining a large language model with a mental health assessment system based on multimodal data, the problems of strong subjectivity and low efficiency of existing assessment methods are solved, achieving a more objective and efficient mental health assessment.
Patent Information
- Application Number
- CN202510700338.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-28
- Publication Date
- 2025-09-12
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Existing mental health assessment methods mainly rely on self-assessment scales, clinical interviews or psychological questionnaires. They have problems such as strong subjectivity, reliance on professional judgment, and low assessment efficiency, making it difficult to meet the automation and real-time requirements of large-scale screening.
A mental health assessment system based on a large language model and multimodal data is used to achieve intelligent identification and classification of depressive states through data collection, preprocessing, feature extraction, multimodal fusion and large language model modeling.
It improves the objectivity, scientificity and identification efficiency of mental health assessment, can more accurately identify individuals with potential psychological risks, and enhances the automation and practicality of mental health management.
Smart Images

Figure CN120636702A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of artificial intelligence and mental health assessment, and in particular to a mental health assessment method and system based on a large language model and multimodal data. Background Art
[0002] As mental health issues receive increasing public attention, depression, a common mental illness, has become a key public health issue in terms of early screening and intervention. Existing depression assessment methods primarily rely on self-rating scales (such as the PHQ-9), clinical interviews, or psychological questionnaires. These methods are subject to high subjectivity, reliance on professional judgment, and low assessment efficiency, making them difficult to meet the automation and real-time requirements of large-scale screening.
[0003] In recent years, artificial intelligence (AI) technologies have been widely applied in medicine and psychology, particularly speech emotion recognition, expression recognition, and natural language processing, offering new insights into the objective detection of depression. However, existing research often processes data from different modalities independently, resulting in a loss of cross-modal semantic connections, limited recognition accuracy, and a lack of effective multimodal fusion mechanisms. While some studies have introduced multimodal models, these models lack the ability to align the temporal order and semantic integration of multi-source data, and lack a unified information representation framework, hindering their generalizability and interpretability.
[0004] Therefore, there is an urgent need for an automated detection method for mental health status that can integrate multimodal information of speech, vision and text and has deep semantic modeling and classification capabilities to improve the accuracy, objectivity and practicality of detection. Summary of the Invention
[0005] (1) Technical issues to be solved
[0006] In view of the above-mentioned shortcomings and deficiencies of the prior art, the present invention provides a psychological health assessment method and system based on a large language model and multimodal data, which integrates multimodal information of speech, vision and text, and improves cross-modal fusion capabilities, time series modeling accuracy and semantic understanding depth through unified embedding structure and large language model modeling, and through algorithm innovation, realizes efficient, accurate and deployable automatic detection of psychological health, thereby improving the objectivity, intelligence and practicality of the assessment.
[0007] The above system collects audio, video and text data programmatically, and uses deep learning models to achieve cross-modal feature extraction, semantic fusion and classification tasks. It can be applied to fields such as emotional computing, human-computer interaction, and behavior analysis.
[0008] (2) Technical solution
[0009] In order to achieve the above objectives, the main technical solutions adopted by the present invention include:
[0010] In a first aspect, an embodiment of the present invention provides a mental health assessment system based on a large language model and multimodal data, comprising:
[0011] A data collection module is used to collect basic attribute information and depression assessment scale data of the subjects, and synchronously collect audio data and video data of the subjects during the assessment process, as well as text records through a standardized interface; the subjects include subjects in a healthy state and subjects in a depressed state;
[0012] A data preprocessing module, used to preprocess audio data, video data and text records respectively;
[0013] The feature extraction module is used to extract the sound source, spectrum, and rhythm features from the pre-processed audio data and convert the sound source, spectrum, and rhythm features into audio pseudo-text; extract facial movements, eye movement trajectories, and head posture from the pre-processed video data and generate video pseudo-text; and perform speech transcription on the pre-processed audio data while combining it with text records for standardization and word segmentation to obtain text features;
[0014] A multimodal fusion module, configured to input the audio pseudo-text, video pseudo-text, and text features into a unified embedding representation structure to obtain a fused multimodal vector;
[0015] The large language model processing module is used to input multimodal vectors into the pre-trained large language model for contextual semantic modeling and output the fused semantic embedding matrix;
[0016] The model training module is used to input the semantic embedding matrix into the pre-built depression binary classification model to train the model and use the trained model to evaluate the user's mental health.
[0017] Optionally, the data acquisition module includes:
[0018] The first data collection unit is used to collect attribute information and total scores of depression assessment scales of a specified number of subjects online using a mini-program;
[0019] The second data acquisition unit is used to collect audio data in m4a format with a sampling frequency of 8kHz;
[0020] The third data acquisition unit is used to collect video data of the subject facing the camera during the self-narration process at a frame rate of not less than 30 frames per second.
[0021] The fourth data collection unit is used to collect text records of the subjects' self-reports during the self-assessment process.
[0022] Optionally, the data preprocessing module includes:
[0023] A first data preprocessing unit compares the total score of the first data collection unit with a specified threshold, and determines the score greater than the specified threshold as a healthy control group, and the rest as a depression group;
[0024] The second data preprocessing unit is used to convert the audio data into audio data in WAV format with a sampling frequency of 16kHz, and obtain preprocessed audio data through endpoint detection, pre-emphasis, framing, and windowing.
[0025] The third data preprocessing unit processes the video data frame by frame, locates the facial area through face detection, performs face alignment and geometric normalization processing, and obtains preprocessed video data.
[0026] The fourth data preprocessing unit performs stop word removal, word segmentation, part-of-speech tagging and normalized coding processing on the text self-description record to obtain a preprocessed text record.
[0027] Optionally, the feature extraction module includes:
[0028] An audio feature extraction unit is used to call an audio processing tool to extract 523 acoustic features from the preprocessed audio data, including sound source features, spectrum features, and prosody features, and convert the sound source, spectrum, and prosody features into audio pseudo-text;
[0029] The visual feature extraction unit is used to call the video analysis tool to process the video frames, extract facial movements, gaze trajectory and head posture features, and perform semantic description of the key frame features in the time series to generate video pseudo text;
[0030] The text feature extraction unit is used to call the speech recognition tool to transcribe the audio content into text, and perform word segmentation, part-of-speech tagging, entity recognition and sentiment word extraction on the transcribed text and the pre-written text records to form text features.
[0031] Optionally, the multimodal fusion module includes:
[0032] An embedding encoding unit is used to input the audio pseudo-text, video pseudo-text and text features into a unified embedding representation structure, wherein the embedding representation structure includes token embedding: converting features of different modalities into vector representations of unified dimensions, position embedding: adding temporal position information to each modal feature, and paragraph embedding: adding modality identifiers to features from different modal sources; thereby obtaining multimodal data;
[0033] A temporal alignment unit, configured to align audio, visual, and text features in a temporal dimension using temporal position information in the position embedding, thereby updating multimodal data;
[0034] A feature fusion unit is used to merge the updated multimodal data in the vector dimension. The merging method includes at least one of the following methods: feature splicing, weighted summation, or attention mechanism to achieve semantic fusion between modalities;
[0035] The fusion output unit is used to output the semantically fused multimodal data as a fused multimodal vector to the large language model processing module.
[0036] Optionally, the large language model processing module includes:
[0037] The context modeling unit receives the multimodal vectors output by the multimodal fusion module and inputs them into a pre-trained large language model for contextual semantic analysis. The large language model uses a deep neural network based on the Transformer structure to support long-range dependency modeling and cross-modal context capture.
[0038] The embedding optimization unit is used to extract high-dimensional semantic features from the output processed by the large language model and generate a fused semantic embedding matrix. The semantic embedding matrix is used to express the subject's emotional and psychological state characteristics at the audio, visual and text levels.
[0039] Optionally, the model training module includes:
[0040] The validation unit was used to train a binary depression classification model using 5-fold cross-validation. The training data was split into 80% training and 20% test sets. The following hyperparameters were optimized using grid search: convolution kernel size: {3, 5, 7}; number of filters: {64, 128, 256}; number of neurons in the fully connected layer: {16, 32, 64, 128}; learning rate: {0.001, 0.0005}. An early stopping mechanism was implemented during training, terminating the training if the AUC on the validation set did not improve for five consecutive rounds.
[0041] Among them, five indicators were selected to evaluate the generalization performance of the depression binary classification model, including receiver operating characteristic curve (ROC), AUC, accuracy, precision, recall rate and F1 value.
[0042] In a second aspect, an embodiment of the present invention further provides a method for psychological health assessment based on a large prediction model and multimodal data, comprising:
[0043] Collecting basic attribute information and depression assessment scale data of the subjects, and simultaneously collecting audio data and video data, as well as text records, of the subjects during the assessment process; the subjects include healthy subjects and depressed subjects;
[0044] Preprocess the audio data, video data and text records separately;
[0045] Extracting sound source, spectrum, and rhythm features from the preprocessed audio data and converting them into audio pseudo-text; extracting facial movements, eye movement trajectories, and head posture from the preprocessed video data and generating video pseudo-text; and performing speech transcription on the preprocessed audio data while combining it with text records for standardization and word segmentation to obtain text features;
[0046] Inputting the audio pseudo-text, video pseudo-text and text features into a unified embedding representation structure to obtain a fused multimodal vector;
[0047] Input the multimodal vector into the pre-trained large language model for contextual semantic modeling, and output the fused semantic embedding matrix;
[0048] The semantic embedding matrix is input into a pre-built depression binary classification model to train the model, and the trained model is used to evaluate the user's mental health.
[0049] Optionally, the collecting of the subject's basic attribute information and depression assessment scale data, and the simultaneous collection of the subject's audio data and video data during the assessment process, as well as text records, include:
[0050] A mini program was used to collect online the attribute information and total scores of the depression assessment scale of a specified number of subjects, collect audio data in m4a format with a sampling frequency of 8kHz, collect video data of the subjects facing the camera during the self-report process, and collect text records of the subjects' self-report during the self-assessment process.
[0051] Optionally, preprocessing is performed on the audio data, video data and text records respectively; including: comparing the total score of the first data collection unit with a specified threshold, determining that the total score greater than the specified threshold is a healthy control group, and the rest are a depression group;
[0052] The audio data is converted into WAV format audio data with a sampling frequency of 16kHz, and preprocessed by endpoint detection, pre-emphasis, framing, and windowing to obtain preprocessed audio data;
[0053] Process the video data frame by frame, locate the facial area through face detection, perform face alignment and geometric normalization, and obtain pre-processed video data;
[0054] The text self-description records are processed by removing stop words, segmenting words, tagging parts of speech and normalizing the coding to obtain the preprocessed text records.
[0055] Optionally, calling an audio processing tool to extract 523 acoustic features from the preprocessed audio data, the features including sound source features, spectrum features, and prosodic features, and converting the sound source, spectrum, and prosodic features into audio pseudo-text;
[0056] Call the video analysis tool to process the video frames, extract facial movements, gaze trajectory and head posture features, and perform semantic description of the key frame features in the time series to generate video pseudo text;
[0057] Call the speech recognition tool to transcribe the audio content into text, and perform word segmentation, part-of-speech tagging, entity recognition, and sentiment word extraction on the transcribed text and pre-written text records to form text features.
[0058] Optionally, the audio pseudo-text, video pseudo-text and text features are input into a unified embedding representation structure to obtain a fused multimodal vector, including:
[0059] Inputting the audio pseudo-text, video pseudo-text and text features into a unified embedding representation structure, the embedding representation structure includes token embedding: converting features of different modalities into vector representations of unified dimensions, position embedding: adding temporal position information to each modal feature, and paragraph embedding: adding modality identifiers to features from different modal sources; thus obtaining multimodal data;
[0060] Using the temporal position information in the position embedding, audio, visual, and text features are aligned in the temporal dimension to update the multimodal data;
[0061] The updated multimodal data are merged in the vector dimension, and the merging method includes at least one of the following methods: feature splicing, weighted summation, or attention mechanism to achieve semantic fusion between modalities; the semantically fused multimodal data is used as the fused multimodal vector.
[0062] Optionally, a multimodal vector output from the multimodal fusion module is received and input into a pre-trained large language model for contextual semantic analysis. The large language model adopts a deep neural network based on the Transformer structure, supports long-distance dependency modeling and cross-modal context capture, and performs high-dimensional semantic feature extraction on the output processed by the large language model to generate a fused semantic embedding matrix. The semantic embedding matrix is used to express the subject's emotional and psychological state characteristics at the audio, visual and text levels.
[0063] (3) Beneficial effects
[0064] This paper innovatively integrates large language model technology to construct a mental health assessment method and system based on large language models and multimodal data. This system integrates objective audio, video, and text data into a unified model, combining deep learning with large language models to automatically identify and analyze depression risk. Compared to traditional subjective assessment methods that rely solely on psychological scales, this paper can effectively improve the objectivity, scientific nature, and recognition efficiency of mental health screening.
[0065] This invention provides a complete multimodal psychological assessment pathway, encompassing data collection, feature extraction, pseudo-text construction, semantic fusion, and classification modeling. With a clear system structure and a highly automated processing flow, it enables comprehensive modeling and accurate identification of psychological states. This system contributes to the shift from empirical assessment to data-driven mental health assessment, providing technical support for early screening and intervention for psychological problems, and possesses practical application value.
[0066] This invention effectively solves the problems existing in existing mental health identification methods, such as strong subjectivity, insufficient single-modal information expression capabilities, and weak multimodal fusion capabilities, and improves the accuracy, stability and scalability of emotion recognition; through pseudo-text construction and large language model semantic enhancement mechanism, it strengthens the system's semantic understanding ability and cross-modal modeling capabilities, and improves the intelligence level of the overall detection system.
[0067] Applying this method to practical mental health management tasks can help relevant institutions improve the speed and accuracy of identifying individuals at risk, optimize assessment and intervention processes, and achieve automation and scalability of mental health services while protecting user privacy. This method has broad engineering feasibility and potential for translational applications. BRIEF DESCRIPTION OF THE DRAWINGS
[0068] Figure 1 A flowchart of a method for mental health assessment based on a large language model and multimodal data provided by one embodiment of the present invention;
[0069] Figure 2 A schematic diagram of the overall architecture of a mental health assessment method and system based on a large language model and multimodal data provided by one embodiment of the present invention;
[0070] Figure 3 This is a schematic diagram of the overall data flow describing the data input and output relationships and collaborative working mechanisms among the modules of the system of the present invention;
[0071] Figure 4 and Figure 5 Schematic diagrams of the ROC and AUC test results in the model training phase. DETAILED DESCRIPTION
[0072] In order to better explain the present invention and facilitate understanding, the present invention is described in detail below through specific implementation methods in conjunction with the accompanying drawings.
[0073] Multimodal data fusion technology leverages information from multiple aspects of an individual, including speech, vision, and text, to objectively reflect their emotional and psychological state. Speech features, as important physiological and behavioral indicators, can reveal an individual's emotional state by analyzing speech rate, intonation, pauses, and other factors. They offer high levels of emotion recognition and the ability to differentiate psychological states. Visual features further enrich emotional expression through behavioral manifestations such as facial expressions, eye movements, and head posture. Text features, based on the content and expression of an individual's language, provide semantic clues to their emotional and cognitive states.
[0074] This invention utilizes advanced multimodal fusion methods, combined with deep learning and large language model technologies, to automatically learn and extract the underlying features and patterns in multimodal data, enabling intelligent identification and classification of depressive states. This system objectively and quantitatively reveals the contribution of various features to depression identification, overcoming the subjectivity and low accuracy of traditional single-modality assessments. It provides strong technical support for the automated and scientific assessment of mental health.
[0075] The purpose of this invention is to synchronously collect audio, video, and text self-description data from subjects, use professional tools to preprocess and extract features from various types of data, construct a unified pseudo-text representation, and then achieve deep fusion of multimodal information through a shared embedding layer and a large language model. Finally, based on a convolutional neural network model, it can accurately classify depression risk. This system is applicable to various application scenarios and provides a scientific, objective, and efficient technical means for mental health management.
[0076] Example 1
[0077] like Figure 1 and Figure 2 As shown, this embodiment provides a psychological health assessment system based on a large language model and multimodal data. The system of this embodiment may include: a data acquisition module, a data preprocessing module, a feature extraction module, a multimodal fusion module, a large prediction model processing module, and a model training module. The system of this embodiment can be integrated into any electronic device to implement psychological assessment of the subject. To better understand the various modules in the above system, the following is combined with Figure 1 and Figure 2 Provide a detailed description of the functions of each module.
[0078] Data Collection Module: This module collects basic attribute information and depression assessment scale data from subjects, as well as audio and video data and text records from the assessment process. This includes both healthy and depressed subjects. For example, multimodal data from subjects can be collected through a customized platform or mini-program tool.
[0079] For example, the data acquisition module may include:
[0080] The first data collection unit is used to collect attribute information and total scores of the depression assessment scale of a specified number of subjects online using a mini-program; for example, the attribute information may include: gender, age, etc.; the depression assessment scale is the PHQ-9 table; and the psychological state groups are divided according to the total score of the PHQ-9.
[0081] The second data acquisition unit is used to collect audio data in m4a format with a sampling frequency of 8kHz; usually, it records the subject's self-narrated voice during the self-assessment process.
[0082] The third data acquisition unit is used to collect video data of the subject facing the camera during the self-narration process, corresponding to the aforementioned audio data acquisition process. At this time, the facial image during the self-narration process is synchronously recorded, and the frame rate of the video data is not less than 30 frames / second.
[0083] The fourth data collection unit is used to collect text records of the subjects' self-reports during the self-assessment process. The text records may include speech transcriptions or subjective self-reports, etc. The above text records are used to explore the subjects' language expression characteristics and emotional tendencies.
[0084] Data preprocessing module: used to preprocess audio data, video data and text records respectively. In other words, it normalizes the collected raw data.
[0085] For example, the data preprocessing module may include: a first data preprocessing unit, a second data preprocessing unit, a third data preprocessing unit, and a fourth data preprocessing unit.
[0086] A first data preprocessing unit compares the total score of the first data collection unit with a specified threshold, and determines the score greater than the specified threshold as a healthy control group, and the rest as a depression group;
[0087] The second data preprocessing unit is used to convert the audio data into audio data in WAV format with a sampling frequency of 16kHz, and obtain preprocessed audio data through endpoint detection, pre-emphasis, framing, and windowing. That is, the audio data is converted into WAV format and the sampling rate is increased to 16kHz. The audio data here refers to the audio data of each subject in the healthy control group and the depression group.
[0088] The third data preprocessing unit processes the video data frame by frame, locates the facial area through face detection, performs face alignment and geometric normalization to ensure the consistency of visual input, and obtains preprocessed video data. The video data here is the video data of each subject in the healthy control group and the depression group.
[0089] The fourth data preprocessing unit performs stop word removal, word segmentation, part-of-speech tagging, and normalized encoding on the text self-report records to obtain preprocessed text records. In this embodiment, the text data can be subjected to stop word removal, word segmentation, part-of-speech tagging, and normalized encoding to adapt to the subsequent semantic modeling of the large prediction model. The text self-report records here are the text self-report records of each subject in the healthy control group and the depression group.
[0090] Feature extraction module: used to extract the sound source, spectrum and rhythm features from the preprocessed audio data, and convert the sound source, spectrum and rhythm features into audio pseudo-text; extract facial movements, eye movement trajectories and head postures from the preprocessed video data and generate video pseudo-text; and perform speech transcription on the preprocessed audio data and standardize and segment it in combination with the text records to obtain text features.
[0091] The audio feature extraction unit of the feature extraction module is used to call an audio processing tool (such as OpenSMILE) to extract 523 acoustic features from the preprocessed audio data. The features include: sound source features, spectrum features, and prosodic features, and convert the sound source, spectrum, and prosodic features into audio pseudo-text, i.e., structured text description;
[0092] In this embodiment, the method for extracting audio data includes: calculating statistics such as maximum value, minimum value, average value, range, standard deviation, skewness and kurtosis for each feature.
[0093] The visual feature extraction unit of the feature extraction module is used to call video analysis tools (such as OpenFace) to process video frames, extract facial movements, gaze tracks, and head posture features, and perform semantic descriptions on key frame features in the time series to generate video pseudotext;
[0094] The text feature extraction unit of the feature extraction module is used to call a speech recognition tool (such as the Whisper model) to transcribe the audio content into text, and perform word segmentation, part-of-speech tagging (such as sentiment word tagging), syntactic analysis, entity recognition, and sentiment word extraction on the transcribed text and pre-written text records to form text features.
[0095] Multimodal fusion module: used to input the audio pseudo-text, video pseudo-text and text features into a unified embedding representation structure to obtain a fused multimodal vector;
[0096] For example, the embedding coding unit of the multimodal fusion module is used to input the audio pseudo-text, video pseudo-text and text features into a unified embedding representation structure (i.e., embedding layer). The embedding representation structure includes token embedding: converting the features of different modalities into vector representations of unified dimensions, position embedding: adding temporal position information to each modal feature, and paragraph embedding: adding modal identifiers to features from different modal sources to keep the temporal alignment; thus, multimodal data is obtained.
[0097] A temporal alignment unit of the multimodal fusion module is configured to align audio, visual, and text features in a temporal dimension using temporal position information in the position embedding to update the multimodal data;
[0098] The feature fusion unit of the multimodal fusion module is used to merge the updated multimodal data in the vector dimension. The merging method includes at least one of the following methods: feature splicing, weighted summation, or attention mechanism to achieve semantic fusion between modalities;
[0099] The fusion output unit of the multimodal fusion module is used to output the semantically fused multimodal data as a fused multimodal vector to the large language model processing module.
[0100] Large language model processing module: used to input multimodal vectors into the pre-trained large language model for contextual semantic modeling and output the fused semantic embedding matrix;
[0101] For example, the context modeling unit of the large language model processing module is used to receive the multimodal vectors output by the multimodal fusion module and input them into the pre-trained large language model for contextual semantic analysis. The large language model uses a deep neural network based on the Transformer structure to support long-range dependency modeling and cross-modal context capture;
[0102] The embedding optimization unit of the large language model processing module is used to extract high-dimensional semantic features from the output processed by the large language model and generate a fused semantic embedding matrix. The semantic embedding matrix is used to express the subject's emotional and psychological state characteristics at the audio, visual and text levels, and fully express the individual's multimodal psychological state information.
[0103] Model training module: used to input the semantic embedding matrix into a pre-built depression binary classification model (i.e., convolutional neural network model) to train the model and use the trained model to conduct mental health assessments on users.
[0104] Specifically, during the training phase, the embedding matrices of healthy and depressed subjects are fed into a convolutional neural network model for training and inference. The CNN model, which includes multiple layers of convolution, activation, pooling, and fully connected layers, uses a sigmoid function to output a binary classification result, determining whether an individual is at risk for depression.
[0105] The depression binary classification model includes:
[0106] Feature input layer: used to receive the semantic embedding matrix output by the large language model processing module. The input data is Z-score normalized before input;
[0107] At least two cascaded multi-scale one-dimensional convolutional layers, each layer is configured with an adjustable convolution kernel kernel_size∈{3,5,7} and a dynamic number of filters filters∈{64,128,256}, and nonlinear transformation is achieved through the ReLU activation function. Each level of convolutional layer is selectively connected to a maximum pooling layer to extract local features, and the final feature map is reduced in dimensionality by global average pooling;
[0108] Fully connected layer: used to perform nonlinear combination of pooled feature vectors. The number of neurons is selected from {16, 32, 64, 128} through grid search.
[0109] Classification output layer: Outputs the binary classification prediction results of depression through the Sigmoid function.
[0110] It is understandable that the above-mentioned model training module is specifically used for:
[0111] A CNN binary depression classification model was constructed. The optimal hyperparameters were selected through grid search. Five-fold cross-validation was used, and the training data was divided into an 80% training set and a 20% test set. Training was performed using the training set data, and the following hyperparameters were optimized through grid search: convolution kernel size: {3, 5, 7}; number of filters: {64, 128, 256}; number of neurons in the fully connected layer: {16, 32, 64, 128}; learning rate: {0.001, 0.0005}. An early stopping mechanism was implemented during training, and training was terminated when the AUC on the validation set did not improve for five consecutive rounds. Predictions were made using the optimized binary depression classification model using the test set data.
[0112] Among them, five indicators were selected to evaluate the generalization performance of the depression binary classification model, including receiver operating characteristic curve (ROC), AUC, accuracy, precision, recall rate and F1 value.
[0113] In this embodiment, a novel multimodal mental health assessment system was constructed using advanced artificial intelligence algorithms, combining objective audio, video, and text data. The system combines voice, facial behavior, and language expression features with psychological assessment scales, and establishes an intelligent recognition model for mental health status based on deep learning and large language model technology. Compared to traditional psychological assessment methods that rely on subjective questionnaires or interviews, the system of this embodiment can more objectively and efficiently identify individuals at risk of mental health problems, improving the scientific nature and operability of mental health management.
[0114] The above embodiments innovatively utilize the powerful semantic understanding capabilities of large language models to achieve deep integration and precise analysis of multimodal depression features, significantly improving the objectivity and accuracy of detection and providing an intelligent solution for mental health assessment.
[0115] This system has the characteristics of high degree of automation, strong modeling capabilities, and flexible deployment. It is suitable for various application scenarios such as medical institutions, psychological counseling platforms, educational scenarios, and remote psychological service systems. It can provide reliable data support and risk warning methods for mental health workers, management agencies, and public decision-making departments. It has good promotion and application prospects and real social value.
[0116] Example 2
[0117] This embodiment provides a mental health assessment system that integrates voice, video and text features, combines psychological scale assessment with multimodal feature modeling, and realizes intelligent assessment of mental health status. The system has good adaptability and scalability, and specifically includes the following functional modules: data acquisition module; data preprocessing and feature extraction module; multimodal modeling and large prediction model fusion module; classification model training and performance evaluation module; and external verification module of the model.
[0118] 1. Data acquisition module, specifically used for:
[0119] (1) The subjects' general information and depression assessment scale PHQ-9 were collected online through the WeChat official account. The subjects' depressive symptoms were assessed based on the total score of the depression assessment scale PHQ-9.
[0120] For example, conducting a scale assessment can be understood as setting students with a total score of <5 on the PHQ-9 scale as a healthy control group, and setting those with a total score of ≥5 on the PHQ-9 scale as a depression group.
[0121] (2) Use mobile devices to synchronously collect the voice data and video data of the subjects during the self-assessment process. The subjects face the camera in a quiet and evenly lit environment and independently perform a structured self-expression task. The expression content can revolve around the following questions, including but not limited to: "What is the most creative thing you have done?", "What do you do when you are sure that you are right, but others disagree with you?", "How hard will you work to achieve your goals?"
[0122] Voice Data Collection: Use your phone's built-in microphone to record audio. It's recommended to save the audio in m4a format, with a sampling rate of 8kHz or 16kHz. The recommended recording duration is between 1 and 5 minutes. The audio signal must be clear and continuous, and background noise and environmental interference should be minimized to ensure the quality of the original data.
[0123] Video Data Collection: Use the front-facing camera of your phone to capture the subject's frontal facial image. The subject should maintain a stable posture and a natural expression during the recording. The camera should be fixed, with a frame rate of at least 30 frames per second and a resolution of at least 720p. Avoid occlusion, uneven lighting, or image jitter to ensure complete extraction of facial key points and behavioral features.
[0124] Text data collection: Automatically transcribe speech content into text through a speech recognition system (such as Whisper) to form a standardized language expression sequence for subsequent text preprocessing and semantic feature extraction.
[0125] In this embodiment, there are no mandatory restrictions on speech content, speaking speed, and emotional tendencies. The purpose is to preserve the psychological expression process of the subject in a natural state as much as possible, so as to improve the authenticity of multimodal feature extraction, the accuracy of semantic analysis, and the individual adaptability of the classification model.
[0126] 2. Data preprocessing and feature extraction module
[0127] This module is used to perform standardization and structured feature extraction on the collected raw voice, video, and text data. It specifically includes the following three submodules:
[0128] (1) Speech preprocessing and feature extraction submodule
[0129] The collected raw audio data (m4a format, 8kHz sampling rate) is first converted to WAV format with a 16kHz sampling rate using multimedia processing tools (such as FFmpeg) to improve speech resolution and match the input requirements of the subsequent feature extraction model. The following processing steps are then performed: endpoint detection: This method uses a combined energy and zero-crossing rate method to determine the start and end times of speech, eliminating silent or noisy segments; pre-emphasis: This method uses a first-order high-pass filter to emphasize high-frequency features and improve the energy balance of the speech signal; and framing and windowing: This method uses a sliding window (e.g., 25ms window length, 10ms frame shift) to divide the continuous signal into multiple short-term stable frames, which are weighted using a Hamming window to reduce spectral leakage.
[0130] Speech feature extraction was performed using the OpenSMILE toolkit, using its emotion recognition-specific feature set (e.g., emoLarge). Multi-dimensional parameters, including sound source features, spectral features, and prosodic features, were extracted. Specifically, these include: sound source features (e.g., pitch, loudness, and jitter), which reflect speaking state and emotional control; spectral features (e.g., MFCC and frequency band energy distribution), which reflect the propagation characteristics of sound waves; and prosodic features (e.g., speaking rate, pauses, and rhythmic changes), which reflect the impact of psychological state on speech rhythm.
[0131] After extraction, each feature is analyzed using tools like Librosa to calculate descriptive statistics such as maximum, minimum, mean, range, standard deviation, skewness, and kurtosis, forming a fixed-length, high-dimensional vector. These features are then categorized by dimension into three main categories: sound source features, spectral features, and rhythmic features, for subsequent modeling.
[0132] In addition, nonlinear dynamic feature analysis methods such as minimum delay time, correlation dimension, phase space reconstruction and other technologies can also be used to extract nonlinear geometric features to assist in judging the stability and dynamic change characteristics of speech emotions.
[0133] (2) Video data preprocessing and visual feature extraction submodule
[0134] The video data is processed frame by frame, and the preprocessing steps include: face detection: using a face detection algorithm based on a deep convolutional neural network (such as Dlib or OpenCV) to locate the face area in each frame and eliminate invalid frames; face alignment: performing affine transformation based on key points such as the corners of the eyes and the tip of the nose to standardize the face to a unified posture; image normalization: graying and unifying the size of the face image, and performing pixel normalization to make the subsequent feature extraction model input consistent.
[0135] Visual feature extraction is completed based on the OpenFace toolkit, including the following three types of features: Facial Action Units (AUs): reflecting the activation of facial muscles, such as AU01 (Inner Brow Raised), AU12 (Smile), etc.; Eye movement trajectories and fixation points: estimating the line of sight direction and fixation distribution through eye movement modeling, used to measure the degree of attention concentration; Head pose: extracting three-dimensional head pose information such as pitch angle, yaw angle, and roll angle to quantify non-verbal behavior patterns.
[0136] The above multi-dimensional video features form a time series in chronological order and are standardized encoded, and further can be converted into visual "pseudo-text" for alignment with other modalities.
[0137] (3) Text data preprocessing and semantic feature extraction sub-module
[0138] Text data comes from speech-to-text results and self-reported text input. The preprocessing steps include: Removing stop words: Using a predefined mental health word list to剔除 common words with no semantic contribution (such as "of", "is", "then", etc.); Word segmentation and词性标注: Using natural language processing tools (such as jieba, HanLP, etc.) to perform Chinese text word segmentation and词性标注; Emotional word extraction and annotation: Matching the《Emotional Word Dictionary》and《Chinese Emotional Lexical Ontology》to identify emotional words, extracting positive, negative, and neutral words and their intensities; Syntactic structure analysis: Extracting basic syntactic patterns such as subject-predicate-object structure and interrogative sentence configuration to assist in modeling language organization ability.
[0139] Text features are further used to construct a structured language feature vector by statistically calculating indicators such as word frequency, TF-IDF, emotional density, sentence length, and negative usage frequency, and can also be converted into "text pseudo-coding" to participate in multi-modal embedding modeling.
[0140] 3. Multi-modal modeling and large prediction model fusion module
[0141] This module is used to perform unified vector quantization encoding and deep semantic modeling on the speech, visual, and text features extracted after preprocessing, mainly including the following two steps:
[0142] (1) Multi-modal feature fusion and embedding representation
[0143] Unify the input of audio pseudo-text, video pseudo-text, and text semantic features into the embedding structure to construct a multi-modal fusion representation. The embedding structure includes the following three levels:
[0144] Token embedding: divide the three types of modal features into pseudo-word units according to semantic units or time segments, and convert them into vector representations of fixed dimensions respectively; position embedding: add the position information of each pseudo-word unit in the sequence, maintain the time order, and support temporal modeling; paragraph embedding: add modal identifiers (such as "audio segment", "video segment", "text segment") to pseudo-texts from different modal sources to preserve the structural differences between modalities.
[0145] The fused embedding sequence represents the subject's psychological state characteristics in terms of language expression, facial behavior and voice features, and is the basis for subsequent semantic modeling and emotion recognition.
[0146] (2) Semantic modeling and fusion embedding output
[0147] The fused embedding sequence is used as input and passed to a large language model for contextual semantic modeling. This language model utilizes a multi-layer encoder structure based on the Transformer architecture, with the following features: a multi-head self-attention mechanism that captures long-range dependencies and contextual associations between multimodal pseudo-texts, extracting cross-modal interaction semantics; residual connections and layer normalization that enhance network stability and improve deep semantic modeling capabilities; and strong deep semantic expression capabilities that can recognize complex psychological cues, such as the matching of speech rhythm and expression, and the relativity of negative emotions and intonation.
[0148] The final output of the large language model is a multimodal fusion semantic embedding matrix, where each vector unit represents the psychological state of a multimodal semantic unit at different time periods. This semantic embedding matrix serves as the input feature of the subsequent classification model to complete the depression / non-depression prediction task.
[0149] 4. Classification model training and performance evaluation module
[0150] This module is used to input the semantic embedding representation after multimodal fusion into the classification model, build an intelligent recognition network for mental health status, and evaluate the performance of the model on the training set and external datasets.
[0151] (1) Model structure design
[0152] This embodiment uses CNN as the classification model architecture. The model structure includes the following four main parts:
[0153] Feature input layer: The input is the semantic embedding vector after multimodal fusion, which is normalized by Z-score. The input dimension is (N, 1), where N is the total dimension of the fused feature.
[0154] Multi-stage convolutional extraction module: This module consists of two or more cascaded one-dimensional convolutional layers, supporting multi-scale receptive field design. Each layer has a kernel size of {3, 5, 7}, a channel number of {64, 128, 256}, and uses a ReLU activation function. A max pooling layer (pool_size = 2) can be added after the convolutional layer to compress local information.
[0155] Fully connected layer: The convolution output is input into the fully connected network after global average pooling. The number of nodes is optimized in {16, 32, 64, 128} through grid search for high-order feature combination.
[0156] Classification output layer: Outputs the probability of depression classification through Sigmoid function. Figure 3 shown.
[0157] (2) Model training strategy
[0158] The model was trained using five-fold cross-validation, with the data divided into 80% training and 20% test sets. The following hyperparameters were optimized using grid search: convolution kernel size: {3, 5, 7}; number of filters: {64, 128, 256}; number of neurons in the fully connected layer: {16, 32, 64, 128}; learning rate: {0.001, 0.0005}. An early stopping mechanism was implemented during training, and training was terminated when the validation set AUC did not improve for five consecutive rounds.
[0159] The semantic embedding representation after multimodal fusion is used as the input feature of the model, and then the classifier variable (the classification output layer of the model) is encoded. The total score of the PHQ-9 scale <5 points is marked as 0 (i.e., no depression), and the total score of the PHQ-9 scale ≥5 points is marked as 1 (depression), thus establishing a depression recognition model.
[0160] (3) Performance evaluation indicators
[0161] This is implemented after the model is tested to evaluate the model's classification performance. After the model is classified, four indicators are obtained: TP, TN, FP, and FN. TP is the number of correctly predicted positive samples, TN is the number of correctly predicted negative samples, FP is the number of incorrectly predicted positive samples, and FN is the number of incorrectly predicted negative samples.
[0162] Based on the confusion matrix data, this embodiment uses five indicators to evaluate the generalization performance of the classifier, including the receiver operating characteristic curve (ROC), the area under the curve (AUC), accuracy, precision, recall, and F1 value. TP is the number of correctly predicted positive samples, TN is the number of correctly predicted negative samples, FP is the number of incorrectly predicted positive samples, and FN is the number of incorrectly predicted negative samples. See the deep learning process for details. Figure 5 .
[0163] 1) ROC and AUC. The ROC curve is a graph plotted with the true positive rate on the y-axis and the false positive rate on the x-axis. The diagonal line corresponds to the random guessing model, and the closer the curve is to the upper left corner, the better the classification performance. The ROC curve is a qualitative metric that is simple, intuitive, and highly readable. The more convex the ROC curve is and the closer it is to the upper left corner, the greater its diagnostic value. The AUC is a quantitative metric that is more accurate and objective. A larger AUC value indicates better classification performance.
[0164] 2) Accuracy: refers to the ratio of the number of correct predictions to the total number of actual predictions. It is the most widely used classification evaluation indicator.
[0165] Calculation formula:
[0166] 3) Precision P (Precision): The proportion of actual positive cases to the total predicted positive cases.
[0167] The calculation formula is:
[0168] 4) Recall (R) refers to the proportion of cases predicted to be positive that are actually positive.
[0169] Calculation formula:
[0170] 5) F1 score: It is the weighted average of precision and recall.
[0171] Calculation formula:
[0172] 5. External validation module of the model
[0173] This module is used to externally validate the established binary depression classification model. Using the same analytical and computational methods, the established binary depression classification model is validated against external datasets to evaluate the model's generalization performance. Cross-validation is performed on external datasets containing diverse population characteristics to evaluate the generalization performance of the depression recognition model. Specific steps include data cleaning and feature extraction, data normalization, and inputting external datasets into the model for prediction. Model performance is also evaluated using multiple metrics such as AUC, accuracy, precision, recall, and F1 score. Furthermore, considering the challenges posed by different data sources, the model's applicability and stability in different practical application scenarios are comprehensively evaluated.
[0174] In this example, a binary depression classification model was first constructed based on the multimodal integration of speech, video, and text, combined with deep learning and large language model technologies. External validation was then conducted on multi-center, heterogeneous datasets to verify the model's stability and adaptability across diverse populations and scenarios. This system, which integrates subjective scale information with objective behavioral data, exhibits high robustness and generalizability, enabling automated identification of individual psychological states and risk warnings, demonstrating promising practical value and widespread adoption.
[0175] Example 3
[0176] This example builds and validates a mental health assessment method based on a large sample of real-world student data, integrating multimodal features from speech, video, and text. This method integrates deep speech features, visual behavioral features, and language expression features. Through deep learning and a large language model, it achieves intelligent recognition of individual depressive states, demonstrating high accuracy, interpretability, and generalizability.
[0177] The method of this embodiment includes the following steps:
[0178] Step 1: Multimodal data collection and annotation
[0179] The PHQ-9 self-rating depression questionnaire (PHQ-9) was collected from 9,484 students at a school. Based on their scores, the students were divided into a healthy control group and a depressed group (the PHQ-9 score was ≥5). Multimodal data was also collected: voice data: Self-assessments were performed using a mobile phone in a quiet environment, with audio files recorded at an 8kHz sampling rate and approximately 3 minutes in length; video data: Frontal facial videos were collected simultaneously during the self-assessment process, with a frame rate ≥30fps and a resolution ≥720p; and text data: The voice transcriptions and subjective self-report text were saved for analysis.
[0180] Step 2: Multimodal preprocessing and feature extraction
[0181] Speech features: FFmpeg is used to complete format conversion, and OpenSMILE is used to extract the emoLarge feature set, including spectrum, sound source and prosody information. Statistics include mean, standard deviation, skewness, kurtosis, etc.; Video features: OpenFace is used to extract facial action units (AUs), head posture (pitch, yaw, roll) and eye movement trajectories to form structured temporal features; Text features: After automatic transcription using Whisper, semantic and emotional features are extracted through Jieba word segmentation, sentiment word tagging, word frequency statistics, etc.; All modal features are encoded as pseudo-text in a unified format and input into a unified embedding structure for vectorization processing.
[0182] Step 3: Multimodal Fusion and Large Language Model Modeling
[0183] The audio pseudo-text, video pseudo-text and text semantic features are uniformly input into the embedding structure to construct a multimodal fusion representation. The embedding structure consists of the following three levels:
[0184] Token embedding: divide the three types of modal features into pseudo-word units according to semantic units or time segments, and convert them into vector representations of fixed dimensions respectively; position embedding: add the position information of each pseudo-word unit in the sequence, maintain the time order, and support temporal modeling; paragraph embedding: add modal identifiers (such as "audio segment", "video segment", "text segment") to pseudo-texts from different modal sources to preserve the structural differences between modalities.
[0185] The fused embedding sequence is used as input and passed to the large language model for contextual semantic modeling. The language model adopts a multi-layer encoder structure based on the Transformer architecture, with the following features: multi-head self-attention mechanism: it can capture the long-distance dependencies and contextual associations between multimodal pseudo-texts and extract cross-modal interaction semantics; residual connection and layer normalization: it enhances network stability and improves deep semantic modeling capabilities; strong deep semantic expression capabilities: it can recognize complex psychological clues, such as the matching degree between speech rhythm and expression, the relativity between negative emotions and intonation, etc. The final output of the large language model is a multimodal fusion semantic embedding matrix, in which each vector unit represents the psychological state representation of different time periods and multimodal semantic units. This semantic embedding matrix is used as the input feature of the subsequent classification model to complete the depression / non-depression prediction task.
[0186] Step 4: Classification model construction and evaluation
[0187] A depression binary classification learning model was constructed using the CNN algorithm. The depression binary classification learning model includes:
[0188] Feature input layer: The input is the semantic embedding vector after multimodal fusion, which is normalized by Z-score. The input dimension is (N, 1), where N is the total dimension of the fused feature.
[0189] Multi-stage convolutional extraction module: This module consists of two or more cascaded one-dimensional convolutional layers, supporting multi-scale receptive field design. Each layer has a kernel size of {3, 5, 7}, a channel number of {64, 128, 256}, and uses a ReLU activation function. A max pooling layer (pool_size = 2) can be added after the convolutional layer to compress local information.
[0190] Fully connected layer: The convolution output is input into the fully connected network after global average pooling. The number of nodes is optimized in {16, 32, 64, 128} through grid search for high-order feature combination.
[0191] Classification output layer: Outputs the probability of depression classification through Sigmoid function. Figure 3 shown.
[0192] The model was trained using five-fold cross-validation. During the training process, the data was divided into 80% training set and 20% test set. The following hyperparameters were optimized through grid search: convolution kernel size: {3, 5, 7}; number of filters: {64, 128, 256}; number of neurons in the fully connected layer: {16, 32, 64, 128}; learning rate: {0.001, 0.0005}. An early stopping mechanism was set during training, and training was terminated when the AUC of the validation set did not improve for five consecutive rounds.
[0193] The semantic embedding representation after multimodal fusion is used as the input feature of the model, and then the classifier variable (the classification output layer of the model) is encoded. The total score of the PHQ-9 scale <5 points is marked as 0 (i.e., no depression), and the total score of the PHQ-9 scale ≥5 points is marked as 1 (depression), thus establishing a depression recognition model.
[0194] To ensure the reliability and stability of the model, reduce the risk of overfitting, and improve the generalization ability of the model, a five-fold cross-validation was used, dividing the data into an 80% training set and a 20% test set for internal validation of the model. True Positive (TP) refers to the number of positive cases predicted as positive, which in this embodiment refers to the number of individuals with actual depression classified as having a depressive tendency. True Negative (TN) refers to the number of negative cases predicted as negative, which in this embodiment refers to the number of individuals with actual non-depressed symptoms classified as non-depressed. False Positive (FP) refers to the number of negative cases predicted as positive, which in this study refers to the number of individuals with actual non-depressed symptoms classified as depressed. False Negative (FN) refers to the number of positive cases predicted as negative, which in this embodiment specifically refers to the number of individuals with actual depression classified as non-depressed. After establishing the classification model using the deep learning algorithm, the generalization performance of the classifier was evaluated using five indicators: receiver operating characteristic (ROC), area under the curve (AUC), accuracy, precision, recall, and F1 value based on the confusion matrix data.
[0195] 1) ROC and AUC. The ROC curve is a graph plotted with the true positive rate on the y-axis and the false positive rate on the x-axis. The diagonal line corresponds to the random guessing model, and the closer the curve is to the upper left corner, the better the classification performance. The ROC curve is a qualitative metric that is simple, intuitive, and highly readable. The more convex the ROC curve is and the closer it is to the upper left corner, the greater its diagnostic value. The AUC is a quantitative metric that is more accurate and objective. A larger AUC value indicates better classification performance.
[0196] 2) Accuracy: refers to the ratio of the number of correct predictions to the total number of actual predictions. It is the most widely used classification evaluation indicator.
[0197] Calculation formula:
[0198] 3) Precision P (Precision): The proportion of actual positive cases to the total predicted positive cases.
[0199] The calculation formula is:
[0200] 4) Recall (R) refers to the proportion of cases predicted to be positive that are actually positive.
[0201] Calculation formula:
[0202] 5) F1 score: It is the weighted average of precision and recall.
[0203] Calculation formula:
[0204] In the constructed binary classification model for depression, AUC = 0.917, accuracy = 0.812, precision = 0.837, recall = 0.800, and F1 score = 0.817. Figure 4 .
[0205] Step 5: External validation of the model
[0206] Independent samples were collected from another school, and the same method was used to complete data collection, preprocessing, feature extraction and model inference.
[0207] The results show: AUC = 0.881, accuracy = 0.751, precision = 0.759, recall = 0.779, F1 score = 0.767. Figure 5 .
[0208] This example builds a comprehensive multimodal mental health recognition model based on a large sample of real-world data. This model leverages the complementary strengths of speech, vision, and text features, achieving excellent accuracy, interpretability, and generalizability. This system provides a replicable and verifiable technical solution for intelligent student mental health management, with significant practical value and social impact.
[0209] Those skilled in the art will appreciate that embodiments of the present invention may be provided as methods, systems, or computer program products. Thus, the present invention may take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0210] The present invention is described with reference to flowcharts and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It should be understood that each process and / or block in the flowcharts and / or block diagrams, as well as combinations of processes and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions.
[0211] It should be noted that, in the claims, any reference signs placed between brackets shall not be construed as limiting the claims. The word "comprising" does not exclude the presence of components or steps not listed in the claims. The word "a" or "an" preceding a component does not exclude the presence of a plurality of such components. The invention may be implemented by means of hardware comprising several different components and by means of a suitably programmed computer. In claims enumerating several means, several of these means may be embodied by one and the same hardware. The use of the words first, second, third etc. is for convenience only and does not indicate any order. These words may be understood as part of the component name.
[0212] In addition, it should be noted that, in the description of this specification, the description of the terms "one embodiment", "some embodiments", "embodiment", "example", "specific example" or "some examples" means that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic expressions of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described can be combined in any one or more embodiments or examples in a suitable manner. In addition, those skilled in the art can combine and combine different embodiments or examples described in this specification and the features of different embodiments or examples, unless they are contradictory.
[0213] Although the preferred embodiments of the present invention have been described, those skilled in the art may make additional changes and modifications to these embodiments after learning the basic creative concept. Therefore, the claims should be interpreted as including the preferred embodiments and all changes and modifications that fall within the scope of the present invention.
[0214] Obviously, those skilled in the art may make various modifications and variations to the present invention without departing from the spirit and scope of the present invention. Thus, if such modifications and variations fall within the scope of the claims and their equivalents, the present invention shall also include such modifications and variations.
Claims
1. A mental health assessment system based on a large language model and multimodal data, characterized by: include: A data acquisition module is used to synchronously collect basic attribute information and depression assessment scale data of subjects through a programmatic interface, and synchronously collect audio data and video data of subjects during the assessment process, as well as text records; the subjects include subjects in a healthy state and subjects in a depressed state; A data preprocessing module, used to preprocess audio data, video data and text records respectively; The feature extraction module is used to extract the sound source, spectrum, and rhythm features from the pre-processed audio data and convert the sound source, spectrum, and rhythm features into audio pseudo-text; extract facial movements, eye movement trajectories, and head posture from the pre-processed video data and generate video pseudo-text; and perform speech transcription on the pre-processed audio data while combining it with text records for standardization and word segmentation to obtain text features; A multimodal fusion module, configured to input the audio pseudo-text, video pseudo-text, and text features into a unified embedding representation structure to obtain a fused multimodal vector; The large language model processing module is used to input multimodal vectors into the pre-trained large language model for contextual semantic modeling and output the fused semantic embedding matrix; The model training module is used to input the semantic embedding matrix into the pre-built depression binary classification model to train the model and use the trained model to evaluate the user's mental health.
2. The system according to claim 1, wherein: The data acquisition module includes: The first data collection unit is used to collect attribute information and total scores of depression assessment scales of a specified number of subjects online using a mini-program; A second data acquisition unit is used to collect audio data with a sampling frequency of 8 kHz; The third data acquisition unit is used to collect video data of the subject facing the camera during the self-narration process at a frame rate of not less than 30 frames per second. The fourth data collection unit is used to collect text records of the subjects' self-reports during the self-assessment process.
3. The system according to claim 2, characterized in that Data preprocessing module, including: A first data preprocessing unit compares the total score of the first data collection unit with a specified threshold, and determines the score greater than the specified threshold as a healthy control group, and the rest as a depression group; The second data preprocessing unit is used to convert the audio data into audio data in WAV format with a sampling frequency of 16kHz, and obtain preprocessed audio data through endpoint detection, pre-emphasis, framing, and windowing. a third data preprocessing unit, which processes the video data frame by frame, locates the facial area through face detection, performs face alignment and geometric normalization processing, and obtains preprocessed video data; The fourth data preprocessing unit performs stop word removal, word segmentation, part-of-speech tagging and normalized coding processing on the text self-description record to obtain a preprocessed text record.
4. The system according to claim 1, wherein: The feature extraction module includes: An audio feature extraction unit is used to call an audio processing tool to extract 523 acoustic features from the preprocessed audio data, including sound source features, spectrum features, and prosody features, and convert the sound source, spectrum, and prosody features into audio pseudo-text; The visual feature extraction unit is used to call the video analysis tool to process the video frames, extract facial movements, gaze trajectory and head posture features, and perform semantic description of the key frame features in the time series to generate video pseudo text; The text feature extraction unit is used to call the speech recognition tool to transcribe the audio content into text, and perform word segmentation, part-of-speech tagging, entity recognition and sentiment word extraction on the transcribed text and the pre-written text records to form text features.
5. The system according to claim 1, wherein: The multimodal fusion module includes: An embedding encoding unit is used to input the audio pseudo-text, video pseudo-text and text features into a unified embedding representation structure, wherein the embedding representation structure includes token embedding: converting features of different modalities into vector representations of unified dimensions, position embedding: adding temporal position information to each modal feature, and paragraph embedding: adding modality identifiers to features from different modal sources; thereby obtaining multimodal data; A temporal alignment unit, configured to align audio, visual, and text features in a temporal dimension using temporal position information in the position embedding, thereby updating multimodal data; A feature fusion unit is used to merge the updated multimodal data in the vector dimension. The merging method includes at least one of the following methods: feature splicing, weighted summation, or attention mechanism to achieve semantic fusion between modalities; The fusion output unit is used to output the semantically fused multimodal data as a fused multimodal vector to the large language model processing module.
6. The system according to claim 1, wherein: The large language model processing module includes: The context modeling unit receives the multimodal vectors output by the multimodal fusion module and inputs them into a pre-trained large language model for contextual semantic analysis. The large language model uses a deep neural network based on the Transformer structure to support long-range dependency modeling and cross-modal context capture. The embedding optimization unit is used to extract high-dimensional semantic features from the output processed by the large language model and generate a fused semantic embedding matrix. The semantic embedding matrix is used to express the subject's emotional and psychological state characteristics at the audio, visual and text levels.
7. The system according to claim 1, wherein: The model training module includes: The validation unit was used to train a binary depression classification model using 5-fold cross-validation. The training data was split into an 80% training set and a 20% test set. The following hyperparameters were optimized using grid search: convolution kernel size: {3, 5, 7}; number of filters: {64, 128, 256}; number of neurons in the fully connected layer: {16, 32, 64, 128}; learning rate: {0.001, 0.0005}. An early stopping mechanism was used during training, terminating the training if the AUC on the validation set did not improve for five consecutive rounds. Among them, five indicators were selected to evaluate the generalization performance of the depression binary classification model, including receiver operating characteristic curve (ROC), AUC, accuracy, precision, recall rate and F1 value.
8. The system according to any one of claims 1 to 7, characterized in that: The system further comprises: The prediction module is used to obtain the audio data, video data, and text records of the user to be analyzed during the evaluation process, and input them into the trained depression binary classification model through the data preprocessing module, feature extraction module, multimodal fusion module, and large language model processing module to obtain the mental health assessment results output by the model.
9. A mental health assessment method based on a large prediction model and multimodal data, characterized in that: include: Collecting basic attribute information and depression assessment scale data of the subjects, and simultaneously collecting audio data and video data, as well as text records, of the subjects during the assessment process; the subjects include healthy subjects and depressed subjects; Preprocess the audio data, video data and text records separately; Extract the sound source, spectrum and rhythm features from the preprocessed audio data, and convert the sound source, spectrum and rhythm features into audio pseudo-text; Extract facial movements, eye movement trajectories, and head postures from pre-processed video data and generate video pseudo-text; perform speech transcription on the pre-processed audio data and perform normalization and word segmentation on the text records to obtain text features; Inputting the audio pseudo-text, video pseudo-text and text features into a unified embedding representation structure to obtain a fused multimodal vector; Input the multimodal vector into the pre-trained large language model for contextual semantic modeling, and output the fused semantic embedding matrix; The semantic embedding matrix is input into a pre-built depression binary classification model to train the model, and the trained model is used to evaluate the user's mental health.
Citation Information
Cited By
User psychological state data monitoring method, system and terminal based on large language model
CN121528445A
Psychological state assessment method and system based on multi-modal data and model fine tuning
CN121658981A