Mental health evaluation and tutoring system based on multi-modal large model

By combining multimodal data collection and deep learning analysis with a psychological counseling module, the system addresses the shortcomings of existing mental health assessment systems in terms of accuracy and personalization, enabling more comprehensive mental health assessment and personalized intervention, and improving the system's real-time performance and accuracy.

CN121910367AInactive Publication Date: 2026-04-24BINGZHI TECHNOLOGY (SHENZHEN) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
BINGZHI TECHNOLOGY (SHENZHEN) CO LTD
Filing Date
2025-11-20
Publication Date
2026-04-24
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Existing mental health assessment and counseling systems based on multimodal large models suffer from insufficient accuracy, real-time performance, and personalization.

Method used

The system employs a multimodal data acquisition module to collect voice, text, facial expression, and behavioral data. Key features are extracted through a data preprocessing module, and in-depth analysis is performed using a large-scale model fusion and analysis module to generate a mental health status assessment report. Personalized suggestions are provided through a psychological counseling and intervention module, and the model is optimized based on user feedback.

Benefits of technology

It achieves a more comprehensive assessment of mental health status, improves the accuracy and reliability of the assessment, can detect potential mental health problems in advance and provide personalized interventions, reduces false alarm rates, and enhances the real-time nature and personalization of the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121910367A_ABST
    Figure CN121910367A_ABST
Patent Text Reader

Abstract

A psychological health assessment and tutoring system based on a multi-modal large model comprises a multi-modal data acquisition module used for a user to collect voice, text, facial expression and behavior data through a terminal equipment authorization system; the data preprocessing module is used for processing the multi-modal data and extracting key emotion and behavior characteristics; the large model fusion and analysis module is used for carrying out deep analysis on the multi-modal data to obtain a psychological health state evaluation report of the user; the psychological tutoring and intervention module is used for the tutoring module to generate personalized psychological health suggestions or intervention schemes, and the user can interact with the system through text and voice dialogue; and the user feedback and model self-optimization module is used for continuously optimizing the model by the system according to user feedback and subsequent use data and providing more accurate service. By integrating voice, text, facial expression and behavior data, more comprehensive mental health state assessment is provided, and the accuracy and reliability of assessment are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of mental health assessment and intervention, specifically involving a mental health assessment and counseling system based on a multimodal large model. Background Technology

[0002] Traditional mental health assessments often rely on regular interviews and questionnaires, which are inefficient, have limited accuracy, and struggle to continuously track changes in mental state. However, with the development of artificial intelligence and deep learning, mental health assessment systems based on multimodal data and large models have emerged. By analyzing users' language, behavior, and physiological characteristics, these systems can more comprehensively and accurately assess mental state and provide interventions. Nevertheless, existing technologies still suffer from limitations in accuracy, real-time performance, and personalization. Therefore, this invention aims to provide a system based on a multimodal large model to address the shortcomings of existing methods. Summary of the Invention

[0003] The purpose of this section is to outline some aspects of embodiments of the present invention and to briefly describe some preferred embodiments. Simplifications or omissions may be made in this section, as well as in the abstract and title of this application, to avoid obscuring the purpose of these documents; however, such simplifications or omissions should not be construed as limiting the scope of the invention.

[0004] In view of the problems of the above-mentioned or existing mental health assessment and counseling systems based on multimodal large models, this invention is proposed.

[0005] To solve the above-mentioned technical problems, the present invention provides the following technical solution: This invention provides a mental health assessment and counseling system based on a multimodal large model, including: a multimodal data acquisition module, used by the user to collect voice, text, facial expression and behavioral data through the authorization system of the terminal device; The data preprocessing module is used to process multimodal data and extract key sentiment and behavioral features; The large-scale model fusion and analysis module is used to perform in-depth analysis of multimodal data and generate a mental health status assessment report for users. The psychological counseling and intervention module is used by the counseling module to generate personalized mental health advice or intervention plans. Users can interact with the system through text and voice dialogue. The user feedback and model self-optimization module is used by the system to continuously optimize the model based on user feedback and subsequent usage data, so as to provide more accurate services.

[0006] As a preferred embodiment of the multimodal large-scale model-based mental health assessment and counseling system of the present invention, the multimodal data acquisition module includes: Voice data unit, used to collect the user's voice through a microphone or mobile phone.

[0007] Text data unit, used to collect user-input text, chat logs, and online conversations.

[0008] The facial expression data unit is used to capture users' facial expressions through a camera and analyze emotional changes.

[0009] Behavioral data unit, used to collect users' movement, usage habits and sleep records through wearable devices or mobile phone sensors.

[0010] As a preferred embodiment of the multimodal large model-based mental health assessment and counseling system described in this invention, the data preprocessing module includes: The speech data unit is used to convert speech into text using speech-to-text technology and extract emotional features from the speech. Text data units are used to perform word segmentation, part-of-speech tagging, and sentiment analysis using natural language processing techniques. The facial expression data unit is used to extract facial expression features and detect emotional states using computer vision technology. The behavioral data unit is used to extract behavioral characteristics that reflect mental health by analyzing users' daily behavior patterns.

[0011] As a preferred embodiment of the multimodal large-scale model-based mental health assessment and counseling system of the present invention, the step of using natural language processing technology for word segmentation, part-of-speech tagging, and sentiment analysis includes: Sentiment analysis is used to identify the sentiment tendency in text. It calculates the number of positive and negative words in the text using a predefined sentiment dictionary and uses a classifier support vector machine to train and predict the features. For dictionary-based sentiment analysis, assuming S is the sentiment score of the text, the expression is: Where wi is the sentiment weight of the i-th word, and ti is the number of times the word appears in the text.

[0012] As a preferred embodiment of the multimodal large model-based mental health assessment and counseling system of the present invention, the step of extracting facial expression features and detecting emotional state through computer vision technology includes: The face detection algorithm is used to locate the facial region, and the detected facial region is cropped out to reduce the amount of computation; the facial image is scaled and rotated to ensure that the facial features are in the same position and size. Convolutional neural networks are used to extract facial expression feature key point detection, facial key point detection algorithms are used to locate facial feature points, facial expression changes are calculated based on the detected key points, the extracted features are input into an emotion classification model, the model is trained using a labeled facial expression dataset, and the user's current emotional state is determined based on the output of the classification model.

[0013] As a preferred embodiment of the multimodal large-scale model-based mental health assessment and counseling system of the present invention, the large-scale model fusion and analysis module includes: A multimodal fusion unit is used to fuse speech, text, facial expression, and behavioral data using a deep learning model; The sentiment analysis and psychological assessment unit is used to analyze users' emotional and psychological states based on pre-trained models and identify potential mental health problems.

[0014] The behavior pattern prediction unit is used to analyze user behavior data using time series models to predict mental health risks.

[0015] As a preferred embodiment of the multimodal large-scale model-based mental health assessment and counseling system of the present invention, the step of using a time-series model to analyze user behavior data and predict mental health risks includes: The collected user behavior data is preprocessed to ensure the consistency and validity of the data input. The data is cleaned and then the behavior data, such as the number of social interactions and the degree of participation in activities, are normalized. Suppose user behavior data can be represented as a time series X={x1,x2,...,xT}, Where xT represents the behavioral characteristics at time step T; The system employs a time-series-based deep learning model, LSTM, to learn from historical data and establish a mapping relationship between behavioral characteristics and mental health status. The model's input is a time series X, and its output is a mental health risk score y. The model formula is as follows: Input gate: Controls how new input affects the hidden state. Among them, i t σ is the activation value of the input gate, σ is the sigmoid activation function, and W is the activation value of the input gate. i U is the input gate weight matrix. i Let b be the input gate weight matrix. i This is the bias term for the input gate; Forget Gate: Controls whether information from the previous state is retained. Among them, f t The activation value of the Forgotten Gate, W fU is the weight matrix of the forget gate. f Let b be the weight matrix of the forget gate. f For the bias term of the forget gate; Output gate: Determines the information to be output from the current hidden state. Where ot is the activation value of the output gate, W o U is the weight matrix of the output gate. o Let b be the weight matrix of the output gate. o This is the bias term for the output gate; Cell state update: Combining the forget gate and the input gate to update the state of memory cells. Among them, C t f represents the cell state at the current time step ttt. t *C t−1 For the memory portion of the previous time step determined by the forget gate, i t *tanh determines the addition of new memories based on the current input and the hidden state of the previous time step. tanh is the hyperbolic tangent function. W c U is the weight matrix for updating cell states. c is the weight matrix for cell state updates, and bc is the bias term for cell state updates; Hidden status update: Among them, h t Let tanh(Ct) be the hidden state at the current time step, and o be the cell state over time. t This is the activation value of the output gate; Output prediction: Among them, y t For model output, W y Let b be the weight matrix of the output layer. y For the bias term of the output layer; After training, the time-series model can predict and output a future mental health risk score y based on real-time user behavior data. t According to the set threshold T r The system assesses the user's mental health status: If y t >T r If this is the case, it indicates that the user's mental health is at high risk, and the system will trigger an alert. If y t ≤T r If the system continues to monitor, it will not trigger an alert.

[0016] As a preferred solution of the mental health assessment and counseling system based on the multimodal large model described in the present invention, wherein: the mental counseling and intervention module includes: Personalized recommendation generation unit: used to automatically generate personalized mental health recommendations based on the analysis results; Emotional guidance dialogue unit, used to create a virtual mental health assistant using dialogue generation technology to interactively counsel users and guide and relieve users' negative emotions; Mental health early warning unit, used to detect serious mental health problems of users and automatically trigger an early warning mechanism.

[0017] As a preferred solution of the mental health assessment and counseling system based on the multimodal large model described in the present invention, wherein: the detection of serious mental health problems of users and automatically triggering an early warning mechanism includes: Using sentiment analysis technology to evaluate the user's emotional state, generate an emotional score, set a threshold according to the user's behavior data, and set a threshold for the emotional score, set a triggering condition, and combine the emotional score and behavior indicators: Serious problem = (emotional score < Te) ∨ (behavior indicator < Tb) Where Te and Tb are the thresholds for emotion and behavior respectively; If it is detected that the user is young, the system lowers the emotional score threshold Te to avoid frequent false alarms; If it is detected that the user is old, the system sets a high emotional threshold Te to ensure the sensitivity of the early warning; If it is detected that the user is an extroverted culture user, the system sets a high emotional threshold Te; If it is detected that the user is an introverted culture user, the system lowers the emotional score threshold Te; By long-term collecting the user's emotional score data, the system determines the score fluctuation range and average emotional level of the user, so as to dynamically adjust the emotional threshold Te and behavior threshold Tb according to individual differences; For users with large emotional fluctuations, appropriately increase the emotional threshold Te to avoid frequent triggering of early warnings; For users with small emotional changes, appropriately lower the emotional threshold Te to enhance the sensitivity of the early warning; As time goes by, the user's emotional reactions and behavior patterns will change. The system uses a rolling window mechanism to continuously analyze the user's recent emotional scores and behavior data to detect the change trend of the score distribution; If the emotional score of the user continues to be low after experiencing long-term stress, the system will automatically lower the threshold Te so as to intervene earlier for intervention; If the user's emotional state gradually returns to normal, the system will also correspondingly increase the threshold Te to reduce over-warning; The system is configured to monitor both emotional and behavioral states simultaneously, and only triggers an alert when both indicators are abnormal, thus reducing the false alarm rate.

[0018] The beneficial effects of this invention are as follows: By integrating voice, text, facial expression, and behavioral data, this invention provides a more comprehensive assessment of mental health status. Compared to the analysis of single-modal data, multimodal fusion can more accurately capture users' emotional changes and mental health status, improving the accuracy and reliability of the assessment.

[0019] This system utilizes advanced time-series models (such as LSTM and Transformer) to analyze user behavioral data and predict future mental health risks. By capturing long-term trends and time dependencies in user behavior, the system can identify potential mental health problems early, enabling earlier intervention and warnings. Based on the analysis of users' multimodal data, the system can automatically generate personalized mental health counseling suggestions, combining cognitive behavioral therapy (CBT) and other psychological intervention techniques to help users improve their emotional state, provide emotional support, and regulate their mental health. By setting risk scoring thresholds, the system can automatically trigger an early warning mechanism when it detects serious mental health problems in users, reminding them to take further action and suggesting they seek professional help to prevent the mental health problems from worsening. Attached Figure Description

[0020] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the following description of the embodiments will be briefly introduced. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0021] Figure 1 This is a schematic diagram of the structure of a mental health assessment and counseling system based on a multimodal large model provided in an embodiment of the present invention. Detailed Implementation

[0022] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, the specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings.

[0023] Many specific details are set forth in the following description in order to provide a full understanding of the invention. However, the invention may also be practiced in other ways different from those described herein, and those skilled in the art can make similar extensions without departing from the spirit of the invention. Therefore, the invention is not limited to the specific embodiments disclosed below.

[0024] Secondly, the term "one embodiment" or "embodiment" as used herein refers to a specific feature, structure, or characteristic that may be included in at least one implementation of the present invention. The phrase "in one embodiment" appearing in different places in this specification does not necessarily refer to the same embodiment, nor is it a single or selective embodiment that is mutually exclusive with other embodiments.

[0025] Example 1 Reference Figure 1 This is one embodiment of the present invention, which provides a mental health assessment and counseling system based on a multimodal large model, including: S1: Multimodal data acquisition module, used by users to collect voice, text, facial expression and behavioral data through the authorization system of terminal devices.

[0026] Preferably, the multimodal data acquisition module includes: Voice data unit, used to collect the user's voice through a microphone or mobile phone.

[0027] Text data unit, used to collect user-input text, chat logs, and online conversations.

[0028] The facial expression data unit is used to capture users' facial expressions through a camera and analyze emotional changes.

[0029] Behavioral data unit, used to collect users' movement, usage habits and sleep records through wearable devices or mobile phone sensors.

[0030] Furthermore, by recording user facial expression videos from the camera, the system categorizes emotions into six types (such as happiness, sadness, anger, surprise, disgust, and calmness), uses a professional emotion recognition dataset for initial model training, and manually annotates the video data with annotation tools to determine the emotion category corresponding to each frame in the video.

[0031] The face detection algorithm is used to extract the user's facial region from video frames and remove background interference; Facial images were uniformly adjusted to 48×48 pixels to reduce the computational load on the model; color images were converted to grayscale to reduce data dimensionality and model computational complexity; images were rotated, flipped, and scaled to increase the diversity of data samples and avoid overfitting.

[0032] A CNN-LSTM hybrid model was constructed. The CNN part was used to extract local features from facial images, such as micro-expression features of eyebrows, eyes, and corners of the mouth. The LSTM part was used to capture the time-series features of expression changes. By combining information from previous and subsequent frames, the model was identified to identify the trend of changes in user emotions. The model was trained using public emotion recognition datasets and collected video data.

[0033] Each video segment is extracted and fed into a CNN-LSTM model; the cross-entropy loss function is used to calculate the error between the predicted sentiment category and the true sentiment category; the Adam optimizer is used to update the model weights. Train the model until the loss function converges, and then determine the model parameters through cross-validation.

[0034] The system captures user facial expression videos in real time via camera, automatically performs face detection and cropping, and inputs video segments per second into the model. The model generates an emotion score Se, which represents the probability distribution of emotion categories (e.g., [0.2, 0.1, 0.5, 0.1, 0.1] corresponding to happiness, sadness, anger, surprise, calmness, etc.). The category with the highest probability is taken as the current emotion label, and a real-time emotion score is generated.

[0035] S2: Data preprocessing module, used to process multimodal data and extract key sentiment and behavioral features.

[0036] Preferably, the data preprocessing module includes: The speech data unit is used to convert speech into text using speech-to-text technology and extract emotional features from the speech. Text data units are used to perform word segmentation, part-of-speech tagging, and sentiment analysis using natural language processing techniques. The facial expression data unit is used to extract facial expression features and detect emotional states using computer vision technology. The behavioral data unit is used to extract behavioral characteristics that reflect mental health by analyzing users' daily behavior patterns.

[0037] Preferably, sentiment analysis is used to identify the sentiment tendency in the text. By using a predefined sentiment dictionary, the number of positive and negative words in the text is calculated, and the features are trained and predicted using a classifier support vector machine. For dictionary-based sentiment analysis, assuming S is the sentiment score of the text, the expression is: Where wi is the sentiment weight of the i-th word, and ti is the number of times the word appears in the text.

[0038] Furthermore, the system needs to construct a predefined sentiment lexicon. The sentiment lexicon contains words representing positive and negative sentiments and their corresponding sentiment weights ωi, for example: Positive words: happiness, excitement, and bliss, with a positive weight ωi>0; Negative words: sadness, anger, and disappointment, with a negative weight ωi<0; For preprocessed words, the system matches the corresponding words and weights ωi from the sentiment dictionary, and then calculates the sentiment score S based on the frequency τi of the words.

[0039] Furthermore, assume that the emotion lexicon includes "happy" (ωhappy=+1.0) and "smooth sailing" (ωsmooth sailing=+0.8); Enter the text: "I am very happy today because things are going well." After word segmentation, the emotional words "happy" and "smoothly" were matched, with "happy" appearing once and "smoothly" appearing once. Calculate the sentiment score using the formula: S=(1.0×1)+(0.8×1)=1.0+0.8=1.8 The text has a positive sentiment score, indicating a positive sentiment tendency.

[0040] S3: Large Model Fusion and Analysis Module, used for in-depth analysis of multimodal data to generate a psychological health status assessment report for users.

[0041] Preferably, the large model fusion and analysis module includes: A multimodal fusion unit is used to fuse speech, text, facial expression, and behavioral data using a deep learning model; The sentiment analysis and psychological assessment unit is used to analyze users' emotional and psychological states based on pre-trained models and identify potential mental health problems.

[0042] The behavior pattern prediction unit is used to analyze user behavior data using time series models to predict mental health risks.

[0043] Preferably, the processed behavioral data is input into a time-series model for training. Psychological health risk labels from historical data are used as supervisory signals. The model predicts future psychological health risks by learning the mapping relationship between behavioral characteristics and psychological health status. After training, the model can predict users' future behavioral data and output a psychological health risk score. Based on the psychological health risk score output by the model, the system assesses the user's psychological health status. If the score is higher than a certain threshold, it indicates that the user faces a high psychological health risk. Combining the predicted psychological health risk score, the system sets a threshold and automatically triggers an early warning mechanism. The system will notify the user or relevant parties and recommend that the user take timely intervention measures.

[0044] Furthermore, the system records user behavior information, such as social interactions, work hours, exercise frequency, sleep quality, and dietary habits, through various sensors, user input, and application data. Common data sources include: wearable devices (recording heart rate, sleep duration, and exercise volume), mobile application data (social media usage time, message sending frequency), and questionnaires (self-reported mood, stress level). The system uses expert-assessed mental health status from historical data as labels, marking them as risk levels, such as: normal, mild risk, moderate risk, and high risk. Through training, the model can learn the mapping relationship between behavioral characteristics and mental health risk.

[0045] The model optimizes prediction accuracy by minimizing a loss function. Commonly used loss functions include mean squared error (MSE) or cross-entropy loss. The system divides the data into training and validation sets. The training set is used to update the model weights, and the validation set is used to evaluate model performance and prevent overfitting. After training, the model can predict users' future mental health risks based on new behavioral data. After the model is trained, the system can predict future mental health risks by inputting new behavioral data Xt and output a mental health risk score R. The score is usually set between 0 and 1, with the closer to 1 indicating a higher mental health risk. R=f(Xt) Where f(Xt) is the prediction function of the time series model for the input feature Xt; To determine whether a user faces serious mental health issues, the system sets a risk threshold θ based on a mental health risk score R. The threshold is set either by analyzing the distribution of historical data or based on recommendations from psychological experts. For example: If R > 0.8, it indicates that the user may face a higher risk to mental health.

[0046] If R ≤ 0.88, it indicates that the user's risk is low or in a normal state.

[0047] S4: Psychological Counseling and Intervention Module, used by the counseling module to generate personalized mental health advice or intervention plans. Users can interact with the system through text and voice dialogue.

[0048] Preferably, the psychological counseling and intervention module includes: Personalized suggestion generation unit: Used to automatically generate personalized mental health suggestions based on the analysis results; The Emotional Guidance Dialogue Unit is used to create a virtual mental health assistant using dialogue generation technology, which interacts with users to provide guidance and support for users' negative emotions. The mental health early warning unit is used to detect serious mental health problems in users and automatically trigger the early warning mechanism.

[0049] Preferably, using sentiment analysis technology, the emotional state of the user is evaluated to generate a sentiment score. Thresholds are set according to the user behavior data, and a threshold for the sentiment score is set, and a triggering condition is set. Combining the sentiment score and the behavior metrics: Serious problem = (sentiment score < Te) ∨ (behavior metric < Tb) where Te and Tb are the thresholds for sentiment and behavior respectively; If it is detected that the user is young, the system lowers the sentiment score threshold Te to avoid frequent false alarms; If it is detected that the user is old, the system sets a high sentiment threshold Te to ensure the sensitivity of the early warning; If it is detected that the user is an extroverted culture user, the system sets a high sentiment threshold Te; If it is detected that the user is an introverted culture user, the system lowers the sentiment score threshold Te; By collecting the user sentiment score data for a long time, the system determines the score fluctuation range and the average sentiment level of the user, and thus dynamically adjusts the sentiment threshold Te and the behavior threshold Tb according to individual differences; For users with large emotional fluctuations, the sentiment threshold Te is appropriately increased to avoid frequent triggering of early warnings; For users with small emotional changes, the sentiment threshold Te is appropriately lowered to enhance the sensitivity of the early warning; As time goes by, the emotional reactions and behavior patterns of users will change. The system uses a rolling window mechanism to continuously analyze the recent sentiment scores and behavior data of users to detect the change trend of their score distributions; If the sentiment score of the user remains low after experiencing long-term stress, the system will automatically lower the threshold Te to intervene earlier; If the emotional state of the user gradually returns to normal, the system will also correspondingly increase the threshold Te to reduce over-warning; The system is configured to monitor both the emotional and behavioral states simultaneously, and only triggers an early warning when both metrics are abnormal, reducing the false alarm rate.

[0050] Furthermore, the system normalizes the sentiment score of the user, sets its range between 0 and 1, where 0 represents negative emotion and 1 represents positive emotion. The normalized sentiment score S 情感 is a quantitative assessment of the user's current emotional state; The system aggregates the behavior data, extracts features and generates behavior metrics. Similar to the sentiment score, the behavior metrics are also processed by normalization, and their range is set between 0 and 1, where 0 represents abnormal behavior and 1 represents normal behavior; According to historical data and expert advice, the system sets two thresholds: Emotional score threshold Te: If the user's emotional score is lower than this threshold, it indicates that the user has emotional problems; Behavioral indicator threshold Tb: If the user's behavioral indicators are lower than this threshold, it indicates that the user's behavior pattern is abnormal; When the emotional score is lower than Te and the behavioral indicators are lower than Tb, the system determines that the user has serious emotional or behavioral problems.

[0051] Furthermore, Emotional score S 情感 : The emotional analysis scores of user A for 7 consecutive days are respectively: 0.8, 0.6, 0.7, 0.5, 0.3, 0.2, 0.4 (lower than the average level).

[0052] Behavioral data B: The number of social interactions of user A is respectively: 10 times, 8 times, 6 times, 5 times, 3 times, 2 times, 4 times (continuously decreasing).

[0053] Young user A: The system detects that he is young, so the emotional threshold is set lower, Te = 0.4, and the behavioral threshold: Tb = 5 (if the social interaction is less than 5 times, it indicates abnormal behavior).

[0054] Joint analysis of emotion and behavior: Data on the 6th day: Emotional score S 情感 = 0.2 (lower than Te = 0.4), and the behavioral indicator B = 2 (lower than Tb = 5).

[0055] Judgment: When both S 情感 < Te and B < Tb are satisfied, the system triggers an alarm.

[0056] Response: The system sends a reminder message to user A and recommends contacting mental health services S5: User feedback and model self-optimization module, which is used for the system to continuously optimize the model according to user feedback and subsequent usage data to provide more accurate services.

[0057] Preferably, the user feedback and model self-optimization module includes: Feedback collection unit, which is used to regularly collect user feedback on the counseling plan and optimize the counseling strategy; Model self-learning unit, which is used to continuously improve the model's understanding of user needs through reinforcement learning or transfer learning to improve the accuracy of evaluation and counseling.

[0058] Furthermore, the basic framework of reinforcement learning. The system models the understanding of user needs as a reinforcement learning problem, which includes the following elements: State S: The current state of the user as observed by the system at each point in time, including sentiment score, behavioral indicators, historical behavioral data, and user feedback; Action A: The system provides users with assessment, feedback, or coaching strategies, including emotion regulation suggestions, health reminders, and behavior adjustment suggestions; Reward R: The quality of user feedback and the improvement of user health status, used to evaluate whether the assessments and guidance provided by the system have a positive impact on users; At each time step t, the system obtains information from state St, executes action At, and receives a reward Rt based on user feedback, then enters the next state St+1. The system optimizes its understanding of user needs by accumulating long-term rewards.

[0059] Furthermore, through transfer learning, the system can adapt to different emotional and behavioral patterns across users and dynamically adjust the emotional threshold Te and behavioral threshold Tb based on individual differences among users. For example, for users with significant emotional fluctuations, the system may set a lower emotional threshold, thereby intervening earlier for emotional regulation. Conversely, for users with more severe behavioral abnormalities, the system will set a higher intensity of behavioral intervention.

[0060] As the system continues to collect more user data, the model will continuously improve its prediction accuracy and guidance effectiveness, for example: Prediction of Emotional Score: The system can predict the trend of a user's emotional changes based on the user's past behavioral patterns and current emotional state, and provide intervention suggestions in advance.

[0061] Prediction of abnormal behavior: By learning user behavior patterns, the system can predict potential abnormal behaviors (such as decreased activity levels or insufficient sleep) and issue timely warnings.

[0062] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the technical solutions of the present invention, and all such modifications or substitutions should be covered within the scope of the claims of the present invention.

Claims

1. A mental health assessment and counseling system based on a multimodal large model, characterized in that, include: The multimodal data acquisition module is used to collect voice, text, facial expression, and behavioral data from users through the authorization system on their terminal devices. The data preprocessing module is used to process multimodal data and extract key sentiment and behavioral features; The large-scale model fusion and analysis module is used to perform in-depth analysis of multimodal data and generate a mental health status assessment report for users. The psychological counseling and intervention module is used by the counseling module to generate personalized mental health advice or intervention plans. Users can interact with the system through text and voice dialogue. The user feedback and model self-optimization module is used by the system to continuously optimize the model based on user feedback and subsequent usage data, so as to provide more accurate services.

2. The mental health assessment and counseling system based on a multimodal large model as described in claim 1, characterized in that, The multimodal data acquisition module includes: Voice data unit, used to collect the user's voice through a microphone or mobile phone. Text data unit, used to collect user-input text, chat logs, and online conversations. The facial expression data unit is used to capture users' facial expressions through a camera and analyze emotional changes. Behavioral data unit, used to collect users' movement, usage habits and sleep records through wearable devices or mobile phone sensors.

3. The mental health assessment and counseling system based on a multimodal large model as described in claim 1, characterized in that, The data preprocessing module includes: The speech data unit is used to convert speech into text using speech-to-text technology and extract emotional features from the speech. Text data units are used to perform word segmentation, part-of-speech tagging, and sentiment analysis using natural language processing techniques. The facial expression data unit is used to extract facial expression features and detect emotional states using computer vision technology. The behavioral data unit is used to extract behavioral characteristics that reflect mental health by analyzing users' daily behavior patterns.

4. The mental health assessment and counseling system based on a multimodal large model as described in claim 3, characterized in that, The use of natural language processing techniques for word segmentation, part-of-speech tagging, and sentiment analysis includes: Sentiment analysis is used to identify the sentiment tendency in text. It calculates the number of positive and negative words in the text using a predefined sentiment dictionary and uses a classifier support vector machine to train and predict the features. For dictionary-based sentiment analysis, assuming S is the sentiment score of the text, the expression is: Where wi is the sentiment weight of the i-th word, and ti is the number of times the word appears in the text.

5. The mental health assessment and counseling system based on a multimodal large model as described in claim 1, characterized in that, The method of extracting facial expression features and detecting emotional state using computer vision technology includes: The face detection algorithm is used to locate the facial region, and the detected facial region is cropped out to reduce the amount of computation; the facial image is scaled and rotated to ensure that the facial features are in the same position and size. Convolutional neural networks are used to extract facial expression feature key point detection, facial key point detection algorithms are used to locate facial feature points, changes in facial expression are calculated based on the detected key points, and the extracted features are input into an emotion classification model. The model is trained using a labeled facial expression dataset, and the user's current emotional state is determined based on the output of the classification model.

6. The mental health assessment and counseling system based on a multimodal large model as described in claim 1, characterized in that, The large model fusion and analysis module includes: A multimodal fusion unit is used to fuse speech, text, facial expression, and behavioral data using a deep learning model; The sentiment analysis and psychological assessment unit is used to analyze users' emotional and psychological states based on pre-trained models and identify potential mental health problems. A behavior pattern prediction unit, which is used to analyze user behavior data using a time series model to predict mental health risks.

7. The mental health assessment and counseling system based on a multimodal large model as described in claim 1, characterized in that, The process of using a time series model to analyze user behavior data and predict mental health risks includes: Preprocessing the collected user behavior data to ensure the consistency and effectiveness of data input, cleaning the data, and then normalizing behavior data such as the number of social interactions and activity participation. Assume that the user behavior data can be represented as a time series X = {x1, x2,..., xT}, where xT represents the behavior characteristics at the T-th time step; The system uses a deep learning model LSTM based on time series analysis. By learning historical data, a mapping relationship between behavior characteristics and mental health status is established. The input of the model is the time series X, and the output is the mental health risk score y. The model formula is: Input gate: Controls how new inputs affect the hidden state Among them, i t σ is the activation value of the input gate, σ is the sigmoid activation function, and W is the activation value of the input gate. i U is the input gate weight matrix. i Let b be the input gate weight matrix. i This is the bias term for the input gate; Forget gate: Controls whether the information of the previous state is retained Among them, f t The activation value of the Forgotten Gate, W f U is the weight matrix of the forget gate. f Let b be the weight matrix of the forget gate. f For the bias term of the forget gate; Output gate: Determines the information output from the current hidden state Where ot is the activation value of the output gate, W o U is the weight matrix of the output gate. o Let b be the weight matrix of the output gate. o This is the bias term for the output gate; Cell state update: Combines the forget gate and the input gate to update the memory unit state Among them, C t f represents the cell state at the current time step ttt. t *C t−1 For the memory portion of the previous time step determined by the forget gate, i t *tanh determines the addition of new memories based on the current input and the hidden state of the previous time step. tanh is the hyperbolic tangent function. W c U is the weight matrix for updating cell states. c is the weight matrix for cell state updates, and bc is the bias term for cell state updates; Hidden state update: Among them, h t Let tanh(Ct) be the hidden state at the current time step, and o be the cell state over time. t This is the activation value of the output gate; Output prediction: Among them, y t For model output, W y Let b be the weight matrix of the output layer. y For the bias term of the output layer; After training, the time-series model can predict and output a future mental health risk score y based on real-time user behavior data. t According to the set threshold T r The system assesses the user's mental health status: If y t >T r If this is the case, it indicates that the user's mental health is at high risk, and the system will trigger an alert. If y t ≤T r If the system continues to monitor, it will not trigger an alert.

8. The mental health assessment and counseling system based on a multimodal large model as described in claim 1, characterized in that, The psychological counseling and intervention module includes: A personalized recommendation generation unit, which is used to automatically generate personalized mental health recommendations based on the analysis results; An emotional guidance dialogue unit, which is used to create a virtual mental health assistant using dialogue generation technology to interactively counsel users and guide and relieve users' negative emotions; A mental health warning unit, which is used to detect serious mental health problems of users and automatically trigger a warning mechanism.

9. The mental health assessment and counseling system based on a multimodal large model as described in claim 8, characterized in that, The process of detecting serious mental health problems of users and automatically triggering a warning mechanism includes: Using sentiment analysis technology to evaluate the user's emotional state, generating an emotional score, setting a threshold according to the user behavior data, and setting a threshold for the emotional score, setting a trigger condition, and combining the emotional score and behavior indicators: Serious problem = (emotional score < Te) ∨ (behavior indicator < Tb) where Te and Tb are the emotional and behavior thresholds respectively; If the detected user is young, the system lowers the emotional score threshold Te to avoid frequent false alarms; If the detected user is old, the system sets a high emotional threshold Te to ensure the sensitivity of the warning; If the detected user is an extroverted culture user, the system sets a high emotional threshold Te; If the detected user is an introverted culture user, the system lowers the emotional score threshold Te; By long-term collecting the user's emotional score data, the system determines the score fluctuation range and average emotional level of the user, and thus dynamically adjusts the emotional threshold Te and behavior threshold Tb according to individual differences; For users with large emotional fluctuations, appropriately increase the emotional threshold Te to avoid frequent triggering of warnings; For users with small emotional changes, appropriately lower the emotional threshold Te to enhance the sensitivity of the warning; Over time, the user's emotional responses and behavior patterns will change. The system uses a rolling window mechanism to continuously analyze the user's recent emotional scores and behavior data to detect the change trend of their score distribution; If the emotional score of the user remains low after experiencing long-term stress, the system will automatically lower the threshold Te to intervene earlier. If the user's emotional state gradually returns to normal, the system will also raise the threshold Te accordingly to reduce excessive warnings; The system is configured to monitor both emotional and behavioral states simultaneously, and only triggers an alert when both indicators are abnormal, thus reducing the false alarm rate.

10. The mental health assessment and counseling system based on a multimodal large model as described in claim 1, characterized in that, The user feedback and model self-optimization module includes: The feedback collection unit is used to regularly collect user feedback on the coaching plan and optimize the coaching strategy. The model self-learning unit is used to continuously improve the model's understanding of user needs through reinforcement learning or transfer learning, thereby increasing the accuracy of assessment and coaching.