Multimodal mental health intelligent assessment and intervention system and method

Through a multimodal intelligent mental health assessment system, combined with scenario simulation and multiple information feature analysis, the authenticity and participation issues of college students' mental health assessment are solved, and more accurate mental health assessment and interactive intervention are achieved.

CN119517413BActive Publication Date: 2025-10-03ZHEJIANG UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411720260.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-11-28
Publication Date
2025-10-03
Estimated Expiration
2044-11-28

AI Technical Summary

Technical Problem

The existing mental health assessment of college students relies on scale questionnaires, which have the problem of artificial falsification of answers, resulting in insufficient authenticity of the test conclusions. In addition, depression and anxiety among college students in psychological counseling rooms are still relatively common.

Method used

A multimodal mental health intelligent assessment system is used, combining scene simulation, image and voice feature extraction, analyzing facial expressions and voice information through convolutional neural networks, combining emotion recognition and intervention modules, providing interactive games and electronic cognitive behavioral therapy, for mental health assessment and intervention.

Benefits of technology

It improves the accuracy of mental health assessments, reduces students' resistance to psychological intervention, enhances participation, broadens application scenarios, and reduces dependence on professional psychologists.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119517413B_ABST
    Figure CN119517413B_ABST
Patent Text Reader

Abstract

The present invention discloses a multimodal intelligent mental health assessment and intervention system and method, wherein the system comprises: a scene simulation module, an image acquisition module, an image feature extraction module, a voice acquisition module, a voice feature extraction module, a mental health recognition module, an emotion recognition module, a picture selection module, an emotion correction module, a video teaching module, an emotion guidance module, an emotion matching module, and an alternative thinking guidance module. The multimodal intelligent mental health assessment and intervention system and method of the present invention, based on multimodal mental health assessment, makes the assessment no longer dependent on scale questionnaires and professional psychologists, broadens the application scenarios, and improves the accuracy of the assessment. At the same time, electronic cognitive behavioral therapy makes psychological intervention no longer boring, reduces students' resistance to psychological intervention, and improves student participation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention specifically relates to a multimodal mental health intelligent assessment and intervention system and method. Background Art

[0002] The mental health of college students is a hot topic at the moment. College students with good mental health can better cope with many pressures in life, study and work, explore their own potential, achieve their goals, and make contributions to society to the best of their ability.

[0003] Many factors can influence mental health. Family, learning, work, and social environments are all protective factors. However, even though many universities have established psychological counseling rooms, depression and anxiety caused by mental illness remain common among college students. The mental health of college students is a social issue that urgently requires solutions.

[0004] The most commonly used method for testing mental health is the scale questionnaire. Since the answers to the scale questionnaire can be falsified, the authenticity of the test conclusions needs to be further verified by the professional judgment of a doctor. Summary of the Invention

[0005] The present invention provides a multimodal mental health intelligent assessment and intervention system and method to solve the above-mentioned technical problems, specifically adopting the following technical solutions:

[0006] A multimodal mental health intelligent assessment and intervention system, comprising:

[0007] Scenario simulation module, used to simulate preset interactive scenarios;

[0008] Image acquisition module, used to collect facial video information of users during scene interaction;

[0009] An image feature extraction module, configured to extract facial image features from the facial video information collected by the image acquisition module;

[0010] The voice collection module is used to collect the user's voice information during the scene interaction process;

[0011] A speech feature extraction module, configured to extract speech features from the speech information collected by the speech collection module;

[0012] a mental health recognition module, configured to obtain a mental health recognition result based on the facial image features extracted by the image feature extraction module and the speech features extracted by the speech feature extraction module;

[0013] an emotion recognition module, configured to identify the emotion category of each image frame in the facial video information collected by the image collection module;

[0014] An image selection module is used to select multiple image frames with negative emotions from all image frames for users to identify;

[0015] An emotion correction module is used to display the selected image to the user, receive the emotion category identified by the user, and display the correct emotion category corresponding to the image frame to the user after the user selects the emotion category of the image frame;

[0016] A video teaching module, which is used to provide users with knowledge about the consequences of negative beliefs through video explanations;

[0017] The emotion guidance module is used to display a number of positive and negative emotion descriptions to the user, receive interruptions made by the user when a negative emotion description is identified, and score the received interruptions based on whether they are correct or timely;

[0018] The emotion matching module displays several negative emotions and corresponding questioning negative emotional thoughts to the user, receives the user's matching operation, and scores the received matching operation based on whether it is correct or not;

[0019] The alternative thinking guidance module is used to show the user several negative thinking modes, receive the user's alternative thinking modes selected for the negative thinking modes, and score the received alternative thinking modes based on whether they are correct or not.

[0020] Furthermore, the speech feature extraction module is used to extract Mel-frequency cepstral coefficients from the speech signal, and use a convolutional neural network to extract speech features from the Mel-frequency cepstral coefficients.

[0021] Furthermore, the image feature extraction module is used to decompose facial video information into multiple image frames, convert each image frame into a grayscale image, and use a convolutional neural network to extract facial image features of the image.

[0022] Furthermore, the mental health recognition module includes a trained multimodal recognition model, which receives the facial image features extracted by the image feature extraction module and the voice features extracted by the voice feature extraction module, and inputs the mental health recognition results.

[0023] Furthermore, the multimodal mental health intelligent assessment and intervention system further comprises:

[0024] The scenario selection module is used for managers to select and determine the interaction scenario.

[0025] Furthermore, in the process of the picture selection module selecting a plurality of image frames with negative emotions from all image frames for user identification, the time interval between the selected plurality of image frames with negative emotions is greater than a preset value.

[0026] Furthermore, the multimodal mental health intelligent assessment and intervention system further comprises:

[0027] The report output module is used to output the user's evaluation report.

[0028] A multimodal mental health intelligent assessment and intervention method, applied to the aforementioned multimodal mental health intelligent assessment and intervention system, comprising:

[0029] S1: simulating a preset interactive scene through the scene simulation module, collecting facial video information of the user during the scene interaction through the image acquisition module, extracting facial image features from the facial video information collected by the image acquisition module through the image feature extraction module, collecting voice information of the user during the scene interaction through the voice acquisition module, extracting voice features from the voice information collected by the voice acquisition module through the voice feature extraction module, and obtaining a mental health recognition result through the mental health recognition module based on the facial image features extracted by the image feature extraction module and the voice features extracted by the voice feature extraction module;

[0030] S2: identifying the emotion category of each image frame in the facial video information collected by the image acquisition module by the emotion recognition module, selecting a plurality of image frames with negative emotions from all the image frames for the user to identify by the image selection module, displaying the selected images to the user by the emotion correction module, receiving the emotion category identified by the user, and displaying the correct emotion category corresponding to the image frame to the user after the user selects the emotion category of the image frame;

[0031] S3: Outputting knowledge related to the consequences of negative beliefs to the user through video explanations in the form of video teaching modules;

[0032] S4: displaying a plurality of positive and negative emotional descriptions to the user through the emotion guidance module, receiving an interruption operation made by the user when a negative emotional description is identified, scoring the received interruption operation based on whether it is correct or timely, and completing this step when the score exceeds a preset value;

[0033] S5: Displaying a number of negative emotions and a number of corresponding questioning negative emotional thoughts to the user through the emotion matching module, receiving a matching operation from the user, scoring the received matching operation based on whether it is correct or not, and completing this step when the score exceeds a preset value;

[0034] S6: presenting several negative thinking patterns to the user through the alternative thinking guidance module, receiving the user's selected alternative thinking pattern for the negative thinking pattern, scoring the received alternative thinking pattern based on whether it is correct or not, and completing this step when the score exceeds a preset value;

[0035] S7: re-execute step S1 and output a new mental health identification result through the mental health identification module.

[0036] Furthermore, steps S1 to S7 are completed in stages at certain time intervals.

[0037] Furthermore, steps S1 to S7 are repeated twice to complete three cycles of testing.

[0038] The benefits of this invention lie in the multimodal intelligent mental health assessment and intervention system and method it provides. Based on multimodal mental health assessment, it eliminates the reliance on questionnaires and professional psychologists, broadens its application scenarios, and improves its accuracy. Furthermore, the combination of electronic cognitive behavioral therapy and interactive games makes psychological intervention less tedious, reduces student resistance to intervention, and increases student engagement. BRIEF DESCRIPTION OF THE DRAWINGS

[0039] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative labor.

[0040] Figure 1 Schematic diagram of the multimodal mental health intelligent assessment and intervention system of the present invention. DETAILED DESCRIPTION

[0041] The following describes in detail embodiments of the present application. Examples of the embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals throughout represent the same or similar elements or elements having the same or similar functions. The embodiments described below with reference to the accompanying drawings are exemplary and are intended to be used to explain the present application, and should not be construed as limiting the present application.

[0042] like Figure 1 Shown is a multimodal mental health intelligent assessment and intervention system of the present application, which includes: a scene simulation module, an image acquisition module, an image feature extraction module, a voice acquisition module, a voice feature extraction module, a mental health recognition module, an emotion recognition module, a picture selection module, an emotion correction module, a video teaching module, an emotion guidance module, an emotion matching module and an alternative thinking guidance module.

[0043] Specifically, the scenario simulation module is used to simulate preset interactive scenarios. Specifically, the multimodal mental health intelligent assessment and intervention system also has a scenario selection module, which is used for managers to select and determine interactive scenarios. In this application, users can be college students who receive multimodal mental health intelligent assessment and intervention, and managers are operators of the system. Administrator members need to select a scenario that may cause college students to feel anxious, such as an interview scenario. The scenario simulation module will provide a scene generated by the Unreal Engine, and users need to answer questions raised by the characters in the scene within the specified time limit and interact.

[0044] The image acquisition module is used to collect facial video information of the user during scene interaction. The voice acquisition module is used to collect voice information of the user during scene interaction, and the voice information includes voice intonation.

[0045] The image feature extraction module is used to extract facial image features from the facial video information collected by the image acquisition module. Specifically, the image feature extraction module is used to decompose the facial video information into multiple image frames, convert each image frame into a grayscale image, and use a convolutional neural network to extract facial image features from the image.

[0046] The speech feature extraction module is used to extract speech features from the speech information collected by the speech collection module. In an embodiment of the present application, the speech feature extraction module is used to extract Mel-frequency cepstral coefficients from the speech signal and use a convolutional neural network to extract speech features from the Mel-frequency cepstral coefficients.

[0047] The mental health recognition module is used to obtain a mental health recognition result based on the facial image features extracted by the image feature extraction module and the speech features extracted by the speech feature extraction module. In the present application, the mental health recognition module includes a trained multimodal recognition model. The multimodal recognition model receives the facial image features extracted by the image feature extraction module and the speech features extracted by the speech feature extraction module and inputs the mental health recognition result.

[0048] The emotion recognition module is used to identify the emotion category of each image frame in the facial video information collected by the image acquisition module. The image selection module is used to select multiple image frames with negative emotions (such as fear, anger, and disgust) from all image frames for the user to identify. The emotion correction module is used to display the selected image to the user, receive the emotion category identified by the user, and display the correct emotion category corresponding to the image frame to the user after the user selects the emotion category of the image frame. The user needs to determine which of the seven basic emotions (happiness, sadness, fear, anger, disgust, surprise, and contempt) the image frame contains, and then compare it with the results provided by the module. Through guidance, the user recalls the emotional reasons for the expression and the negative beliefs generated during the evaluation process.

[0049] It is understandable that consecutive image frames with negative emotions are relatively similar. To avoid such a situation, in an embodiment of the present application, when the image selection module selects multiple image frames with negative emotions from all image frames for user identification, the time interval between the selected multiple image frames with negative emotions is greater than a preset value.

[0050] The video teaching module uses video explanations to teach users about the consequences of negative beliefs. This interactive video explains the consequences of negative beliefs. Users learn to understand the consequences of negative beliefs through interactive simulations, taking on the perspective of an observer.

[0051] The emotion guidance module displays several positive and negative emotional descriptions to the user, receives interruptions from the user when a negative emotional description is identified, and scores the user based on the accuracy and timeliness of the interruption. Specifically, several positive and negative beliefs are randomly displayed in the form of text and images. The user is required to identify the negative belief within a short period of time and press the interrupt button in a timely manner. The module then receives the user's interruption. If the user interrupts untimely or incorrectly, the module will deduct points; otherwise, points will be awarded.

[0052] The emotion matching module displays a number of negative emotions and a number of corresponding questioning negative emotional thoughts to the user, receives the user's pairing operation, and scores according to whether the received pairing operation is correct or not. Specifically, in the present application, the emotion matching module interacts with the user in a game manner and receives the negative beliefs identified by the user. In the game scene, the user needs to click on furniture and objects that may contain clues, find all the information about negative beliefs (such as a note in a drawer saying "I am not experienced enough to do this well"), and its corresponding evidence of questioning negative beliefs (such as "I got full marks in last week's report" written on the back of the picture frame). Such information and evidence are regarded as a pair. When the user finds a pair of matching information and evidence, the module will prompt "One negative belief has been cleared", and the number of negative beliefs cleared and their total number will be displayed in the upper left corner of the screen.

[0053] The Alternative Thinking Guidance Module is used to present several negative thinking patterns to the user, receive user-selected alternative thinking patterns, and score the received alternative thinking patterns based on their accuracy. In an embodiment of the present application, the Alternative Thinking Guidance Module interacts with the user through a game, guiding the user to learn alternative thinking. Specifically, the game visualizes negative beliefs as different monsters, each representing a common negative thought (such as "I'm not good enough" or "No one likes me"). These monsters have different characteristics and attack methods, requiring the user to defeat them using positive thinking cards. Users can use positive thinking cards to attack the monsters. Each card features a specific alternative thinking pattern, expressed in concise and easy-to-understand text and images. The card design allows users to intuitively see how to replace negative thinking with positive, rational thinking when faced with negative thinking. When a user uses a positive thinking card to reduce the monster's health bar to zero, the user is considered to have defeated the negative thinking pattern. The module records whether the user has selected the correct alternative thinking card in each round of battle. A higher accuracy rate indicates that the user's mastery of alternative thinking is improving.

[0054] In the embodiments of this application, the multimodal mental health intelligent assessment and intervention system further includes a report output module. This module is used to output the user's assessment report. The system displays the assessment results in a visual report format, using colors and icons to distinguish changes, allowing users to clearly identify areas for improvement and areas requiring further effort.

[0055] Alternatively, the entire system could be ported to an embedded hardware device and developed into an app for iOS or Android, allowing users to implement mental health assessment and intervention in their daily lives. By using fun and serious games, users' resistance to mental health assessments could be reduced, improving assessment accuracy and intervention effectiveness.

[0056] This application also discloses a multimodal mental health intelligent assessment and intervention method, which is applied to the aforementioned multimodal mental health intelligent assessment and intervention system, specifically comprising:

[0057] Day 1: Use the scene simulation module to simulate the preset interactive scene, use the image acquisition module to collect the user's facial video information during the scene interaction process, use the image feature extraction module to extract facial image features from the facial video information collected by the image acquisition module, use the voice acquisition module to collect the user's voice information during the scene interaction process, use the voice feature extraction module to extract voice features from the voice information collected by the voice acquisition module, and use the mental health recognition module to obtain the mental health recognition result based on the facial image features extracted by the image feature extraction module and the voice features extracted by the voice feature extraction module.

[0058] Day 2: The emotion recognition module identifies the emotion category of each image frame in the facial video information collected by the image acquisition module. The picture selection module selects multiple image frames with negative emotions from all image frames for the user to identify. The emotion correction module displays the selected images to the user, receives the emotion category identified by the user, and displays the correct emotion category corresponding to the image frame to the user after the user selects the emotion category of the image frame.

[0059] Day 3: Through the video teaching module, users are provided with knowledge about the consequences of negative beliefs through video explanations.

[0060] Day 4: Use the emotion guidance module to present several positive and negative emotion descriptions to the user, receive interruptions from the user when a negative emotion description is identified, score the interruptions based on whether they are correct and timely, and complete the step when the score exceeds the preset value.

[0061] Day 5: Use the emotion matching module to present several negative emotions and corresponding questioning negative emotional thoughts to the user, receive the user's matching operation, score the received matching operation based on whether it is correct or not, and complete this step when the score exceeds the preset value.

[0062] Day 6: Use the alternative thinking guidance module to present several negative thinking patterns to the user, receive the user's selected alternative thinking patterns for the negative thinking patterns, score the received alternative thinking patterns based on whether they are correct or not, and complete this step when the score exceeds the preset value.

[0063] Day 7: Re-evaluate the assessment from Day 1 for a second evaluation, and output new mental health recognition results through the mental health recognition module. Specifically, users need to complete tasks in the same scenario, and the system will conduct a mental health assessment based on facial expressions and voice intonation. The system will then present a comparison of the assessment reports from Days 1 and 7. The system will display these comparison results in a visual report, using colors and icons to distinguish changes, allowing users to clearly see areas for improvement and areas that require further effort. At the same time, users will receive a system-wide loop completion reward badge and unlock a second loop and new assessment scenarios. Users will also be ranked within the server based on the number of days they check in. In addition, when a user's overall mental health score improves, the user will receive points based on the improvement score. Points can be redeemed within the system for mental health education resources (such as professional articles and relaxing audio) or virtual items (such as personalized avatars or themed backgrounds) to enhance the interactive experience.

[0064] The aforementioned seven days constitute one evaluation cycle. In an embodiment of the present application, preferably, the aforementioned seven-day cycle is repeated twice, completing three cycles of testing, for a total of 21 days. At the end of the third cycle, the user will receive a detailed evaluation report for comparison, which includes the results of each evaluation. According to the psychological 21-day effect, after three cycles, the user's mental health will be significantly improved.

[0065] The above shows and describes the basic principles, main features and advantages of the present invention. Those skilled in the art should understand that the above embodiments do not limit the present invention in any form, and any technical solutions obtained by equivalent replacement or equivalent transformation fall within the scope of protection of the present invention.

Claims

1. A multimodal mental health intelligent assessment and intervention system, characterized by: Include: Scenario simulation module, used to simulate preset interactive scenarios; Image acquisition module, used to collect facial video information of users during scene interaction; An image feature extraction module, configured to extract facial image features from the facial video information collected by the image acquisition module; The voice collection module is used to collect the user's voice information during the scene interaction process; A speech feature extraction module, configured to extract speech features from the speech information collected by the speech collection module; a mental health recognition module, configured to obtain a mental health recognition result based on the facial image features extracted by the image feature extraction module and the speech features extracted by the speech feature extraction module; an emotion recognition module, configured to identify the emotion category of each image frame in the facial video information collected by the image collection module; An image selection module is configured to select a plurality of image frames with negative emotions from all image frames for user identification, wherein the time interval between the selected plurality of image frames with negative emotions is greater than a preset value during the process of the image selection module selecting the plurality of image frames with negative emotions from all image frames for user identification; An emotion correction module is used to display the selected image frame to the user, receive the emotion category identified by the user, and display the correct emotion category corresponding to the image frame to the user after the user selects the emotion category of the image frame; A video teaching module, which is used to provide users with knowledge about the consequences of negative beliefs through video explanations; The emotion guidance module is used to display a number of positive and negative emotion descriptions to the user, receive interruptions made by the user when a negative emotion description is identified, and score the received interruptions based on whether they are correct or timely; The emotion matching module displays several negative emotions and corresponding questioning negative emotional thoughts to the user, receives the user's matching operation, and scores the received matching operation based on whether it is correct or not; The alternative thinking guidance module is used to show the user several negative thinking modes, receive the user's alternative thinking modes selected for the negative thinking modes, and score the received alternative thinking modes based on whether they are correct or not.

2. The multimodal mental health intelligent assessment and intervention system according to claim 1, characterized in that: The speech feature extraction module is used to extract Mel-frequency cepstral coefficients from the speech signal, and use a convolutional neural network to extract speech features from the Mel-frequency cepstral coefficients.

3. The multimodal mental health intelligent assessment and intervention system according to claim 1, characterized in that: The image feature extraction module is used to decompose facial video information into multiple image frames, convert each image frame into a grayscale image, and use a convolutional neural network to extract facial image features of the image.

4. The multimodal mental health intelligent assessment and intervention system according to claim 1, characterized in that: The mental health recognition module includes a trained multimodal recognition model, which receives the facial image features extracted by the image feature extraction module and the voice features extracted by the voice feature extraction module, and outputs a mental health recognition result.

5. The multimodal mental health intelligent assessment and intervention system according to claim 1, characterized in that: The multimodal mental health intelligent assessment and intervention system further comprises: The scenario selection module is used for managers to select and determine the interaction scenario.

6. The multimodal mental health intelligent assessment and intervention system according to claim 1, characterized in that: The multimodal mental health intelligent assessment and intervention system further comprises: The report output module is used to output the user's evaluation report.

7. A multimodal mental health intelligent assessment and intervention method, applied to the multimodal mental health intelligent assessment and intervention system according to any one of claims 1 to 6, characterized in that: Include: S1: simulating a preset interactive scene through the scene simulation module, collecting facial video information of the user during the scene interaction through the image acquisition module, extracting facial image features from the facial video information collected by the image acquisition module through the image feature extraction module, collecting voice information of the user during the scene interaction through the voice acquisition module, extracting voice features from the voice information collected by the voice acquisition module through the voice feature extraction module, and obtaining a mental health recognition result through the mental health recognition module based on the facial image features extracted by the image feature extraction module and the voice features extracted by the voice feature extraction module; S2: identifying, by the emotion recognition module, the emotion category of each image frame in the facial video information collected by the image acquisition module, selecting, by the picture selection module, a plurality of image frames with negative emotions from all the image frames for the user to identify; in the process of selecting, by the picture selection module, a plurality of image frames with negative emotions from all the image frames for the user to identify, a time interval between the selected plurality of image frames with negative emotions is greater than a preset value; displaying the selected image frames to the user by the emotion correction module, receiving the emotion category identified by the user, and displaying to the user the correct emotion category corresponding to the image frame after the user selects the emotion category of the image frame; S3: Outputting knowledge related to the consequences of negative beliefs to the user through video explanations in the form of video teaching modules; S4: Displaying a plurality of positive and negative emotional descriptions to the user through the emotion guidance module, receiving an interruption operation made by the user when a negative emotional description is identified, scoring the received interruption operation based on whether it is correct or timely, and completing this step when the score exceeds a preset value; S5: Displaying a number of negative emotions and a number of corresponding questioning negative emotional thoughts to the user through the emotion matching module, receiving a matching operation from the user, scoring the received matching operation based on whether it is correct or not, and completing this step when the score exceeds a preset value; S6: presenting several negative thinking patterns to the user through the alternative thinking guidance module, receiving the user's selected alternative thinking pattern for the negative thinking pattern, scoring the received alternative thinking pattern based on whether it is correct or not, and completing this step when the score exceeds a preset value; S7: re-execute step S1 and output a new mental health identification result through the mental health identification module.

8. The multimodal mental health intelligent assessment and intervention method according to claim 7, characterized in that: Steps S1 to S7 are completed in stages at certain time intervals.

9. The multimodal mental health intelligent assessment and intervention method according to claim 7, characterized in that: Repeat steps S1 to S7 twice to complete three cycles of testing.

Citation Information

Patent Citations

  • Multi-modal feature consistency mental health abnormity identification method and system

    CN116230234A

  • Multi-modal data-driven student psychological health education and monitoring system

    CN116665900A