Multi-mode mental health intelligent assessment and intervention system for old people
By integrating voice, video, physiological and semantic data, the multimodal mental health intelligent assessment and intervention system solves the problems of subjectivity in existing mental health assessment methods and lack of systematicity in electronic intervention. It enables real-time assessment and personalized intervention of mental health for the elderly, improving the accuracy of detection and user participation.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-29
- Publication Date
- 2026-04-07
AI Technical Summary
Existing mental health assessment methods are highly subjective and difficult to track continuously, while electronic intervention methods lack systematicity, interest, and personalization, making it difficult to motivate older adults to participate actively in the long term.
A multimodal intelligent assessment and intervention system for mental health is adopted, including a scenario module, a multimodal signal acquisition module, a unified fusion and assessment module, and an intelligent recommendation module. Through the multimodal fusion of voice, video, physiological and semantic data, real-time assessment and personalized intervention of mental state are carried out.
It improves the authenticity and accuracy of mental health testing, enables early assessment and personalized intervention of the mental health of the elderly, avoids user resistance, and provides an intelligent assessment platform that is social, fun, and emotionally adaptive.
Smart Images

Figure CN121812083A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of mental health assessment for the elderly, and specifically relates to a multimodal intelligent mental health assessment and intervention system for the elderly. Background Technology
[0002] With the accelerating aging of the population, mental health problems among the elderly are becoming increasingly prominent, such as loneliness, depression, and cognitive decline. According to the National Health Commission, by around 2035, China's population aged 60 and above will exceed 400 million, accounting for over 30% of the total population. Globally, 310 million people aged 65 and above suffer from depression, with nearly one-third of the elderly in China experiencing depression, and more than half experiencing significant loneliness.
[0003] Currently, common methods for assessing mental health still primarily rely on questionnaires and doctor interviews, which suffer from issues such as strong subjectivity, susceptibility to concealment, and difficulty in sustained follow-up. Existing electronic intervention methods are mostly information pushes or simple games, lacking systematicity, engagement, and personalization, making it difficult to motivate older adults to participate actively in the long term. Summary of the Invention
[0004] This invention provides a multimodal intelligent mental health assessment and intervention system for the elderly, addressing the technical problem of inaccurate traditional early warning methods mentioned above. The specific technical solution is as follows:
[0005] A multimodal intelligent mental health assessment and intervention system for the elderly includes:
[0006] The scenario module is used to allow users to make corresponding assessments and interventions.
[0007] The multimodal signal acquisition module is used to collect multimodal data of speech, video, physiological and semantic data when the user selects the corresponding scene module for intervention;
[0008] The scene module includes:
[0009] The cognitive-behavioral intervention module is used to identify and adjust negative automatic thoughts in elderly users;
[0010] The breathing training intervention module is used to guide and regulate the user's breathing rhythm and monitor the user's physiological information during the process.
[0011] The memory intervention module is used to create memory scenarios and guide users to recount positive experiences in their lives through recollection.
[0012] The group art intervention module is used to provide a painting scenario for the user and other participating users;
[0013] The unified fusion and evaluation module is used to perform temporal alignment, feature extraction, feature normalization, and feature fusion processing on the collected voice, video, physiological and semantic data. The data is then input into a pre-trained classification model and outputs a comprehensive mental health report and personalized intervention suggestions for the day.
[0014] The intelligent recommendation module recommends the order in which scenario modules can be selected based on the user's previous comprehensive mental health report.
[0015] Furthermore, the cognitive-behavioral intervention module includes a cognitive restructuring unit;
[0016] The cognitive restructuring unit includes: multiple negative thought interfaces containing negative thought information of the user;
[0017] Each negative thinking interface has multiple positive alternative thinking cards for users to choose from, which can guide users into a positive thinking mode.
[0018] The multimodal signal acquisition module records the user's reading speed, pitch, intonation changes, and facial emotions when the user selects multiple negative thought interfaces and corresponding positive alternative thought cards.
[0019] The unified fusion and assessment module performs feature fusion and psychological assessment based on the negative thinking interface selected by the user and the corresponding information collected by the multimodal signal acquisition module.
[0020] Furthermore, the breathing training intervention module includes:
[0021] A breathing rhythm guidance unit is used to help users maintain even breathing through recorded voice and visual animation;
[0022] The physiological monitoring unit is used to receive the user's heart rate variability (HRV), skin conductance response (GSR), and respiratory rate from the multimodal signal acquisition module;
[0023] The unified fusion and evaluation module records the information corresponding to the breathing rhythm guidance unit and the physiological monitoring unit on the time axis, performs feature fusion, and outputs the relaxation effect.
[0024] Furthermore, the recall intervention module includes:
[0025] The emotional narrative unit is used to allow users to tell their stories based on the nostalgic scene animations displayed, and to collect voice and facial expression information of users in real time during the narration process;
[0026] The semantic analysis unit is used to transcribe speech information into text, analyze the ratio of positive to negative words, and calculate the emotional stability of the user's narrative.
[0027] The unified fusion and evaluation module generates and records speech emotion curves, narrative duration, and semantic tendency scores.
[0028] Furthermore, the group arts intervention module includes:
[0029] The collaborative drawing unit is used to record the number of brush strokes, color usage, and contribution to the artwork by the user.
[0030] The group emotion analysis unit is used to collect the facial expressions and voice features of each user and calculate the consistency of group emotions. When the overall group emotion is positive, it automatically adjusts the scene lighting and plays soft music to create a relaxing atmosphere.
[0031] The unified integration and evaluation module records individual interaction frequency, number of times spoken, and average group sentiment deviation value.
[0032] Furthermore, the unified fusion and evaluation module encrypts and stores all collected multimodal data locally, and deletes video and audio data after feature extraction, retaining only the extracted feature vectors for subsequent model optimization.
[0033] Furthermore, the unified fusion and evaluation module uses the video timeline as a baseline to perform time alignment on speech, video, physiological and semantic speech frames, visual image frames, physiological signal windows and text paragraphs to ensure that features at the same moment correspond accurately.
[0034] For signals with low sampling frequencies, the unified fusion and evaluation module performs time synchronization on the signal through interpolation or aggregation.
[0035] The unified fusion and evaluation module standardizes the features of all time-aligned modalities to make them comparable under a unified dimension.
[0036] Furthermore, the unified fusion and evaluation module establishes a dedicated feature extraction module for each modality, mapping speech, video, physiological and semantic features to a unified feature space;
[0037] The unified fusion and evaluation module uses a multimodal fusion engine to weightedly integrate features within a unified feature space;
[0038] When fusing features, the multimodal fusion engine dynamically adjusts the weights of each modality to adaptively reflect the contribution of different psychological dimensions based on the module attributes.
[0039] Furthermore, the features fused by the multimodal fusion engine are input into the evaluation engine to calculate scores for five psychological dimensions: intensity of positive thinking, degree of breathing relaxation, emotional stability, positivity of memory, and social activity.
[0040] The unified integration and assessment module records the daily changes of five psychological dimension indicators and generates a comprehensive mental health report for the day;
[0041] The evaluation engine automatically adjusts the weights of each psychological dimension indicator based on the input expert experience information and historical samples, and optimizes the parameters of the multimodal fusion engine based on data from long-term user use.
[0042] Furthermore, the unified integration and evaluation module sets individualized baselines for users and tracks the trends of various psychological dimension indicators of users through a sliding time window;
[0043] When any psychological dimension indicator shows a decrease greater than the preset threshold over several consecutive days, or when the physiological signal deviates significantly from the individualized baseline, the unified fusion and evaluation module automatically identifies it as a potential abnormality.
[0044] When a potential anomaly occurs, the Unified Integration and Assessment module pauses the current scene task, plays a reassuring voice prompt, and simultaneously sends an alert to caregivers or family members.
[0045] The multimodal mental health intelligent assessment and intervention system for the elderly provided by this invention is an intelligent psychological assessment and intervention platform that integrates multimodal data collection, analysis and serious games. Through multi-dimensional data such as voice, image, physiological signals and semantics, combined with scenario intervention, it realizes real-time assessment and personalized intervention of mental state.
[0046] This invention can improve the authenticity and accuracy of mental health testing for the elderly, enabling early assessment and intervention. Through four intervention modules, it constructs an intelligent system platform with social, engaging, and emotionally adaptive features, achieving a closed loop of psychological assessment, intervention adjustment, and post-adjustment feedback reporting. Users can conduct mental health assessments within daily, enjoyable tasks, subtly accepting psychological interventions and avoiding resistance or defensiveness. Attached Figure Description
[0047] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0048] Figure 1 This is a schematic diagram of the modules of the multimodal mental health intelligent assessment and intervention system for the elderly proposed in this application;
[0049] Figure 2This is a schematic diagram of the overall process framework of the multimodal mental health intelligent assessment and intervention system for the elderly proposed in this application. It shows the overall data analysis process and the data processing model of the specific application of the unified fusion and assessment module.
[0050] Figure 3 This is a schematic diagram of the scenario modules of the multimodal mental health intelligent assessment and intervention system for the elderly proposed in this application;
[0051] Figure 4 This is a schematic diagram of the negative thinking interface of the cognitive behavior intervention module of the multimodal mental health intelligent assessment and intervention system for the elderly proposed in this application;
[0052] Figure 5 This is a schematic diagram of the positive alternative thought card of the cognitive behavior intervention module of the multimodal mental health intelligent assessment and intervention system for the elderly proposed in this application;
[0053] Figure 6 This is a schematic diagram of the interface of the breathing training intervention module of the multimodal mental health intelligent assessment and intervention system for the elderly in this application;
[0054] Figure 7 This is a schematic diagram of the interface of the memory intervention module of the multimodal mental health intelligent assessment and intervention system for the elderly proposed in this application;
[0055] Figure 8 This is a schematic diagram of the group art intervention module of the multimodal mental health intelligent assessment and intervention system for the elderly proposed in this application. Detailed Implementation
[0056] The embodiments of this application are described in detail below. Examples of the embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and intended to explain this application, and should not be construed as limiting this application.
[0057] like Figure 1 and Figure 2The illustration shows a multimodal mental health intelligent assessment and intervention system for the elderly, as described in this application. It includes a scenario module for users to conduct corresponding assessments and interventions. Specifically, the scenario module includes a cognitive behavioral intervention module, a breathing training intervention module, a memory intervention module, and a group art intervention module. The cognitive behavioral intervention module identifies and adjusts the negative automatic thoughts of elderly users; the breathing training intervention module guides and regulates the user's breathing rhythm and monitors the user's physiological information during this process; the memory intervention module establishes a memory scenario, guiding users to recount positive experiences in their lives; and the group art intervention module provides a scene for the user and other participating users to create art. On the system's display, the cognitive behavioral intervention module, breathing training intervention module, memory intervention module, and group art intervention module are named: Mind Nursery, Tranquil Lakeside, Memory Loft, and Co-creation Studio, respectively. The scenario module records the start and exit times of the cognitive behavioral intervention module, breathing training intervention module, memory intervention module, and group art intervention module, verifying whether all interactive content of the intervention modules has been completed, thereby ensuring the accuracy of the assessment and intervention.
[0058] The multimodal mental health intelligent assessment and intervention system for the elderly also includes a multimodal signal acquisition module. This module collects multimodal data on speech, video, physiological, and semantic aspects from users when they select appropriate scenario modules for intervention. Specifically, the multimodal signal acquisition module collects speech through a microphone and extracts acoustic features such as Mel-frequency cepstral coefficients; it collects facial expression and behavioral data through a camera and extracts visual features; it collects users' physiological information through wearable devices; and it extracts emotional and cognitive features by analyzing the semantic content of users' interactions.
[0059] The multimodal mental health intelligent assessment and intervention system for the elderly also includes a unified fusion and assessment module. This module performs temporal alignment, feature extraction, feature normalization, and feature fusion processing on the collected voice, video, physiological, and semantic data. The data is then input into a pre-trained classification model, outputting a comprehensive daily mental health report and personalized intervention suggestions. In other words, the unified fusion and assessment module integrates voice, facial expression, text, and behavioral data collected by the multimodal signal acquisition module, inputs it into a pre-trained classification model (such as SVM or neural networks), and outputs a mental health level assessment result and personalized intervention suggestions.
[0060] The multimodal mental health intelligent assessment and intervention system for the elderly also includes an intelligent recommendation module. This module recommends the order of scenario module selections based on the user's previous comprehensive mental health report; that is, it recommends a second intervention assessment based on personalized intervention suggestions, suggesting the user select scenario modules in a specific order.
[0061] As a specific implementation method, the cognitive behavioral intervention module (Thinking Nursery) guides older adults to identify and challenge negative thoughts based on cognitive behavioral intervention, and to learn alternative thoughts through interactive tasks. It includes a cognitive restructuring unit, which comprises multiple negative thought interfaces containing information about the user's negative thoughts. Each negative thought interface corresponds to multiple positive alternative thought cards for the user to choose from, guiding them into a positive thinking mode. A multimodal signal acquisition module records the user's speech rate, pitch, intonation changes, and facial expressions while reading, as the user selects multiple negative thought interfaces and corresponding positive alternative thought cards. A unified fusion and assessment module performs feature fusion and psychological assessment based on the negative thought interfaces selected by the user and the corresponding information collected by the multimodal signal acquisition module.
[0062] As a specific implementation method, the breathing training intervention module (Quiet Lakeside) guides elderly individuals in deep breathing exercises and assesses anxiety through physiological signal monitoring. It includes a breathing rhythm guidance unit and a physiological monitoring unit. The breathing rhythm guidance unit helps users maintain even breathing through recorded voice and visual animations, while the physiological monitoring unit receives data from wearable devices on the user's heart rate variability (HRV), skin conductance response (GSR), and respiratory rate. A unified fusion and assessment module records the corresponding information from the breathing rhythm guidance unit and the physiological monitoring unit on a timeline, thereby performing feature fusion and outputting the relaxation effect.
[0063] As a specific implementation method, the memory intervention module (Memory Loft) utilizes memory intervention to evoke positive emotions and alleviate loneliness and depression through media such as old photos and music. It includes an emotional narrative unit and a semantic analysis unit. The emotional narrative unit allows users to recount their experiences based on a constructed nostalgic scene animation, and collects voice and facial expression information in real time during the narration. The semantic analysis unit transcribes the voice information into text, analyzes the ratio of positive to negative words, and calculates the emotional stability of the user's narrative. The unified fusion and evaluation module generates and records the voice emotion curve, narrative duration, and semantic tendency score.
[0064] As a specific implementation method, the group art intervention module (co-creation studio) combines group art intervention to support multiple people collaborating to complete painting tasks. Through communication and artistic creation, it enhances social interaction and emotional expression. It includes a collaborative painting unit and a group emotion analysis unit. The collaborative painting unit records the number of brushstrokes, color usage, and contribution to the artwork. The group emotion analysis unit collects each user's facial expressions and voice characteristics to calculate group emotional consistency. When the overall group emotion leans towards positive, it automatically adjusts the scene lighting and plays soft music to create a relaxing atmosphere. The unified integration and evaluation module records individual interaction frequency, number of times spoken, and the average group emotion deviation value.
[0065] In the specific operational process:
[0066] The intelligent recommendation module guides users to complete the four psychological intervention modules in the scenario module in sequence, such as... Figure 3 As shown, after entering the system, the main interface displays four virtual scene entrances: "Mind Nursery," "Tranquil Lakeside," "Memory Attic," and "Co-creation Studio." Users can unlock modules randomly in sequence, and must complete all modules to end the day's task. The scene selection module records the start and exit times of each module to verify whether all interactive content has been completed. The module interface uses large icons and voice prompts to suit the operating habits of elderly users. Before entering a module, the system will prompt "Please complete the experience in a comfortable position" to ensure the stability of the collected data. The intelligent recommendation module automatically recommends scene order based on the user's previous psychological reports.
[0067] When a user selects the "Mind Garden" interface, they enter the virtual scene where three negative thoughts appear in the garden. The system uses voice and facial expression capture modules to record the user's speaking speed, pitch, intonation, and facial expressions (such as frowning and smiling) while reading. When a user clicks to enter a negative thought interface, they must select a "positive alternative thought card" to water, fertilize, and remove pests from the plant. Each card contains a positive alternative thought (such as "I can help others," or "My children are busy with work; I can proactively contact them"). Figure 4 and Figure 5 As shown, the system records the user's selected type, reaction time, voice energy, and improvement in facial expression in the background.
[0068] The system's backend records the number of positive thinking cards selected by the user, their reaction speed, and completion rate. Faster and more accurate selections indicate higher mental flexibility. The system also analyzes changes in the user's vocal energy and the degree of improvement in facial expressions during the positive thinking learning process to assess the effectiveness of cognitive restructuring.
[0069] When a user selects the "Tranquil Lakeside" option, the user enters the "Tranquil Lakeside" scene, such as... Figure 6 As shown, the ripples on the lake surface are synchronized with the system's voice prompts, guiding the user to perform breathing exercises according to a rhythm. The breathing rhythm guidance unit helps users maintain even breathing through voice and visual animation. The physiological monitoring unit collects heart rate variability (HRV), skin conductance response (GSR), and respiratory rate via wearable devices to assess the relaxation effect. The system records data in the background and does not display real-time indicators on the screen to avoid causing anxiety or distraction.
[0070] The system records the duration of breathing training, rhythm matching rate, and the degree of improvement in physiological indicators. If a user can stably complete rhythmic breathing for a relatively long period and their heart rate variability significantly improves, the system determines that their level of relaxation is high.
[0071] When the user selects the "Memory Loft" option, such as... Figure 7 As shown, users enter the "Memory Loft" virtual scene. The system guides users through a voice assistant to select historical photos provided by the system, or to recall old photos uploaded by their children. Through these recollections, users can recount positive life experiences, such as family time, work achievements, and dreams from their youth. The emotional narrative unit receives and collects voice and facial expressions in real time, displaying animations of relevant nostalgic items (such as old photo albums, phonographs, and diaries) to enhance the sense of context. The semantic analysis unit transcribes the speech into text, analyzes the ratio of positive to negative words, and calculates the emotional stability of the narrative. The system backend records the voice emotion curve, narrative duration, and semantic tendency score.
[0072] The system backend tracks the duration, fluency, and emotional fluctuations of user narratives. The semantic analysis module identifies the ratio of positive to negative words and calculates narrative stability. A user's ability to express themselves steadily for a relatively long period, primarily using positive words, indicates a positive recall experience.
[0073] When a user selects the co-creation studio's click interface, such as Figure 8 Users enter the "Co-creation Studio" scene and collaborate with other elderly users to complete paintings or collages. The collaborative painting unit records the number of brushstrokes, color usage, and contribution to the artwork. The group emotion analysis unit collects each user's facial expressions and voice characteristics through a camera and microphone, calculating the consistency of group emotions. When the overall group emotion leans towards positive, the system automatically adjusts the scene lighting or plays soft music to create a relaxing atmosphere. The system backend records individual interaction frequency, number of times spoken, and the average group emotion deviation value.
[0074] In addition to recording the duration of collaborative painting, the number of brushstrokes, and the frequency of communication, the system's backend also analyzes the richness, saturation, and diversity of colors in the artwork. Color richness is considered one of the indicators of emotional expression: paintings using a wide variety of colors with high saturation often reflect a more positive psychological state; while works using a single color with low saturation suggest low mood or insufficient motivation. The system also analyzes team collaboration parameters, such as the contribution of roles in the painting, group synchronicity, and the frequency of verbal communication, to assess social activity and emotional consistency.
[0075] In addition, the system simultaneously records and analyzes the usage duration and interaction depth of each module, including: the average dwell time and task completion time for each module; the number of clicks within the module, the duration of voice interaction, and the frequency of task interruptions; the distribution of continuous usage time and rest intervals. These indicators collectively reflect the user's level of psychological engagement and fatigue, and can serve as an important reference for emotional stability and participation.
[0076] After the user completes the four modules mentioned above, the system automatically launches the unified fusion and assessment module. This module performs temporal alignment and feature extraction on all visual, speech, physiological, and semantic data. After normalizing the features, it performs multimodal feature fusion to generate a comprehensive mental health report for the day. The report includes five main indicators: intensity of positive thinking, degree of breathing relaxation, emotional stability, positivity of memories, and social activity. The system generates a visual report for the day (such as a radar chart or change curve) for the user and caregiver to view. Based on the analysis results, the system generates task suggestions for the next day in the background, such as extending the time spent at the meditation lake, adding memory themes, or adjusting the number of participants in the drawing interaction. These suggestions are automatically pushed to the user the following day by the intelligent recommendation module and are not immediately prompted on the same day.
[0077] In other words, this system generates a personalized training plan for the next day based on daily psychological assessment results and long-term trends. For example, when the system detects a low score in positive thinking, it increases the frequency of the "Mind Nursery" module; when the relaxation index is insufficient, it extends the training time of "Tranquil Lakeside"; when recall positivity or social activity declines, the system automatically recommends the "Memory Loft" or "Co-creation Studio" modules to encourage users to participate in more emotional and social interactions. The recommended content is automatically updated in the background and presented naturally upon login the next day, avoiding interruption to the daily user experience of the elderly.
[0078] This system supports continuous training for 21 days, generating a periodic report every 7 days to observe psychological trends. During this process, the unified fusion and evaluation module locally encrypts and stores all collected multimodal data, and deletes video and audio data after feature extraction, retaining only the extracted feature vectors for subsequent model optimization. In other words, all collected data is locally encrypted and stored, and sensitive data (video and audio) is deleted immediately after feature extraction, retaining only the feature vectors, thus ensuring user privacy.
[0079] As a specific implementation scheme, the unified fusion and evaluation module uses the video timeline as a baseline to perform time alignment on speech, video, physiological and semantic speech frames, visual image frames, physiological signal windows, and text segments, ensuring accurate correspondence of features at the same moment. For signals with low sampling frequencies, the unified fusion and evaluation module performs time synchronization on these signals through interpolation or aggregation. The module then standardizes the features of all time-aligned modalities to make them comparable under a unified dimension. In other words, since different modalities have different data sampling frequencies, the system uses the video timeline as a reference to perform time alignment on speech frames, physiological signal windows, and text segments, ensuring accurate correspondence of features at the same moment. For signals with low sampling frequencies, the system synchronizes them through interpolation or aggregation. Subsequently, the features of all modalities are standardized to make them comparable under a unified dimension, reducing the impact of individual differences and extreme values.
[0080] Furthermore, the unified fusion and evaluation module establishes a dedicated feature extraction module for each modality, mapping speech, video, physiological, and semantic features to a unified feature space. The unified fusion and evaluation module then uses a multimodal fusion engine to weightedly integrate features within the unified feature space. During feature fusion, the multimodal fusion engine dynamically adjusts the weights of each modality to adaptively reflect the contributions of different psychological dimensions based on module attributes. For example, it increases the weight of physiological signals in "Quiet Lakeside" and emphasizes semantic and facial features in "Memory Attic," thus adaptively reflecting the contributions of different psychological dimensions according to module attributes.
[0081] Furthermore, the features fused by the multimodal fusion engine are input into the assessment engine to calculate scores for five psychological dimensions: intensity of positive thinking, degree of breathing relaxation, emotional stability, positivity of memory, and social activity. The unified fusion and assessment module records the daily changes of the five psychological dimension indicators and generates a comprehensive daily mental health report. Based on the input expert experience information and historical samples, the assessment engine automatically adjusts the weights of each psychological dimension indicator and optimizes the parameters of the multimodal fusion engine based on long-term user data. This gives the system continuous learning capabilities, ensuring that the overall score is optimized in real time to accurately reflect the true psychological state of elderly users.
[0082] As a specific implementation method, the unified fusion and assessment module sets an individualized baseline for each user and tracks the trends of various psychological indicators through a sliding time window. When any psychological indicator shows a decrease greater than a preset threshold over several consecutive days, or when physiological signals (such as a sharp drop in heart rate variability or a sharp increase in skin conductance) deviate significantly from the individualized baseline, the unified fusion and assessment module automatically identifies it as a potential anomaly. When a potential anomaly occurs, the unified fusion and assessment module pauses the current scenario task, plays a reassuring voice prompt, and simultaneously sends an alert to caregivers or family members. This ensures the safety of elderly users. All data is immediately de-identified and encrypted after analysis, retaining only statistical features for subsequent model optimization. In other words, the system has an anomaly detection function; when significant negative emotions or physiological stress responses (such as a sharp drop in HRV or a sharp increase in GSR) are detected, it can trigger a prompt to caregivers or pause the task to ensure the safety of elderly users. The system also supports multi-role login modes for families, elderly care institutions, and medical institutions, enabling data sharing and remote psychological support.
[0083] This solution provides a multimodal intelligent mental health assessment and intervention system for the elderly. It integrates multimodal data collection and analysis with a sophisticated gaming platform for intelligent psychological assessment and intervention. Through multi-dimensional data including voice, images, physiological signals, and semantics, combined with scenario-based intervention, it achieves real-time assessment and personalized intervention of mental state. This invention improves the authenticity and accuracy of mental health detection for the elderly, enabling early assessment and intervention. Through four intervention modules, it constructs an intelligent system platform with social, engaging, and emotionally adaptive features, achieving a closed loop of psychological assessment, intervention adjustment, and post-intervention feedback reporting. Users can conduct mental health assessments within daily, engaging tasks, subtly accepting psychological intervention and avoiding resistance or defensiveness.
[0084] This solution provides a multimodal mental health intelligent assessment and intervention system for the elderly, enabling non-invasive and continuous monitoring of mental state and avoiding subjective bias in scales. It improves assessment accuracy through multimodal data fusion and combines cognitive-behavioral intervention, recall intervention, and serious play intervention mechanisms involving group art intervention. Personalized intervention enhances assessment and adjustment effectiveness. The system is scalable and adaptable to elderly individuals with different cognitive and physical abilities.
[0085] In summary, the multimodal mental health intelligent assessment and intervention system for the elderly provided in this solution integrates multimodal data collection, analysis, and serious games into an intelligent psychological assessment and intervention platform. Through multi-dimensional data such as voice, images, physiological signals, and semantics, combined with scenario-based intervention, it achieves real-time assessment and personalized intervention of mental state. This invention can improve the authenticity and accuracy of mental health detection for the elderly, enabling early assessment and intervention. Through four intervention modules, it constructs an intelligent system platform with social, engaging, and emotionally adaptive features, achieving a closed loop of psychological assessment, intervention adjustment, and post-adjustment feedback reporting. Users can conduct mental health assessments within daily engaging tasks and subtly accept psychological intervention, avoiding resistance or defensiveness.
[0086] The foregoing has shown and described the basic principles, main features, and advantages of the present invention. Those skilled in the art should understand that the above embodiments do not limit the present invention in any way, and all technical solutions obtained by equivalent substitution or equivalent transformation fall within the protection scope of the present invention.
Claims
1. A multimodal intelligent mental health assessment and intervention system for the elderly, characterized in that, include: The scenario module is used to allow users to make corresponding assessments and interventions. The multimodal signal acquisition module is used to acquire multimodal data of speech, video, physiological and semantic data when the user selects the corresponding scene module for intervention; The scene module includes: The cognitive-behavioral intervention module is used to identify and adjust negative automatic thoughts in elderly users; The breathing training intervention module is used to guide and regulate the user's breathing rhythm and monitor the user's physiological information during the process. The memory intervention module is used to create memory scenarios and guide users to recount positive experiences in their lives through recollection. The group art intervention module is used to provide a painting scenario for the user and other participating users; The unified fusion and evaluation module is used to perform temporal alignment, feature extraction, feature normalization, and feature fusion processing on the collected voice, video, physiological and semantic data. The data is then input into a pre-trained classification model and outputs a comprehensive mental health report and personalized intervention suggestions for the day. The intelligent recommendation module recommends the order in which scenario modules can be selected based on the user's previous comprehensive mental health report.
2. The multimodal mental health intelligent assessment and intervention system for the elderly according to claim 1, characterized in that, The cognitive behavioral intervention module includes a cognitive restructuring unit; The cognitive reconstruction unit includes: multiple negative thought interfaces containing negative thought information of the user; Each of the aforementioned negative thinking interfaces corresponds to multiple positive alternative thinking cards for users to choose from in order to guide them into a positive thinking mode; The multimodal signal acquisition module records the user's reading speed, pitch, intonation changes, and facial emotions when the user selects multiple negative thought interfaces and corresponding positive alternative thought cards. The unified fusion and evaluation module performs feature fusion and psychological evaluation based on the negative thinking interface selected by the user and the corresponding information collected by the multimodal signal acquisition module.
3. The multimodal mental health intelligent assessment and intervention system for the elderly according to claim 1, characterized in that, The breathing training intervention module includes: A breathing rhythm guidance unit is used to help users maintain even breathing through recorded voice and visual animation; The physiological monitoring unit is used to receive the user's heart rate variability (HRV), skin conductance response (GSR), and respiratory rate collected by the multimodal signal acquisition module. The unified fusion and evaluation module records the information corresponding to the breathing rhythm guidance unit and the physiological monitoring unit on the time axis, performs feature fusion, and outputs the relaxation effect.
4. The multimodal mental health intelligent assessment and intervention system for the elderly according to claim 1, characterized in that, The memory intervention module includes: The emotional narrative unit is used to allow users to tell their stories based on the nostalgic scene animations displayed, and to collect voice and facial expression information of users in real time during the narration process; The semantic analysis unit is used to transcribe speech information into text, analyze the ratio of positive to negative words, and calculate the emotional stability of the user's narrative. The unified fusion and evaluation module generates and records the voice emotion curve, narrative duration, and semantic tendency score.
5. The multimodal mental health intelligent assessment and intervention system for the elderly according to claim 1, characterized in that, The group art intervention module includes: The collaborative drawing unit is used to record the number of brush strokes, color usage, and contribution to the artwork by the user. The group emotion analysis unit is used to collect the facial expressions and voice features of each user and calculate the consistency of group emotions. When the overall group emotion is positive, it automatically adjusts the scene lighting and plays soft music to create a relaxing atmosphere. The unified integration and evaluation module records individual interaction frequency, number of times spoken, and average group sentiment shift value.
6. The multimodal mental health intelligent assessment and intervention system for the elderly according to claim 1, characterized in that, The unified fusion and evaluation module encrypts and stores all collected multimodal data locally, and deletes video and audio data after feature extraction, retaining only the extracted feature vectors for subsequent model optimization.
7. The multimodal mental health intelligent assessment and intervention system for the elderly according to claim 1, characterized in that, The unified fusion and evaluation module uses the video timeline as a baseline to perform time alignment on speech, video, physiological and semantic speech frames, visual image frames, physiological signal windows and text paragraphs to ensure that features at the same moment correspond accurately. For signals with low sampling frequencies, the unified fusion and evaluation module performs time synchronization on the signal through interpolation or aggregation. The unified fusion and evaluation module standardizes the features of all time-aligned modalities to make them comparable under a unified dimension.
8. The multimodal mental health intelligent assessment and intervention system for the elderly according to claim 7, characterized in that, The unified fusion and evaluation module establishes a dedicated feature extraction module for each modality, mapping speech, video, physiological and semantic features to a unified feature space; The unified fusion and evaluation module uses a multimodal fusion engine to weightedly integrate features within the unified feature space; The multimodal fusion engine dynamically adjusts the weights of each modality when fusing features, so as to adaptively reflect the contribution of different psychological dimensions according to the module attributes.
9. The multimodal mental health intelligent assessment and intervention system for the elderly according to claim 8, characterized in that, The features fused by the multimodal fusion engine are input into the evaluation engine to calculate scores for five psychological dimensions: intensity of positive thinking, degree of breathing relaxation, emotional stability, positivity of memory, and social activity. The unified integration and assessment module records the daily changes of the five psychological dimension indicators and generates a comprehensive mental health report for the day. The evaluation engine automatically adjusts the weights of each psychological dimension indicator based on the input expert experience information and historical samples, and optimizes the parameters of the multimodal fusion engine based on data from long-term user use.
10. The multimodal mental health intelligent assessment and intervention system for the elderly according to claim 1, characterized in that, The unified integration and evaluation module sets individualized baselines for users and tracks the trends of various psychological dimension indicators of users through a sliding time window; When any psychological dimension indicator shows a decrease greater than a preset threshold over several consecutive days, or when the physiological signal deviates significantly from the individualized baseline, the unified fusion and evaluation module automatically identifies it as a potential anomaly. When a potential anomaly occurs, the unified fusion and evaluation module pauses the current scene task, plays a reassuring voice prompt, and simultaneously sends an alarm to caregivers or family members.