Intelligent evaluation and intervention system for whole process of mental health

By combining gamified general testing, multimodal screening, and adaptive intervention modules, the system achieves closed-loop management of the entire mental health assessment process, improving assessment accuracy and intervention response speed. It solves the problems of inaccurate assessment results and poor intervention effects in existing technologies and expands the system's deployment scenario coverage.

CN120932892APending Publication Date: 2025-11-11ZHONGKE XINHE (BEIJING) TECHNOLOGY CO LTD
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202511097612.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-06
Publication Date
2025-11-11

AI Technical Summary

Technical Problem

Existing mental health assessment systems suffer from problems such as limited assessment dimensions, insufficient data integration depth, lack of static intervention plans and real-time response capabilities, fragmented system functions, and lack of a closed-loop process. These issues lead to inaccurate assessment results and poor intervention effects, and the deployment scenarios and hardware integration limit the service coverage.

Method used

The system employs a gamified general testing module, a multimodal fine screening module, and an adaptive intervention module. Through multimodal data fusion and dynamic feedback mechanisms, it achieves closed-loop management throughout the entire process. The gamified general testing module collects multimodal data through various gamified tasks, the multimodal fine screening module performs refined data fusion analysis, and the adaptive intervention module adjusts the intervention plan based on real-time physiological data.

Benefits of technology

It has achieved a significant improvement in assessment accuracy and intervention response speed, built a closed-loop service chain for the entire process, adapted to individual differences and scenario changes, and solved the scenario coverage problem of traditional systems.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120932892A_ABST
    Figure CN120932892A_ABST
Patent Text Reader

Abstract

The invention discloses a psychological health full-process intelligent evaluation and intervention system, and mainly relates to the technical field of psychological health evaluation and intervention. Comprising a gamification general measurement module used for collecting multi-modal initial data of a user through multiple gamification tasks; the multi-modal fine screening module is connected with the gamification general survey module and is used for collecting multi-dimensional fine data of the user when the general survey risk score reaches a preset threshold value; the self-adaptive intervention module is connected with the multi-modal fine screening module and used for matching a corresponding intervention scheme according to the risk level and continuously collecting physiological data of the user in the intervention process to dynamically adjust the intervention scheme; and the data interaction and control module is connected with the gamification general measurement module, the multi-mode fine screening module and the self-adaptive intervention module and is used for realizing data transmission and cooperative control among the modules. The method has the beneficial effects that the evaluation precision is remarkably improved, and meanwhile, the intervention response speed is increased.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of mental health assessment and intervention technology, specifically a full-process intelligent assessment and intervention system based on multimodal data fusion and dynamic feedback mechanisms. Background Technology

[0002] With the surge in global demand for mental health services, digital mental health services have become a core trend in the industry. According to the World Health Organization, approximately one billion people worldwide are affected by mental health issues, while traditional offline psychological counseling suffers from limitations such as uneven resource distribution and low service efficiency. Against this backdrop, intelligent assessment and intervention systems based on artificial intelligence, the Internet of Things, and big data technologies have become a research hotspot. For example, China's "Anxin Ai" system has improved its assessment accuracy to 92% through multimodal data fusion (drawing, voice, video, etc.), and Tsinghua University's depression detection system achieves an individual detection accuracy of 0.909 through multimodal analysis of voice, text, and video. These technologies, by integrating multi-source data, have initially achieved dynamic assessment of mental states, but existing solutions still have significant shortcomings.

[0003] However, the existing technology still has the following problems: 1. Limited Assessment Dimensions and Insufficient Data Integration: Traditional mental health assessments primarily rely on self-report scales or single physiological signals (such as heart rate), making results susceptible to subjective bias or data noise. For example, in rural schools in central and western China, outdated equipment leads to an error rate exceeding 35%, and single EEG signal analysis cannot comprehensively reflect the dynamic changes in mental state. Furthermore, while existing systems attempt multimodal data collection (such as voice + text), they lack deep integration mechanisms. For instance, they merely perform simple weighted processing of multi-source data, failing to fully explore the semantic relationships and temporal features between the data. 2. Static Intervention Programs and Lack of Real-Time Response: Existing intervention systems typically have pre-set fixed programs (such as standardized cognitive behavioral therapy) and lack dynamic adjustments based on users' real-time physiological feedback. For example, a mental health app's breathing training module only guides breathing at a fixed frequency without optimizing training intensity based on real-time changes in the user's heart rate variability (HRV). This static design leads to significant differences in intervention effectiveness, and some users experience resistance due to program mismatch. 3. Fragmented System Functions and Lack of a Closed-Loop Process: Existing technologies generally suffer from a disconnect between the "assessment-intervention" stages. For example, some systems only generate reports after completing the assessment without automatically triggering the intervention module; or the intervention process lacks retrospective verification of the assessment results, leading to a break in the service chain. Furthermore, data silos are severe; data from different modules (such as general assessment, detailed screening, and intervention) do not interact in real time, making it difficult to form a closed-loop management system from risk identification to precise intervention.

[0004] The specific reasons for the above problems are as follows: 1. Limitations of Technical Architecture Design: Existing systems mostly adopt modular and independent designs, lacking cross-module collaboration mechanisms. For example, the data interfaces between the general testing module and the detailed screening module are incompatible, causing risk scoring to fail to automatically trigger the detailed screening process. Furthermore, data fusion algorithms are mostly based on traditional statistical models (such as random forests), making it difficult to handle the high-dimensional features and nonlinear relationships of multimodal data. 2. Lack of dynamic feedback mechanisms: Intervention programs often rely on expert experience and rules, failing to establish adaptive adjustment models based on real-time physiological data. For example, while a certain VR psychological training system can collect eye-tracking data, it does not integrate gaze patterns with EEG data. The correlation between wave power variations and the inability to accurately identify the user's cognitive load and adjust the training difficulty accordingly makes this static rule base design ill-suited to individual differences and changing scenarios. 3. Constraints on Deployment Scenarios and Hardware Integration: Existing systems are mostly deployed in fixed locations (such as medical institutions) and rely on specialized equipment (such as large-scale EEG monitoring devices), resulting in limited service coverage. For example, in remote areas, unstable network connections make it difficult to use cloud-based mental health apps. In addition, the size and cost of hardware devices limit their application in emergency scenarios (such as disaster relief).

[0005] Therefore, there is an urgent need for a full-process intelligent assessment and intervention system based on multimodal data fusion and dynamic feedback mechanisms to solve the above problems. Summary of the Invention

[0006] The purpose of this invention is to provide an intelligent assessment and intervention system for the entire process of mental health, which significantly improves assessment accuracy while enhancing intervention response speed.

[0007] To achieve the above objectives, the present invention employs the following technical solution: On the one hand, it provides a smart assessment and intervention system for the entire process of mental health, including: a gamified general assessment module, a multimodal fine screening module, an adaptive intervention module, and a data interaction and control module; The gamified general assessment module is used to collect multimodal initial data from users through various gamified tasks, and to process and analyze the multimodal initial data to generate a general assessment risk score; The multimodal fine screening module is connected to the gamified general assessment module. It is used to collect multi-dimensional fine data of users when the risk score of the general assessment reaches a preset threshold, and to perform fusion processing on the multi-dimensional fine data to output a psychological state quantitative report and risk level. The adaptive intervention module is connected to the multimodal screening module and is used to match the corresponding intervention plan according to the risk level, continuously collect the user's physiological data during the intervention process, and dynamically adjust the intervention plan according to the physiological data. The data interaction and control module is connected to the gamified general testing module, the multimodal fine screening module, and the adaptive intervention module, respectively, to realize data transmission and collaborative control between the modules, so as to complete the closed loop of the entire process from assessment to intervention.

[0008] Preferably, the gamified general testing module includes a first data acquisition unit, a first data processing unit, and a general testing result generation unit; The first data acquisition unit includes a camera, a microphone, and a touch screen. The camera is used to acquire facial image sequences when the user plays the "Emotional Matching" game, the microphone is used to acquire voice signals when the user plays the "Voice Story Relay" game, and the touch screen is used to record touch coordinate sequences when the user plays the "Attention Maze" game. The first data processing unit is connected to the first data acquisition unit and is used to process the facial image sequence using an improved convolutional neural network to obtain an emotion perception ability score, extract features from the speech signal using Mel frequency cepstral coefficients and process it through a classification model to obtain a speech emotion score, and perform trajectory smoothing and feature extraction on the touch coordinate sequence to obtain an attention concentration index. The general test result generation unit is connected to the first data processing unit and is used to perform fusion processing on the emotion perception ability score, voice emotion score and attention concentration index using a weighted voting model to generate the general test risk score.

[0009] Preferably, the improved convolutional neural network is a VGG16 convolutional neural network with an added attention mechanism module and an introduced temporal difference network. The attention mechanism module focuses on emotion-related muscle regions such as the corner of the eye and the zygomaticus major muscle, and the temporal difference network analyzes the dynamic changes of facial images in multiple consecutive frames. The classification model is a sentiment classifier built based on a Gaussian mixture model, and the model parameters are iteratively optimized through the EM algorithm. In the weighted voting model, the weight of the emotion perception ability score is 40%, the weight of the voice emotion score is 30%, and the weight of the attention concentration index is 30%.

[0010] Preferably, the multimodal fine screening module includes a second data acquisition unit, a second data processing unit, and a fine screening result output unit; The second data acquisition unit includes an EEG headband, a VR device, and an AI virtual assistant. The EEG headband is used to collect EEG signals from the user's prefrontal cortex. The VR device is used to simulate social scenarios and collect the user's eye movement data and physiological signals. The AI ​​virtual assistant is used to initiate structured interviews and collect the user's voice responses. The second data processing unit is connected to the second data acquisition unit and is used to preprocess and calculate the features of the EEG signal to obtain the wave power ratio, to perform synchronous analysis of the eye movement data and physiological signals to identify stress response, and to perform sentiment analysis on the text converted from the speech response to obtain the proportion of negative words. The fine screening result output unit is connected to the second data processing unit. Based on a random forest model trained from multi-source data, it processes the wave power ratio, stress response identification results, and negative word proportion, and outputs the psychological state quantitative report and risk level.

[0011] Preferably, the preprocessing of the EEG signal employs wavelet transform to decompose the signal in order to remove electromyographic and power frequency interference, and the feature calculation is achieved through power spectral density analysis; In the synchronous analysis of the eye movement data and physiological signals, the eye movement features include fixation duration and saccade amplitude, and the physiological signals include heart rate variability. Avoidance fixation patterns are identified by a hidden Markov model, and stress response is judged by combining the time-domain and frequency-domain indices of heart rate variability. The sentiment analysis of the text is based on a BERT pre-trained model. After word segmentation and stop word removal, word vectors are generated, negative words and their intensity are identified, and the proportion of negative words is calculated. At the same time, semantic dependency analysis is combined to identify the causal chain of sentiment.

[0012] Preferably, the adaptive intervention module includes an intervention plan matching unit and an intervention plan adjustment unit; The intervention plan matching unit is used to call the corresponding intervention plan from the preset rule base according to the risk level. The rule base contains intervention rules corresponding to three risk levels: low, medium and high. Each rule contains triggering conditions and execution parameters, and the execution parameters can be fine-tuned according to the user's age, gender and other demographic characteristics. The intervention program adjustment unit is connected to the intervention program matching unit and is used to collect the user's heart rate variability and brain wave power in real time during the intervention process, set a baseline value, and upgrade and adjust the intervention program when the increase in heart rate variability and the decrease in brain wave power within a preset time do not reach the preset threshold.

[0013] Preferred intervention options for low-risk cases include a combination of pet robot companionship and interaction with jujube seed tea; for medium-risk cases, the combination of VR breathing guidance training, lavender tea, and timed interaction with a pet robot; and for high-risk cases, the combination of low-intensity ultrasound stimulation and remote video intervention by a professional psychological counselor. The preset time is 5 minutes, and the preset threshold is a heart rate variability increase of ≥10% and an EEG power decrease of ≥15%.

[0014] Preferably, the data interaction and control module includes a data transmission unit and a collaborative control unit; The data transmission unit adopts edge computing technology to ensure that the delay from data acquisition to processing output is less than a preset value. At the same time, it realizes data interaction and remote consultation data transmission with the cloud through 4G / 5G network. The collaborative control unit is used to control the activation and switching of the gamified general testing module, the multimodal fine screening module, and the adaptive intervention module, so as to realize the linkage process of general testing triggering fine screening and fine screening results driving the matching of intervention schemes.

[0015] Preferably, the preset value is 2 seconds; When the risk assessment score is greater than or equal to 60, the collaborative control unit triggers the multimodal fine screening module to start. After the fine screening results are output, it immediately drives the adaptive intervention module to match the corresponding intervention plan.

[0016] Preferably, the system is deployed on a mobile vehicle, which adopts a modular design and is treated with shock absorption and sound insulation.

[0017] Compared with the prior art, the beneficial effects of the present invention are as follows: 1. Closed-loop management throughout the entire process: Through the seamless collaboration of three major modules—gamified general testing, multimodal fine screening, and adaptive intervention—a complete service chain of "assessment-intervention-feedback" is constructed to achieve dynamic tracking of risk levels and precise matching of intervention plans.

[0018] 2. Multimodal deep fusion: An improved convolutional neural network (including attention mechanism) is used to analyze facial expressions, combined with Mel frequency cepstral coefficients to extract speech emotion features, and a weighted voting model is used to achieve deep fusion of multi-source data, which significantly improves the accuracy of the evaluation.

[0019] 3. Dynamic adaptive intervention: The intensity of intervention is adjusted based on real-time HRV and EEG beta wave power. For example, the intervention plan is automatically upgraded when the user's stress response does not reach the preset threshold, forming a closed-loop control of "monitoring-evaluation-adjustment".

[0020] 4. Mobile and modular deployment: The system is integrated into a shock-absorbing and sound-insulating mobile vehicle, which can be deployed in a short time, effectively solving the problem of scene coverage of traditional systems. Attached Figure Description

[0021] Figure 1 This is a schematic diagram of the structure of the intelligent assessment and intervention system for the entire process of mental health in this invention. Detailed Implementation

[0022] The present invention will be further illustrated below with reference to specific embodiments. It should be understood that these embodiments are for illustrative purposes only and are not intended to limit the scope of the invention. Furthermore, it should be understood that after reading the teachings of this invention, those skilled in the art can make various alterations or modifications to the invention, and these equivalent forms also fall within the scope defined in this application.

[0023] like Figure 1 As shown, this embodiment provides a smart assessment and intervention system for the entire process of mental health, including: a gamified general assessment module, a multimodal fine screening module, an adaptive intervention module, and a data interaction and control module; The gamified general testing module is used to collect multimodal initial data from users through various gamified tasks, and to process and analyze the multimodal initial data to generate a general testing risk score; The multimodal fine screening module is connected to the gamified general testing module. It is used to collect multi-dimensional fine data of users when the risk score of the general testing reaches a preset threshold, and to perform fusion processing on the multi-dimensional fine data to output a psychological state quantitative report and risk level. The adaptive intervention module is connected to the multimodal screening module and is used to match the corresponding intervention plan according to the risk level, continuously collect the user's physiological data during the intervention process, and dynamically adjust the intervention plan based on the physiological data. The data interaction and control module is connected to the gamified general testing module, the multimodal fine screening module, and the adaptive intervention module, respectively, to realize data transmission and collaborative control between the modules, so as to complete the closed loop of the entire process from assessment to intervention.

[0024] The gamified general testing module includes a first data acquisition unit, a first data processing unit, and a general testing result generation unit. The first data acquisition unit includes a camera, a microphone, and a touch screen. The camera is used to acquire facial image sequences when the user plays the "Emotional Matching" game, the microphone is used to acquire voice signals when the user plays the "Voice Story Relay" game, and the touch screen is used to record touch coordinate sequences when the user plays the "Attention Maze" game. The first data processing unit is connected to the first data acquisition unit and is used to process the facial image sequence using an improved convolutional neural network to obtain an emotion perception ability score, extract features from the speech signal using Mel frequency cepstral coefficients and process it through a classification model to obtain a speech emotion score, and perform trajectory smoothing and feature extraction on the touch coordinate sequence to obtain an attention concentration index. The general test result generation unit is connected to the first data processing unit and is used to perform fusion processing on the emotion perception ability score, voice emotion score and attention concentration index using a weighted voting model to generate a general test risk score. The improved convolutional neural network is a VGG16 convolutional neural network with a newly added attention mechanism module and an introduced temporal difference network. The attention mechanism module focuses on emotion-related muscle regions such as the caudal muscles and zygomaticus major muscles, while the temporal difference network analyzes the dynamic changes of facial images in multiple consecutive frames. The classification model is a sentiment classifier built based on a Gaussian mixture model, and the model parameters are iteratively optimized using the EM algorithm; In the weighted voting model, the weight of emotion perception ability score is 40%, the weight of voice emotion score is 30%, and the weight of attention concentration index is 30%. The specific implementation of the game beta testing module includes the following steps: 1. First data acquisition unit: Acquisition of initial multimodal data: The first data acquisition unit collects users' physiological and behavioral data synchronously during gamified tasks using three types of hardware devices: Camera: In the "Emotion Matching" game, the camera captures a sequence of facial images of the user at a frame rate of 30fps, focusing on capturing the dynamic changes of emotion-related muscles such as the contraction of the corner muscles of the eyes (such as crow's feet when smiling) and the activity of the zygomaticus major muscle (such as the degree of upward movement of the corners of the mouth), providing raw data for subsequent expression recognition. Microphone: In the "Voice Story Relay" game, it collects 30 seconds of user voice signals and records features such as tone frequency (e.g., fluctuations of 250-500Hz in an anxious state) and pause duration (e.g., a percentage of >20% in a stressed state). The sampling accuracy ensures that it can distinguish between normal and abnormal voice patterns. Touchscreen: In the "Attention Maze" game, the user's touch coordinate sequence is recorded at a sampling rate of 100Hz to capture the movement trajectory of the virtual character (such as turning frequency and path offset), reflecting the user's executive function and level of attention concentration; 2. First data processing unit: Analysis and scoring of multimodal data: The first data processing unit performs stratified processing on the collected raw data and outputs three types of core scores. The specific process is as follows: Generation of emotion perception ability scores (corresponding to the "emotion matching" task): Data preprocessing: For the 30fps facial image sequence collected by the camera, infrared supplementary lighting is used to eliminate the interference of ambient light, ensuring stable image brightness (gray value fluctuation < 5%), providing high-quality input for subsequent feature extraction; Improved convolutional neural network processing: The "VGG16 convolutional neural network with a newly added attention mechanism module and temporal difference network" is adopted. The specific process is as follows: Attention mechanism module (SE module): Through adaptive weight allocation, the feature extraction weights of emotion-related regions such as the angular muscle and zygomatic major muscle are increased by 20% (compared with the traditional VGG16), suppressing the noise interference of irrelevant regions (such as the forehead and mandible), and enhancing the recognition of emotion features; Temporal difference network (TDN): Analyze the dynamic changes of three consecutive frames of images. For example, calculate the rate of the upward curvature of the corners of the mouth when smiling (pixel movement distance / time difference) and the amplitude change of the contraction of the glabella muscle when frowning, and output the probability distributions of six types of emotions: happiness, anger, sadness, fear, surprise, and neutrality (such as the probability of "happiness" is 85%, "neutrality" is 10%, etc.); Scoring quantization: Generate an emotion perception ability score (0 - 100 points) through two indicators: Emotion recognition delay: Calculate the time difference from image input to emotion probability output, requiring < 300ms (to ensure real-time performance); Cross-category confusion rate: Statistically calculate the proportion of cross-category errors such as misjudging "sadness" as "fear" and "anger" as "surprise", requiring < 15% (to ensure recognition accuracy); Comprehensive score: Based on the weighting of delay and confusion rate (delay weight 60%, confusion rate 40%), generate an emotion perception ability score (such as when the delay is 200ms and the confusion rate is 10%, the score is 85 points); Generation of speech emotion score (corresponding to the "speech story continuation" task): Data preprocessing: For the speech signal collected by the microphone, Wiener filtering is used to remove ambient noise (such as vehicle engine noise and external conversations), and the signal-to-noise ratio is increased to more than 60dB; then the effective speech segment is intercepted through the double-threshold method (energy threshold + zero-crossing rate threshold) (eliminating the silent and noisy parts and retaining the user's effective speech); Feature quantization: Extract three core features: Fundamental frequency (F0) fluctuation range: Calculate the difference between the maximum and minimum values of the fundamental frequency in the speech signal. The normal range is less than ±30Hz, and in an anxious state, it is greater than ±50Hz (such as if a user's fundamental frequency fluctuates by ±60Hz, it is determined as abnormal); Speech rate: Statistically calculate the number of syllables per second (syllables / second). The normal range is 6 - 8, and in a depressive state, it is < 4 (such as if a user's speech rate is 3 syllables / second, it is determined as a depressive tendency); Pause duration percentage: Calculate the ratio of the total pause duration to the total duration in a 30-second voice message. Normal is <10%, stress is >20% (e.g., if a user's pause duration is 25%, it is considered high stress).

[0025] Classification model processing: A limited "Gaussian mixture model (GMM) based emotion classifier" is adopted. The model parameters are iteratively optimized through the EM algorithm (the number of iterations is ≥50 to ensure convergence). The model distinguishes between three modes: "calm", "anxious" and "depressed", and outputs a voice emotion score (e.g., when the matching degree of "anxious" mode is 88%, the score is 70). Generation of attention concentration index (corresponding to the "Attention Maze" task): Data preprocessing: The 100Hz touch coordinate sequence collected from the touch screen is processed by Kalman filtering algorithm to remove hand shaking interference (filter window size is 5 sampling points) and the trajectory curve is smoothed (ensuring that the trajectory offset error is less than 1mm). Feature extraction: Calculate three types of core features: trajectory offset: the Euclidean distance between the touch coordinates and the optimal path in the maze, normally <5mm (if a user's average offset is 6mm, it is judged as distraction); Turning frequency: The number of turns per unit path length (meters). Normally <3 times / meter, but more than 5 times / meter when attention is distracted (e.g., if a user turns 6 times in a 1-meter path, it is considered as a lack of concentration). Time efficiency: The ratio of actual completion time to theoretical optimal time, normally less than 1.5 times (e.g., if a user's actual time is twice the optimal time, it is considered that the execution function is insufficient). Overall score: The "Analytic Hierarchy Process (AHP)" is used to calculate the attention concentration index from 0 to 100 points by weighting "offset 40% + turning frequency 30% + time efficiency 30%". 3. General survey result generation unit, weighted fusion and risk score output: The general assessment result generation unit adopts the "weighted voting model" as defined in claim 3, which integrates the above three types of scores to generate a general assessment risk score. The specific process is as follows: Weighting: Based on ROC curve analysis of 1000 clinical cases (true positive rate ≥85%, false positive rate ≤15%), the weights of emotion perception ability score (40%), voice emotion score (30%), and attention concentration index (30%) were determined to ensure high discrimination. Total score calculation: The three categories of scores (each from 0 to 100 points) are summed according to their weights, that is: The general risk assessment score is calculated as follows: Emotional perception ability score × 40% + Voice emotion score × 30% + Attention concentration index × 30%. Triggering condition: Set a threshold of 60 points (calculated by the Youden index, taking into account both sensitivity and specificity). When the total score is ≥60 points, the multimodal fine screening module (the multimodal fine screening module in claim 1) is triggered.

[0026] Efficiency control: By using edge computing technology (local data processing, avoiding cloud transmission delays), we ensure that the total delay from data collection to score output is less than 2 seconds, improving the user experience; 4. Improved VGG16 convolutional neural network: By adding an SE module (increasing the weight of emotion-related muscle regions by 20%) and TDN (analyzing dynamic changes in continuous frames), the problem of insufficient accuracy of traditional models in micro-expression recognition is solved (such as reducing the confusion rate between "sorrow" and "fear" from 25% to 12%).

[0027] GMM sentiment classifier: Improves speech pattern discrimination by optimizing parameters (such as mean and covariance matrix) through the EM algorithm (e.g., the recognition accuracy of "anxiety" and "depression" is increased from 80% to 88%).

[0028] Weighted voting model: The reliability of the risk score is ensured by weighting (40%, 30%, 30%) and setting a threshold (60 points) through clinical data validation.

[0029] The multimodal fine screening module includes a second data acquisition unit, a second data processing unit, and a fine screening result output unit; The second data acquisition unit includes an EEG headband, a VR device, and an AI virtual assistant. The EEG headband is used to collect EEG signals from the user's prefrontal cortex, the VR device is used to simulate social scenarios and collect the user's eye movement data and physiological signals, and the AI ​​virtual assistant is used to initiate structured interviews and collect the user's voice responses. The second data processing unit is connected to the second data acquisition unit and is used to preprocess and perform feature calculations on the EEG signals to obtain... Wave power ratio, synchronous analysis of eye movement data and physiological signals to identify stress response, and sentiment analysis of the text converted from speech response to obtain the proportion of negative words; The fine screening result output unit is connected to the second data processing unit, and uses a random forest model trained based on multi-source data to process... The wave power ratio, stress response identification results, and negative word ratio are processed to output a quantitative report of the psychological state and risk level. The preprocessing of EEG signals employs wavelet transform decomposition to remove electromyographic and power frequency interference, and feature calculation is achieved through power spectral density analysis. In the synchronous analysis of eye movement data and physiological signals, eye movement features include fixation duration and saccade amplitude, and physiological signals include heart rate variability. Avoidance fixation patterns are identified using a hidden Markov model, and stress responses are judged by combining time-domain and frequency-domain indices of heart rate variability. Sentiment analysis of text is based on a BERT pre-trained model. After word segmentation and stop word removal, word vectors are generated, negative words and their intensity are identified, and the proportion of negative words is calculated. At the same time, semantic dependency analysis is combined to identify the causal chain of sentiment. The multimodal fine screening module is implemented by following these steps: 1. Second data acquisition unit: Acquiring multi-dimensional and detailed data: The second data acquisition unit simultaneously collects users' physiological, behavioral, and linguistic data using three types of specialized equipment: EEG headband: Employing an 8-channel dry electrode design, positioned at FP1, FP2, F3, F4, C3, C4, P3, and P4 according to the international 10-20 system, it acquires prefrontal EEG signals at a sampling rate of 256Hz, with an impedance <10kΩ to ensure signal quality meets subsequent requirements. Wave power ratio calculation requirements; VR device: Using HTC ViveProEye, eye movement data (including fixation point and pupil diameter changes) is collected at 120Hz in simulated social scenarios (such as a 10-person conference room presentation). At the same time, heart rate and skin conductance data are collected at 10Hz through the built-in heart rate sensor and skin conductance response (GSR) module for stress response analysis. AI Virtual Assistant: Generates natural speech based on TTS technology (such as "Please describe something that makes you feel stressed"), and collects user voice responses through a dual-microphone array with a sampling rate of 44.1kHz and a dynamic range of 96dB to ensure clear and intelligible speech; 2. Second data processing unit: Multi-dimensional data fusion analysis: The second data processing unit performs multi-dimensional analysis on the collected raw data and outputs three types of core indicators. The specific process is as follows: Wave power ratio calculation (corresponding to EEG signal): (1) Data preprocessing: Bandpass filtering: Using a Butterworth 4th order filter, the 1-30Hz frequency band is preserved, and DC drift and high-frequency noise are removed; Independent component analysis (ICA): Separates electrooculogram artifacts (such as spikes produced by blinking) to ensure the purity of EEG signals (artifact residue rate <5%). (2) Feature extraction: Short-time Fourier Transform (STFT): Window length 0.5 seconds, overlap rate 50%, converts EEG signals to the frequency domain; Power spectral density (PSD) calculation: In Band (8-13Hz) and Integrating each band (13-30Hz) separately yields the following results: Wave power and Wave power ; Indicator Calculation: Wave power ratio equals The normal range is 0.8-1.2, and this ratio is usually <0.7 in an anxious state (e.g., a user). (Indicates an anxiety tendency) Stress response recognition (corresponding to eye movements and physiological signals in VR scenes): (1) Eye movement feature extraction: Fixation duration: The percentage of time spent fixing on threatening stimuli (such as a frowning face). Normally <30%, high anxiety >50%; Sagging amplitude: the average distance between adjacent fixation points, normally ranging from 3-8° of visual angle, which usually decreases under stress. (2) Physiological signal analysis: Heart rate variability (HRV): Calculate the standard deviation (SDNN) of adjacent RR intervals. The normal range is 50-100 ms, and <40 ms under stress. Skin conductance response (SCR): Calculate the rate of change in conductivity; normal baseline fluctuation <0.5 μS, stress condition >1.0 μS. Hidden Markov Model (HMM) Recognition: State definitions: including "relaxed" (state 1), "mild stress" (state 2), and "high stress" (state 3); Training parameters: The HMM was trained based on 100 clinical cases, with the transition probability matrix A, emission probability matrix B, and initial state probability π set. Stress determination: When the probability of output state 3 is >70% and the duration is >30 seconds, it is determined that there is a significant stress response; Negative word percentage calculation (corresponding to voice responses collected by the AI ​​virtual assistant): Speech-to-text conversion: Using the Wav2Vec2.0 model, with a word error rate (WER) of <5%, the user's speech is converted into text; (1) Text preprocessing: Word segmentation: Using the jieba word segmentation tool, combined with a mental health dictionary (containing 5000+ negative emotion words, such as "despair" and "collapse"); Stop word removal: Remove function words without emotional tendency such as "的" and "了". (2)Sentiment analysis: Word vector generation: Use the BERT pre-trained model to convert the tokenized text into 12×64-dimensional word vectors. Negative word recognition: Match negative emotion words in the domain dictionary through cosine similarity calculation (threshold 0.85). Ratio calculation: Number of negative words / Total number of words × 100%. For example, if the total number of words in a user's answer is 100 and the number of negative words is 15, the ratio is 15%. 3. Fine screening result output unit, risk level determination: The fine screening result output unit performs fusion processing on the above three types of indicators based on the random forest model, and outputs a psychological state quantification report and risk level: Feature vectorization: Combine the wave power ratio (continuous value), stress response recognition result (0-1 binary), and negative word ratio (percentage) into a three-dimensional feature vector [0.6, 1, 15%].

[0030] Random forest model: Training parameters: Trained based on 500 cases of clinical data, with 100 trees, a maximum depth of 5 layers, and a minimum sample split number of 5. Risk level mapping: Low risk (level 1): Prediction probability < 30%, corresponding feature vector range (α / β > 0.8, stress response = 0, negative word ratio < 10%); Medium risk (level 2): Prediction probability 30% - 70%, corresponding feature vector range (α / β ∈ [0.6, 0.8], stress response = 1, negative word ratio ∈ [10%, 20%]); High risk (level 3): Prediction probability > 70%, corresponding feature vector range (α / β < 0.6, stress response = 1, negative word ratio > 20%); Quantification report generation: Output a structured report including scores for each dimension (such as 70 points for the EEG dimension, 85 points for the stress response dimension, and 60 points for the language dimension), comprehensive risk level (such as "medium risk"), and recommended intervention measures (such as "recommended VR breathing training"). EEG signal processing: Wavelet transform decomposition: Use the db4 wavelet basis to decompose the signal into 5 layers, corresponding to the δ (1-4Hz), θ (4-8Hz), α (8-13Hz), β (13-30Hz), and γ (30-50Hz) frequency bands respectively; Power spectral density calculation: Use the Welch method, with a window length of 0.5 seconds and an overlap rate of 50% to ensure a frequency domain resolution < 2Hz; Eye movement and physiological signal synchronization analysis: Time calibration: Ensure eye-tracking data is aligned with physiological signal timestamps through hardware clock synchronization (error <1ms); Multi-feature fusion: The fixation duration, saccade amplitude, RMSSD index of HRV (root mean square of the difference between adjacent RR intervals), and SCR change rate are input into HMM, and the parameters are iteratively optimized through the Baum-Welch algorithm (convergence threshold 1e-6). Text sentiment analysis: BERT pre-training: Fine-tuning the BERT-base-chinese model using a mental health corpus (containing 100,000 psychological counseling dialogues) to improve the recognition accuracy of negative emotion words; Semantic dependency analysis: Using the HanLP tool, dependency syntactic trees of sentences are constructed to identify emotional causal chains (such as "because of high work pressure (reason), I feel hopeless (result)"), providing a targeted basis for intervention programs.

[0031] The adaptive intervention module is implemented by following these steps: The adaptive intervention module includes an intervention program matching unit and an intervention program adjustment unit; The intervention plan matching unit is used to call the corresponding intervention plan from the preset rule base according to the risk level. The rule base contains intervention rules corresponding to three risk levels: low, medium and high. Each rule contains triggering conditions and execution parameters, and the execution parameters can be fine-tuned according to the user's age, gender and other demographic characteristics. The intervention program adjustment unit is connected to the intervention program matching unit. It is used to collect the user's heart rate variability and brain wave power in real time during the intervention process, set baseline values, and upgrade and adjust the intervention program when the increase in heart rate variability and the decrease in brain wave power do not reach the preset threshold within the preset time. The intervention program corresponding to low risk is a combination of pet robot companionship and interaction with jujube seed tea; the intervention program corresponding to medium risk is a combination of VR breathing guidance training, lavender tea, and timed interaction with pet robot; and the intervention program corresponding to high risk is a combination of low-intensity ultrasound stimulation and remote video intervention by a professional psychological counselor. The preset time is 5 minutes, and the preset threshold is a heart rate variability increase of ≥10% and an EEG power decrease of ≥15%.

[0032] The data interaction and control module is implemented using the following steps: The data interaction and control module includes a data transmission unit and a collaborative control unit; The data transmission unit adopts edge computing technology to ensure that the delay from data acquisition to processing output is less than the preset value. At the same time, it realizes data interaction and remote consultation data transmission with the cloud through 4G / 5G network. The collaborative control unit is used to control the activation and switching of the gamified general testing module, the multimodal fine screening module, and the adaptive intervention module, so as to realize the linkage process of general testing triggering fine screening and fine screening results driving the matching of intervention plans; The default value is 2 seconds; When the risk assessment score in the general test is greater than or equal to 60, the collaborative control unit triggers the multimodal fine screening module to start. After the fine screening results are output, it immediately drives the adaptive intervention module to match the corresponding intervention plan.

[0033] The above describes the specific embodiments of the present invention. However, the present invention is not limited to the embodiments described. Those skilled in the art can make various equivalent modifications or substitutions without departing from the spirit of the present invention. All such equivalent modifications or substitutions are included within the scope defined by the claims of this application.

Claims

1. A comprehensive intelligent assessment and intervention system for mental health, characterized in that, include: Gamified general testing module, multimodal fine screening module, adaptive intervention module, and data interaction and control module; The gamified general assessment module is used to collect multimodal initial data from users through various gamified tasks, and to process and analyze the multimodal initial data to generate a general assessment risk score; The multimodal fine screening module is connected to the gamified general assessment module. It is used to collect multi-dimensional fine data of users when the risk score of the general assessment reaches a preset threshold, and to perform fusion processing on the multi-dimensional fine data to output a psychological state quantitative report and risk level. The adaptive intervention module is connected to the multimodal screening module and is used to match the corresponding intervention plan according to the risk level, continuously collect the user's physiological data during the intervention process, and dynamically adjust the intervention plan according to the physiological data. The data interaction and control module is connected to the gamified general testing module, the multimodal fine screening module, and the adaptive intervention module, respectively, to realize data transmission and collaborative control between the modules, so as to complete the closed loop of the entire process from assessment to intervention.

2. The intelligent assessment and intervention system for the entire process of mental health as described in claim 1, characterized in that, The gamified general testing module includes a first data acquisition unit, a first data processing unit, and a general testing result generation unit; The first data acquisition unit includes a camera, a microphone, and a touch screen. The camera is used to acquire facial image sequences when the user plays the "Emotional Matching" game, the microphone is used to acquire voice signals when the user plays the "Voice Story Relay" game, and the touch screen is used to record touch coordinate sequences when the user plays the "Attention Maze" game. The first data processing unit is connected to the first data acquisition unit and is used to process the facial image sequence using an improved convolutional neural network to obtain an emotion perception ability score, extract features from the speech signal using Mel frequency cepstral coefficients and process it through a classification model to obtain a speech emotion score, and perform trajectory smoothing and feature extraction on the touch coordinate sequence to obtain an attention concentration index. The general test result generation unit is connected to the first data processing unit and is used to perform fusion processing on the emotion perception ability score, voice emotion score and attention concentration index using a weighted voting model to generate the general test risk score.

3. The intelligent assessment and intervention system for the entire process of mental health as described in claim 2, characterized in that, The improved convolutional neural network is a VGG16 convolutional neural network with a newly added attention mechanism module and an introduced temporal difference network. The attention mechanism module focuses on emotion-related muscle regions, and the temporal difference network analyzes the dynamic changes of facial images in multiple consecutive frames. The classification model is a sentiment classifier built based on a Gaussian mixture model, and the model parameters are iteratively optimized through the EM algorithm. In the weighted voting model, the weight of the emotion perception ability score is 40%, the weight of the voice emotion score is 30%, and the weight of the attention concentration index is 30%.

4. The intelligent assessment and intervention system for the entire process of mental health as described in claim 1, characterized in that, The multimodal fine screening module includes a second data acquisition unit, a second data processing unit, and a fine screening result output unit; The second data acquisition unit includes an EEG headband, a VR device, and an AI virtual assistant. The EEG headband is used to collect EEG signals from the user's prefrontal cortex. The VR device is used to simulate social scenarios and collect the user's eye movement data and physiological signals. The AI ​​virtual assistant is used to initiate structured interviews and collect the user's voice responses. The second data processing unit is connected to the second data acquisition unit and is used to preprocess and calculate the features of the EEG signal to obtain the wave power ratio, to perform synchronous analysis of the eye movement data and physiological signals to identify stress response, and to perform sentiment analysis on the text converted from the speech response to obtain the proportion of negative words. The fine screening result output unit is connected to the second data processing unit. Based on a random forest model trained from multi-source data, it processes the wave power ratio, stress response identification results, and negative word proportion, and outputs the psychological state quantitative report and risk level.

5. The intelligent assessment and intervention system for the entire process of mental health as described in claim 4, characterized in that, The preprocessing of the EEG signal employs wavelet transform decomposition to remove electromyographic and power frequency interference, and the feature calculation is achieved through power spectral density analysis. In the synchronous analysis of the eye movement data and physiological signals, the eye movement features include fixation duration and saccade amplitude, and the physiological signals include heart rate variability. Avoidance fixation patterns are identified by a hidden Markov model, and stress response is judged by combining the time-domain and frequency-domain indices of heart rate variability. The sentiment analysis of the text is based on a BERT pre-trained model. After word segmentation and stop word removal, word vectors are generated, negative words and their intensity are identified, and the proportion of negative words is calculated. At the same time, semantic dependency analysis is combined to identify the causal chain of sentiment.

6. The intelligent assessment and intervention system for the entire process of mental health as described in claim 1, characterized in that, The adaptive intervention module includes an intervention plan matching unit and an intervention plan adjustment unit; The intervention plan matching unit is used to call the corresponding intervention plan from the preset rule base according to the risk level. The rule base contains intervention rules corresponding to three risk levels: low, medium and high. Each rule contains triggering conditions and execution parameters, and the execution parameters are fine-tuned according to the user's demographic characteristics. The intervention program adjustment unit is connected to the intervention program matching unit and is used to collect the user's heart rate variability and brain wave power in real time during the intervention process, set a baseline value, and upgrade and adjust the intervention program when the increase in heart rate variability and the decrease in brain wave power within a preset time do not reach the preset threshold.

7. The intelligent assessment and intervention system for the entire process of mental health as described in claim 6, characterized in that, The intervention program corresponding to low risk is a combination of pet robot companionship and interaction with jujube seed tea; the intervention program corresponding to medium risk is a combination of VR breathing guidance training, lavender tea, and timed interaction with pet robot; and the intervention program corresponding to high risk is a combination of low-intensity ultrasound stimulation and remote video intervention by a professional psychological counselor. The preset time is 5 minutes, and the preset threshold is a heart rate variability increase of ≥10% and an EEG power decrease of ≥15%.

8. The intelligent assessment and intervention system for the entire process of mental health as described in claim 1, characterized in that, The data interaction and control module includes a data transmission unit and a collaborative control unit; The data transmission unit adopts edge computing technology to ensure that the delay from data acquisition to processing output is less than a preset value. At the same time, it realizes data interaction and remote consultation data transmission with the cloud through 4G / 5G network. The collaborative control unit is used to control the activation and switching of the gamified general testing module, the multimodal fine screening module, and the adaptive intervention module, so as to realize the linkage process of general testing triggering fine screening and fine screening results driving the matching of intervention schemes.

9. The intelligent assessment and intervention system for the entire process of mental health as described in claim 8, characterized in that, The preset value is 2 seconds; When the risk assessment score is greater than or equal to 60, the collaborative control unit triggers the multimodal fine screening module to start. After the fine screening results are output, it immediately drives the adaptive intervention module to match the corresponding intervention plan.

10. A fully intelligent assessment and intervention system for mental health according to any one of claims 1-9, characterized in that, The system is deployed on a mobile vehicle, which is modular in design and features shock absorption and sound insulation.

Citation Information

Cited By

  • Man-machine conversation psychological stress identification and intervention method, system, medium and product

    CN121306436A

  • Video type psychological and physiological index analysis method and system

    CN121943315A