Intelligent psychological intervention system based on multi-modal fusion

Through a multimodal fusion intelligent psychological intervention system, combined with voice emotion, keyboard dynamics and physiological signals, dynamic assessment and personalized intervention of psychological states are realized, the rigid problem of traditional evaluation is solved, the accuracy and real-time nature of psychological state recognition is improved, and the dependence of artificial intervention is reduced.

CN120337162AInactive Publication Date: 2025-07-18JIANGSU ZHUODUN INFORMATION TECH CO LTD

Patent Information

Application Number
CN202510828771.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-20
Publication Date
2025-07-18
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Traditional psychological assessment relies on a single standardized scale, which is difficult to fully reflect the complexity and dynamics of individual psychological states. It lacks continuous tracking of digital chemotherapy efficacy indicators, cannot dynamically adjust intervention strategies, and multimodal data has not achieved joint modeling.

Method used

An intelligent psychological intervention system based on multimodal fusion is used to realize time series alignment and feature weighting through speech emotion analysis, keyboard dynamics monitoring and physiological signal acquisition, combined with a cross-modal Transformer model, and dynamically adjust the intervention strategy using reinforcement learning algorithms, integrating an adaptive intervention engine and clinical decision support system to realize multi-dimensional state space definition and personalized intervention.

Benefits of technology

It significantly improves the accuracy and real-time nature of psychological state recognition, shortens the recognition timeliness of high-risk signals, solves the rigid problem of traditional evaluation, improves the individual intervention effect and reduces artificial dependence.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120337162A_ABST
    Figure CN120337162A_ABST
Patent Text Reader

Abstract

The invention discloses an intelligent psychological intervention system based on multi-modal fusion, which is characterized in that a three-dimensional evaluation system is constructed by integrating speech sentiment analysis, keyboard dynamics monitoring and physiological signal acquisition, and time sequence alignment and feature weighted fusion of multi-source data are realized by adopting a cross-modal Transform model. The core of the system comprises an adaptive intervention engine which defines a multi-dimensional state space based on a hierarchical reinforcement learning architecture, optimizes an intervention strategy through a PPO algorithm, and realizes dynamic emotion interaction in AR and VR scenes in combination with a digital twin training module; according to the clinical decision support system, physiological behavior characteristics and psychological assessment trends are integrated by using a multi-time scale risk prediction model, and a personalized early warning threshold system is constructed, so that the psychological state recognition accuracy is improved, the intervention intensity self-adaptive adjustment response time is shortened, and the high-risk signal early warning timeliness reaches the minute level; and the problems of evaluation hysteresis and strategy stiffness of traditional psychological intervention are obviously improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of mental health treatment applications, and particularly to an intelligent psychological intervention system based on multimodal fusion. Background Art

[0002] In current society, mental health problems are becoming increasingly prominent, with extensive and profound impacts. According to the report of the World Health Organization (WHO), the incidence of mental disorders worldwide continues to rise, bringing heavy disease burdens and economic pressures to individuals, families, and even the entire society. For example, research continuously reveals "the increasing prominence of mental health problems among college students" and "the severity of mental health problems among adolescents". As the future backbone of society, the mental health status of these groups deserves particular attention. Therefore, providing timely and effective psychological assessment and intervention services is of crucial significance for enhancing national well-being and promoting social harmony and stability.

[0003] Common traditional psychological assessments largely rely on standardized psychological scales, such as the Symptom Checklist-90 and the Self-Rating Anxiety Scale. As a result, the assessment dimensions are relatively fixed and single, making it difficult to comprehensively reflect the complexity and dynamics of an individual's mental state. Moreover, the clinical path is fixed, unable to dynamically adjust the intervention strategy based on individual feedback. At the same time, the lack of continuous tracking of digital efficacy indicators prevents the cross-modal joint modeling of voice emotions, physiological data, and behavioral data, failing to meet the work requirements of mental health treatment applications. Therefore, an intelligent psychological intervention system based on multimodal fusion is proposed. Summary of the Invention

[0004] The present invention provides the following technical solutions: An intelligent psychological intervention system based on multimodal fusion, comprising: A multimodal assessment system, an adaptive intervention engine, a clinical decision support system, and a standardized clinical path implementation module. The multimodal assessment system includes a data acquisition layer and a fusion analysis layer. The multimodal assessment system is used to collect and fuse the user's voice emotion data, keyboard dynamics data, and physiological signal data; The data acquisition layer includes a voice emotion analysis unit, a keyboard dynamics monitoring unit, and a physiological signal integration unit. The fusion analysis layer is used to align time series data through a cross-modal Transformer model and weight multimodal features through a multi-head attention mechanism; The adaptive intervention engine is used to dynamically adjust the intervention strategy through a reinforcement learning algorithm. The clinical decision support system is used to predict the psychological risk level and provide joint intervention suggestions. The standardized clinical path implementation module is used to execute digital intervention processes and biofeedback linkages; An adaptive intervention engine, including a reinforcement learning tuning module and a digital twin training module. The reinforcement learning tuning module is used to define a multi-dimensional state space including the PSQI index and the SDS score, and optimize the intervention strategy through the PPO algorithm. The digital twin training module is used to achieve dynamic emotional interaction and cognitive remodeling through AR and VR scenarios; A clinical decision support system, including a risk prediction model and a combined drug and psychology algorithm. The standardized clinical pathway implementation module internally integrates a PM+ digital process engine, an intelligent session management unit, and a biofeedback linkage unit.

[0005] Preferably, the voice emotion analysis unit identifies changes in emotional states by extracting acoustic feature parameters from the user's voice call records and combining them with a deep learning model. The acoustic feature parameters include fundamental frequency, speech rate, and intonation changes. The keyboard dynamics monitoring unit assists in evaluating the mental state by capturing the user's input behavior characteristics. The behavior characteristics include recording the keystroke interval time, input error rate, and input pressure value. The physiological signal integration unit continuously collects indicators of the activity of the autonomic nervous system through wearable devices, including heart rate variability, skin conductance response, and body temperature fluctuation data. When persistent sympathetic nerve excitation characteristics are detected, a corresponding biofeedback training plan is automatically matched and the intervention intensity level is adjusted.

[0006] Preferably, the fusion analysis layer integrates multi-source data through a combination of feature-level fusion and decision-level fusion. That is, first, the time series alignment and standardization processing of each modality data are performed, then a hierarchical feature expression space is constructed, and finally, the weight coefficients of each modality feature are dynamically allocated through a multi-head attention mechanism to form a comprehensive mental health state assessment vector.

[0007] Preferably, the reinforcement learning tuning module adopts a hierarchical reinforcement learning architecture. The upper policy network of the hierarchical reinforcement learning architecture is responsible for long-term intervention planning, and the lower execution network of the hierarchical reinforcement learning architecture processes immediate intervention adjustments.

[0008] Preferably, the risk prediction model uses a multi-time scale analysis method to integrate short-term physiological behavior characteristics and long-term psychological assessment trends, and constructs a personalized risk warning threshold system. When the risk prediction model detects a risk level transition, a hierarchical response mechanism is automatically activated.

[0009] Preferably, the combined drug and psychology algorithm is used to establish a drug response prediction model. By analyzing the correlation between historical medication records and changes in psychological indicators, personalized drug dose adjustment suggestions are generated. At the same time, in combination with the implementation progress of cognitive behavioral therapy, the collaborative cooperation plan of drug intervention and psychological intervention is optimized. The combined drug and psychology algorithm uses a Bayesian network to associate the SSRI medication dose with the HRV improvement rate and dynamically adjusts the rules.

[0010] Preferably, the PM + digital process engine incorporates a modular intervention protocol library. The modular intervention protocol library supports the automatic combination of standardized technical modules for cognitive restructuring, behavioral activation, and mindfulness training according to user characteristics, while retaining the personalized adjustment interface for human therapists, realizing the configuration of an intervention path that combines standardization and personalization. The intelligent session management unit is used to automatically annotate key dialogue nodes and extract constructive dialogue content.

[0011] Preferably, the biofeedback linkage unit is used to establish a dynamic mapping relationship between physiological signals and intervention content, and automatically trigger the switching of preset intervention content when physiological index changes are detected. The intervention content includes breathing training guidance, relaxation music playback, and virtual scene adjustment. The biofeedback linkage unit internally integrates a desk pressure sensor, which continuously monitors the changes in the user's physiological indicators. When it detects that the writing force fluctuation of the user exceeds the preset threshold, the biofeedback linkage unit automatically triggers relaxation intervention measures such as breathing training, and simultaneously adjusts the difficulty and progress of relaxation training in the VR scene through HRV data.

[0012] Preferably, a cross-modal emotion transfer sub-module is added inside the digital twin training module. The cross-modal emotion transfer sub-module is used to construct a user-specific digital emotional twin, and map voice emotion features, facial micro-expression features, and physiological signal features to the emotional expression of the virtual image through a deep generative adversarial network. An emotion regulation guidance mechanism is implanted in the virtual training scene of the digital twin training module. When negative emotions of the user are detected, the virtual image will actively display a preset positive emotional expression pattern.

[0013] Preferably, an intelligent narrative therapy unit is provided inside the standardized clinical pathway implementation module. The intelligent narrative therapy unit is used to analyze the narrative content generated by the user during the intervention process through natural language processing technology, including voice dialogue records, written expression materials, and interactive discourse in the virtual scene. The intelligent narrative therapy unit automatically identifies the core conflict points, resource points, and transformation points in the narrative by constructing a narrative analysis model based on a knowledge graph, and generates personalized narrative reconstruction suggestions. The intelligent narrative therapy unit internally integrates augmented reality technology, which is used to visually present the analysis results as a dynamic narrative map to assist the therapist in guiding the user to complete the narrative reconstruction process.

[0014] In summary, compared with the prior art, the present invention provides an intelligent psychological intervention system based on multi-modal fusion, which has the following beneficial effects: 1. The present invention integrates multimodal data of speech emotion, keyboard dynamics and physiological signals, and combines the cross-modal Transformer model to achieve time series alignment and feature weighting, breaking through the single evaluation mode of traditional scale reliance, significantly improving the accuracy and real-time performance of identifying psychological states such as depression and anxiety, and constructing a dynamic tuning mechanism based on reinforcement learning, defining the state space with multidimensional indicators of PSQI sleep index and SDS score, and realizing real-time adaptation of cognitive behavioral intervention through the digital twin training module, solving the rigidity problem of fixed clinical pathways and improving the effect of individual intervention; 2. The present invention can obtain objective indicators of continuous quantitative HRV and keyboard input delay by integrating the PM+ digital process engine and the biofeedback linkage unit, and integrate multi-source heterogeneous data such as voice emotion, eye movement behavior, and physiological signals through a multi-head attention mechanism to avoid the misjudgment problem caused by modal splitting of the multimodal system. At the same time, it can explore the implicit correlation between emotional fluctuations and suicidal thoughts, achieve early warning of high-risk signals, and analyze multimodal data streams in real time through risk prediction models and drug and psychological joint algorithms, automatically trigger graded warnings, and shorten the identification time of high-risk cases from hours to minutes, reducing manual dependence. BRIEF DESCRIPTION OF THE DRAWINGS

[0015] Figure 1 It is a schematic diagram of the system structure of the present invention.

[0016] Figure 2 It is a schematic diagram of the structure of the multimodal evaluation system of the present invention. DETAILED DESCRIPTION

[0017] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0018] Embodiment 1 See also Figure 1 The present invention provides a technical solution, an intelligent psychological intervention system based on multimodal fusion, comprising: Multimodal evaluation system, adaptive intervention engine, clinical decision support system and standardized clinical pathway implementation module. The multimodal evaluation system includes data collection layer and fusion analysis layer. Please refer to Figure 2 ,The multimodal evaluation system is used to collect and fuse the user’s voice emotion data, keyboard dynamics data and physiological signal data; The data acquisition layer includes a voice emotion analysis unit, a keyboard dynamics monitoring unit, and a physiological signal integration unit. The fusion analysis layer is used to align time series data through a cross-modal Transformer model and weight multi-modal features through a multi-head attention mechanism. The voice emotion analysis unit extracts acoustic feature parameters from the user's voice call records and combines them with a deep learning model to identify changes in emotional states. The acoustic feature parameters include fundamental frequency, speech rate, and intonation changes. The keyboard dynamics monitoring unit assists in the assessment of mental states by capturing the user's input behavior characteristics. The behavior characteristics include recording the keystroke interval time, input error rate, and input pressure value. The physiological signal integration unit continuously collects indicators of the activity of the autonomic nervous system through wearable devices, including heart rate variability, skin conductance response, and body temperature fluctuation data. When persistent sympathetic nerve excitation characteristics are detected, it automatically matches the corresponding biofeedback training plan and adjusts the intervention intensity level; Please refer to Figure 2 , the fusion analysis layer integrates multi-source data through a combination of feature-level fusion and decision-level fusion. That is, first, the time series alignment and normalization processing are performed on each modal data, then a hierarchical feature expression space is constructed, and finally, the weight coefficients of each modal feature are dynamically allocated through a multi-head attention mechanism to form a comprehensive mental health status assessment vector; The adaptive intervention engine is used to dynamically adjust the intervention strategy through a reinforcement learning algorithm. The clinical decision support system is used to predict the mental risk level and provide joint intervention suggestions. The standardized clinical path implementation module is used to execute the digital intervention process and biofeedback linkage. The adaptive intervention engine includes a reinforcement learning tuning module and a digital twin training module. The reinforcement learning tuning module adopts a hierarchical reinforcement learning architecture. The upper-level policy network of the hierarchical reinforcement learning architecture is responsible for long-term intervention planning, and the lower-level execution network of the hierarchical reinforcement learning architecture processes immediate intervention adjustments; The reinforcement learning tuning module is used to define a multi-dimensional state space containing the PSQI index and SDS score, and optimize the intervention strategy through the PPO algorithm. The digital twin training module is used to achieve dynamic emotional interaction and cognitive remodeling through AR and VR scenarios; Clinical decision support system, including a risk prediction model and a combined drug and psychology algorithm. The internal integration of the standardized clinical pathway implementation module has a PM+ digital process engine, an intelligent conversation management unit, and a biofeedback linkage unit. The risk prediction model uses a multi-time scale analysis method to integrate short-term physiological behavior characteristics and long-term psychological assessment trends, and constructs a personalized risk warning threshold system. When the risk prediction model detects a risk level transition, it automatically activates a hierarchical response mechanism. The combined drug and psychology algorithm is used to establish a drug response prediction model. By analyzing the correlation between historical medication records and changes in psychological indicators, it generates personalized drug dose adjustment suggestions. At the same time, in combination with the implementation progress of cognitive behavioral therapy, it optimizes the coordination plan of drug intervention and psychological intervention. The combined drug and psychology algorithm uses a Bayesian network to associate the SSRI medication dose with the HRV improvement rate and dynamically adjusts the rules. The PM+ digital process engine has a built-in modular intervention protocol library. The modular intervention protocol library supports the automatic combination of standardized technical modules for cognitive restructuring, behavior activation, and mindfulness training according to user characteristics, while retaining the personalized adjustment interface of the human therapist to achieve the configuration of an intervention path that combines standardization and personalization. The intelligent conversation management unit is used to automatically label key conversation nodes and extract constructive conversation content; The biofeedback linkage unit is used to establish a dynamic mapping relationship between physiological signals and intervention content. When specific physiological index changes are detected, it automatically triggers the switching of preset intervention content. The intervention content includes breathing training guidance, relaxation music playback, and virtual scene adjustment. The biofeedback linkage unit internally integrates a desk pressure sensor. The desk pressure sensor real-time monitors the changes in the user's physiological indicators. When it detects that the writing force fluctuation of the user exceeds the preset threshold, the biofeedback linkage unit automatically triggers relaxation intervention measures such as breathing training, and at the same time adjusts the relaxation training difficulty and progress in the VR scene through HRV data; The workflow of the multimodal evaluation system is as follows: In the voice emotion acquisition channel, the system first establishes a deep connection with the user's communication device and continuously obtains voice call records. The voice emotion analysis unit adopts a segmented feature extraction strategy, cutting the continuous voice stream into analysis units with a duration of 30 seconds. For each voice segment, the system first performs acoustic feature deconstruction, extracting the fundamental frequency trajectory, speech rate change curve, and intonation fluctuation pattern through speech signal processing algorithms. The fundamental frequency analysis uses the dynamic warping algorithm to capture the subtle fluctuations in the vocal cord vibration frequency, the speech rate calculation is based on the statistical distribution of lexical interval times, and the intonation recognition relies on the Fourier transform analysis of the pitch contour. These underlying features are then input into a pre-trained deep learning model, which adopts a spatio-temporal convolutional network architecture and can learn the non-linear mapping relationship between acoustic features and emotion labels. The model training data covers millions of voice samples, including eight basic emotion categories such as happy, sad, angry, and neutral, and the cross-scenario generalization ability of emotion recognition is ensured through the transfer learning mechanism. The keyboard dynamics monitoring channel adopts a dual-mode data acquisition scheme. In the desktop scenario, the system deploys a lightweight keyboard hook program to record the timestamps of each keystroke event with millisecond-level precision. For the mobile scenario, it connects to intelligent peripherals through the Bluetooth protocol to obtain touch screen pressure sensor data. The collected feature dimensions include three core indicators: the key press interval time is constructed by calculating the time difference sequence of adjacent keys, the input error rate is calculated based on the backspace key usage frequency and the number of text modifications, and the input pressure value is obtained from the keyboard pressure sensor or the touch screen pressure sensing array. These raw data are processed by a sliding window to generate a feature vector reflecting the user's cognitive load. When abnormal patterns such as the standard deviation of the key press interval exceeding the threshold or the error rate continuously rising are detected, the system automatically triggers the mental state warning mechanism. The physiological signal integration channel constructs a relay for wearable device access, supporting protocol adaptation for mainstream smart bracelets, electrocardiogram patches, and other devices. The data acquisition focuses on the activity indicators of the autonomic nervous system, obtaining the time-domain features of heart rate variability (HRV) through photoplethysmography technology, capturing the changes in sweat gland activity using skin conductance sensors, and continuously monitoring the fingertip temperature fluctuations with an infrared array sensor. These physiological signals are uploaded to the edge computing node at a sampling frequency of once per second for real-time preprocessing. The system has a built-in sympathetic excitation detection model. When the LF / HF ratio of HRV exceeds the warning threshold for 10 consecutive minutes, or the skin conductance level shows a stepwise upward trend, it automatically determines that the user is in a stress state. At this time, the biofeedback linkage unit will start the intervention intensity evaluation process, and dynamically adjust the guiding rhythm of breathing training or the immersion depth of the VR scene within the intensity range of 1-5 according to the degree of deviation of the physiological indicators from the baseline. After the three-channel data acquisition is completed, it enters the cross-modal fusion analysis stage. The system first performs a time axis alignment operation, using the dynamic time warping algorithm to compensate for the sampling rate differences of each modal data. The feature normalization module uses the quantile normalization method to eliminate the numerical differences caused by different dimensions.Subsequently, a four-dimensional feature tensor is constructed, which includes speech emotion features, keyboard dynamics features, physiological signal features, and temporal context encoding. This tensor is input into a pre-trained cross-modal Transformer model, and the feature weight allocation is achieved through the multi-head attention mechanism. The model training adopts a contrastive learning strategy to minimize the distance between the feature vectors of positive sample pairs (multi-modal data in the same time window) and maximize the distance between negative sample pairs (random combinations of different time windows). Finally, a state vector containing four dimensions of valence, arousal, cognitive load, and physiological stress is generated, providing a basis for real-time decision-making for the upper-level intervention engine. When high-risk signals such as persistent sympathetic nerve excitation are detected, the system initiates an emergency response process. The physiological signal integration unit first increases the data sampling frequency to 10 Hz and performs signal denoising through the Kalman filter algorithm. The keyboard dynamics monitoring unit switches to the high-sensitivity mode and shortens the keystroke interval analysis window to 5 seconds. The speech emotion analysis unit activates the emotion intensity tracking sub-module and focuses on monitoring the expression frequency of high-risk emotions such as anger and despair. The multi-modal data is input into the risk prediction model through the emergency fusion channel. This model adopts a lightweight gradient boosting machine architecture and can complete the risk level determination within 200 milliseconds. When the prediction probability exceeds the preset threshold, the system automatically activates the biofeedback linkage unit, dynamically configures the intervention plan according to the current physiological indicators. For example, for users with low HRV, resonance breathing training is preferentially initiated, and for users with high skin conductance, progressive muscle relaxation guidance is adopted. At the same time, cognitive restructuring prompts are pushed through the intelligent conversation management unit to form a multi-modal collaborative intervention effect; The specific implementation process of the adaptive intervention engine is as follows: First, in the environmental perception stage, the system first constructs a state space containing multi-dimensional indicators. The reinforcement learning optimization module integrates real-time data from the multi-modal evaluation system, including standardized evaluation results such as the PSQI sleep quality index and the SDS self-rating depression scale score, and also incorporates dynamic monitoring indicators such as the fundamental frequency change of speech emotion, the keyboard input error rate, and the HRV time-domain characteristics. These data are encoded into a 12-dimensional state vector, covering key psychological dimensions such as emotional valence, cognitive load, physiological stress, and sleep quality. The state space adopts a sliding time window mechanism to retain the historical data trajectory of the most recent 72 hours, providing a temporal context reference for the policy network. In the policy generation stage, a hierarchical reinforcement learning architecture is adopted, including the collaborative work of the upper-level policy network and the lower-level execution network. The upper-level policy network is responsible for formulating weekly intervention plans. Based on the deep Q-learning algorithm, its input is the state representation abstracted by time. The network focuses on key state dimensions through the attention mechanism. For example, when it detects that the SDS score continues to rise, it automatically raises the priority of the cognitive behavioral therapy (CBT) module. The output is a macro plan containing intervention goals, technology combinations, and frequency arrangements, such as "Focus on implementing emotion recognition training this week, 2 times a day, 15 minutes each time". The lower-level execution network then processes hourly immediate adjustments, using the Actor-Critic framework. Among them, the Actor network generates specific intervention actions according to the real-time state, such as adjusting the guiding rhythm of breathing training or switching the immersion mode of the VR scene; the Critic network guides policy iteration through value evaluation. When it detects an improvement in the user's physiological indicators, it gives a positive reward signal. In the execution feedback stage, a multi-channel effect evaluation mechanism is established. The system synchronously collects three types of response data: Physiological signal changes are monitored in real time through wearable devices, focusing on indicators such as the HRV improvement rate and the skin conductance decline amplitude; Behavioral performance data includes the stability of keyboard dynamics characteristics and the interaction duration in the virtual scene; Subjective feedback is collected through an embedded scale, such as the intervention satisfaction score and the emotion self-rating scale. These data are standardized and then form the input of the reward function, where the weight of physiological improvement accounts for 40%, the behavioral stability accounts for 35%, and the subjective feedback accounts for 25%. When the cumulative reward value exceeds the threshold, a policy retention mechanism is triggered; if the evaluation is lower than the baseline for 3 consecutive times, a policy rollback process is started. In the continuous optimization stage, the policy evolution is achieved relying on the digital twin training module. This module constructs a user-specific digital twin, and simulates the user's response mode to the intervention plan through a generative adversarial network (GAN). In the AR / VR training scenario, the virtual avatar can reproduce the typical emotional response characteristics of the user, such as specific speech intonation patterns or physiological stress thresholds. The intervention strategy is first tested under stress in the digital twin environment, and the optimal parameter combination is explored through Monte Carlo tree search. For example, when testing the coping strategy for an anxiety scenario, the system will simulate typical reactions such as abnormal fluctuations in the user's HRV and an increase in the keyboard input error rate to verify the robustness of the intervention plan.The strategies verified by digital twin will be incorporated into the candidate strategy pool and undergo online A / B testing through the multi-armed bandit algorithm. The ultimately winning strategy will be updated to the main intervention engine. In terms of special scenario handling, the system sets up an emergency intervention channel. When the risk prediction model of the clinical decision support system detects high-risk signals such as suicidal ideation, the adaptive intervention engine immediately switches to the crisis response mode. At this time, the upper-layer strategy network freezes the regular intervention plan, and the lower-layer execution network activates the pre-trained emergency intervention protocol. The intervention intensity is dynamically adjusted according to the physiological index risk level. For example, for users with extremely low HRV, the guiding frequency of breathing training is increased from the regular 0.2 Hz to 0.5 Hz. At the same time, the VR scene is switched to the safe island environment, and a warning notice is automatically sent to the artificial therapist terminal. After the crisis intervention ends, the system generates a strategy optimization report through the post-analysis module and supplements it to the digital twin training library.

[0019] Example Two The difference between this example and the above Example One is that a cross-modal emotion transfer sub-module is added inside the digital twin training module. The cross-modal emotion transfer sub-module is used to construct a user-specific digital emotion twin, and maps speech emotion features, facial micro-expression features, and physiological signal features to the emotional expression of the virtual image through a deep generative adversarial network. An emotion regulation guidance mechanism is implanted in the virtual training scenario of the digital twin training module. When detecting the user's negative emotion, the virtual image will actively display the preset positive emotion expression mode; And an intelligent narrative therapy unit is provided inside the standardized clinical path implementation module. The intelligent narrative therapy unit is used to analyze the narrative content generated by the user during the intervention through natural language processing technology, including voice conversation records, written expression materials, and interactive discourses in the virtual scenario; the intelligent narrative therapy unit automatically identifies the core conflict points, resource points, and transformation points in the narrative by constructing a narrative analysis model based on a knowledge graph, and generates personalized narrative reconstruction suggestions. The intelligent narrative therapy unit integrates augmented reality technology inside, and the augmented reality technology is used to visually present the analysis results as a dynamic narrative graph to assist the therapist in guiding the user to complete the narrative reconstruction process.

[0020] This solution combines multi-modal data of speech emotion, keyboard dynamics features, and physiological signals, and uses a cross-modal Transformer model to achieve time series alignment and feature weighting, breaking through the single evaluation mode relying on traditional scales, significantly improving the accuracy and real-time performance of recognizing psychological states such as depression and anxiety. And a dynamic tuning mechanism is constructed based on reinforcement learning, defining the state space with multi-dimensional indicators such as the PSQI sleep index and SDS score, and realizing the real-time adaptation of cognitive behavior intervention through the digital twin training module, solving the rigidity problem of fixed clinical paths and improving the individual intervention effect.

[0021] This solution integrates the PM + digital process engine with the biofeedback linkage unit to obtain objective indicators of continuous quantified HRV and keyboard input latency. It also fuses multi-source heterogeneous data such as speech emotion, eye movement behavior, and physiological signals through the multi-head attention mechanism, avoiding misjudgment problems caused by modal fragmentation in multi-modal systems. At the same time, it can explore the implicit relationship between emotional fluctuations and suicidal ideation, achieve early warning of high-risk signals, and through the risk prediction model and the combined algorithm of medication and psychology, analyze the multi-modal data stream in real-time, automatically trigger hierarchical warnings, shorten the identification time of high-risk cases from the hour level to the minute level, and reduce the dependence on manual work.

[0022] It should be noted that in this article, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements not only includes those elements but also includes other elements not expressly listed, or also includes elements inherent to such process, method, article or device.

[0023] Although the embodiments of the present invention have been shown and described, it will be understood by those of ordinary skill in the art that various changes, modifications, substitutions and variations can be made in these embodiments without departing from the principles and spirit of the present invention, and the scope of the present invention is defined by the appended claims and their equivalents.

Claims

1. An intelligent psychological intervention system based on multimodal fusion, characterized in that, Including: A multimodal assessment system, an adaptive intervention engine, a clinical decision support system, and a standardized clinical pathway implementation module. The multimodal assessment system includes a data acquisition layer and a fusion analysis layer. The multimodal assessment system is used to collect and fuse the user's voice emotion data, keyboard dynamics data, and physiological signal data; The data acquisition layer includes a voice emotion analysis unit, a keyboard dynamics monitoring unit, and a physiological signal integration unit. The fusion analysis layer is used to align time series data through a cross-modal Transformer model and weight multimodal features through a multi-head attention mechanism; The adaptive intervention engine is used to dynamically adjust the intervention strategy through a reinforcement learning algorithm. The clinical decision support system is used to predict the psychological risk level and provide joint intervention suggestions. The standardized clinical pathway implementation module is used to execute the digital intervention process and biofeedback linkage; The adaptive intervention engine includes a reinforcement learning tuning module and a digital twin training module. The reinforcement learning tuning module is used to define a multi-dimensional state space including the PSQI index and SDS score, and optimize the intervention strategy through the PPO algorithm. The digital twin training module is used to achieve dynamic emotional interaction and cognitive remodeling through AR and VR scenarios; The clinical decision support system includes a risk prediction model and a drug and psychology joint algorithm. The standardized clinical pathway implementation module internally integrates a PM+ digital process engine, an intelligent session management unit, and a biofeedback linkage unit.

2. The intelligent psychological intervention system based on multimodal fusion according to claim 1, characterized in that: The voice emotion analysis unit extracts acoustic feature parameters from the user's voice call record and combines a deep learning model to identify changes in the emotional state. The acoustic feature parameters include fundamental frequency, speech rate, and intonation changes. The keyboard dynamics monitoring unit realizes auxiliary assessment of the psychological state by capturing the user's input behavior characteristics. The behavior characteristics include recording the key press interval time, input error rate, and input pressure value. The physiological signal integration unit continuously collects indicators of the activity of the autonomic nervous system through a wearable device, including heart rate variability, skin conductance response, and body temperature fluctuation data. When persistent sympathetic nerve excitation characteristics are detected, an appropriate biofeedback training plan is automatically matched and the intervention intensity level is adjusted.

3. The intelligent psychological intervention system based on multimodal fusion according to claim 1, characterized in that: The fusion analysis layer integrates multi-source data through a combination of feature-level fusion and decision-level fusion. That is, first, the time series alignment and standardization processing of each modal data are performed, then a hierarchical feature expression space is constructed, and finally, the weight coefficients of each modal feature are dynamically allocated through a multi-head attention mechanism to form a comprehensive mental health state assessment vector.

4. The intelligent psychological intervention system based on multimodal fusion according to claim 1, wherein: The reinforcement learning tuning module adopts a hierarchical reinforcement learning architecture. The upper-level policy network of the hierarchical reinforcement learning architecture is responsible for long-term intervention planning, and the lower-level execution network of the hierarchical reinforcement learning architecture processes immediate intervention adjustments.

5. The intelligent psychological intervention system based on multimodal fusion according to claim 1, characterized in that: The risk prediction model uses a multi-time scale analysis method to integrate short-term physiological behavior characteristics and long-term psychological assessment trends, and constructs a personalized risk warning threshold system. When the risk prediction model detects a risk level transition, a hierarchical response mechanism is automatically activated.

6. The intelligent psychological intervention system based on multi-modal fusion according to claim 1, wherein: The drug and psychology combined algorithm is used to establish a drug response prediction model. By analyzing the correlation between historical medication records and changes in psychological indicators, it generates personalized drug dosage adjustment suggestions. At the same time, in combination with the implementation progress of cognitive behavioral therapy, it optimizes the collaborative cooperation plan of drug intervention and psychological intervention. The drug and psychology combined algorithm uses a Bayesian network to associate the SSRI dosage with the HRV improvement rate and dynamically adjusts the rules.

7. The intelligent psychological intervention system based on multi-modal fusion according to claim 1, characterized in that: The PM+ digital process engine is built with a modular intervention protocol library. The modular intervention protocol library supports automatically combining standardized technical modules of cognitive restructuring, behavioral activation, and mindfulness training according to user characteristics, while retaining the personalized adjustment interface for human therapists to achieve the configuration of an intervention path that combines standardization and personalization. The intelligent session management unit is used to automatically label key dialogue nodes and extract constructive dialogue content.

8. The intelligent psychological intervention system based on multimodal fusion according to claim 1, characterized in that: The biofeedback linkage unit is used to establish a dynamic mapping relationship between physiological signals and intervention content. When a change in physiological indicators is detected, it automatically triggers the switching of preset intervention content. The intervention content includes breathing training guidance, relaxation music playback, and virtual scene adjustment. The biofeedback linkage unit is internally integrated with a desk pressure sensor. The desk pressure sensor continuously monitors the changes in the user's physiological indicators. When it detects that the writing force fluctuation of the user exceeds the preset threshold, the biofeedback linkage unit automatically triggers the relaxation intervention measure of breathing training and adjusts the relaxation training difficulty and progress in the VR scene through HRV data.

9. The intelligent psychological intervention system based on multimodal fusion according to claim 1, wherein: The digital twin training module is internally equipped with a cross-modal emotion transfer sub-module. The cross-modal emotion transfer sub-module is used to construct a user-specific digital emotion twin and map voice emotion features, facial micro-expression features, and physiological signal features to the emotion expression of the virtual image through a deep generative adversarial network. An emotion regulation guidance mechanism is implanted in the virtual training scene of the digital twin training module. When negative emotions of the user are detected, the virtual image will actively display a preset positive emotion expression pattern.

10. The intelligent psychological intervention system based on multimodal fusion according to claim 1, characterized in that: The standardized clinical path implementation module is internally provided with an intelligent narrative therapy unit. The intelligent narrative therapy unit is used to analyze the narrative content generated by the user during the intervention through natural language processing technology, including voice dialogue records, written expression materials, and interactive discourse in the virtual scene. The intelligent narrative therapy unit automatically identifies the core conflict points, resource points, and transformation points in the narrative by constructing a narrative analysis model based on a knowledge graph and generates personalized narrative reconstruction suggestions. The intelligent narrative therapy unit is internally integrated with augmented reality technology. The augmented reality technology is used to visually present the analysis results as a dynamic narrative map to assist the therapist in guiding the user to complete the narrative reconstruction process.

Citation Information

Patent Citations

  • Teenager mental health data analysis and early warning system

    CN117912710A

  • System platform for psychological assessment and emotion feedback

    CN120124085A

  • Assessing adherence fidelity to behavioral interventions using interactivity and natural language processing

    US20180317840A1

Cited By

  • Depression assessment titration optimization method and system based on professional labels

    CN120636707A

  • Multi-modal medical information intelligent integration and decision support system for acupuncture rehabilitation

    CN120766879A

  • Digital human speech synthesis method and system based on multi-modal speech feature fusion

    CN120833777A

  • AI large model psychological evaluation and dredging voice interaction method for micro hyperbaric oxygen chamber

    CN120859497A

  • Wound memory integration evaluation and treatment system

    CN120998425A