A psychological support system and a psychological support and medical closed-loop intelligent management system
By constructing multimodal patient profiles and using dynamic content adaptation modules, the problem of insufficient personalization and real-time performance of existing AI tools in the management of patients with serious diseases is solved. This enables personalized remote psychological support and closed-loop medical management, and improves the real-time and accuracy of emotional intervention.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- AFFILIATED HOSPITAL OF CHENGDU UNIV (CHENGDU INST OF TRAUMATOLOGY & ORTHOPEDICS)
- Filing Date
- 2025-12-29
- Publication Date
- 2026-05-22
AI Technical Summary
Existing AI tools lack the ability to generate personalized content, integrate with medical systems, and model and intervene in family support systems throughout the entire course of a patient's illness. This results in fragmented emotion recognition and intervention outcomes, making it impossible to achieve real-time, personalized, and precise remote psychological support.
It employs a multimodal patient profiling module, which integrates voice, vision, and physiological data through the Transformer architecture to generate personalized AI emotion support videos. Combined with a dynamic content adaptation module, it enables real-time adjustment of video content and integrates product recommendations and closed-loop medical intervention.
It enables personalized, real-time remote psychological support and closed-loop medical management, improves the real-time nature and accuracy of emotional intervention, and supports joint intervention by patients and their family members.
Smart Images

Figure CN122073147A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of intelligent disease management technology, specifically relating to a psychological support system and an intelligent management system for psychological support and medical closed-loop. Background Technology
[0002] With the development of precision medicine and digital medicine, the treatment pathways and efficacy for major diseases (such as malignant tumors, neurodegenerative diseases, severe post-traumatic sequelae, etc.) have been continuously improved, and the survival period of patients has been significantly extended. However, clinical practice shows that the diagnosis itself often brings a strong psychological impact. In addition, patients face problems such as physical function degeneration, family role imbalance, and increased economic pressure during the long-term treatment process after surgery, chemotherapy, and rehabilitation. As a result, patients and their families often exhibit the following characteristic difficulties in the whole process of "diagnosis-treatment-rehabilitation": (1) severe emotional fluctuations (such as depression, fear, and anger); (2) difficulty in understanding medical information and insufficient understanding of treatment pathways; (3) inability to judge the rehabilitation products or life adjustment plans that are suitable for them; (4) family members are unable to effectively participate in care decisions and lack emotional synchronization.
[0003] Currently, providing emotional support and rehabilitation guidance to patients and their families after the initial diagnosis of a disease still primarily relies on offline, manual methods, such as verbal guidance from doctors, health education leaflets, psychological counseling, and hospitalization procedure guidance. However, the development of technologies such as AI-generated content (AIGC), multimodal sensory interaction, digital therapy (DTx), and patient profiling modeling offers new possibilities for building intelligent, personalized, and contextualized emotional support and rehabilitation assistance systems.
[0004] Currently, AI tools have been explored in areas such as psychological counseling, AI video generation, and health follow-up. For example: (1) Some chronic disease management apps (such as Woebot, Wysa, etc.) have developed conversational robot modules based on cognitive behavioral therapy (CBT) to alleviate emotional problems through text communication; (2) AI character videos (such as Synthesia, HeyGen, etc.) are used for health education; (3) Follow-up and rehabilitation platforms (such as DXY's "DXY Doctor Pro" and Tencent's "WeDoctor") provide functions such as drug delivery, consultation and follow-up, and information push.
[0005] However, existing technologies have not yet truly enabled the application of such AI tools in the entire course of patient management, especially in terms of emotional and behavioral interventions. For example: (1) There is a lack of personalized content generation capabilities. Current AI systems generally lack dynamic perception and modeling of users' emotional states, cultural contexts, and behavioral feedback, resulting in highly homogenized video content that cannot meet the individual psychological support needs of patients, especially in the early stages of diagnosis of serious diseases where there is a significant lack of adaptation; (2) There is a lack of deep integration mechanism between AI-generated content and the medical system. Currently, AI content is mainly used in static scenarios such as preoperative education and health knowledge explanation. It has not yet formed a system docking mechanism with individual rehabilitation paths, intervention rhythms, product recommendations, and follow-up visits. AI-generated content cannot be adjusted in real time according to changes in the patient's condition or medical advice, nor can it be embedded in the patient management platform as an intervention node for execution, making it difficult to... (3) Lack of modeling and intervention capabilities for family support systems. Traditional follow-up and rehabilitation platforms usually take patients as the only core modeling object and lack assessment and response mechanisms for their family members—especially primary caregivers—in terms of emotional state, care pressure, cognitive level, etc. They cannot provide intelligent suggestions, video content or auxiliary decision support for family joint care, and it is difficult to form a complete support chain for joint adaptation of patients and families; (4) Broken information interaction chain and lack of closed-loop feedback mechanism. Existing systems generally lack a unified data flow structure covering patients, family members and medical platforms, resulting in emotion recognition, intervention results and user feedback being scattered in different subsystems, which cannot be integrated and processed in real time. Especially when there are significant changes in emotional state (such as an increase in the risk of depression), the system cannot realize automatic early warning, intervention plan adjustment or medical staff active intervention, which seriously restricts the real-time, personalized and accurate nature of remote intervention.
[0006] Therefore, developing an AI system that integrates personalized video emotion guidance, intelligent recommendation, and closed-loop medical intervention is a challenge in this field. Summary of the Invention
[0007] In view of the shortcomings of existing technologies, this invention provides a psychological support system and a closed-loop intelligent management system for psychological support and medical care, with the aim of achieving real-time, personalized and precise remote intervention.
[0008] This invention provides a psychological support system, which includes an AI emotion support video module, wherein the AI emotion support video module integrates the following modules: The data acquisition module is configured to collect patient information; The multimodal patient profile construction module is configured to build a patient profile from the patient information through a multimodal fusion model, wherein the multimodal fusion model includes the following parameters: voice emotion fluctuation rate V-Vol, visual situational adaptation index VSAI, and stress sensitivity index SSI. The AI-generated content-driven video generation module is configured to generate videos based on patient profiles, patient disease stage information, user identity, and preset video themes, using large-scale pre-trained language models and virtual human generation engines. The dynamic content adaptation module is configured to optimize the video generated by the AI-generated content-driven video generation module based on the patient's signals.
[0009] Preferably, in the multimodal patient profile construction module, the process of establishing a patient profile includes: using a multimodal fusion model with a Transformer architecture to extract and fuse features from different modal data such as text, speech, image, and physiology; using a hybrid emotion recognition model of bidirectional long short-term memory network and convolutional neural network to obtain structured labels and feature vectors containing emotion category and emotion intensity information; and recording them in time series form to form a patient profile.
[0010] Preferably, in the multimodal patient profile construction module, the multimodal fusion model includes: concatenating parameters and location encoding information as additional bias terms into attention calculation, fusing multimodal features through weighted summation, and dynamically adjusting loss weights using a joint loss function and a dynamic weight adjustment strategy.
[0011] Preferably, in the multimodal patient profile construction module, the method for attention calculation includes: in: Let be the multimodal feature vector, where Q, K, and V represent the query vector, key vector, and value vector obtained from the multimodal feature vector through different linear transformations, respectively, and T represents the position encoding vector. Let be the dimension of the key vector. For position encoding, t This represents the position index of each element in the sequence. The context vector for aggregation; And / or, the weighted summation method includes: The weights are calculated as follows: in, , h m Representing the characteristics of speech, images, or physiological signals, Represents the weight matrix. Represents the bias vector; And / or, the joint loss function includes: in: Let the emotion category loss function be... The weights for the loss of the emotion category, The VSAI regression loss function is used. The weights for the VSAI regression loss, For SSI multi-task loss function, The weights for the SSI multi-task loss are... Let V-Vol be the prediction loss function. Weights for predicting loss for V-Vol. For parameterized regularization loss function, The weights for the parameter regularization loss; And / or, the dynamic weight adjustment strategy includes: in or is the importance coefficient, i and j are the feature numbers, and t is the time step.
[0012] Preferably, the For the focus loss function, the For Huber's loss function, the The loss function is a combination of the classification cross-entropy loss function and the regression mean squared error loss function. This is the mean squared error loss function.
[0013] Preferably, in the multimodal patient profile construction module, the V-Vol is used to quantify the emotional fluctuation features in the user's speech. The V-Vol is obtained by the following steps: collecting the user's real-time speech information and converting it into a discrete digital speech sequence, extracting at least one speech feature, characterizing the volatility of the speech feature, extracting the speech feature through a hybrid structure model, and calculating the magnitude of the change of the speech feature over time. And / or, the VSAI is used to quantify the degree of dynamic adaptation of the user's facial emotion expression to the visual context, VSAI>0.7; And / or, the SSI is divided into intensity prediction and classification prediction. By integrating multimodal data analysis of speech, image and physiological signals, the intensity of the user's stress response to the triggering element (SSI intensity) and the stress classification result (SSI class) are predicted. The multimodal data analysis process includes multimodal feature extraction process and multimodal feature fusion process.
[0014] Preferably, in the multimodal patient profile construction module, the speech features in V-Vol include: fundamental frequency, energy, speech rate, and 13-dimensional Mel frequency cepstral coefficients. The volatility of the speech features is characterized by the coefficient of variation. The hybrid structure model consists of convolutional neural network layers and long short-term memory network layers. The hybrid structure model uses a loss function and volatility score to optimize the model parameters. And / or, the formula for calculating V-Vol includes the following: Where h represents the hidden state output by the hybrid structure model. This represents the stability correction term, where t represents the time step and T represents the total number of time steps. And / or, prior to AI-driven content generation, VSAI's calculation formula includes the following: in, Sigmoid function and For model parameters, User facial image features for initial visual content. For visual context-related text features, For feature splicing operations; And / or, after AI-generated content-driven video generation, VSAI's calculation formula includes the following: Where E(t) is the video emotion labeling function, P is the user expression distribution, and Sim is the function for calculating the cosine similarity between P and E(t). P is improved using a multi-frame weighted fusion model, and Sim is improved using an attention-enhanced matching model. The time series smoothing process is performed using the exponential moving average algorithm, where the smoothing coefficient is... ; And / or, in the SSI, the triggering elements include visual triggers, auditory triggers, and time strategies, and the stress classification results By The input is another fully connected layer, which is then processed by the Softmax function to obtain the result. And / or, in the SSI, the process of multimodal feature extraction is optimized by multimodal early fusion technology, and the process of multimodal feature fusion is optimized by dynamic threshold adaptive adjustment strategy; And / or, The calculation formulas include the following: in, This is the weight matrix. To convert audio feature vectors and physiological feature vectors Concatenate them into a longer vector. It is the bias vector; And / or, The calculation formulas include the following: in, This is the weight matrix. For bias vectors, Category labels representing the classification results, This indicates the category of "no stress". This indicates the category of "significant stress".
[0015] Preferably, the data acquisition module interfaces with the hospital information system through an interface conforming to the HL7 FHIR standard, thereby accessing patients and collecting their information. And / or, in the data acquisition module, the information collected includes: structured data such as age, gender, social role, diagnosis results, pathological stage, and treatment plan; natural language processing technology is used to extract and supplement information from unstructured text in electronic medical records; and speech recognition technology is used to collect voice data from patients and their families. And / or, in the multimodal patient profile construction module, a two-stage training strategy and an adaptive learning algorithm are used to optimize the model; And / or, the AIGC-driven video generation module includes: using a prompting engineering approach, inputting patient profile feature vectors, disease stage information, user identity and preset video themes into a large-scale pre-trained language model to generate a video script; and using 3D modeling and animation rendering technology to construct a detailed library of facial expressions and body movements for each character through the anthropomorphic role-playing of virtual doctors and patients. And / or, in the AIGC-driven video generation module, the user identity includes the patient himself / herself, the patient's parents, the patient's spouse, and the patient's children, and the preset video theme includes psychological counseling and emotional acceptance; And / or, in the dynamic content adaptation module, the patient signal includes: user video viewing time, likes, comments, and other behavioral data; And / or, in the dynamic content adaptation module, the optimization method includes: establishing a dynamic content adjustment model based on reinforcement learning, which automatically triggers the regeneration and rendering process of the video script when the user's emotional state changes, thereby realizing dynamic adjustment of the content.
[0016] This invention provides a closed-loop intelligent management system for psychological support and medical care, which includes the following modules: AI emotion support video module in the psychological support system according to any one of claims 1-7; The product and resource matching and recommendation module integrates an intelligent evaluation module, a personalized product combination recommendation module, and a supply chain system integration module. The medical intervention and monitoring module integrates a medical institution collaboration platform access module, an AI-assisted medical decision-making module, and a three-party interaction module; The feedback monitoring and follow-up support module integrates a multimodal emotion monitoring module, a dynamic service recommendation and video update module, and a system optimization module; The data acquired by the data acquisition module in the AI emotion support video module is transmitted to the multimodal emotion monitoring module. The data acquired by the data acquisition module in the AI emotion support video module, the patient profile acquired by the multimodal patient profile construction module, and the video content usage effect data acquired by the dynamic content adaptation module are respectively input into the intelligent evaluation module and the medical institution collaboration platform access module. The product recommendation feedback data obtained by the supply chain system docking module is also transmitted to the medical institution collaboration platform access module. The three-way interaction module also exchanges information with the dynamic content adaptation module, the personalized product combination recommendation module, and the system optimization module. The emotional changes obtained by the multimodal emotion monitoring module are also transmitted as patient signals to the AI emotion support video module.
[0017] Preferably, in the product and resource matching and recommendation module, The intelligent assessment module is configured to use an intelligent assessment model to determine the patient's rehabilitation needs and the priority of resource matching for various needs; The personalized product combination recommendation module is configured to generate a personalized product combination package containing rehabilitation training equipment, psychological adjustment client, nutritional supplement package, and assistive device through a hybrid recommendation model. The supply chain system integration module is configured to integrate personalized product bundles with the supply chain management system through a standardized interface. And / or, in the medical intervention and monitoring module, The medical institution collaboration platform access module is configured to synchronize patient data to the medical collaboration platform in real time through a secure data transmission protocol; The AI-assisted medical decision-making module is configured to provide personalized adjustment suggestions to medical staff through the AI-assisted decision-making model, patient profiles obtained by the multimodal patient profile construction module, and user feedback data, and to enable medical staff to formulate adjustment instructions. The three-way interaction module is configured to upload medical staff's adjustment instructions and user feedback data to the system through a real-time interaction mechanism between AI, patients, and medical staff, thereby achieving information sharing among the three parties.
[0018] And / or, in the feedback monitoring and follow-up support module, The multimodal emotion monitoring module is configured to use a multimodal fusion algorithm to display emotion change trends in the form of visual charts and monitor users' emotion changes in real time. The dynamic service recommendation and video update module is configured to use time-series prediction models to predict users' future emotional changes and rehabilitation needs based on data from the multimodal emotion monitoring module, and dynamically adjust personalized product recommendation strategies based on the prediction results. The system optimization module is configured to optimize the parameters in the system through parameter feedback loops and personalized weight update strategies.
[0019] Preferably, in the system optimization module, the parameter feedback loop strategy includes: user interaction with the system, the system collecting user feedback data and updating model parameters and attention weights, fine-tuning the updated model for the next emotion recognition and recommendation, and the feedback data including user viewing of recommended content and performing emotion self-assessment; And / or, the personalized weight update strategy includes: calculating the feedback loss based on user feedback data, and updating the parameter-related weights in the model using gradient descent, wherein the feedback loss is as follows: Based on genuine user feedback, For model prediction feedback, is the time decay factor, i represents the data point number, and N represents the total number of data points.
[0020] This invention provides a personalized dynamic psychological support system for newly diagnosed patients and their families by optimizing model parameters and data flow. The system includes an AI-powered emotional support video module that integrates parameters such as voice emotion fluctuation rate (V-Vol), visual situational adaptation index (VSAI), and stress sensitivity index (SSI) to achieve visualized and personalized emotional support. Based on this, the invention also provides a personalized dynamic psychological support and medical closed-loop management system. This system consists of four modules: an AI-powered emotional support video module, a product and resource matching and recommendation module, a medical intervention and monitoring module, and a feedback monitoring and follow-up support module. This system achieves a continuous intervention closed loop centered on the patient, and has broad application prospects in the management of patients throughout their entire disease course, especially in the management of emotional and behavioral interventions.
[0021] Obviously, based on the above description of the present invention, and according to common technical knowledge and conventional methods in the field, various other modifications, substitutions, or alterations can be made without departing from the basic technical concept of the present invention.
[0022] The following detailed embodiments further illustrate the above-described content of the present invention. However, this should not be construed as limiting the scope of the present invention to the following examples. All technologies implemented based on the above-described content of the present invention fall within the scope of the present invention. Attached Figure Description
[0023] Figure 1 Flowchart for extracting voice emotion fluctuation rate (V-Vol); Figure 2 A flowchart for the Visual Context Adaptation Index (VSAI); Figure 3 The flowchart for calculating the Stress Sensitivity Index (SSI); Figure 4 Flowchart of a parameter-guided multimodal attention fusion framework; Figure 5 This is a flowchart illustrating a personalized dynamic psychological support and medical closed-loop intelligent management system in Example 2. Figure 6 This is a flowchart of the parameter feedback mechanism. Detailed Implementation
[0024] In the following embodiments and experimental examples, the algorithms for data acquisition, transmission, storage, and processing steps not specifically described, as well as the hardware structures and circuit connections not specifically described, can all be implemented using the content already disclosed in the prior art.
[0025] Example 1: A Personalized Dynamic Psychological Support System The personalized dynamic psychological support system in this embodiment consists of an AI emotion support video module, which integrates a data acquisition module, a multimodal patient profile construction module, an AI-generated content (AIGC) driven video generation module, and a dynamic content adaptation module.
[0026] 1. The data acquisition module is configured to deeply interface with the hospital information system (HIS) through an interface conforming to the HL7 FHIR standard, to access patients and collect information.
[0027] The patient data collection method involves automatically collecting patient disease information from the hospital information system using API interface technology, including structured data such as diagnosis results, pathological stages, and treatment plans. Simultaneously, Natural Language Processing (NLP) technology is used to extract information from the unstructured text in the electronic medical records, supplementing it with disease-related details.
[0028] For the collection of information on patients and their family members, in addition to structured data such as age, gender, and social role, speech recognition technology is used to collect voice data from patients and their families. This data is combined with questionnaires (such as PHQ-9 and GAD-7) and historical medical records to construct a multi-dimensional data collection system. During the voice collection process, far-field speech enhancement technology is used to eliminate environmental noise interference and improve the quality of voice data. For questionnaire data, anti-cheating mechanisms and logical verification rules are used to ensure the authenticity and validity of the data.
[0029] 2. The multimodal patient profile construction module is configured to deeply integrate the structured and unstructured data collected by the data acquisition module through multimodal data fusion technology, and establish a dynamic and personalized patient profile through a hybrid emotion recognition model of bidirectional long short-term memory network (BiLSTM) and convolutional neural network (CNN).
[0030] The multimodal data fusion technology employs a Transformer-based multimodal fusion model to extract and fuse features from different modalities, including text, speech, images, and physiological data. The hybrid emotion recognition model uses PHQ-9 and GAD-7 questionnaire data as text feature input, employing NLP techniques for word segmentation, part-of-speech tagging, and semantic analysis. Speech data is input into the model after acoustic feature extraction using Mel-frequency cepstral coefficients (MFCC). Through joint training, the model achieves accurate assessment of the patient's current psychological state, outputting structured labels and feature vectors containing information such as emotion category (e.g., anxiety, depression, calmness) and emotion intensity, recorded in time-series format to form a dynamic and personalized patient profile.
[0031] During the acquisition of multimodal data, issues such as time asynchrony and data gaps often arise due to differences in equipment and acquisition time. Therefore, synchronous preprocessing is fundamental to ensuring the accuracy of subsequent analysis. First, timestamp alignment is performed using a specific algorithm, with the timestamps of the speech data as a benchmark, to calibrate the timestamps of the image frame sequences and physiological signals, ensuring temporal consistency across modalities. Next, missing value handling is performed; for a small number of missing data points, linear interpolation is used to fill in the gaps, while data with significant missing values is discarded. Then, filtering algorithms are used to reduce noise and remove interference from environmental noise. Next, standardization is applied to normalize the data to a specific range, eliminating dimensional differences. Finally, to increase data diversity and improve model generalization ability, data augmentation techniques are employed, such as adding background noise to the speech data and randomly flipping the images.
[0032] Extraction of specific features from multimodal data: Speech Feature Extraction: Speech data contains rich emotional information. The spectral features of the speech are extracted using the MFCC (Mel-frequency cepstral coefficients) algorithm, while prosodic information such as fundamental frequency and energy is extracted using prosodic features. These two types of features are concatenated and input into a network composed of BiLSTM and CNN for temporal feature modeling and deep feature extraction, resulting in a speech feature representation. .
[0033] Image feature extraction: First, a face detection algorithm is used to locate the face region in the image. Then, the face image is input into the MobileNetV2 network to extract facial expression features. Finally, the facial expression features are further processed using a Transformer network to obtain the image feature representation. .
[0034] Physiological signal processing: For physiological signals, such as skin conductance response (GSR) data, filtering is first performed to remove noise interference, followed by normalization. For physiological signals related to heart rate variability (HRV), relevant features are extracted. The processed GSR and HRV features are concatenated and then subjected to attention pooling to obtain the physiological signal feature representation. .
[0035] Parameter-guided multimodal fusion mechanism: In multimodal data fusion technology, accurate quantification of the features of each modality is crucial for achieving intelligent emotion recognition. By defining the core parameters of speech, vision, and physiological signals, a standardized feature representation system is constructed, providing a quantitative basis for subsequent cross-modal fusion and dynamic adaptation. Table 1 below shows the design framework of the system's core parameters.
[0036] Table 1 (1) Voice emotion fluctuation rate (V-Vol) Voice emotion volatility (V-Vol) aims to accurately characterize a user's emotional stability by quantifying the fluctuations in pitch, intensity, and speech rate during speech, thereby helping to determine the degree of anxiety or depression. Specifically, it obtains numerical indicators reflecting the amplitude and frequency of emotional fluctuations through multi-dimensional feature analysis and mathematical modeling of speech signals. The voice emotion volatility extraction process is as follows: Figure 1 As shown.
[0037] (1.1) Data Acquisition: Use a high-fidelity microphone at a sampling frequency (Unit: Hz) Real-time acquisition of user speech is performed, converting the analog speech signal into a discrete digital speech sequence x[n], where This represents the length of the speech sequence. To ensure the validity and consistency of the data, the microphone needs to be calibrated during the acquisition process, and appropriate gain parameters need to be set to avoid signal overload or weakness.
[0038] (1.2) Feature extraction: (1.2.1) Basic speech feature extraction The OpenSMILE tool was used to extract Prosodic features from the acquired speech signal x[n], including fundamental frequency (F0), energy (RMS Energy), speech rate, and 13-dimensional Mel-Frequency Cepstral Coefficients (MFCCs). The specific calculation process is as follows: (1.2.2) Fundamental frequency (F0) extraction: The fundamental frequency of the speech signal is calculated using the autocorrelation function or the cepstrum method. F 0[n], its core principle is to find the fundamental frequency period of the periodicity of the speech signal. T 0, thus obtaining the fundamental frequency. F 0 = 1 / T 0. In actual calculations, the following formula can be used: in, L To analyze window length, TThis represents the possible range of values for the fundamental frequency period.
[0039] (1.2.3) Energy (RMS Energy) Calculation: The root mean square energy is used to measure the strength of a speech signal. The calculation formula is as follows: in, M Calculate the window length for energy.
[0040] (1.2.4) Speech rate calculation: The speech rate S is determined by statistically counting the number of effective vocal frames or words in the speech signal per unit time. Assuming that within the time interval [ t 1, t Within [2], the effective number of sound frames is N f Then the speech rate S = N f / ( t 2- t 1).
[0041] (1.2.5) 13-dimensional MFCC extraction: MFCC is a speech feature based on the characteristics of human hearing. Its calculation process mainly includes pre-emphasis, framing, windowing, short-time Fourier transform (STFT), Mel filter bank, logarithmic operation, and discrete cosine transform (DCT). The specific steps are as follows: Pre-emphasis: The speech signal x[n] is pre-emphasized to enhance high-frequency components. The formula is as follows: , where α is the pre-weighting coefficient, which is usually between 0.9 and 1.0.
[0042] Framing: dividing the pre-emphasized speech signal Divided into frames with a length of Frame shift is short frame sequences ,in , This represents the total number of frames.
[0043] Windowing: For each frame Apply Hamming window , obtain the windowed frame The Hamming window formula is , .
[0044] Short-Time Fourier Transform (STFT): Perform STFT on each windowed frame z_m[n] to obtain its spectrum. The formula is .
[0045] Mel filter bank: This filters the spectrum after STFT. The Mel spectrum is obtained through the Mel filter bank. The formula is ,in Let i be the frequency response of the i-th Mel filter. , This represents the number of Mel filters.
[0046] Logarithmic operation: on the Mel spectrum Taking the logarithm, we get .
[0047] Discrete Cosine Transform (DCT): for Perform DCT and extract the first 13 coefficients as MFCC features, i.e. .
[0048] (1.3) Feature framing and statistical calculation The extracted speech features are processed using a 1-second sliding window frame segmentation method, with each frame containing F sampling points. For each feature in each frame, its mean is calculated. and standard deviation The formulas are as follows: in, Let be the feature value of the i-th sampling point in this frame.
[0049] (1.4) Characterization of volatility The coefficient of variation (CV) is used to characterize the volatility of each feature. The coefficient of variation is defined as the ratio of the standard deviation to the mean, and the formula is: By calculating the coefficients of variation of features such as fundamental frequency, energy, speech rate, and 13-dimensional MFCC, a vector containing the volatility of multiple features is obtained. .
[0050] (1.5) Modeling method (1.5.1) Model Structure A CNN-LSTM hybrid model is constructed to model the extracted speech features. This model consists of Convolutional Neural Network (CNN) layers and Long Short-Term Memory (LSTM) layers. The CNN layers are used to extract local spatial features of the speech features, while the LSTM layers are used to capture the temporal series information of the speech features.
[0051] CNN layers: Multiple convolutional layers with different kernel sizes are used to perform convolution operations on the input feature matrix. Assume the input feature matrix is... (15-dimensional features × 50 frames), convolution kernel size is With a stride of s and padding of p, the output feature map after the convolutional layer is... The calculation formula is: in, The convolution kernel weight matrix is... This is a bias term.
[0052] STM layer: LSTM cells pass through the forget gate Input gate Output gate and cell state To process time series information, the calculation formula is as follows: in, Here, is the sigmoid function, and tanh is the hyperbolic tangent function. This indicates element-wise multiplication. This is the weight matrix. For bias terms, This indicates that the hidden state from the previous moment will be restored. and the input at the current moment Then, the parts are assembled.
[0053] (1.5.2) Model Training and Output The feature vector obtained after feature extraction The model is trained using a CNN-LSTM hybrid architecture, with the input being the CNN-LSTM input. The mean squared error (MSE) is used as the loss function during training, as shown in the formula: Where N is the number of training samples, For real labels, These are the model's predicted values. The model parameters are optimized using stochastic gradient descent (SGD) or its improved versions, such as the Adam algorithm.
[0054] The model output is a continuous numerical volatility. This indicates the degree of emotional fluctuation in speech. Simultaneously, by detecting abrupt changes in the volatility sequence, the number of abrupt changes is calculated. The final volatility score is formed by combining the volatility V. The formula is: in, and These are weighting coefficients used to balance the continuous volatility V (reflecting the overall strength of sentiment fluctuations) with the number of abrupt change points. The impact of (the frequency of sudden mood changes) on the final score, combined with the quantitative requirements of voice emotion features and multimodal parameter collaborative logic, The recommended value range is 0.6 to 0.8. The suggested value range is 0.2 to 0.4. The specific value needs to be determined through iterative optimization using controlled experiments based on voice emotion annotation data to ensure consistent scoring. Accurately depict the emotional fluctuations in speech.
[0055] The speech volatility is obtained by calculating the magnitude of changes in speech features over time, calculating the relative differences between speech features at adjacent time steps, and summing and averaging them.
[0056] Note: In the above formula, The hidden states output by the LSTM layer in the CNN-LSTM hybrid architecture model mentioned earlier "Consistent" refers to the temporal features of the speech modality after model processing; here, it is uniformly expressed as... ; For numerical stability correction terms (take the smallest positive number, such as...) The magnitude (t) is used to avoid calculation anomalies caused by a denominator of zero; t represents the time step, and T represents the total number of time steps.
[0057] Through the above process, the Voice Emotional Volatility (V-Vol) module can more accurately and comprehensively quantify the emotional fluctuation characteristics in a user's voice.
[0058] (2) Visual Context Adaptation Index (VSAI) The Visual Situation Adaptation Index (VSAI) quantifies the dynamic adaptation between a user and a visual context, providing a basis for content adaptation decisions in intelligent interaction systems. Before AIGC video generation, an auxiliary VSAI is generated based on the user's facial image features and visual context-related text features. After AIGC video generation, the core VSAI is generated by quantifying the dynamic matching degree between the user's facial emotional expression and the emotions guided by the AIGC video.
[0059] 1. VSAI Assist (Applicable before AIGC video generation) Before AIGC video generation, it is achieved through the fusion calculation of image features and text features, that is, the user's facial image features. Text features related to visual context (As described in the preset scenario) Concatenate the data, input it to a fully connected layer and activate it using the Sigmoid function, then output a value in the range of 0-1: in, Sigmoid function and For model parameters, For initial visual content (such as static images and text on the system's introductory interface, preset science materials, etc.), the user's facial image features are used. For visual context-related text features, This is a feature splicing operation.
[0060] Output The range of values reflects the user's degree of adaptation to the initial visual context before video generation. This auxiliary VSAI is only used in the transitional stage before AIGC video generation, and its main functions include: providing initial visual feature data for multimodal patient profiling, and assisting the AIGC video generation module in determining initial parameters (such as adjusting the video theme style based on the user's adaptation to the initial images and text).
[0061] 2. Core VSAI (applicable after AIGC video generation) After AIGC video generation, the dynamic matching degree between user facial emotion expression and AIGC video-guided emotion is quantified. Specifically, the cosine similarity between the user's facial expression distribution P and the video emotion label function E(t) is calculated, and after normalization and time-series smoothing, the result is: in, Distribution of user facial expressions With video emotion tagging function cosine similarity, Time series smoothing is performed using the exponential moving average algorithm (smoothing coefficient). ). The two VSAI approaches described above are connected through the scenarios of "initial assessment - video generation - dynamic adaptation," together covering the entire process.
[0062] The core VSAI is implemented through the following technical paths: Real-time emotion perception: Analyzing user facial expression distribution based on multi-frame facial images; Contextual semantic mapping: Constructing a time-series model of sentiment labels for video content; Dynamic matching evaluation: The cosine similarity algorithm is used to calculate the emotional compatibility index, and the final output is a continuous value in the range of [0,1], which is used to drive adaptive interaction logic such as video switching and repeat playback.
[0063] The process of Visual Context Adaptation Index (VSAI) is as follows: Figure 2 As shown.
[0064] (2.1) Image acquisition and face localization system (2.1.1) Multi-task cascaded convolutional network (MTCNN) A three-stage cascaded structure is used to achieve face detection and key point localization: Proposal Network (P-Net): Generates candidate face regions using a fully convolutional network. Each detection box is represented as: in The coordinates of the top left corner Width and height, This is the confidence score.
[0065] Refine Network (R-Net): For The output proposal is subjected to bounding box regression and non-maximum suppression (NMS) to optimize the detection box coordinates: in This is the regression offset.
[0066] Output Network (O-Net): Outputs the coordinates of 5 facial key points. Each key point is represented as: (2.1.2) Image preprocessing pipeline Pose normalization: based on binocular key points Calculate the rotation angle : Affine transformation is used to correct the face to a standard pose: in The coordinates are the center coordinates of the face.
[0067] Region clipping and scaling: Expand the detection bounding box by 1.5 times its side length to form the ROI region, and scale it to a fixed size. (e.g., 224×224), and then perform pixel normalization: (2.2) Lightweight Facial Expression Recognition Model (FER) (2.2.1) Feature Extraction Network Architecture The MobileNetV2-FER model is adopted, and its core structure includes: Depthwise separable convolution: For input images Perform layer-by-layer feature extraction, and output the feature map of layer l: in For convolution kernel, This indicates a convolution operation.
[0068] Global Average Pooling (GAP): Compresses the last layer of feature maps into a 128-dimensional feature vector. in The size of the feature map.
[0069] (2.2.2) Calculation of Facial Expression Probability Distribution Seven types of facial expression probability distributions are generated using a fully connected layer and a Softmax function: in: Weight matrix For bias vector Softmax function definition: P Satisfying the probability normalization condition .
[0070] (2.2.3) Emotion matching scoring model (2.2.3.1) Video emotion tag timeline modeling Define the video emotion labeling function E(t) as a piecewise continuous probability distribution sequence, where each time window... Corresponding to specific emotion distributions: Each emotion tag It is a 7-dimensional unit vector that satisfies .
[0071] Example: The emotion tag for a 60-second video is defined as: (2.2.3.2) Cosine similarity measure Calculate the user's facial expression distribution P and the target emotion in the video. Cosine similarity: because For a unit vector, the formula simplifies to: (2.2.3.3) VSAI value normalization Map the cosine similarity to the interval [0,1]: when hour, (Exact match); when hour, (Complete mismatch).
[0072] (2.2.3.4) Time series smoothing Using exponential moving averages (EMA) to reduce the impact of noise: smoothness coefficient initial value .
[0073] (2.3) Extended Mathematical Model (2.3.1) Multi-frame weighted fusion model Considering the temporal continuity of facial expression changes, the probability distributions of the first K frames are weighted and fused: The weights satisfy the decay property: Where τ is the time decay constant, which controls the influence weight of historical frames.
[0074] (2.3.2) Attention-enhanced matching model Introducing attention weights for emotion categories Highlighting key emotional match: Weight parameters It can be trained using user historical data or manually set using task objectives (such as assigning higher weight to the emotion of "happiness").
[0075] (2.4) System performance indicators Table 2 (3) Stress sensitivity index (SSI) The Stress Sensitivity Index (SSI) accurately quantifies the intensity of a user's stress response during an "emotional trigger test segment" in a video by integrating multimodal data analysis of speech, images, and physiological signals. Its core functionality focuses on: Capture significant stress behaviors triggered by sudden changes in visuals, keyword stimuli, and other unexpected scenarios; Output binary classification decision results (significant stress / no significant stress) and continuous intensity values in the 0-1 interval; This provides a quantitative basis for personalized intervention strategies such as dynamically adjusting the intensity of video content and inserting emotional buffering segments.
[0076] The calculation process for the Stress Sensitivity Index (SSI) is as follows: Figure 3 As shown.
[0077] (3.1) Test segment triggering mechanism design (3.1.1) Trigger element design The system embeds standardized "emotion-triggered test segments" (fixed duration) into the video content. It includes three types of triggering elements: Visual triggers: high-contrast images with instantaneous brightness changes exceeding 40%, or specific semantic keywords (such as "hospitalization" or "complications"); Auditory triggers: a sudden increase in speech rate of more than 30% or a change in tone of voice with a fundamental frequency jump of 50Hz; Time strategy: Test segment start time follow Evenly distributed to avoid interference at the beginning and end of the video.
[0078] (3.1.2) Mathematical Modeling of Trigger Signals The trigger time window is defined as: (3.2) Multimodal feature extraction and discrimination (3.2.1) Speech signal feature analysis (temporal change detection) Key features are extracted using a 1-second sliding window (0.5-second step size): Fundamental frequency mutation index: Energy mutation index: (3.2.2) Image signal feature analysis (expression dynamic monitoring) Probability distribution of 7 types of facial expressions based on the output of the FER model Construct a two-dimensional detection: KL divergence mutation: Expression category transfer marker: (3.2.3) Physiological signal characteristic analysis (GSR skin conductance response) Baseline of skin conductance response (GSR) The average value of the first 5 seconds of the test segment : Mutation amount is defined as: when Time markers for physiological mutations : (3.2.4) Multimodal feature fusion strategy The three types of signals are normalized and then concatenated. T ×3 timing matrix (T = 20 time steps, 0.5 seconds / step): (3.3) BiLSTM-Attention Model Architecture (3.3.1) Bidirectional Long Short-Term Memory Network (BiLSTM) Forward LSTM cell computation: For the Sigmoid function, For element-wise multiplication, (for learnable parameters) Total hidden state: (3.3.2) Attention mechanism Attention score calculation: Context vector aggregation: (3.3.3) Output layer design SSI (Stress Sensitivity Index) consists of two parts: intensity prediction and classification prediction, as defined below: First, the speech features Physiological signal characteristics After concatenation, the input is fed into a fully connected layer, and the stress intensity value is obtained through the Sigmoid function: in, This is the weight matrix. To extract speech features and physiological signal characteristics Concatenate them into a longer vector. This is the bias vector.
[0079] The stress intensity values are then input into another fully connected layer, and processed by the Softmax function to obtain the stress classification results (binary classification: no stress / significant stress): in, This is the weight matrix. For bias vectors, Category labels representing the classification results, This indicates the category of "no stress". This indicates the category of "significant stress".
[0080] This definition uses a progressive process of "intensity prediction → classification prediction" to form a complete SSI calculation logic, with the intensity prediction result as the classification input.
[0081] Joint loss function in, It is a loss function used to measure the intensity of stress.
[0082] (3.4) Extended Model and Optimization Strategy (3.4.1) Multimodal early fusion technology In the feature extraction stage, the speech fundamental frequency, original image pixels, and original physiological signal values are directly fused to construct... A high-dimensional feature matrix (D is the total feature dimension) enhances the model's ability to capture subtle stress signals.
[0083] (3.4.2) Adaptive adjustment of dynamic threshold Dynamically update mutation thresholds based on user historical data: ( As the initial threshold, For adjustment coefficients, (Historical standard deviation) (3.5) System performance indicators Table 3 (4) Parameter-guided cross-modal attention The three parameters VSAI, SSI, and V-Vol are concatenated with the positional encoding information and incorporated as an additional bias term into the attention calculation. In this way, when calculating attention weights, the model considers not only the relationships between modal features but also the user's current state information, enabling the model to focus more on key information related to the user's state.
[0084] in, This represents a multimodal parameter vector consisting of the Visual Context Adaptation Index (VSAI), the Stress Sensitivity Index (SSI), and the Voice Emotional Volatility Rate (V-Vol). Q, K, and V represent the query vector, key vector, and value vector obtained from the multimodal feature vector through different linear transformations, respectively, and are used for the calculation of attention weights. T represents the positional encoding vector, which is used to model the temporal information of the input features; The dimension of the key vector is used to scale the dot product of Q and K to prevent the value from being too large and affecting the computational stability of softmax.
[0085] That is, the aggregated context vector obtained through the Concat (concatenation) operation. Position encoding is used to capture timing information. t It represents the position index of each element in the sequence, which is used in conjunction with (position encoding) to capture the order information of multimodal features in the time or sequence dimension, ensuring that the model can recognize the sequential relationship of different elements in the sequence; (W) p () is the weight matrix, used to weight the context vector. Perform a linear transformation. (5) Multimodal weighted fusion After obtaining the features of each modality and the attention weights guided by parameters, the multimodal features are fused through a weighted summation. The weights are calculated based on the features of each modality and the core parameters, enabling the model to dynamically adjust the importance of different modalities according to the user's state. For example, when a user's SSI value is high, indicating a state of stress, they may pay more attention to the emotional information conveyed by physiological and speech signals, and the weight of the corresponding modality will be increased accordingly.
[0086] The weights are calculated as follows: in, P , which is a multimodal feature vector; h m They represent h audio , h image , h physio , h audio Representing speech features, h image Representing image features, h physio Indicates physiological signal characteristics; Represents the weight matrix. This represents the bias vector, used to adjust the baseline value of the weights.
[0087] (6) Multi-task learning framework (6.1) Joint Loss Function To simultaneously optimize multiple tasks, including sentiment classification, VSAI prediction, SSI prediction, and V-Vol prediction, a joint loss function is designed. Different loss function forms are adopted to address the characteristics of different tasks. Sentiment classification uses focus loss to address class imbalance; VSAI prediction uses Huber loss for better robustness to outliers; SSI prediction combines classification cross-entropy loss and regression mean squared error loss; and V-Vol prediction uses mean squared error loss. Furthermore, to prevent overfitting, parameter regularization loss is introduced. By adjusting the weights of the losses for each task, collaborative optimization of multiple tasks is achieved.
[0088] in: Emotion categorization loss (focus loss) VSAI regression loss (Huber loss) SSI multi-task loss (classification + regression). V-Vol is the predicted loss (mean squared error). : Parameter regularization loss.
[0089] (6.2) Dynamic weight adjustment To enable the model to dynamically adjust its training focus based on the difficulty and importance of different tasks, a dynamic weight adjustment strategy was designed. Based on the loss values of each task in the previous training step, the weights of the losses for each task in the current step are calculated using an exponential function. Tasks with larger loss values receive increased weights, thus gaining more attention in subsequent training and achieving adaptive adjustment of the model's training.
[0090] in or is the importance coefficient, i and j are the feature numbers, and t is the time step.
[0091] (7) Training process and optimization strategies (7.1) Two-stage training strategy To improve the efficiency and effectiveness of model training, a two-stage training strategy is adopted. In the pre-training stage, the multimodal data is first aligned, and then the feature extraction network and parameter calculation module of each modality are trained independently, enabling the model to initially learn the features and parameter calculation methods of each modality. In the joint fine-tuning stage, the entire network is trained jointly. Through parameter-guided attention training and end-to-end optimization, the model parameters are further adjusted, enabling the model to better integrate multimodal data and core parameters, thereby improving overall performance.
[0092] (7.2) Adaptive learning algorithm To adapt to the characteristics of different tasks and data and improve model training efficiency, an adaptive learning rate scheduler was designed. This scheduler maintains a counter and a record of the best loss for each task. During training, the counter is updated based on the loss changes for each task. When the loss of a task does not decrease within a certain number of steps, it is considered that the task is trapped in a local optimum or the training difficulty is high. At this point, the learning rate for that task is reduced to help the model escape local optima or to fine-tune the parameters. In this way, the learning rate for different tasks is dynamically adjusted, improving the model's training performance.
[0093] Based on this, this embodiment constructs a parameter-guided multimodal attention fusion framework, the process of which is as follows: Figure 4 As shown, traditional multimodal fusion methods often involve simply concatenating or weighting data features, lacking awareness and utilization of the user's real-time state. However, the three core parameters introduced in this solution quantify the user's state from three dimensions: visual, physiological / psychological, and vocal. During model operation, these parameters participate in the multimodal data fusion decision-making process, enabling the model to dynamically adjust its focus on different modalities based on the user's real-time state, thereby improving the effectiveness of emotion recognition and recommendation.
[0094] 3. The AIGC-driven video generation module is configured to automatically generate personalized video scripts based on patient profiles, using large-scale pre-trained language models (such as the GPT series) combined with a virtual human generation engine. Employing a prompt engineering approach, it inputs patient profile feature vectors, disease stage information, user identity (such as the patient, their parents, spouse, or children), and preset video themes (such as psychological counseling or emotional acceptance) into the GPT model to generate logically coherent and emotionally resonant video scripts.
[0095] The videos feature anthropomorphic characters such as virtual doctors and patients undergoing rehabilitation. Utilizing 3D modeling and animation rendering techniques, a detailed database of facial expressions and body movements is built for each character. During video rendering, multimodal synthesis technology is employed, combined with a text-to-speech (TTS) emotional prosody control algorithm to generate corresponding speech based on the script's emotional tone. A deep learning-based facial expression animation synchronization algorithm achieves precise matching between speech and facial expressions; for example, adjusting the character's facial expressions such as raising the corners of their mouth and frowning in real time based on the intonation of the speech, enhancing the video's realism and emotional impact.
[0096] 4. The dynamic content adaptation module is configured to optimize the video generated by the AIGC-driven video generation module based on the patient's signals.
[0097] A dynamic content adjustment model based on reinforcement learning (RL) is established, using patient feedback (such as behavioral data like video viewing time, likes, and comments) and emotional changes (obtained through multimodal emotion monitoring) as reward signals to optimize video content generation strategies. When the patient's emotional state changes, the model automatically triggers the regeneration and rendering process of the video script, achieving dynamic content adjustment.
[0098] Video content is pushed via mobile apps or mini-programs, employing a personalized recommendation algorithm based on user behavior prediction. This algorithm combines time series analysis and collaborative filtering techniques to analyze factors such as the patient's historical viewing behavior, current time, and geographical location, predicting video content the patient might be interested in and pushing it accurately at the appropriate time, supporting on-demand playback and personalized recommendations.
[0099] Example 2: A Personalized Dynamic Psychological Support and Medical Closed-Loop Intelligent Management System The system in this embodiment is a four-module closed-loop service system for AI support for newly diagnosed patients and their families. The four modules are: AI emotion support video module, product and resource matching and recommendation module, medical intervention and monitoring module, and feedback monitoring and follow-up support module.
[0100] like Figure 5As shown, an intelligent closed-loop service platform for patient treatment and rehabilitation is constructed through multimodal data collection, AI-driven personalized emotion support video generation, intelligent product recommendation, dynamic medical intervention and feedback monitoring. The AI emotion support video module is the core, and the other three modules form auxiliary support.
[0101] I. AI Emotion Support Video Module The AI emotion support video module was constructed according to the method in Example 1.
[0102] II. Product and Resource Matching Recommendation Module The product and resource matching and recommendation module integrates an intelligent evaluation module, a personalized product combination recommendation module, and a supply chain system integration module.
[0103] 1. The intelligent assessment module is configured to determine the patient's rehabilitation needs and the priority of resource matching for various needs based on the intelligent assessment model. At the same time, the intelligent assessment model dynamically adjusts the assessment parameters by continuously learning from the patient's feedback data on the effects of using the product, thereby improving the accuracy of prediction.
[0104] The intelligent assessment model is as follows: the Barthel Index and other daily living ability scales are used to assess the patient's daily living ability. Based on the patient's disease stage, daily living ability assessment results, and patient profiles obtained by the multimodal patient profile construction module, an intelligent assessment model integrating XGBoost and random forest algorithms is constructed.
[0105] The method for determining the resource matching priority of various patient needs based on the intelligent assessment model is as follows: The intelligent assessment model conducts multi-dimensional analysis of the patient's rehabilitation needs, including physical function recovery needs, psychological adjustment needs, nutritional supplementation needs, etc., and uses the analytic hierarchy process (AHP) to determine the resource matching priority of various needs.
[0106] 2. The personalized product bundle recommendation module is configured to generate personalized product bundles containing rehabilitation training equipment, a psychological adjustment app, nutritional supplements, assistive devices, etc., through a hybrid recommendation model. For the recommendation results, an online learning mechanism based on user feedback is employed. Based on patient behavior data such as clicks, purchases, and reviews of recommended products, the weight parameters of the recommendation algorithm are adjusted in real time to optimize the recommendation effect.
[0107] The hybrid recommendation model is based on a hybrid recommendation algorithm, which is obtained by weighted fusion of content recommendation (CB) and collaborative filtering (CF) algorithms. In content-based recommendation, textual features such as product functionalities, applicable disease types, and usage scenarios are extracted, and the similarity to the patient's rehabilitation needs (obtained by the intelligent assessment module) is calculated using the TF-IDF algorithm. In collaborative filtering recommendation, the product usage history of other patients with similar disease stages, living abilities, and emotional states to the target patient is analyzed to uncover potential product preferences.
[0108] 3. The supply chain system integration module is configured to seamlessly integrate the personalized products required by patients (obtained by the personalized product combination recommendation module) with the supply chain management system through a standardized RESTful API interface.
[0109] A microservice architecture is used to design the interface interaction process, ensuring the stability and efficiency of data transmission. For real-time inventory queries, message queue technology (such as RabbitMQ) is employed to achieve asynchronous data interaction between the system and the supply chain system, reducing query latency. During automatic order placement, a distributed transaction processing mechanism (such as Seata) is introduced to ensure the consistency and integrity of order data. By integrating a logistics tracking API, product delivery status information is obtained and pushed to patients' mobile devices in real time, enabling full tracking of the delivery process.
[0110] III. Medical and Nursing Intervention and Monitoring Module The medical intervention and monitoring module integrates a medical institution collaboration platform access module, an AI-assisted medical decision-making module, and a three-party interaction module.
[0111] 1. The medical institution collaboration platform access module is configured to synchronize patient data and status to the medical collaboration platform in real time through secure data transmission protocols (such as HTTPS, SSL / TLS).
[0112] WebSocket technology is used to achieve bidirectional real-time data communication, ensuring that medical staff can promptly obtain patients' latest emotional and usage feedback, as well as dynamically updated health profiles. During data synchronization, data encryption technologies (such as the AES encryption algorithm) are used to encrypt sensitive information to protect patient privacy and security; at the same time, a strict access control mechanism is set up to assign different data viewing and operation permissions according to the roles and responsibilities of medical staff.
[0113] 2. The AI-assisted medical decision-making module is configured to provide medical staff with personalized video content adjustment suggestions (such as increasing the explanation time of specific topics, adjusting the performance style of the role), product combination optimization solutions (such as replacing unsuitable products, supplementing relevant products), and intervention frequency settings (determining the follow-up interval based on the patient's recovery progress) through AI-assisted decision-making models, patient profiles (obtained by the multimodal patient profile construction module), video content usage effects (obtained by the dynamic content adaptation module), and product recommendation feedback (obtained by the supply chain system docking module), assisting medical staff in formulating precise adjustment instructions.
[0114] 3. The three-party interaction module is configured to upload medical staff's adjustment instructions and patients' feedback data to the system through a real-time interaction mechanism between AI, patients, and medical staff, thereby achieving efficient collaboration and information sharing among the three parties.
[0115] Based on the intervention strategies obtained by the AI-assisted medical decision-making module, medical staff make remote adjustment instructions. These instructions are transmitted in real time to the corresponding modules of the system via message queues. For example, instructions to adjust video content parameters (such as script modification or character action adjustment) are transmitted in real time to the system's dynamic content adaptation module via message queues, and instructions to adjust product plans (adding or deleting products) are transmitted in real time to the system's personalized product combination recommendation module via message queues.
[0116] Patient feedback on mobile devices (such as suggestions on video content and product usage issues) is synchronized to the system optimization module in the system's feedback monitoring and follow-up support module. After analyzing and processing the feedback data, the system uses it to optimize the AI model on the one hand, and pushes key information to the medical intervention and monitoring module on the other hand, forming a dynamic closed-loop management.
[0117] IV. Feedback Monitoring and Follow-up Support Module The feedback monitoring and follow-up support module integrates a multimodal emotion monitoring module, a dynamic service recommendation and video update module, and a system optimization module.
[0118] 1. The multimodal emotion monitoring module is configured to monitor patients' emotional changes in real time through a multimodal fusion algorithm and display the trend of emotional changes in the form of visual charts, providing a basis for subsequent service adjustments.
[0119] The multimodal fusion algorithm integrates questionnaire data and voice emotion data at the feature level, and uses DS evidence theory to make decisions on the fused features, enabling real-time monitoring of patient emotion changes. The questionnaire data and voice emotion data are obtained as follows: patients periodically complete psychological assessment questionnaires (PHQ-9, GAD-7, etc.) through the system. The system automatically recognizes the questionnaire content using optical character recognition (OCR) technology, and combines NLP technology for semantic understanding and sentiment analysis. Simultaneously, the system collects the patient's voice emotion data and employs a deep feature fusion-based voice emotion recognition algorithm, fusing MFCC features, spectrogram features, etc., and inputting them into an improved ResNet neural network model to achieve high-precision recognition of voice emotions.
[0120] 2. The dynamic service recommendation and video update module is configured to predict patients' future emotional changes and rehabilitation needs trends using time-series prediction models (such as LSTM and Prophet) based on data from the multimodal emotion monitoring module. According to the prediction results, the personalized product recommendation strategy (personalized product combination recommendation module) is dynamically adjusted, such as recommending products needed for the next rehabilitation stage in advance. Simultaneously, the AI emotion support video module is triggered to update video content, generating new video scripts that match the patient's future emotional state and rehabilitation needs, and promptly pushing them to the patient's mobile device, thus achieving personalized and dynamic service optimization.
[0121] 3. The system optimization module is configured to upload all patient data to the cloud platform via a secure data transmission channel, using distributed storage technology (such as the Hadoop Distributed File System HDFS) for storage management. Big data analytics (such as Spark) are used to evaluate service effectiveness, establishing an evaluation indicator system across multiple dimensions, including video viewing effects (such as completion rate and repeat viewing counts), product usage feedback (such as purchase conversion rate and user reviews), and the degree of improvement in patient emotions.
[0122] Machine learning algorithms (such as cluster analysis and regression analysis) are used to mine evaluation data, identify problems and optimization opportunities in the service process, such as identifying the preferences of certain types of patients for specific video topics or the reasons for poor recommendation results of certain products. In this way, AI model parameters are optimized, recommendation algorithm strategies are adjusted, and system functional modules are improved to continuously enhance the overall intelligence level of the system.
[0123] The system optimization module uses a parameter feedback mechanism. (Flowchart shown) Figure 6 As shown.
[0124] (1.1) Parameter feedback loop In practical applications, the model collects feedback data through user interaction, forming a closed-loop parameter feedback loop. Users interact with the system, such as watching recommended content and performing emotion self-assessments. The system collects this feedback data, including user behavior data (such as viewing time and skip counts), emotion self-assessment data, and physiological change data. Based on the feedback data, the model parameters are updated, and attention weights are adjusted. The updated model is then fine-tuned and used for the next emotion recognition and recommendation, thereby continuously optimizing model performance to better adapt to user needs.
[0125] (1.2) Personalized weight update Based on user feedback data, a feedback loss is calculated, and the weights related to the core parameters in the model are updated using gradient descent. The feedback loss reflects the difference between the model's predictions and the user's actual feedback. By minimizing the feedback loss, the model can better adapt to the user's personalized needs, enabling personalized weight updates and improving the model's personalized service capabilities.
[0126] The feedback loss is defined as: Real user feedback Model prediction feedback, : Time decay factor, i represents the data point number, and N represents the total number of data points.
[0127] Through the above embodiments, the system of the present invention achieves the following: (1) Parameter-guided multimodal fusion: Breaking through the traditional multimodal fusion mode, the three core parameters VSAI, SSI and V-Vol are deeply integrated into the multimodal fusion process as the core input of the attention mechanism to realize dynamic weight allocation, so that the model can flexibly adjust the attention to different modal data according to the user's state.
[0128] (2) Physiological-psychological cross-modal association: By using the SSI index, the physiological and psychological responses of users to stimulating content are quantified, and a cross-modal mapping relationship between physiological signals and psychological emotions is established, providing more comprehensive information for emotion recognition.
[0129] (3) Adaptive learning framework: through dynamic weight adjustment and adaptive learning rate.
[0130] In summary, the system of the present invention has the following beneficial effects: 1. Visualized and personalized emotion support This invention, through multimodal parameter fusion, can accurately capture a patient's emotional state in the early stages of diagnosis, including parameters such as vocal emotion fluctuation rate and visual situational adaptation index, analyzing the patient's emotions from multiple dimensions such as voice and facial expressions. Based on these accurate emotion recognition results, the system automatically generates visual support content tailored to the individual characteristics of each patient, including customized videos. For example, for patients with lower levels of education who have difficulty understanding complex medical information, the system generates easy-to-understand videos that combine text and images and incorporate familiar life scenarios, helping patients quickly accept their emotions and adapt to the changes in their lives after illness. This is an effect that traditional static, single-information presentation methods cannot achieve.
[0131] 2. Alleviate pressure on medical resources Traditional methods rely on in-person emotional support and rehabilitation guidance from mental health / medical professionals, which is highly manpower-intensive, time- and space-constrained, and inefficient. This invention's AI-powered emotional support system operates 24 / 7, automatically responding to patient needs. Patients do not need to make appointments or wait for professionals; they can receive emotional support and rehabilitation guidance anytime, anywhere. Taking common emotional counseling scenarios as an example, the system can quickly process a large number of patient inquiries. Preliminary estimates suggest that it can reduce the burden of manual counseling by approximately 70%-80%, significantly alleviating the strain on medical resources and allowing professionals to focus their energy on more complex and urgent medical matters.
[0132] 3. Full-cycle tracking and personalized intervention to improve rehabilitation compliance. The system leverages continuous collection and analysis of multimodal data to track patients throughout their entire recovery cycle. From the initial diagnosis to the later stages of recovery, intervention strategies are dynamically adjusted based on the patient's emotional changes, physical condition, and treatment progress at different stages. During chemotherapy, if the system detects that the patient is experiencing low mood or decreased adherence to rehabilitation due to physical discomfort, it automatically adjusts rehabilitation recommendations, such as reducing the intensity of rehabilitation training, increasing the delivery of content that soothes emotions, and encouraging the patient to persist with treatment through personalized videos. Studies have shown that personalized intervention can improve patient adherence to rehabilitation by approximately 30%–40%, effectively promoting the patient's recovery process.
[0133] 4. Construct a digital closed-loop platform This invention organically integrates AI-generated content, medical services, and rehabilitation products to form a complete digital closed-loop platform. Patients access personalized rehabilitation videos generated by AI within the system. These videos recommend suitable rehabilitation products based on the patient's condition. After purchasing the products, the patient's usage data is fed back to the system, allowing medical personnel to adjust treatment plans based on the product usage data and the patient's recovery progress. Simultaneously, the system is closely integrated with medical follow-up, synchronizing patient rehabilitation data with medical staff in real time. This achieves precise matching and collaborative optimization of medical services and rehabilitation product recommendations, providing patients with a one-stop, end-to-end service.
[0134] 5. Dynamic identification, real-time recommendation, and continuous optimization The innovative parameter system and model structure enable the system to possess powerful dynamic recognition capabilities, allowing it to perceive changes in patients' emotions and behaviors in real time. Based on real-time data, the system rapidly analyzes and infers, enabling real-time recommendations of rehabilitation products and pathways. Furthermore, by continuously collecting patient feedback and usage data, the system constantly optimizes its model and recommendation strategies. If the system detects that a certain type of rehabilitation product is ineffective for a specific patient group, it automatically adjusts the recommendation algorithm, reducing the recommendation of such products and increasing the recommendation of more effective products, ensuring that the system always maintains a high level of efficient and accurate service.
[0135] 6. Excellent interpretability and modular architecture, facilitating implementation. The system employs an interpretable design, clearly explaining the basis and logic behind the generated emotion analysis results, rehabilitation suggestions, and product recommendations to patients, medical staff, and their families. For example, when recommending rehabilitation products, the system details the reasons for the recommendation, including the product's suitability for the patient's condition and data on the effectiveness of similar products used in the past. Simultaneously, the modular architecture design allows the system's functional modules to work relatively independently yet collaboratively, facilitating flexible configuration and deployment according to different medical scenario needs. Whether it's a large general hospital or a primary healthcare institution, the system can be quickly integrated and applied.
[0136] 7. Adapting to the continuous evolution of patients' emotions and behaviors. Online learning and behavioral feedback mechanisms are key features of this system. As patients' treatment progresses and their emotions and behaviors change, the system continuously collects multimodal data and updates its model in real time using online learning algorithms. If a patient experiences new emotional problems during the mid-recovery phase, the system can promptly identify and adjust intervention strategies, continuously providing emotional support and rehabilitation guidance tailored to the patient's current state, ensuring the system's effectiveness and adaptability remain consistently online.
[0137] As can be seen from the above embodiments and experimental examples, the present invention provides a personalized dynamic psychological support system for newly diagnosed patients and their families. This system includes an AI-powered emotional support video module, which integrates parameters such as voice emotion fluctuation rate (V-Vol), visual situational adaptation index (VSAI), and stress sensitivity index (SSI) to achieve visualized and personalized emotional support. Based on this, the present invention also provides a personalized dynamic psychological support and medical closed-loop management system. This system consists of four modules: an AI-powered emotional support video module, a product and resource matching and recommendation module, a medical intervention and monitoring module, and a feedback monitoring and follow-up support module. This system achieves a continuous intervention closed loop centered on the patient, and has broad application prospects in the management of patients throughout their entire disease course, especially in the management of emotional and behavioral interventions.
Claims
1. A psychological support system, characterized in that, It includes an AI emotion-supporting video module, which integrates the following modules: The data acquisition module is configured to collect patient information; The multimodal patient profile construction module is configured to build a patient profile from the patient information through a multimodal fusion model, wherein the multimodal fusion model includes the following parameters: voice emotion fluctuation rate V-Vol, visual situational adaptation index VSAI, and stress sensitivity index SSI. The AI-generated content-driven video generation module is configured to generate videos based on patient profiles, patient disease stage information, user identity, and preset video themes, using large-scale pre-trained language models and virtual human generation engines. The dynamic content adaptation module is configured to optimize the video generated by the AI-generated content-driven video generation module based on the patient's signals.
2. The dynamic psychological support system according to claim 1, characterized in that, The process of building a patient profile in the multimodal patient profile construction module includes: using a multimodal fusion model with a Transformer architecture to extract and fuse features from different modalities of data such as text, speech, images, and physiology; using a hybrid emotion recognition model of bidirectional long short-term memory network and convolutional neural network to obtain structured labels and feature vectors containing emotion category and emotion intensity information; and recording them in time series form to form a patient profile.
3. The psychological support system according to claim 2, characterized in that, In the multimodal patient profile construction module, the multimodal fusion model includes: concatenating parameters and location encoding information as additional bias terms into attention calculation; fusing multimodal features through weighted summation; and dynamically adjusting loss weights using a joint loss function and a dynamic weight adjustment strategy.
4. The psychological support system according to claim 3, characterized in that, In the multimodal patient profile construction module, the method for attention calculation includes: in: Let be the multimodal feature vector, where Q, K, and V represent the query vector, key vector, and value vector obtained from the multimodal feature vector through different linear transformations, respectively, and T represents the position encoding vector. Let be the dimension of the key vector. For position encoding, t This represents the position index of each element in the sequence. The context vector for aggregation; And / or, the weighted summation method includes: The weights are calculated as follows: in, , h m Representing the characteristics of speech, images, or physiological signals, Represents the weight matrix. Represents the bias vector; And / or, the joint loss function includes: in: For the emotion category loss function, The weights for the loss of the emotion category, The VSAI regression loss function is used. The weights for the VSAI regression loss, For SSI multi-task loss function, The weights for the SSI multi-task loss are... Let V-Vol be the prediction loss function. Weights for predicting loss for V-Vol. Let the parameter be the regularization loss function. The weights for the parameter regularization loss; And / or, the dynamic weight adjustment strategy includes: in or is the importance coefficient, i and j are the feature numbers, and t is the time step.
5. The psychological support system according to claim 1, characterized in that, In the multimodal patient profile construction module, the V-Vol is used to quantify the emotional fluctuation features in the user's speech. The V-Vol is obtained by the following steps: collecting the user's real-time speech information and converting it into a discrete digital speech sequence, extracting at least one speech feature, characterizing the volatility of the speech feature, extracting the speech feature through a hybrid structure model, and calculating the magnitude of the change of the speech feature in the time series. And / or, the VSAI is used to quantify the degree of dynamic adaptation of the user's facial emotion expression to the visual context, VSAI>0.7; And / or, the SSI is divided into intensity prediction and classification prediction. By integrating multimodal data analysis of speech, image and physiological signals, the intensity of the user's stress response to the triggering element (SSI intensity) and the stress classification result (SSI class) are predicted. The multimodal data analysis process includes multimodal feature extraction process and multimodal feature fusion process.
6. The psychological support system according to claim 5, characterized in that, In the multimodal patient profile construction module, the speech features in V-Vol include: fundamental frequency, energy, speech rate, and 13-dimensional Mel frequency cepstral coefficients. The volatility of the speech features is characterized by the coefficient of variation. The hybrid structure model consists of convolutional neural network layers and long short-term memory network layers. The hybrid structure model uses a loss function and volatility score to optimize the model parameters. And / or, the formula for calculating V-Vol includes the following: Where h represents the hidden state output by the hybrid structure model. This represents the stability correction term, where t represents the time step and T represents the total number of time steps. And / or, prior to AI-driven content generation, VSAI's calculation formula includes the following: in, Sigmoid function and For model parameters, User facial image features for initial visual content. For visual context-related text features, For feature splicing operations; And / or, after AI-generated content-driven video generation, VSAI's calculation formula includes the following: Where E(t) is the video emotion labeling function, P is the user expression distribution, and Sim is the function for calculating the cosine similarity between P and E(t). P is improved using a multi-frame weighted fusion model, and Sim is improved using an attention-enhanced matching model. The time series smoothing process is performed using the exponential moving average algorithm, where the smoothing coefficient is... ; And / or, in the SSI, the triggering elements include visual triggers, auditory triggers, and time strategies, and the stress classification results By The input is another fully connected layer, which is then processed by the Softmax function to obtain the result. And / or, in the SSI, the process of multimodal feature extraction is optimized by multimodal early fusion technology, and the process of multimodal feature fusion is optimized by dynamic threshold adaptive adjustment strategy; And / or, The calculation formulas include the following: in, This is the weight matrix. To convert audio feature vectors and physiological feature vectors Concatenate them into a longer vector. It is the bias vector; And / or, The calculation formulas include the following: in, This is the weight matrix. For bias vectors, Category labels representing the classification results, This indicates the category of "no stress". This indicates the category of "significant stress".
7. The psychological support system according to claim 1, characterized in that, The data acquisition module connects to the hospital information system through an interface conforming to the HL7 FHIR standard, allowing it to access patients and collect their information. And / or, in the data acquisition module, the information collected includes: structured data such as age, gender, social role, diagnosis results, pathological stage, and treatment plan; natural language processing technology is used to extract and supplement information from unstructured text in electronic medical records; and speech recognition technology is used to collect voice data from patients and their families. And / or, in the multimodal patient profile construction module, a two-stage training strategy and an adaptive learning algorithm are used to optimize the model; And / or, the AIGC-driven video generation module includes: using a prompting engineering approach, inputting patient profile feature vectors, disease stage information, user identity and preset video themes into a large-scale pre-trained language model to generate a video script; and using 3D modeling and animation rendering technology to construct a detailed library of facial expressions and body movements for each character through the anthropomorphic role-playing of virtual doctors and patients. And / or, in the AIGC-driven video generation module, the user identity includes the patient himself / herself, the patient's parents, the patient's spouse, and the patient's children, and the preset video theme includes psychological counseling and emotional acceptance; And / or, in the dynamic content adaptation module, the patient signal includes: user video viewing time, likes, comments, and other behavioral data; And / or, in the dynamic content adaptation module, the optimization method includes: establishing a dynamic content adjustment model based on reinforcement learning, which automatically triggers the regeneration and rendering process of the video script when the user's emotional state changes, thereby realizing dynamic adjustment of the content.
8. A closed-loop intelligent management system for psychological support and medical care, characterized in that, It includes the following modules: AI emotion support video module in the psychological support system according to any one of claims 1-7; The product and resource matching and recommendation module integrates an intelligent evaluation module, a personalized product combination recommendation module, and a supply chain system integration module. The medical intervention and monitoring module integrates a medical institution collaboration platform access module, an AI-assisted medical decision-making module, and a three-party interaction module; The feedback monitoring and follow-up support module integrates a multimodal emotion monitoring module, a dynamic service recommendation and video update module, and a system optimization module; The data acquired by the data acquisition module in the AI emotion support video module is transmitted to the multimodal emotion monitoring module. The data acquired by the data acquisition module in the AI emotion support video module, the patient profile acquired by the multimodal patient profile construction module, and the video content usage effect data acquired by the dynamic content adaptation module are respectively input into the intelligent evaluation module and the medical institution collaboration platform access module. The product recommendation feedback data obtained by the supply chain system docking module is also transmitted to the medical institution collaboration platform access module. The three-way interaction module also exchanges information with the dynamic content adaptation module, the personalized product combination recommendation module, and the system optimization module. The emotional changes obtained by the multimodal emotion monitoring module are also transmitted as patient signals to the AI emotion support video module.
9. The intelligent management system for psychological support and medical closed-loop treatment according to claim 8, characterized in that: In the product and resource matching and recommendation module The intelligent assessment module is configured to use an intelligent assessment model to determine the patient's rehabilitation needs and the priority of resource matching for various needs; The personalized product combination recommendation module is configured to generate a personalized product combination package containing rehabilitation training equipment, psychological adjustment client, nutritional supplement package, and assistive device through a hybrid recommendation model. The supply chain system integration module is configured to integrate personalized product bundles with the supply chain management system through a standardized interface. And / or, in the medical intervention and monitoring module, The medical institution collaboration platform access module is configured to synchronize patient data to the medical collaboration platform in real time through a secure data transmission protocol; The AI-assisted medical decision-making module is configured to provide personalized adjustment suggestions to medical staff through the AI-assisted decision-making model, patient profiles obtained by the multimodal patient profile construction module, and user feedback data, and to enable medical staff to formulate adjustment instructions. The three-way interaction module is configured to upload medical staff's adjustment instructions and user feedback data to the system through a real-time interaction mechanism between AI, patients, and medical staff, thereby achieving information sharing among the three parties. And / or, in the feedback monitoring and follow-up support module, The multimodal emotion monitoring module is configured to use a multimodal fusion algorithm to display emotion change trends in the form of visual charts and monitor users' emotion changes in real time. The dynamic service recommendation and video update module is configured to use time-series prediction models to predict users' future emotional changes and rehabilitation needs based on data from the multimodal emotion monitoring module, and dynamically adjust personalized product recommendation strategies based on the prediction results. The system optimization module is configured to optimize the parameters in the system through parameter feedback loops and personalized weight update strategies.
10. The intelligent management system for psychological support and medical closed-loop treatment according to claim 9, characterized in that: In the system optimization module, the parameter feedback loop strategy includes: user interaction with the system, the system collecting user feedback data and updating model parameters and attention weights, fine-tuning the updated model for the next emotion recognition and recommendation, and the feedback data including user viewing of recommended content and self-evaluation of emotions; And / or, the personalized weight update strategy includes: calculating the feedback loss based on user feedback data, and updating the parameter-related weights in the model using gradient descent, wherein the feedback loss is as follows: Based on genuine user feedback, For model prediction feedback, is the time decay factor, i represents the data point number, and N represents the total number of data points.