Virtual doctor empathetic interaction method and system based on multi-modal stress features

CN122552162APending Publication Date: 2026-08-11GUOXING CHUANGDA (BEIJING) TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-06-11
Publication Date
2026-08-11

AI Technical Summary

Technical Problem

然而,现有技术中,多数系统仅依赖语音或文本单模态,难以全面捕捉用户的应激状态,尤其对愤怒、恐惧、崩溃等过激情绪的识别准确性低;并且采用固定阈值判断情绪,无法区分用户天生的语速快、音量大等性格特征与真实的情绪波动,导致误报率高

Benefits of technology

1.本发明同步采集心电、皮电、面部动作单元、语音声学及文本语义五类信号,针对过激情绪构建敏感特征集;通过前30秒静息状态建立用户个性化基线(心率、语速、音量、表情强度等)并动态更新,有效区分性格特征与真实情绪波动,相比现有技术召回率提升至87%以上、误报率降低40%以上。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122552162A_ABST
    Figure CN122552162A_ABST
Patent Text Reader

Abstract

This invention discloses a virtual doctor empathic interaction method and system based on multimodal stress characteristics, belonging to the fields of artificial intelligence and digital healthcare. It includes real-time acquisition of the user's electrocardiogram (ECG), skin conductance, facial video, and voice signals through wearable sensors, cameras, and microphones; extraction of high-frequency components of heart rate and skin conductance, action unit intensity, and voice acoustic features; and dynamic adjustment of stress judgment thresholds. A sliding window is used to jointly judge and output stress levels L0-L3, and corresponding control command sequences are invoked to drive a 3D virtual doctor to perform graded interventions. After interaction, the user's cloud storage profile is updated, and proactive care is achieved based on historical stress score trends. This invention achieves accurate detection and layered empathic intervention for excessive emotions and can be widely applied in scenarios such as online consultations, mental health management, and chronic disease follow-up, significantly improving the interactive security and personalization level of virtual doctors.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of artificial intelligence and digital healthcare technology, and in particular to a virtual doctor empathic interaction method and system based on multimodal stress characteristics. Background Technology

[0002] With the development of artificial intelligence and digital healthcare, virtual doctors are widely used in remote consultations, mental health management, and home monitoring of chronic diseases. However, most existing systems rely solely on voice or text-based monomodal responses, making it difficult to fully capture a user's stress state, especially with low accuracy in recognizing extreme emotions such as anger, fear, and breakdown. Furthermore, using fixed thresholds to judge emotions fails to distinguish between a user's natural personality traits such as fast speech and loud volume and genuine emotional fluctuations, resulting in a high false alarm rate.

[0003] Furthermore, normal emotional reactions when users describe their serious illness (such as "tumor" or "not long to live") are often misjudged as pathological overreactions, triggering unnecessary interventions. When intervention is needed, the lack of a tiered, dynamically adjustable intervention mechanism makes it difficult to implement differentiated interactive controls based on the stress level.

[0004] Therefore, there is an urgent need for a virtual doctor empathic interaction method that can accurately detect excessive emotions, combine individual baseline calibration, dynamically adapt to semantic shocks, and perform tiered interventions. Summary of the Invention

[0005] The purpose of this invention is to solve the problems existing in the prior art, and to propose a virtual doctor empathic interaction method and system based on multimodal stress characteristics.

[0006] To achieve the above objectives, the present invention adopts the following technical solution: A virtual doctor empathic interaction method based on multimodal stress characteristics, including a virtual doctor generated based on a 3D animation model, the method comprising the following steps: S1: The user's electrocardiogram (ECG) and skin conductance signals are collected in real time through wearable sensors, facial video sequences are collected through a camera, and voice signals are collected through a microphone; R-wave detection is performed on the ECG to obtain heart rate and heart rate variability, low-pass filtering is performed on the skin conductance signals to obtain phase and high-frequency components, key point detection is performed on the facial video sequences to obtain motion unit intensity, and acoustic analysis is performed on the voice signals to obtain fundamental frequency, energy, and speech rate; S2: During the first 30 seconds of the user's idle state, the signal feature values ​​extracted in S1 are stored as individual baselines in non-volatile memory; in subsequent interactions, the normalized deviation of the current feature value relative to the baseline is calculated every 200 milliseconds, and the formula for calculating the normalized deviation is as follows:

[0007] The baseline standard deviation is the sample standard deviation of this feature value within 30 seconds before rest; S3: The user's voice signal is automatically transcribed into text by speech recognition and input into the BERT classification model based on medical dialogue corpus fine-tuning, and the semantic impact categories S0, S1, S2 and S3 are output. The stress determination threshold is dynamically adjusted according to a preset impact coefficient table, wherein the impact coefficient table is as follows: S0 corresponds to 0, S1 corresponds to 0.2, S2 corresponds to 0.5, and S3 corresponds to 0.8; the impact coefficient table is stored in a read-only memory. S4: Employ a sliding window of 3 seconds, updating the normalized deviation of each modality within the window every 200 milliseconds; when one of the following conditions is met within two consecutive windows, output the corresponding stress level through comparison and timing logic: L1: Single-modal normalization deviation ≥ 0.15 and duration exceeding 5 seconds; L2: Normalized deviation of at least two modes ≥ 0.15, or normalized deviation of a single mode ≥ 0.15 for more than 15 seconds; L3: Normalized deviation ≥ 0.15 for at least three modalities, or the probability of the text classification model outputting aggressive and self-harming keywords ≥ 0.8, or heart rate exceeding 1.4 times the resting heart rate and high-frequency skin conductance exceeding the baseline by 2 times; If conditions L1, L2, and L3 are not met in two consecutive windows, then the stress level L0 is output, which represents the normal state. S5: Based on the stress level, retrieve the corresponding control instruction sequence from the read-only memory to drive the virtual doctor's 3D rendering engine. If it is L1: Send an audio gain control signal to reduce the speech rate of the speech synthesizer by 20% and call the confirmation question template (template ID: C001) stored in the speech script library. If it is L2: Execute four sub-steps sequentially, each sub-step corresponding to a script template ID and an animation trigger signal: (1) Call the verification statement template (template ID: V001) and send the virtual doctor's head tilt animation command at the same time; (2). Call the second type of rhetoric (template ID: A002) in the external attribution rhetoric library and send the virtual doctor's palm-up gesture command; (3) Call the control transfer statement template (template ID: G003) and send a command to the UI to display the "Skip / Pause" button; (4). Call the breathing guidance statement template (template ID: B004) and start the breathing ball animation on the screen; after execution, delay for 10 seconds and call S4 assessment again. If the level is still L2, repeat sub-step (4) once. If it is upgraded to L3, jump to L3 instruction sequence. If it is L3: immediately interrupt all medical question and answer threads, and output a security protocol instruction sequence: stop speech synthesis, forcibly reset the virtual doctor's expression to neutral, and send a short message containing timestamp and location to the user's preset emergency contact through the wireless communication module; S6: After each interaction, the stress level sequence, trigger modality combination, instruction sequence identifier, and level change 10 seconds after intervention for this session are encrypted and stored in the user's private cloud storage file. S7: Map L0, L1, L2, and L3 to stress scores of 0, 1, 2, and 3, respectively; calculate the linear regression slope based on the stress scores of the most recent 10 sessions in the cloud storage archive; when the slope is >0.1, automatically preset the initial speech rate of the virtual doctor to 0.9 times the normal value when the next session starts, and call the caring audio clip (audio ID: W001).

[0008] As a preferred embodiment, the individual baseline in S2 is updated using an exponentially weighted moving average, and the update is performed only when the stress level output in S4 is L0 and the semantic impact level output in S3 is S0 or S1. The update formula is: New baseline = 0.3 × current window feature value + 0.7 × historical baseline; The historical baseline is stored in non-volatile memory and is not lost when the device is powered off.

[0009] As a preferred embodiment, the specific correspondence between the speech template IDs and animation trigger signals of the four sub-steps corresponding to L2 in S5 is stored in a read-only memory: Sub-step (1): Template ID: V001 corresponds to the text "I noticed a change in your voice / expression / heartbeat, are you feeling unwell?"; the animation instruction is for the virtual doctor's eyebrows to symmetrically rise by 0.2 arcs; Sub-step (2): Template ID: A002 corresponds to the text "Some issues during the consultation may cause you stress, but this is not your problem."; the animation instruction is for the virtual doctor to open both palms upwards. Sub-step (3): Template ID: G003 corresponds to the text "Next, you can say 'skip' or 'pause' at any time, and I will follow your instructions completely."; the animation instruction is to display a clickable button on the UI; Sub-step (4): Template ID: B004 corresponds to the text "Please follow the breathing ball on the screen to take three deep breaths: inhale... exhale...", and at the same time send the start command for the breathing ball animation.

[0010] As a preferred embodiment, the method further includes a pre-buffering instruction sequence: when the semantic impact of the output in step S3 is S3 and the medical information to be output contains preset high-impact keywords, the keywords include but are not limited to "tumor," "malignancy," and "mortality risk," the following pre-buffering instruction sequence is executed before executing step S5 to output the information, in the following order: ①. Call the preview template (template ID: P001) and send a virtual doctor's eye-forward pause gesture animation; ②. Start the deep breathing animation and wait for the user to confirm by replying "Continue" via touchscreen or voice. ③. Call the speech synthesizer to output the core information with parameters of reducing the speech rate by 30% and the volume by 6dB; ④. Immediately invoke the supporting template (template ID: S001) and send the virtual doctor nodding animation.

[0011] As a preferred embodiment, the cloud storage archive shall include at least the following structured fields: user anonymity identifier, timestamps of each session, total session duration, trigger time and trigger mode combination of L1 / L2 / L3 events, instruction sequence identifier for each intervention, and change in stress level before and after intervention; The system periodically scans the fields and counts the high-frequency trigger modal combinations for each user. When the number of L2 events corresponding to preset keywords related to pain and death is detected to be ≥3 times, the system automatically extends the question interval to 1.5 times the normal value at the beginning of the user's subsequent conversation.

[0012] As a preferred embodiment, the method further includes a virtual doctor embodied parameter linkage step: based on the stress level output by S4, the corresponding facial expression control parameters, speech synthesis parameters, and UI theme color values ​​are read from the register and sent to the 3D rendering engine. L0: Facial expression parameters (AU12 intensity 0.2), speech rate coefficient 1.0, volume gain 0dB, UI color value #E0F0FF; L1: Expression parameters (AU1+2 intensity 0.3), speech rate coefficient 0.9, volume gain -2dB, UI color value #F5F0E0; L2: Facial expression parameters (AU4 intensity 0.4), speech rate coefficient 0.7, volume gain -4dB, UI color value #FFE0B5; L3: Expression parameters (no AU activation), speech rate coefficient 0.5, volume gain -6dB, UI color value #FFF0D0.

[0013] As a preferred embodiment, the normalized deviation threshold of 0.15 in S4, the duration thresholds of 5 seconds and 15 seconds, and the heart rate threshold of 1.4 times the resting heart rate in L3 are predetermined by the following method: collecting multimodal data from 200 volunteers in simulated medical consultations, with three clinical psychologists independently labeling the stress levels, and maximizing the F1 score on the validation set using a grid search method to determine the final parameter values; the parameter values ​​are stored in non-volatile memory and cannot be modified online.

[0014] As a preferred embodiment, the breathing guidance animation in sub-step (4) of S5 has the following parameters: breathing bulb expansion time 4 seconds, contraction time 4 seconds, repeated 3 times; This parameter was determined by testing the vestibular sympathetic response of 50 users, resulting in an average heart rate decrease of 12%.

[0015] As a preferred embodiment, the semantic impact coefficient table (S0:0, S1:0.2, S2:0.5, S3:0.8) is obtained by collecting text and synchronous physiological signals from 1000 real doctor-patient dialogues, calculating the average stress physiological deviation under each impact level, and using the least squares method to fit the optimal coefficients so that the variance of physiological deviation within the same level is minimized; the coefficient table is stored in read-only memory.

[0016] A virtual doctor empathic interaction system based on multimodal stress characteristics, wherein the virtual doctor is presented as a 3D animated model, the system includes: Sensor interface unit: includes ECG / skin conductance analog front end, camera driver, and microphone array driver, used to acquire raw physical signals; Signal processing unit: includes ECG R-wave detector, skin conductance filter, facial key point detection network, and speech acoustic feature extractor, used to extract feature values ​​from raw signals; Baseline storage and update unit: includes non-volatile memory and an exponentially weighted moving average calculator; Text classification unit: A BERT-based medical impact classification model, running on GPU or NPU; Stress level determination unit: includes a comparator, a timer and logic gates, used to implement the threshold comparison and time accumulation in S4 of claim 1; Policy instruction library: Multiple control instruction sequences stored in read-only memory, including L1 instruction set, L2 sub-step instruction set (including a dialogue template ID and animation template ID mapping table), L3 security protocol instruction set, and pre-buffered instruction set; Interactive control unit: Receives stress level signals, reads corresponding instructions from the strategy instruction library, and outputs audio gain control, animation trigger signals, and communication module call signals; 3D rendering engine unit: Receives facial expression parameters, gesture parameters, and UI color values ​​to generate the virtual doctor's facial animation, body movements, and interface background in real time; Cloud storage client: Uploads encrypted session files to a remote server; Vertical trend analysis unit: Configured to read cloud storage files, calculate the linear regression slope of the stress scores of the most recent 10 conversations, and output speech rate preset adjustment instructions when the slope is >0.1.

[0017] Compared with the prior art, the beneficial effects of the present invention are as follows: 1. This invention simultaneously collects five types of signals: electrocardiogram, skin conductance, facial motion unit, speech acoustics, and text semantics, and constructs a sensitive feature set for excessive emotions; it establishes a personalized baseline for users (heart rate, speech rate, volume, facial expression intensity, etc.) by observing the first 30 seconds of resting state and updates it dynamically, effectively distinguishing personality traits from real emotional fluctuations. Compared with existing technologies, the recall rate is increased to over 87% and the false positive rate is reduced by over 40%.

[0018] 2. This invention designs a four-level stress response strategy from L0 to L3: L0 is normal questioning and answering, L1 is slowing down the speech rate, L2 sequentially executes a four-step degrading dialogue tree of emotion verification → externalization of attribution → transfer of control → breathing guidance, and L3 interrupts the consultation and initiates a safety protocol (including user confirmation and emergency contact); after intervention, a reassessment is performed 10 seconds later to form a closed loop. The L2 breathing guidance parameters, tested on 50 people, resulted in an average heart rate decrease of 12%, effectively preventing the escalation of emotions. Through the four-level graded response and the degrading dialogue tree, it is ensured that excessive emotions are relieved in a timely manner.

[0019] 3. This invention encrypts and stores the stress level sequence, trigger modality, and intervention effect to the user's cloud storage file after each interaction; by calculating the linear regression slope of the stress score of the most recent 10 conversations, when the slope is >0.1, it automatically reduces the initial speech speed of the virtual doctor and plays caring audio, thus creating a longitudinal stress profile and proactive care, achieving a deeper understanding of the user the more it is used. Attached Figure Description

[0020] Figure 1 This is a framework diagram of the virtual doctor empathic interaction system based on multimodal stress characteristics proposed in this invention; Figure 2 This is a flowchart of the virtual doctor empathic interaction method based on multimodal stress characteristics proposed in this invention; Figure 3 This is a schematic diagram of the control instruction sequence invoked according to the stress level in this invention. Detailed Implementation

[0021] The technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without creative effort are within the scope of protection of the present invention.

[0022] It should be understood that, when used in this specification and the appended claims, the terms "comprising" and "including" indicate the presence of the described features, integrals, steps, operations, elements and / or components, but do not exclude the presence or addition of one or more other features, integrals, steps, operations, elements, components and / or collections thereof.

[0023] It should also be understood that the terminology used in this specification is for the purpose of describing particular embodiments only and is not intended to limit the invention. As used in this specification and the appended claims, the singular forms “a,” “an,” and “the” are intended to include the plural forms unless the context clearly indicates otherwise.

[0024] It should also be further understood that the term "and / or" as used in this specification and the appended claims refers to any combination of one or more of the associated listed items and all possible combinations, and includes such combinations.

[0025] The technical solution is as follows, for reference. Figure 1 A virtual doctor empathic interaction system based on multimodal stress characteristics, wherein the virtual doctor is presented as a 3D animated model, the system includes: Sensor interface unit: includes ECG / skin conductance analog front end, camera driver, and microphone array driver, used to acquire raw physical signals; Signal processing unit: includes ECG R-wave detector, skin conductance filter, facial key point detection network, and speech acoustic feature extractor, used to extract feature values ​​from raw signals; Baseline storage and update unit: includes non-volatile memory and an exponentially weighted moving average calculator; Text classification unit: A BERT-based medical impact classification model, running on GPU or NPU; Stress level determination unit: includes comparators, timers and logic gates, used to implement threshold comparison and time accumulation in S4; Policy instruction library: Multiple control instruction sequences stored in read-only memory, including L1 instruction set, L2 sub-step instruction set (including dialogue template ID and animation template ID mapping table), L3 security protocol instruction set (including user confirmation sub-logic), and pre-buffered instruction set; Interactive control unit: Receives stress level signals, reads corresponding instructions from the strategy instruction library, and outputs audio gain control, animation trigger signals, and communication module call signals; 3D rendering engine unit: Receives facial expression parameters, gesture parameters, and UI color values ​​to generate the virtual doctor's facial animation, body movements, and interface background in real time; Cloud storage client: Uploads encrypted session files to a remote server; Vertical trend analysis unit: Configured to read cloud storage files, calculate the linear regression slope of the stress scores of the most recent 10 conversations, and output speech rate preset adjustment instructions when the slope is >0.1.

[0026] Data Privacy Protection Statement: User anonymity identifiers in cloud storage files are encrypted using HMAC-SHA256 one-way encryption. Emergency contact phone numbers and location information are only stored with explicit user authorization, and each transmission requires user confirmation or (except in extreme cases where confirmation is automatically skipped). The system complies with the Personal Information Protection Act and HIPAA regulations regarding de-identified data.

[0027] Reference Figures 1 to 3 A virtual doctor empathic interaction method based on multimodal stress characteristics, including a virtual doctor generated based on a 3D animation model, includes the following steps: S1: Real-time acquisition of the user's electrocardiogram (ECG) and electrodermal signal (EDS) through wearable sensors, facial video sequences through a camera, and speech signals through a microphone; R-wave detection of the ECG signal to obtain heart rate and heart rate variability (HRV, using the root mean square SSD of the difference between adjacent RR intervals); low-pass filtering of the EDS signal to obtain phase and high-frequency components; key point detection of the facial video sequence to obtain action unit (AU) intensity, the action units are defined based on the FACS facial motion coding system, this method uses AU1 (inner corner of eyebrow raised), AU2 (outer corner of eyebrow raised), AU4 (frowning), and AU12 (corner of mouth raised); acoustic analysis of the speech signal to obtain fundamental frequency, energy, and speech rate.

[0028] S2: During the first 30 seconds of the user's idle state, store the signal feature values ​​extracted in S1 as individual baselines in non-volatile memory; in subsequent interactions, calculate the normalized deviation of the current feature value relative to the baseline every 200 milliseconds. The formula for calculating the normalized deviation is as follows:

[0029] The baseline standard deviation is the sample standard deviation of the feature value within 30 seconds before rest.

[0030] If the quality of any modal signal does not meet the requirements within the first 30 seconds (e.g., ECG signal-to-noise ratio <20dB, camera does not detect a face or microphone is muted), the system will prompt "Please remain relaxed and face the camera, we will re-collect a 30-second baseline"; if the second collection still fails, a general population preset baseline (based on the average of 500 healthy adults) will be used, and the "baseline abnormal" field will be marked after the session ends.

[0031] S3: The user's voice signal is automatically transcribed into text via speech recognition and input into the BERT classification model, which is fine-tuned based on medical dialogue corpus. The model outputs semantic impact categories S0, S1, S2, and S3. The stress judgment threshold is dynamically adjusted according to the preset impact coefficient table, which is as follows: S0 corresponds to 0, S1 corresponds to 0.2, S2 corresponds to 0.5, and S3 corresponds to 0.8. The impact coefficient table is stored in read-only memory.

[0032] S4: A sliding window of 3 seconds is used, and the normalized deviation of each modality within the window is updated every 200 milliseconds; when one of the following conditions is met within two consecutive windows, the corresponding stress level is output through a comparator and a timer: L1: Normalized deviation of a single modality ≥ 0.15 and lasting for more than 5 seconds; L2: Normalized deviation of at least two modalities ≥ 0.15, or normalized deviation of a single modality ≥ 0.15 lasting for more than 15 seconds; L3: Normalized deviation of at least three modalities ≥ 0.15, or the probability of the text classification model outputting aggressive and self-harming keywords ≥ 0.8, or the heart rate exceeds 1.4 times the resting heart rate and the high-frequency component of skin conductance exceeds twice the baseline. If conditions L1, L2, and L3 are not met in two consecutive windows, then the stress level L0 is output, which is the normal state.

[0033] S5: Based on the stress level, it retrieves the corresponding control instruction sequence from read-only memory to drive the virtual doctor's 3D rendering engine. If it is L0: The virtual doctor maintains a neutral and friendly expression (AU12 intensity 0.2), speech rate coefficient 1.0, volume gain 0dB, UI color value #E0F0FF, does not trigger any intervention commands, and only executes the normal medical question and answer process.

[0034] If it is L1: Send an audio gain control signal to reduce the speech rate of the speech synthesizer by 20% and call the confirmation question template (template ID: C001) stored in the speech script library. If it is L2: Execute four sub-steps sequentially, each sub-step corresponding to a script template ID and an animation trigger signal: (1) Call the verification statement template (template ID: V001) and send the virtual doctor's head tilt animation command; (2) Call the second type of rhetoric in the external attribution rhetoric library (template ID: A002) and send the virtual doctor's palm-up gesture command; (3) Call the control transfer statement template (template ID: G003) and send the UI display "skip / pause" button command; (4) Call the breathing guidance statement template (template ID: B004) and start the breathing ball animation on the screen; after execution, delay for 10 seconds and call the S4 assessment again. If the level is still L2, repeat sub-step (4) once. If it is upgraded to L3, jump to the L3 instruction sequence; If it is L3: Immediately interrupt all medical Q&A threads, output the security protocol instruction sequence: stop speech synthesis, forcibly reset the virtual doctor's expression to neutral, and then execute the following confirmation logic: First, confirm via virtual doctor's voice: "I will contact your emergency contact. If you wish to cancel, please say 'cancel' within 10 seconds." If the user replies "cancel", the event will be recorded but no SMS will be sent; if there is no cancellation instruction within 10 seconds or the user confirms "send", an SMS containing a timestamp and location will be sent to the user's preset emergency contact via the wireless communication module.

[0035] When the heart rate exceeds 1.6 times the resting heart rate and the high-frequency component of skin conductance exceeds 3 times the baseline, skip the confirmation and send the text message directly.

[0036] S6: After each interaction, the stress level sequence, trigger modality combination, instruction sequence identifier, and level change in 10 seconds after intervention (intervention effect is defined as the number of stress levels reduced before and after intervention, such as from L2 to L0 as a decrease of 2 levels) of this session are encrypted and stored in the user's exclusive cloud storage file.

[0037] S7: Map L0, L1, L2, and L3 to stress scores of 0, 1, 2, and 3, respectively; calculate the linear regression slope based on the stress scores of the most recent 10 conversations in the cloud storage archive; when the slope is >0.1, automatically preset the virtual doctor's initial speech rate to 0.9 times the normal value when the next conversation starts, and call the caring audio clip (audio ID: W001).

[0038] Furthermore, the individual baseline in S2 is updated using an exponentially weighted moving average (EWMA) with a smoothing factor α = 0.3 (this value was determined by optimizing the time-varying characteristics of 200 users), and the update is performed only when the stress level output by S4 is L0 and the semantic impact of the output by S3 is S0 or S1; the update formula is: new baseline = 0.3 × current window feature value + 0.7 × historical baseline; the historical baseline is stored in non-volatile memory and is not lost after the device is powered off.

[0039] Furthermore, the specific correspondence between the speech template IDs and animation trigger signals of the four sub-steps corresponding to L2 in S5 is stored in read-only memory, including: Sub-step (1): Template IDV001 corresponds to the text "I noticed a change in your voice / expression / heartbeat, are you feeling unwell?"; the animation instruction is for the virtual doctor's eyebrows to symmetrically rise by 0.2 arcs; Sub-step (2): The text corresponding to the dialogue template IDA002 is "Some issues during the consultation may cause you stress, but this is not your problem."; the animation instruction is for the virtual doctor to open both palms upwards. Sub-step (3): Template IDG003 corresponds to the text "Next, you can say 'skip' or 'pause' at any time, and I will completely follow your instructions."; the animation instruction is to display a clickable button in the UI; Sub-step (4): The template IDB004 corresponds to the text "Please follow the breathing ball on the screen to take three deep breaths: inhale... exhale...", and at the same time send the start command for the breathing ball animation.

[0040] Furthermore, the method also includes a pre-buffering instruction sequence: when the semantic impact of the output in step S3 is S3 and the medical information to be output contains preset high-impact keywords, the keywords include but are not limited to "tumor," "malignancy," and "mortality risk," the following pre-buffering instruction sequence is executed before executing step S5 to output the information, in sequence: ① Call the preview template (template ID: P001) and send a virtual doctor's eye-looking pause gesture animation; ② Start the deep breathing animation and wait for the user to reply "continue" via touch screen or voice confirmation; ③ Call the speech synthesizer to output the core information with parameters of 30% reduced speech speed and 6dB reduced volume; ④ Immediately call the support template (template ID: S001) and send a virtual doctor's nodding animation.

[0041] Furthermore, the cloud storage archive contains at least the following structured fields: user anonymity identifier, timestamps of each session, total session duration, trigger time and trigger modality combination of L1 / L2 / L3 events, instruction sequence identifier for each intervention, and change in stress level before and after intervention. The system periodically scans the fields and counts the high-frequency trigger modality combinations for each user. When the number of L2 events corresponding to preset keywords related to pain and death is detected to be ≥3 times, the system automatically extends the question interval to 1.5 times the normal value at the beginning of the user's subsequent sessions.

[0042] Furthermore, the method also includes a virtual doctor embodied parameter linkage step: based on the stress level output by S4, the corresponding facial expression control parameters, speech synthesis parameters, and UI theme color values ​​are read from the register and sent to the 3D rendering engine. L0: Expression parameters (AU12 intensity 0.2), speech rate coefficient 1.0, volume gain 0dB, UI color value #E0F0FF; L1: Expression parameters (AU1+2 intensity 0.3), speech rate coefficient 0.9, volume gain -2dB, UI color value #F5F0E0; L2: Expression parameters (AU4 intensity 0.4), speech rate coefficient 0.7, volume gain -4dB, UI color value #FFE0B5; L3: Expression parameters (no AU activation), speech rate coefficient 0.5, volume gain -6dB, UI color value #FFF0D0.

[0043] It is worth noting that the normalized deviation threshold of 0.15 in S4, the duration thresholds of 5 seconds and 15 seconds, and the heart rate threshold of 1.4 times the resting heart rate in L3 were predetermined in the following way: multimodal data of 200 volunteers in simulated medical consultation were collected, and the stress levels were independently labeled by three clinical psychologists. The F1 score was maximized on the validation set using a grid search method, and the final parameter values ​​were determined. The parameter values ​​are stored in non-volatile memory and cannot be modified online.

[0044] Among them, the breathing guidance animation in L2 sub-step (4) of S5 has the following parameters: breathing bulb expansion time 4 seconds, contraction time 4 seconds, repeated 3 times; these parameters were determined by testing the vestibular sympathetic response of 50 users, resulting in an average heart rate decrease rate of 12%.

[0045] The semantic impact coefficient table (S0:0, S1:0.2, S2:0.5, S3:0.8) was obtained by collecting text and synchronous physiological signals from 1000 real doctor-patient dialogues, calculating the average stress physiological deviation under each impact level, and using the least squares method to fit the optimal coefficients so that the variance of physiological deviation within the same level is minimized. This coefficient table is stored in read-only memory.

[0046] Example 1 (First-time user, triggering L2 intervention) Scenario: A 45-year-old male patient is using a virtual doctor for the first time, complaining of "chest pain and insomnia".

[0047] step: 1. Baseline Acquisition: The system prompts "Please relax for 30 seconds," and collects resting heart rate (72 bpm), HRV (35 ms), speech rate (4.0 words / second), volume (65 dB), fundamental frequency (120 Hz), and AU4 intensity (0.1) to establish an individual baseline. If the signal is normal within 30 seconds, the baseline is valid; if a modality is missing, it is handled according to the backup plan (such as using a universal baseline).

[0048] 2. Normal consultation: The virtual doctor is in L0 state, with an AU12 intensity of 0.2, normal speaking speed, light blue UI, and performs routine medical Q&A.

[0049] 3. Stress Trigger: The user describes, "I checked online, it might be a heart attack, I'm afraid I'll suddenly die..." (The speech frequency rises to 145Hz, speech rate is 5.0 words / second; heart rate is 88 bpm; text impact S2). The system detects a monomodal (speech + physiological) deviation, lasting 6 seconds, and outputs L1. The virtual doctor reduces the speech rate by 20% and calls template C001: "You seem a bit anxious, let's talk slowly." 4. Upgrade L2: The user's voice trembles, the AU4 intensity is increased to 0.6, the heart rate is 95 bpm, the voice fundamental frequency variation coefficient is twice the baseline, the number of activated modalities reaches 2, and L2 output is applied.

[0050] 5. Execute L2 downgrade dialogue tree: (1) Call V001: "I noticed a change in your voice / expression / heartbeat. Are you feeling unwell?" accompanied by an animation of raised eyebrows.

[0051] (2) Call A002: "Some questions during the consultation may make you feel stressed, but this is not your problem." Accompanied by a palm-up gesture.

[0052] (3) Call G003: "You can say 'skip' or 'pause' at any time", and the UI will display the button.

[0053] (4) Call B004: "Please follow the breathing ball on the screen to take three deep breaths", breathing ball animation (expand for 4 seconds, contract for 4 seconds, repeat 3 times).

[0054] 6. Post-intervention assessment: After 10 seconds, the heart rate dropped to 82 bpm, AU4 intensity was 0.3, and the stress level decreased from L2 to L1. The system recorded that this intervention was effective (level decreased by 1 level) and stored it in the cloud storage archive.

[0055] 7. Session End: Encrypt and store session files, and update individual baselines only using L0 time period data.

[0056] Example 2 (Pre-buffering and Vertical Trend Early Warning) Scenario: User B has 10 historical conversations. The stress scores for the last 3 conversations are 0.2, 0.6 and 1.1 respectively. The linear slope is calculated to be 0.45 > 0.1.

[0057] 1. Proactive Care: When the next conversation starts, the system will automatically preset the virtual doctor's speaking speed to 0.9 times the normal value and play the care audio W001: "In our recent conversations, I noticed that you seem to be under some pressure. Would you like to talk about it?" 2. Pre-buffering: When a user asks, "Are my test results bad?" and the system needs to output "tumor" information, S3 impact triggers pre-buffering: Invoking P001, the virtual doctor looks straight ahead and pauses with a gesture: "I will inform you of an important result, please be prepared." The deep breathing animation will start, and the user will be asked to confirm "Continue" via voice.

[0058] Output the following message with a 30% reduction in speech rate and a 6dB reduction in volume: "Pathology report indicates the presence of malignant tumor cells." Immediately invoke S001: "Regardless of the outcome, I will continue to assist you." and send a nodding animation.

[0059] 3. Effect: Although the user experienced emotional fluctuations, due to pre-buffering preparation, the level above L2 was not triggered.

[0060] Experimental data for parameter optimization (supporting the determination basis of the above thresholds, breathing bulb parameters, and coefficient tables); Thresholds of 0.15, 5 seconds, and 15 seconds: 200 volunteers simulated consultations, and three psychologists independently labeled the stress levels. Grid search yielded the optimal parameter combination for F1 score: deviation 0.15 (recall 0.87, precision 0.92), 5 seconds (false positive rate ≤5%), and 15 seconds (the maximum Youden index for distinguishing between mild and moderate stress).

[0061] Breathing bulb parameters: In a test with 50 users, inflating / contracting for 4 seconds each, repeated 3 times, resulted in an average decrease in heart rate of 12%, which was better than other parameter combinations (such as 2 seconds / 2 seconds decrease of 5%, 6 seconds / 6 seconds decrease of 9%).

[0062] Impact coefficient: In 1000 real dialogues, S1 corresponds to the mean physiological deviation increment of 0.21 times the baseline difference, S2 corresponds to 0.53 times, and S3 corresponds to 0.79 times. The least squares fitting yielded coefficients [0, 0.2, 0.5, 0.8] that minimized the within-group variance.

[0063] Based on the above, this invention can be deployed on smartphones, tablets, smart speakers, or dedicated medical terminals as the core interactive module of virtual doctor applications. It is widely used in scenarios such as online consultations, mental health screening, chronic disease management, and postoperative rehabilitation follow-up, and has clear market value and industrial applicability.

[0064] The above description is only a preferred embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any equivalent substitutions or modifications made by those skilled in the art within the scope of the technology disclosed in the present invention, based on the technical solution and inventive concept of the present invention, should be covered within the scope of protection of the present invention.

Claims

1. A virtual doctor empathic interaction method based on multimodal stress characteristics, characterized in that, The method, which includes generating a virtual doctor based on a 3D animated model, comprises the following steps: S1: The user's electrocardiogram (ECG) and skin conductance signals are collected in real time through wearable sensors, facial video sequences are collected through a camera, and voice signals are collected through a microphone; R-wave detection is performed on the ECG to obtain heart rate and heart rate variability, low-pass filtering is performed on the skin conductance signals to obtain phase and high-frequency components, key point detection is performed on the facial video sequences to obtain motion unit intensity, and acoustic analysis is performed on the voice signals to obtain fundamental frequency, energy, and speech rate; S2: During the first 30 seconds of the user's idle state, the signal feature values ​​extracted in S1 are stored as individual baselines in non-volatile memory; in subsequent interactions, the normalized deviation of the current feature value relative to the baseline is calculated every 200 milliseconds, and the formula for calculating the normalized deviation is as follows: The baseline standard deviation is the sample standard deviation of this feature value within 30 seconds before rest; S3: The user's voice signal is automatically transcribed into text by speech recognition and input into the BERT classification model based on medical dialogue corpus fine-tuning, and the semantic impact categories S0, S1, S2 and S3 are output. The stress determination threshold is dynamically adjusted according to a preset impact coefficient table, wherein the impact coefficient table is as follows: S0 corresponds to 0, S1 corresponds to 0.2, S2 corresponds to 0.5, and S3 corresponds to 0.8; the impact coefficient table is stored in a read-only memory. S4: Employ a sliding window of 3 seconds, updating the normalized deviation of each modality within the window every 200 milliseconds; when one of the following conditions is met within two consecutive windows, output the corresponding stress level through comparison and timing logic: L1: Single-modal normalization deviation ≥ 0.15 and duration exceeding 5 seconds; L2: Normalized deviation of at least two modes ≥ 0.15, or normalized deviation of a single mode ≥ 0.15 for more than 15 seconds; L3: Normalized deviation ≥ 0.15 for at least three modalities, or the probability of the text classification model outputting aggressive and self-harming keywords ≥ 0.8, or heart rate exceeding 1.4 times the resting heart rate and high-frequency skin conductance exceeding the baseline by 2 times; If conditions L1, L2, and L3 are not met in two consecutive windows, then the stress level L0 is output, which represents the normal state. S5: Based on the stress level, retrieve the corresponding control instruction sequence from the read-only memory to drive the virtual doctor's 3D rendering engine: If it is L1: Send an audio gain control signal to reduce the speech rate of the speech synthesizer by 20% and call the confirmation question template (template ID: C001) stored in the speech script library. If it is L2: Execute four sub-steps sequentially, each sub-step corresponding to a script template ID and an animation trigger signal: (1). Call the verification statement template (template ID: V001) and send the virtual doctor's head tilt animation command at the same time; (2). Call the second type of rhetoric (template ID: A002) in the external attribution rhetoric library and send the virtual doctor's palm-up gesture command; (3) Call the control transfer statement template (template ID: G003) and send a command to the UI to display the "Skip / Pause" button; (4). Call the breathing guidance statement template (template ID: B004) and start the breathing ball animation on the screen; after execution, delay for 10 seconds and call S4 assessment again. If the level is still L2, repeat sub-step (4) once. If it is upgraded to L3, jump to L3 instruction sequence. If it is L3: immediately interrupt all medical question and answer threads, and output a security protocol instruction sequence: stop speech synthesis, forcibly reset the virtual doctor's expression to neutral, and send a short message containing timestamp and location to the user's preset emergency contact through the wireless communication module; S6: After each interaction, the stress level sequence, trigger modality combination, instruction sequence identifier, and level change 10 seconds after intervention for this session are encrypted and stored in the user's private cloud storage file. S7: Map L0, L1, L2, and L3 to stress scores of 0, 1, 2, and 3, respectively; calculate the linear regression slope based on the stress scores of the most recent 10 sessions in the cloud storage archive; when the slope is >0.1, automatically preset the initial speech rate of the virtual doctor to 0.9 times the normal value when the next session starts, and call the caring audio clip (audio ID: W001).

2. The virtual doctor empathic interaction method based on multimodal stress characteristics according to claim 1, characterized in that, The individual baseline in S2 is updated using an exponentially weighted moving average, and the update is performed only when the stress level output in S4 is L0 and the semantic impact level output in S3 is S0 or S1. The update formula is: New baseline = 0.3 × current window feature value + 0.7 × historical baseline; The historical baseline is stored in non-volatile memory and is not lost when the device is powered off.

3. The virtual doctor empathic interaction method based on multimodal stress characteristics according to claim 1, characterized in that, The specific correspondence between the speech template IDs and animation trigger signals of the four sub-steps corresponding to L2 in S5 is stored in a read-only memory: Sub-step (1): Template ID: V001 corresponds to the text "I noticed a change in your voice / expression / heartbeat, are you feeling unwell?"; the animation instruction is for the virtual doctor's eyebrows to symmetrically rise by 0.2 arcs; Sub-step (2): Template ID: A002 corresponds to the text "Some issues during the consultation may cause you stress, but this is not your problem."; the animation instruction is for the virtual doctor to open both palms upwards. Sub-step (3): Template ID: G003 corresponds to the text "Next, you can say 'skip' or 'pause' at any time, and I will follow your instructions completely."; the animation instruction is to display a clickable button in the UI; Sub-step (4): Template ID: B004 corresponds to the text "Please follow the breathing ball on the screen to take three deep breaths: inhale... exhale...", and at the same time send the start command for the breathing ball animation.

4. The virtual doctor empathic interaction method based on multimodal stress characteristics according to claim 1, characterized in that, The method further includes a pre-buffering instruction sequence: when the semantic impact of the output in step S3 is S3 and the medical information to be output contains preset high-impact keywords, the keywords include but are not limited to "tumor", "malignant" and "mortality risk", before executing step S5 to output the information, the following pre-buffering instruction sequence is executed in sequence: ①. Call the preview template (template ID: P001) and send a virtual doctor's eye-forward pause gesture animation; ②. Start the deep breathing animation and wait for the user to confirm by replying "Continue" via touchscreen or voice. ③. Call the speech synthesizer to output the core information with parameters of reducing the speech rate by 30% and the volume by 6dB; ④. Immediately invoke the supporting template (template ID: S001) and send the virtual doctor nodding animation.

5. The virtual doctor empathic interaction method based on multimodal stress characteristics according to claim 1, characterized in that, The cloud storage archive contains at least the following structured fields: user anonymity identifier, timestamps of each session, total session duration, trigger time and trigger mode combination of L1 / L2 / L3 events, instruction sequence identifier for each intervention, and change in stress level before and after intervention; The system periodically scans the fields and counts the high-frequency trigger modal combinations for each user. When the number of L2 events corresponding to preset keywords related to pain and death is detected to be ≥3 times, the system automatically extends the question interval to 1.5 times the normal value at the beginning of the user's subsequent conversation.

6. The virtual doctor empathic interaction method based on multimodal stress characteristics according to claim 1, characterized in that, The method also includes a virtual doctor embodied parameter linkage step: based on the stress level output by S4, the corresponding facial expression control parameters, speech synthesis parameters, and UI theme color values ​​are read from the register and sent to the 3D rendering engine. L0: Expression parameters (AU12 intensity 0.2), speech rate coefficient 1.0, volume gain 0dB, UI color value #E0F0FF; L1: Expression parameters (AU1+2 intensity 0.3), speech rate coefficient 0.9, volume gain -2dB, UI color value #F5F0E0; L2: Facial expression parameters (AU4 intensity 0.4), speech rate coefficient 0.7, volume gain -4dB, UI color value #FFE0B5; L3: Expression parameters (no AU activation), speech rate coefficient 0.5, volume gain -6dB, UI color value #FFF0D0.

7. The virtual doctor empathic interaction method based on multimodal stress characteristics according to claim 1, characterized in that, The normalized deviation threshold of 0.15 in S4, the duration thresholds of 5 seconds and 15 seconds, and the heart rate threshold of 1.4 times the resting heart rate in L3 are predetermined by the following method: collecting multimodal data from 200 volunteers in simulated medical consultations, with three clinical psychologists independently labeling the stress levels, and maximizing the F1 score on the validation set using a grid search method to determine the final parameter values; the parameter values ​​are stored in non-volatile memory and cannot be modified online.

8. The virtual doctor empathic interaction method based on multimodal stress characteristics according to claim 1, characterized in that, The breathing guidance animation in sub-step (4) of S5 has the following parameters: breathing bulb expansion time 4 seconds, contraction time 4 seconds, repeated 3 times. This parameter was determined by testing the vestibular sympathetic response of 50 users, resulting in an average heart rate decrease of 12%.

9. The virtual doctor empathic interaction method based on multimodal stress characteristics according to claim 1, characterized in that, The semantic impact coefficient table (S0:0, S1:0.2, S2:0.5, S3:0.8) was obtained by collecting text and synchronous physiological signals from 1000 real doctor-patient dialogues, calculating the average stress physiological deviation under each impact level, and using the least squares method to fit the optimal coefficients so that the variance of physiological deviation within the same level is minimized; the coefficient table is stored in read-only memory.

10. A virtual doctor empathic interaction system based on multimodal stress characteristics, used to implement the method described in any one of claims 1-9, characterized in that, The virtual doctor is presented as a 3D animated model, and the system includes: Sensor interface unit: includes ECG / skin conductance analog front end, camera driver, and microphone array driver, used to acquire raw physical signals; Signal processing unit: includes ECG R-wave detector, skin conductance filter, facial key point detection network, and speech acoustic feature extractor, used to extract feature values ​​from raw signals; Baseline storage and update unit: includes non-volatile memory and an exponentially weighted moving average calculator; Text classification unit: A BERT-based medical impact classification model, running on GPU or NPU; Stress level determination unit: includes a comparator, a timer and logic gates, used to implement the threshold comparison and time accumulation in S4 of claim 1; Policy instruction library: Multiple control instruction sequences stored in read-only memory, including L1 instruction set, L2 sub-step instruction set (including a dialogue template ID and animation template ID mapping table), L3 security protocol instruction set, and pre-buffered instruction set; Interactive control unit: Receives stress level signals, reads corresponding instructions from the strategy instruction library, and outputs audio gain control, animation trigger signals, and communication module call signals; 3D rendering engine unit: Receives facial expression parameters, gesture parameters, and UI color values ​​to generate the virtual doctor's facial animation, body movements, and interface background in real time; Cloud storage client: Uploads encrypted session files to a remote server; Vertical trend analysis unit: Configured to read cloud storage files, calculate the linear regression slope of the stress scores of the most recent 10 conversations, and output speech rate preset adjustment instructions when the slope is >0.1.