Diagnosis and treatment and communication two-dimensional linkage medical training scene generation method and system
By generating emotional state vectors through multimodal data fusion, the synergy between diagnosis and communication strategies is driven, solving the problem of the separation between diagnosis and communication in the existing medical training system, and realizing the dynamic adjustment of training scenarios and the improvement of comprehensive capabilities.
Patent Information
- Application Number
- CN202511819379.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-04
- Publication Date
- 2026-02-17
AI Technical Summary
The existing medical training system lacks a unified linkage mechanism for diagnosis and treatment pathways and communication strategies, resulting in a large gap between the training scenarios and the real clinical environment. This makes it impossible to effectively develop comprehensive coping abilities and may also solidify erroneous perceptions.
By acquiring multimodal data of doctor-patient interactions in real time, we extract speech prosody features, facial motion unit vectors, and semantic sentiment tags. We then use a multimodal emotion fusion model to generate emotion state vectors, drive the synergy between the diagnosis and treatment decision-making path and the communication strategy path, detect key event triggering conditions in real time to reconstruct the scenario, and dynamically adjust the training scenario.
It enhanced the realism and interactivity of the training scenarios, improved trainees' comprehensive coping abilities, increased the accuracy of diagnosis and treatment decisions and communication and collaboration skills, and reduced mechanical interactions.
Smart Images

Figure CN121545790A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of machine learning technology, and in particular to a method and system for generating medical training scenarios that link diagnosis and communication in two dimensions. Background Technology
[0002] In the field of medical education and clinical simulation training, medical training systems typically design the development of diagnostic and treatment skills and doctor-patient communication abilities as two separate modules. However, this fragmented training model fails to realistically reflect the complex reality of the intertwined and synergistic interaction between diagnostic and treatment processes and communication behaviors in clinical practice, resulting in a gap between the training scenario and the real medical environment. Although some training systems have attempted to incorporate multimodal data processing and artificial intelligence technologies, these technologies remain limited to single-dimensional training objectives, failing to delve into the deep connections between the two dimensions of diagnosis and treatment and communication in terms of medical history context, dialogue strategies, and emotional states, and thus failing to achieve synergistic interaction between the two.
[0003] More importantly, the current medical training system lacks a unified linkage mechanism for diagnosis and treatment pathways and communication strategies. This technical deficiency results in training scenarios generally lacking the complexity, dynamism, and interactivity of real clinical environments. This not only fails to effectively train trainees' comprehensive coping abilities in complex situations, but may also solidify their erroneous perceptions of "emphasizing diagnosis and treatment while neglecting communication" or "communication being disconnected from diagnosis and treatment," which limits the improvement of medical personnel's overall quality.
[0004] Therefore, existing medical training systems need to be improved to overcome the shortcomings of current technologies. Summary of the Invention
[0005] To overcome the problems existing in related technologies, one of the objectives of this invention is to provide a method for generating medical training scenarios that links diagnosis and treatment with communication in two dimensions. This method can improve the authenticity and clinical relevance of training scenarios, thereby optimizing the effectiveness and efficiency of medical training.
[0006] A method for generating medical training scenarios that integrate diagnosis and communication in two dimensions includes: Real-time acquisition of input data during doctor-patient interactions, including voice signals, facial expression image sequences, and text dialogue content; Speech prosody features are extracted based on speech signals; facial action unit vectors are obtained based on facial expression image sequences; semantic role labeling and sentiment polarity classification are performed on text dialogue content to obtain semantic sentiment labels; The speech prosody features, facial action unit vectors, and semantic sentiment labels are input into the multimodal sentiment fusion model to obtain the fused sentiment state vector; Based on the emotional state vector, the diagnosis and treatment decision path and the communication strategy path are driven synchronously; the diagnosis and treatment decision path is based on medical knowledge graph for diagnostic reasoning, and the communication strategy path is based on the emotional state vector to generate a dialogue intent sequence. During the interaction, preset key event triggering conditions are detected in real time. When any triggering condition is met, the dynamic scene reconstruction mechanism is activated to generate a set of scene adjustment parameters. The parameter set is adjusted according to the scenario to dynamically modify the medical training scenario.
[0007] In a preferred embodiment of the present invention, the dynamic modification of the medical training scenario further includes: Input the dialogue intent sequence and historical interaction data of the training system into the template matching module, calculate the matching utility value of the candidate dialogue templates, select the template with the highest utility value, and generate AI simulated dialogue output; The deviation between the AI-simulated dialogue output and the target emotional state is periodically evaluated. Based on the deviation feedback signal obtained from the evaluation, the feature weighting parameters of the multimodal emotion fusion model and the reward function weight configuration of the template matching module are updated.
[0008] In a preferred embodiment of the present invention, the multimodal emotion fusion model adopts a structure that combines attention mechanism with long short-term memory network to weightedly fuse features of each modality and map the fused high-level features to a three-dimensional psychological space.
[0009] In a preferred embodiment of the present invention, the template matching module is a template matching model constructed based on a deep Q-network; The template matching module is used for: Based on the dialogue intent sequence and the interaction state of historical data, calculate the matching utility value of each dialogue template in the preset candidate dialogue template set; Based on the calculated matching utility value, the optimal template combination is selected from the candidate dialogue template set; Based on the optimal template combination, the AI-simulated patient's dialogue response is generated and output.
[0010] In a preferred embodiment of the present invention, calculating the matching utility value includes: A deep Q-network is used to estimate the long-term expected reward for selecting each candidate dialogue template in a given interaction state, and this long-term expected reward value is used as the matching utility value of the template.
[0011] In a preferred embodiment of the present invention, the method further includes: dynamically adjusting the output of the diagnosis and treatment decision path and the communication strategy path through a two-way feedback mechanism; The two-way feedback mechanism is achieved through dynamic weight allocation. The weight values are calculated based on the empathy and trust sub-features extracted from the emotional state vector to balance the contribution ratio of treatment decisions and communication strategies.
[0012] In a preferred embodiment of the present invention, during the simulated doctor-patient interaction process, at least one preset key event triggering condition is detected in parallel and in real time. When any of the aforementioned key event triggering conditions is detected, a dynamic scene reconstruction instruction is triggered. In response to the dynamic scene reconstruction instruction, the dynamic scene reconstruction mechanism is executed to generate a set of scene adjustment parameters for adjusting the current training scene.
[0013] The key event triggering conditions include at least one of the following: Diagnostic path deviation condition: The deviation between the current diagnostic decision path and the expected path exceeds the first preset threshold; Conditions for sudden change in emotional state: The change in the emotional state vector of the user or simulated patient obtained by multimodal data fusion exceeds the second preset threshold. Abnormal user input condition: The user's input pattern is detected to deviate from the historical normal pattern or the preset typical pattern.
[0014] In a preferred embodiment of the present invention, the scenario adjustment parameters include instructions for adjusting the patient's past medical history, social network, diversity of dialogue template library, and weights of various behaviors in AI behavior strategy.
[0015] In a preferred embodiment of the present invention, the speech prosodic features include the fundamental frequency mean, variance, and rate of change; The facial motion unit vector includes the intensity values of multiple facial motion units and the probability distribution of basic expressions; The semantic sentiment tags include sentiment intensity values and sentiment polarity categories.
[0016] The second objective of this invention is to provide a training and assessment system that integrates diagnosis and treatment simulation with doctor-patient communication evaluation. This system is used to implement the medical training scenario generation method described above, which links diagnosis and treatment with communication in two dimensions. The beneficial effects of this invention are as follows: This invention provides a method and system for generating medical training scenarios that integrate diagnosis and communication in a dual-dimensional manner. The method includes: real-time acquisition of input data during doctor-patient interaction; extraction of speech prosodic features, facial motion unit vectors, and semantic sentiment tags from the input content; inputting the extracted content into a multimodal emotion fusion model to obtain a fused emotion state vector; synchronously driving the diagnosis and treatment decision path and the communication strategy path based on the emotion state vector; the diagnosis and treatment decision path using a medical knowledge graph for diagnostic reasoning, and the communication strategy path generating a dialogue intent sequence based on the emotion state vector; real-time detection of preset key event triggering conditions during the interaction, activating a dynamic scenario reconstruction mechanism and generating a scenario adjustment parameter set when any triggering condition is met; and dynamically modifying the medical training scenario according to the scenario adjustment parameter set. By generating emotion state vectors through multimodal fusion of speech, facial, and text, the matching degree with the actual emotions of physicians is improved, thereby solving the problem of "mechanical interaction" in existing training systems and making the interactive experience of the training scenario closer to real doctor-patient communication. Furthermore, by driving dual-path collaboration through emotion state vectors, the "accuracy of diagnosis and treatment decisions" of trained physicians was effectively improved in the experiment, strengthening the training of physicians' collaborative abilities in diagnosis and treatment and communication. Attached Figure Description
[0017] Figure 1 This is a flowchart of a medical training scenario generation method that integrates diagnosis and communication in an embodiment of the present invention. Detailed Implementation
[0018] Preferred embodiments of the invention will now be described in more detail with reference to the accompanying drawings. While preferred embodiments of the invention are shown in the drawings, it should be understood that the invention may be implemented in various forms and should not be limited to the embodiments set forth herein. Rather, these embodiments are provided so that the invention will be thorough and complete, and will fully convey the scope of the invention to those skilled in the art.
[0019] The terminology used herein is for the purpose of describing particular embodiments only and is not intended to be limiting of the invention. The singular forms “a,” “the,” and “the” used in this invention and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term “and / or” as used herein refers to and includes any or all possible combinations of one or more of the associated listed items.
[0020] It should be understood that although the terms "first," "second," "third," etc., may be used in this invention to describe various information, this information should not be limited to these terms. These terms are only used to distinguish information of the same type from one another. For example, without departing from the scope of this invention, first information may also be referred to as second information, and similarly, second information may also be referred to as first information. Thus, features defined as "first" or "second" may explicitly or implicitly include one or more of that feature. In the description of this invention, "a plurality of" means two or more, unless otherwise explicitly specified.
[0021] The current medical training system lacks a unified linkage mechanism for diagnosis and treatment pathways and communication strategies. This technical deficiency results in training scenarios generally lacking the complexity, dynamism, and interactivity of real clinical environments. This not only fails to effectively train trainees' comprehensive coping abilities in complex situations, but may also solidify their erroneous perceptions of "emphasizing diagnosis and treatment while neglecting communication" or "communication being disconnected from diagnosis and treatment," which limits the improvement of medical personnel's overall quality.
[0022] Based on this, this application provides a method for generating medical training scenarios that integrates diagnosis and treatment with communication.
[0023] Example like Figure 1 As shown, this application provides a method for generating medical training scenarios that integrate diagnosis and communication in two dimensions, including: S100: Real-time acquisition of input data during doctor-patient interaction, including voice signals, facial expression image sequences, and text dialogue content; In this step, multimodal input data during the doctor-patient interaction process is acquired, including voice tone, facial expression image sequences and text dialogue content, and interaction timestamps are recorded to maintain the synchronization of multi-source data.
[0024] More specifically, this step involves simultaneously acquiring voice, video, and text data from doctor-patient interactions using a multimodal data acquisition system. Voice signals are acquired using a high-fidelity microphone array with a sampling rate of 48kHz and 16-bit quantization precision. Facial expression image sequences are acquired using an RGB-D camera with a resolution of 1280×720 and a frame rate of 30fps, including RGB images and depth information. Text dialogue content is transcribed in real-time using an ASR system employing a Transformer-based model, achieving a recognition accuracy of ≥95%. The timestamp synchronization mechanism uses the PTP precision clock protocol to ensure that the time deviation of multi-source data is ≤1ms. The system achieves synchronous acquisition and storage of data from various modalities through a multi-threaded data acquisition framework, establishing a unified timeline index. This step constructs a high-precision, multi-dimensional doctor-patient interaction dataset, providing high-quality input for subsequent emotion recognition and scene generation.
[0025] For example, a multimodal data acquisition system was deployed in the clinical skills training center of a tertiary hospital, equipped with a four-microphone array, six RGB-D cameras, and a speech recognition terminal. In a simulated consultation scenario, trainee doctors interacted with standardized patients, and the system simultaneously acquired 120 minutes of multimodal data. Testing showed that the time synchronization error for each modality was 0.8ms, the speech recognition accuracy reached 96.3%, and the facial expression capture completeness reached 98.5%. Experimental data demonstrates that this acquisition system can effectively support subsequent emotion recognition and scene generation tasks.
[0026] S200: Extract speech prosody features based on speech signals; obtain facial action unit vectors based on facial expression image sequences; perform semantic role labeling and sentiment polarity classification on text dialogue content to obtain semantic sentiment labels; Specifically, the prosodic features include the fundamental frequency mean, variance, and rate of change; The facial motion unit vector includes the intensity values of multiple facial motion units and the probability distribution of basic expressions; The semantic sentiment tags include sentiment intensity values and sentiment polarity categories.
[0027] In practical applications, endpoint detection and fundamental frequency extraction are performed on the speech signal, key point localization and micro-expression recognition are performed on the image sequence, and semantic role labeling and sentiment polarity classification are performed on the text content, generating speech prosody features, facial action unit vectors and semantic sentiment tags respectively.
[0028] The speech signal processing employs a dual endpoint detection algorithm based on energy and zero-crossing rate, combined with a deep neural network for fundamental frequency estimation. The specific formula is as follows: in This represents the fundamental frequency at time t. It is a voice signal. The window function is used. The extracted speech prosodic features include 12 parameters such as fundamental frequency mean, variance, and rate of change.
[0029] Facial expression analysis employs a keypoint localization algorithm based on a 3D deformation model, combined with OpenFace 2.0 to extract facial action unit (AU) intensity values. Micro-expression recognition uses a spatiotemporal convolutional network to output the probability distributions of six basic expressions. The final generated facial action unit vector contains 20 AU intensity values and 6 expression probability values.
[0030] Text analysis employs a BERT-BiLSTM-CRF hybrid model for semantic role labeling, combined with a sentiment lexicon and deep learning methods for sentiment polarity classification. The sentiment polarity classification loss function is: in For real labels, To predict probabilities, the generated semantic sentiment labels include sentiment intensity values (0-1) and sentiment categories (positive, neutral, negative).
[0031] This step enables refined extraction of multimodal features, providing rich feature representations for subsequent emotion fusion.
[0032] S300. Input the speech prosody features, facial action unit vectors and semantic sentiment labels into the multimodal emotion fusion model to obtain the fused emotion state vector; For example, tested on a dataset of 1000 doctor-patient dialogues, the accuracy rate for voice endpoint detection reached 98.7%, with a fundamental frequency estimation error of less than 5 Hz. The facial keypoint localization error was 2.3 pixels, and the micro-expression recognition accuracy was 89.2%. The F1 score for text sentiment classification reached 92.5%. Experimental results show that this feature extraction method can effectively capture emotional cues in doctor-patient interactions.
[0033] S3: Input the speech prosody features, facial action unit vectors and semantic emotion tags into the multimodal emotion fusion model to obtain the fused emotion state vector, which includes a three-dimensional psychological dimension representation of pleasure, arousal and control.
[0034] Furthermore, the multimodal emotion fusion model adopts a structure that combines attention mechanism with long short-term memory network to weightedly fuse features of each modality and map the fused high-level features to a three-dimensional psychological space.
[0035] In one implementation, multimodal emotion fusion employs an attention-gated LSTM network, with the following structure: in Let be the input features at time t. In hidden state, For the sigmoid function, This is the Hadamard product. The formula for calculating attention weights is: The final output emotion state vector Mapped to three-dimensional psychological space: Through the multimodal emotion fusion model, this application can achieve deep fusion of multimodal emotional information, and the generated three-dimensional emotion state vector can more comprehensively reflect the emotional changes in doctor-patient interaction.
[0036] For example, training the model on the IEMOCAP dataset and using 5-fold cross-validation, the fusion model achieved an average recognition accuracy of 88.6%, significantly outperforming single-modal methods (speech 72.3%, face 76.5%, text 79.8%). On the doctor-patient dialogue test set, the Pearson correlation coefficient for pleasure prediction was 0.83, arousal was 0.79, and control was 0.75, indicating that the model can effectively capture the complex emotional states in doctor-patient interactions.
[0037] S400. Based on the emotional state vector, the diagnosis and treatment decision path and the communication strategy path are driven synchronously; the diagnosis and treatment decision path is based on the medical knowledge graph to perform diagnostic reasoning, and the communication strategy path is based on the emotional state vector to generate a dialogue intent sequence. More preferably, this step also includes: dynamically adjusting the output of the diagnosis and treatment decision-making path and the communication strategy path through a two-way feedback mechanism; The two-way feedback mechanism is implemented through dynamic weight allocation. The weight values are calculated based on the empathy and trust sub-features extracted from the emotional state vector to balance the contribution ratio of diagnostic decisions and communication strategies. The diagnostic decision-making path employs a knowledge graph-based reasoning method, defining a medical knowledge graph. ,in A collection of medical entities This is a set of relationships between entities. The diagnostic reasoning process is represented as: in As a diagnostic hypothesis, Clinical manifestations, For characteristic function, For parameter vectors. Edge weights in a knowledge graph. Calculated using the PageRank algorithm: in The damping coefficient is... Let j be the in-degree of node j. This represents the total number of nodes.
[0038] The communication strategy path adopts a hierarchical reinforcement learning framework, defining a state space. Includes information such as emotion state vectors and dialogue history, action space It includes various communication strategies. The reward function is defined as: in These are the weighting coefficients. The variable represents the change, and TaskProgress measures the task completion rate. Policy optimization uses the PPO algorithm, with the objective function being: in , This is the dominant function.
[0039] The two-way feedback mechanism is achieved by dynamically adjusting the path weights: in For dynamic weights, and This is derived from an emotional state vector. This application achieves collaborative decision-making between diagnosis and communication through a dual-pathway approach, balancing medical judgment and emotional exchange through a two-way feedback mechanism. This dual-pathway collaborative decision-making and two-way feedback allows trainees to simultaneously develop clinical reasoning and doctor-patient communication skills in a single training session, aligning with real-world workflows. S500: Real-time detection of preset key event triggering conditions during interaction; when any triggering condition is met, activation of dynamic scene reconstruction mechanism to generate scene adjustment parameter set. Furthermore, during the simulated doctor-patient interaction process, at least one preset key event triggering condition is detected in parallel and in real time; When any of the aforementioned key event triggering conditions is detected, a dynamic scene reconstruction instruction is triggered. In response to the dynamic scene reconstruction instruction, the dynamic scene reconstruction mechanism is executed to generate a set of scene adjustment parameters for adjusting the current training scene.
[0040] The key event triggering conditions include at least one of the following: Diagnostic path deviation condition: The deviation between the current diagnostic decision path and the expected path exceeds the first preset threshold; Conditions for sudden change in emotional state: The change in the emotional state vector of the user or simulated patient obtained by multimodal data fusion exceeds the second preset threshold. Abnormal user input condition: The user's input pattern is detected to deviate from the historical normal pattern or the preset typical pattern.
[0041] Specifically, the critical event detection employs a multi-indicator fusion anomaly detection method, defining the diagnostic path offset: in For the current diagnostic path, This is the expected path. The degree of emotional abrupt change is defined as: User input anomaly pattern detection uses an LSTM-based anomaly detection model.
[0042] The dynamic scene reconstruction mechanism includes a scene parameter adjustment module, which generates a set of adjustment parameters. , where each parameter Adjustments are made to the corresponding medical record background, communication template, or AI behavior strategy. The advantage of this step is that it enables the dynamic evolution of the training scenario, enhancing the complexity and realism of the training by triggering scenario reconstruction through real-time detection of key events.
[0043] For example, during a simulated consultation, scene reconstruction is triggered when the diagnostic path deviation exceeds a threshold of 0.4 (ΔD=0.45) or the emotional abrupt change exceeds 0.3 (ΔE=0.32). This detection can improve accuracy and reduce false positive rate. After scene reconstruction, adjustments to the medical record background include adding patient history (such as a history of diabetes) and changing social relationships (such as adding family conflicts). The communication template library is updated to include adding reassuring dialogue templates, and the AI behavior strategy is adjusted to increase the weight of empathetic behavior by 30%.
[0044] S600. Adjust the parameter set according to the scenario to dynamically modify the medical training scenario. More specifically, the scenario adjustment parameters include instructions for adjusting the weights of various behaviors in the AI behavior strategy, such as the patient's medical history, social network, diversity of dialogue template library, and the overall behavior strategy.
[0045] The modification of medical record background information adopts a reasoning method based on medical knowledge graph, which updates the medical record content by adding new entities and relationship types; the adjustment of social relationship network adopts a random graph model, which defines the probability of connection between nodes through probability matrix; the update of communication template library adopts a template matching method based on vector space, which updates templates with similarity below the threshold.
[0046] In a preferred embodiment, the dynamic modification of the medical training scenario further includes: S700: Input the dialogue intent sequence and the historical interaction data of the training system into the template matching module, calculate the matching utility value of the candidate dialogue templates, select the template with the highest utility value, and generate AI simulated dialogue output; Furthermore, the template matching module is a template matching model built based on a deep Q-network; The template matching module is used for: Based on the dialogue intent sequence and the interaction state of historical data, calculate the matching utility value of each dialogue template in the preset candidate dialogue template set; Based on the calculated matching utility value, the optimal template combination is selected from the candidate dialogue template set; Based on the optimal template combination, the AI-simulated patient's dialogue response is generated and output.
[0047] Furthermore, calculating the matching utility value includes: A deep Q-network is used to estimate the long-term expected reward for selecting each candidate dialogue template in a given interaction state, and this long-term expected reward value is used as the matching utility value of the template.
[0048] Specifically, the template matching module uses a deep Q-network (DQN) to define the state space. Includes information such as dialogue intent sequences, historical interaction data, and emotional state vectors, action space This is a set of candidate dialogue templates. The Q network structure is as follows: in For state-action feature embedding, To extract local features for convolutional neural networks, Perform nonlinear transformations on the multilayer perceptron. The training objective function is: in As a discount factor, Here are the target network parameters. The formula for calculating the matching utility value is: in Each item measures intent matching, sentiment matching, and contextual consistency, respectively.
[0049] This step enables intelligent matching of dialogue templates. By learning patterns from historical interaction data through a deep Q-network, it generates more natural AI dialogue output that better meets the needs of the scenario.
[0050] S800 periodically evaluates the deviation between the AI simulated dialogue output and the target emotional state. Based on the deviation feedback signal obtained from the evaluation, it updates the feature weighting parameters of the multimodal emotion fusion model and the reward function weight configuration of the template matching module.
[0051] In this embodiment, through a closed loop of "interaction-evaluation-update," the system can automatically identify its shortcomings in emotional response or dialogue strategies (such as being ineffective at comforting highly anxious patients) and adjust the model parameters accordingly. This transforms a static system into a dynamic, self-learning intelligent agent, significantly extending the system's effective lifespan and reducing the cost of subsequent manual maintenance and iterative upgrades.
[0052] By optimizing the feature weights of the multimodal emotion fusion model, the system can more accurately understand and respond to the complex emotions of trainees or simulated patients, reduce the average emotional state bias, improve the training effect of empathy and emotion regulation in doctor-patient communication, and make simulated dialogue no longer a mechanical question and answer, but an interaction full of emotional intelligence.
[0053] By adjusting the reward weights for template matching, the system can guide AI to focus on different objectives at different training stages. For example, in the early stages of training, the focus can be on intent matching (improving α) to solidify the process; in the later stages, the focus can be on emotion matching (improving β) to enhance communication skills. This approach makes the training scenario no longer "one-size-fits-all," but rather allows for dynamic adjustment of the challenge focus based on training objectives and trainee performance, achieving a personalized and progressive training path.
[0054] Example 2 This embodiment provides a training and assessment system that integrates diagnosis and treatment simulation with doctor-patient communication assessment. This system is used to implement the medical training scenario generation method described above, which links diagnosis and treatment with communication in two dimensions.
[0055] Specifically, the system includes: A multimodal data acquisition module is used to simultaneously acquire voice, facial image, and text data; The feature extraction module is used to extract speech prosody features, facial action unit vectors, and semantic sentiment labels from various modal data; The multimodal emotion fusion module is used to fuse multimodal features into an emotion state vector; The dual-path decision engine, comprising a treatment decision submodule and a communication strategy submodule, is used to generate diagnostic reasoning results and dialogue intent sequences based on emotional state vectors, and supports two-way feedback adjustment. The dynamic scene reconstruction module is used to detect key events and trigger scene adjustments, generating a set of scene adjustment parameters; The resource adjustment module is used to modify the medical record background, update the dialogue template library, and adjust the AI behavior strategy according to the scenario by adjusting the parameter set. The dialogue generation module includes a template matching unit based on a deep Q-network, which is used to select the optimal dialogue template and generate AI responses; The evaluation and optimization module is used to periodically evaluate the quality of the system output and update the parameters of the emotion fusion model and template matching module.
[0056] The following describes the working process of this system: First, trainees (doctors) log into the training system and select the "Acute Abdominal Pain" training case. The system initializes and loads basic medical records (e.g., "Patient, female, 28 years old, sudden onset of right lower abdominal pain for 2 hours").
[0057] Trainees enter the simulated consultation room and begin a conversation with an AI virtual patient on the screen.
[0058] The multimodal data acquisition module is activated, and the microphone array simultaneously records the student's voice ("Hello, where do you feel unwell?"); the camera captures a sequence of facial expression images of the student (such as a concerned look in the eyes and a focused expression); the automatic speech recognition (ASR) system converts the dialogue into text in real time.
[0059] All data streams are tagged with a unified, high-precision timestamp to ensure strict alignment of audio, video, and text during subsequent analysis.
[0060] Raw data is fed into the feature extraction module in real time for parallel processing: speech signals are analyzed to extract prosodic features (e.g., a smooth tone and a stable speaking rate). Facial image sequences are analyzed to generate facial motion unit vectors (e.g., the degree of eyebrow extension and mouth movement). Text content is analyzed to output semantic sentiment labels (e.g., greetings with a sentiment polarity of "neutral").
[0061] The three types of features mentioned above are fed into the multimodal emotion fusion module. This module uses a neural network model to fuse the scattered features into a comprehensive emotional state vector, such as [pleasure: 0.5, arousal: 0.6, control: 0.7], which represents the overall emotional atmosphere of the current interaction.
[0062] This emotional state vector is simultaneously input into a dual-path decision engine. Based on this emotional vector and the input clinical manifestation (abdominal pain), reasoning is performed within a medical knowledge graph. It may calculate the initial probability distribution of "acute appendicitis," "ovarian cyst torsion," and "gastroenteritis," and prepare key symptom questions to be asked later (such as "Has the pain shifted?").
[0063] Based on the same emotion vector, the current communication tone is determined. For example, if the patient's initial emotion shows moderate anxiety (arousal level 0.6), the module may generate a dialogue intent sequence of "empathy-structured questioning".
[0064] The two sub-modules are not independent. For example, if the diagnostic pathway deduces an extremely high risk of ectopic pregnancy and requires urgent questioning about the patient's pregnancy and childbirth history, it will send a signal to the communication pathway: "It is necessary to break through conventional social etiquette and directly inquire about sensitive information." The communication pathway will then adjust its strategy, potentially generating an "urgent-protective inquiry" intention, and feed back to the diagnostic pathway: "The patient may become resistant as a result; please prepare alternative diagnoses."
[0065] During the training, the dynamic scene reconstruction module monitors interactions in real time. It continuously calculates metrics such as: Diagnostic path deviation: Whether the trainee has ignored key symptoms, causing the diagnostic path to deviate from the optimal path.
[0066] Emotional state abrupt change: When a student suddenly asks "Are you married?", the AI patient's emotional vector may change instantly (pleasure decreases, arousal level soars), and the abrupt change exceeds the threshold.
[0067] Triggering Reconstruction and Generating Parameters: Once any of the above conditions is met (e.g., a sudden change in emotion), the module immediately activates the dynamic scene reconstruction mechanism and generates a set of scene adjustment parameters. For example: {"Add Background": "The patient is unmarried and extremely sensitive to privacy", "Communication Template Weight Adjustment": "Increase the priority of defensive response templates", "AI Strategy Adjustment": "Increase the resistance coefficient to direct questioning"}.
[0068] The resource adjustment module receives this parameter set and takes effect immediately, adding a "social and psychological background" to the virtual patient in the background; it can also increase or raise the template weight for "responding to privacy inquiries".
[0069] This makes AI patients more likely to avoid or react emotionally to similar direct questions in subsequent conversations.
[0070] The "dialogue intent sequence" (such as "empathy-clarification") generated by the communication strategy submodule and the current interaction history data are fed into the dialogue generation module. Its core template matching unit, based on a deep Q-network, calculates the "matching utility value" of all candidate response templates according to the current state (intent, emotion, and updated context). Then, it selects the template combination with the highest utility value to generate a natural language response. For example, faced with a student's abrupt privacy question, the AI patient might answer: "Doctor, I don't think this question is related to my stomachache... I mainly have severe pain in my stomach." (This response conforms to a "defensive" strategy and also implies emotion).
[0071] The AI's response is presented to the trainees in the form of voice and virtual avatar animation. The system then enters the next interaction cycle, restarting, forming a real-time closed loop of "collection-analysis-decision-adjustment-response".
[0072] When the consultation ends (or the preset time is reached), the evaluation and optimization module is activated.
[0073] The evaluation and optimization module retrieves full-link data from the entire conversation and periodically evaluates the quality of the AI response. The core is to calculate the deviation between the "actual emotion evoked by AI" and the "expected target emotion," while also evaluating the fluency and relevance of the dialogue.
[0074] Based on the deviation feedback signal generated by the evaluation, this module initiates the parameter optimization process, including: Update the feature weighting parameters of the multimodal sentiment fusion model: for example, if text semantics is found to have a consistently low weight in predicting the "control" dimension, automatically increase its weight.
[0075] Update the reward function weight configuration of the template matching module: For example, if the evaluation finds that the system generally scores low on “emotional matching”, the weight of “emotional matching degree (β)” in the reward function is automatically increased to guide DQN to be more inclined to select emotionally compatible templates in the future.
[0076] The optimized model parameters are updated in the system. When the next trainee begins a new training session, they will be faced with a slightly more intelligent system that has learned from the previous round of interaction. Thus, the system achieves a two-level optimization loop, from "real-time adaptation" in a single session to "long-term self-evolution" across sessions.
[0077] Furthermore, it should be noted that the use of terms such as "first" and "second" to define components is merely for the purpose of distinguishing the corresponding components. Unless otherwise stated, these terms have no special meaning and therefore should not be construed as limiting the scope of protection of this application. The above are merely preferred embodiments of the present invention and are not intended to limit the present invention. For those skilled in the art, the present invention can have various modifications and variations. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.
Claims
1. A method for generating medical training scenarios that integrates diagnosis and communication in two dimensions, characterized in that, include: Real-time acquisition of input data during doctor-patient interactions, including voice signals, facial expression image sequences, and text dialogue content; Extracting prosodic features from speech signals; Facial action unit vectors are obtained from facial expression image sequences; semantic role labeling and sentiment polarity classification are performed on text dialogue content to obtain semantic sentiment labels; The speech prosody features, facial action unit vectors, and semantic sentiment labels are input into the multimodal sentiment fusion model to obtain the fused sentiment state vector; Based on emotional state vectors, the diagnosis and treatment decision-making path and communication strategy path are driven simultaneously; The diagnostic decision-making path is based on a medical knowledge graph for diagnostic reasoning, and the communication strategy path generates a dialogue intent sequence based on the emotional state vector. During the interaction, preset key event triggering conditions are detected in real time. When any triggering condition is met, the dynamic scene reconstruction mechanism is activated to generate a set of scene adjustment parameters. The parameter set is adjusted according to the scenario to dynamically modify the medical training scenario.
2. The method for generating medical training scenarios that link diagnosis and treatment with communication in two dimensions according to claim 1, characterized in that: The dynamically modified medical training scenario, following this, also includes: The dialogue intent sequence and historical interaction data from the training system are input into the template matching module (a template matching module based on a deep Q-network), the matching utility value of the candidate dialogue templates is calculated, and the template with the highest utility value is selected to generate AI simulated dialogue output. The deviation between the AI-simulated dialogue output and the target emotional state is periodically evaluated. Based on the deviation feedback signal obtained from the evaluation, the feature weighting parameters of the multimodal emotion fusion model and the reward function weight configuration of the template matching module are updated.
3. The method for generating medical training scenarios that link diagnosis and treatment with communication in two dimensions as described in claim 1 or 2, characterized in that: The multimodal emotion fusion model adopts a structure that combines attention mechanism with long short-term memory network to weightedly fuse features from various modalities and map the fused high-level features to a three-dimensional psychological space.
4. The method for generating medical training scenarios that link diagnosis and treatment with communication as described in claim 2, characterized in that: The template matching module is a template matching model built based on a deep Q-network; The template matching module is used for: Based on the dialogue intent sequence and the interaction state of historical data, calculate the matching utility value of each dialogue template in the preset candidate dialogue template set; Based on the calculated matching utility value, the optimal template combination is selected from the candidate dialogue template set; Based on the optimal template combination, the AI-simulated patient's dialogue response is generated and output.
5. The method for generating medical training scenarios that link diagnosis and treatment with communication in two dimensions according to claim 4, characterized in that: Calculating the matching utility value includes: A deep Q-network is used to estimate the long-term expected reward for selecting each candidate dialogue template in a given interaction state, and this long-term expected reward value is used as the matching utility value of the template.
6. The method for generating medical training scenarios that link diagnosis and treatment with communication in any one of claims 1-5, characterized in that: It also includes: dynamically adjusting the output of the diagnosis and treatment decision-making path and communication strategy path through a two-way feedback mechanism; The two-way feedback mechanism is achieved through dynamic weight allocation. The weight values are calculated based on the empathy and trust sub-features extracted from the emotional state vector to balance the contribution ratio of treatment decisions and communication strategies.
7. The method for generating medical training scenarios that link diagnosis and treatment with communication in any one of claims 1-5, characterized in that: During the simulated doctor-patient interaction process, at least one preset key event triggering condition is detected in parallel and in real time. When any of the aforementioned key event triggering conditions is detected, a dynamic scene reconstruction instruction is triggered. In response to the dynamic scene reconstruction instruction, the dynamic scene reconstruction mechanism is executed to generate a set of scene adjustment parameters for adjusting the current training scene. The key event triggering conditions include at least one of the following: Diagnostic path deviation condition: The deviation between the current diagnostic decision path and the expected path exceeds the first preset threshold; Conditions for sudden change in emotional state: The change in the emotional state vector of the user or simulated patient obtained by multimodal data fusion exceeds the second preset threshold. Abnormal user input condition: The user's input pattern is detected to deviate from the historical normal pattern or the preset typical pattern.
8. The method for generating medical training scenarios that link diagnosis and treatment with communication as described in claim 1, characterized in that: The scenario adjustment parameters include instructions for adjusting patient medical history, social networks, the diversity of dialogue template libraries, and the weights of various behaviors in AI behavior strategies.
9. The method for generating medical training scenarios that link diagnosis and treatment with communication in two dimensions according to claim 1, characterized in that: The prosodic features include the fundamental frequency mean, variance, and rate of change; The facial motion unit vector includes the intensity values of multiple facial motion units and the probability distribution of basic expressions; The semantic sentiment tags include sentiment intensity values and sentiment polarity categories.
10. A medical training scenario generation system that integrates diagnosis and treatment with communication, characterized in that, A method for generating medical training scenarios that integrates diagnosis and communication in two dimensions as described in any one of claims 1 to 9.