Medical training evaluation method and system based on intelligent simulated patient

By building a high-precision 3D virtual patient model through multimodal interaction technology, the problems of single interaction and lack of personalization in existing medical simulation training systems are solved, highly simulated and personalized clinical training experience and evaluation are achieved, and the immersion and authenticity of medical education are enhanced.

CN120656734APending Publication Date: 2025-09-16CHONGQING MEDICAL UNIVERSITY

Patent Information

Application Number
CN202510721170.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-30
Publication Date
2025-09-16

AI Technical Summary

Technical Problem

Existing medical simulation training systems have limited functions and are unable to realistically simulate the physiological and psychological responses of patients at different stages of their illness. They lack personalization and are unable to meet the learning needs of different students. In addition, their single interaction method prevents them from providing a comprehensive clinical experience.

Method used

It uses multimodal interaction technology, integrates speech recognition, visual recognition and natural language processing, builds a high-precision 3D virtual patient model, interacts through voice, facial expressions and body language, combines with medical knowledge database for semantic parsing and behavioral analysis, and generates personalized medical training evaluation reports.

Benefits of technology

It provides a highly simulated and interactive clinical training experience, supports customizable virtual patient characteristics, covers multiple medical education links, enhances immersion and realism, meets personalized learning needs, and provides targeted training suggestions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120656734A_ABST
    Figure CN120656734A_ABST
Patent Text Reader

Abstract

The invention discloses a medical training evaluation method and system based on an intelligent simulation patient, and relates to the field of artificial intelligence technology and medical simulation, and the method comprises a virtual patient model construction step, a multi-modal perception step, an intelligent interaction step and an intelligent evaluation step. By integrating voice recognition, natural language understanding, facial expression recognition, gesture recognition and virtual reality technologies, a high-simulation and high-interactivity intelligent patient model is constructed for medical students or clinical medical staff to perform diagnosis and treatment training and skill operation assessment. The virtual patient model supports custom information such as age, gender, race, medical history, living habits and the like, different virtual patient images and different pathophysiological states are generated, and personalized learning requirements of users are met. In addition, targeted training suggestions are provided, the training difficulty and content can be adjusted according to the performance of the user, and different training requirements are met.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the fields of artificial intelligence technology and medical simulation, and in particular to a medical training evaluation method and system based on intelligent simulated patients. Background Art

[0002] 1. The importance of traditional medical experimental teaching and the shortcomings of simulation training In medical education and clinical practice, experimental teaching in traditional medicine plays a crucial role in cultivating the clinical skills and practical abilities of medical students and medical staff. However, existing simulation training has many limitations. Traditional simulated patients usually have relatively simple functions and can only simulate some common and simple clinical scenarios. They cannot realistically simulate the physiological and psychological reactions of patients at different stages of the disease, making it difficult to create a highly realistic clinical diagnosis and treatment scenario. In addition, traditional simulation training lacks personalization and often adopts a unified teaching model, which makes it difficult to meet the learning needs of different students and provide targeted feedback on students' learning effects and progress. This has, to a certain extent, affected the teaching effect and the cultivation of students' practical ability.

[0003] (2) Development of multimodal interaction technology and its current application in the medical field Multimodal interaction technology emerged in the 1990s. With the recent advancement of computing power, multimodal interaction technology has evolved from simple technology integration to the integration of multiple information interaction methods such as voice, vision, and touch, providing users with a more natural, flexible, and efficient interactive experience. In the medical field, multimodal interaction technology has gradually been applied. The development of multimodal interaction technology models varies significantly across different organ systems, with a particular focus on disease diagnosis. For example, multimodal interaction technology can provide doctors with more comprehensive patient information by integrating medical images (such as CT and MRI) with textual data (such as medical histories), enabling doctors to more comprehensively assess the patient's condition, thereby assisting clinical diagnosis and helping to develop personalized treatment plans. Furthermore, in rehabilitation therapy, computer vision technology can be used to capture the patient's movements and gestures, or combined with wearable devices and sensors, to provide real-time guidance and assessment for rehabilitation training. However, the application of this technology in medical training is still in the exploratory stage and a comprehensive system has not yet been established. The scope and depth of its application need to be further expanded.

[0004] (3) Limitations of existing technologies Most existing medical simulation systems only support a single interaction method, such as simple voice Q&A or touch operation. These systems are clearly inadequate in terms of interaction and functionality. In real clinical scenarios, students need to access information through multiple senses, such as observing patients' facial expressions and body language. Voice or touchscreen interaction alone cannot provide a comprehensive clinical experience. Their accuracy and stability also need to be improved, lacking authenticity and immersion, and therefore cannot meet the complex information exchange requirements of clinical diagnosis and treatment. This makes it difficult to simulate the diversity and uncertainty of real clinical conditions, resulting in students lacking the ability to solve complex problems in actual clinical scenarios. Summary of the Invention

[0005] The purpose of the present invention is to overcome the deficiencies of the prior art and to provide a medical training evaluation method and system based on intelligent simulated patients.

[0006] The object of the present invention is achieved through the following technical solutions: In one aspect, the present invention provides a medical training evaluation method based on an intelligent simulated patient, comprising the following steps: S1. Generate a virtual patient model based on a plurality of preset disease cases or standard cases, wherein the virtual patient model includes a case data module and a three-dimensional simulation module, and construct a virtual diagnosis and treatment scene based on the virtual patient model; S2. The user uses voice and visual acquisition devices to interactively consult with the virtual patient model in the virtual diagnosis and treatment scenario and make corresponding treatment decisions; S3. Acquire voice and visual information when the user interacts with the virtual patient model, and convert the voice information during the interaction into text information; the visual information includes facial expression information, body language information, and eye contact information; construct a medical knowledge database, and perform semantic analysis on the text information based on the medical knowledge database; and perform behavioral analysis on the user's emotional state, focus, and gestures based on the visual information; S4. Compare the user's semantic parsing results and behavioral analysis results with standard diagnosis and treatment data to generate a personalized medical training and evaluation report for the user.

[0007] Furthermore, the step S1 specifically includes: Based on Unity / Unreal Engine technology, 3D modeling and animation drive are carried out to build 3D virtual patient images through high-precision 3D character modeling; Extracting data on patients' basic conditions, symptom responses, and disease progression from multiple disease cases or standard cases, and using a generative adversarial network to construct a case data module that supports customizable age, gender, race, and pathological characteristics under different disease states; 3D modeling of the face of the 3D virtual patient image to construct a three-dimensional simulation module; Combining the case data module and the three-dimensional simulation module, different virtual diagnosis and treatment scenarios are constructed, including emergency scenarios and chronic disease management scenarios.

[0008] Furthermore, the 3D modeling of the face of the 3D virtual patient image and the construction of a three-dimensional simulation module specifically include: Adopting the generative adversarial network architecture, we build an expression synthesis model and generate corresponding expression images based on emotion parameters. The generated expression images are synthesized and combined with 3D modeling technology to perform 3D modeling of the simulated patient's face.

[0009] Furthermore, the process of acquiring voice information and visual information when the user interacts with the virtual patient model and converting the voice information during the interaction into text information specifically includes: By configuring a high-sensitivity, low-noise microphone array, the voice information during user interaction is captured, the collected voice information is processed for noise reduction and echo elimination, the processed voice information is input into the automatic speech recognition system, and the deep learning algorithm is used to convert the voice information into text data.

[0010] Furthermore, the construction of a medical knowledge database and the semantic analysis of the text information based on the medical knowledge database specifically include: Medical terms are stored in a structured manner, including disease names, symptoms, and treatment plans. Relationship graphs are used to connect medical terms and build a medical knowledge database. Combined with the medical knowledge database, the graph neural network model is used to identify and semantically parse the medical terms in the information in this article.

[0011] Furthermore, the behavioral analysis of the user's emotional state, focus, and gestures based on visual information specifically includes: By configuring a high-resolution camera, the system can capture the user's facial expressions, gestures, and eye direction in real time. The histogram equalization method is used to optimize the collected visual information and enhance the image quality; Use target detection algorithms to analyze and identify the user's emotional state, focus, and gestures in real time.

[0012] Furthermore, it also includes: Based on the visual information extracted during the user interaction process, the emotional change parameters of the virtual patient model are generated, and the expression synthesis model generates new expression images based on the emotional change parameters; an expression transition algorithm is designed and used to fuse the new expression images, and 3D modeling technology is combined to perform 3D modeling of the simulated patient's face, and update the stereo simulation module; Based on the consultation and treatment decisions made during the user interaction process, skeletal animation, facial drive technology, and emotion recognition model are used in combination with contextual simulation to adjust the symptom response of the virtual patient model in real time and update the case data module.

[0013] Furthermore, the step S4 specifically includes: Using machine learning algorithms, the system automatically analyzes the user's diagnosis, treatment, and assessment process based on the user's semantic analysis and behavioral analysis results. It evaluates the fluency and professionalism of the voice conversation, as well as the rationality of the diagnosis and treatment process and sequence, and provides personalized training recommendations. Combined with a deep learning model, the user's diagnosis, treatment and evaluation process is scored, mainly including the accuracy of the assessment, the rationality of treatment decisions, language fluency, tone and emotional management, the interactivity of facial expressions and body language, and the user's humanistic care ability, and finally a personalized training report is generated.

[0014] Another aspect of the present invention provides a medical training and evaluation system based on an intelligent simulated patient, comprising: Virtual patient model construction module: used to generate a virtual patient model based on multiple preset disease cases or standard cases. The virtual patient model includes a case data module and a three-dimensional simulation module, and a virtual diagnosis and treatment scene is constructed based on the virtual patient model; Multimodal perception module: This module allows users to conduct interactive consultations with virtual patient models in virtual diagnosis and treatment scenarios and make corresponding treatment decisions through voice and visual acquisition devices. Intelligent interaction module: used to obtain voice and visual information when the user interacts with the virtual patient model, and convert the voice information during the interaction into text information; the visual information includes facial expression information, body language information, and eye contact information; build a medical knowledge database, and perform semantic analysis on the text information based on the medical knowledge database; and perform behavioral analysis on the user's emotional state, focus, and gestures based on the visual information; Intelligent evaluation module: used to compare the user's semantic parsing results and behavioral analysis results with standard diagnosis and treatment data to generate personalized medical training and evaluation reports for the user.

[0015] Furthermore, it also includes: Real-time adjustment module for the virtual patient model: This module is used to generate emotion change parameters for the virtual patient model based on visual information extracted during user interaction. The expression synthesis model generates new expression images based on these emotion change parameters. The module also designs and utilizes an expression transition algorithm to fuse the new expression images, and combines 3D modeling technology to create a 3D model of the simulated patient's face, thereby updating the stereoscopic simulation module. Based on the consultation and treatment decisions made during the user interaction process, skeletal animation, facial drive technology, and emotion recognition model are used in combination with contextual simulation to adjust the symptom response of the virtual patient model in real time and update the case data module.

[0016] The beneficial effects of the present invention are: 1) Integration and comprehensiveness of multimodal interaction: This invention integrates multiple multimodal interaction technologies, such as speech recognition, visual recognition, natural language processing, expression and emotion simulation, into a single system, and constructs a 3D virtual patient image, creating a highly realistic and interactive intelligent simulated patient model that can simultaneously process multiple information such as speech, expression, and body movements, providing users with a more comprehensive and realistic clinical training experience.

[0017] 2) Personalization and Customization: This invention utilizes a virtual patient management system that supports customization of age, gender, ethnicity, medical history, lifestyle, and other information, generating diverse virtual patient profiles and pathophysiological states to meet the user's personalized learning needs. Furthermore, the system provides targeted training recommendations, adjusting the difficulty and content of training based on the user's performance to meet diverse training needs.

[0018] 3) Expansion of application scenarios and functions: This invention covers multiple aspects of medical education, including clinical skills training, communication skills training, situational simulation training, etc. The system can simulate the physiological and psychological reactions at different stages of the disease, and support full-process training from consultation, diagnosis to treatment recommendations. While training users' basic clinical skills, it can also promote the formation of their holistic clinical thinking.

[0019] 4) Emphasis on immersion and realism: This invention utilizes high-precision voice acquisition and recognition, coupled with natural language processing, to enable the system to mimic the natural and fluent human voice under various physiological and pathological conditions. Integrating specific speech contexts, the system simulates and adjusts the patient's facial expressions and body language, intuitively presenting them through a 3D virtual avatar. The system also captures the user's facial and body language, providing real-time feedback and enabling natural interaction across the evolving disease pathology. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] Figure 1 A flowchart of a medical training evaluation method based on intelligent simulated patients; Figure 2This is an architectural diagram of a medical training evaluation method based on intelligent patient simulation; Figure 3 This is a structural diagram of a medical training and evaluation system based on intelligent simulated patients; Figure 4 This is the structural block diagram of the multimodal perception module. DETAILED DESCRIPTION

[0021] The following will clearly and completely describe the technical solutions of the present invention in conjunction with the embodiments. Obviously, the embodiments described are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative work shall fall within the scope of protection of the present invention.

[0022] See Figures 1-4 , the present invention provides a technical solution: On the one hand, an embodiment of the present invention provides a medical training evaluation method based on intelligent simulated patients, the flow chart and architecture diagram of which are respectively as follows: Figure 1 and Figure 2 As shown, the following steps S1-S4 are included: S1. Generate a virtual patient model based on a plurality of preset disease cases or standard cases, wherein the virtual patient model includes a case data module and a three-dimensional simulation module, and construct a virtual diagnosis and treatment scene based on the virtual patient model; In a preferred embodiment, step S1 specifically includes: S11. 3D modeling and animation driven by Unity / Unreal Engine technology, and construction of 3D virtual patient images through high-precision 3D character modeling; S12. Extract the basic condition, symptom response and disease progression data of patients from multiple disease cases or standard cases, and use generative adversarial networks to build a case data module. The case data module supports customized age, gender, race and pathological characteristics under different disease states; Generative adversarial networks (GANs) can be used to generate realistic virtual patient data, such as cardiovascular diseases, respiratory diseases, mental illnesses, etc.

[0023] S13, performing 3D modeling on the face of the 3D virtual patient image to construct a three-dimensional simulation module; specifically comprising: Adopting the generative adversarial network architecture, we build an expression synthesis model and generate corresponding expression images based on emotion parameters. The generated expression images are synthesized and combined with 3D modeling technology to perform 3D modeling of the simulated patient's face.

[0024] S14. Combine the case data module and the three-dimensional simulation module to construct different virtual diagnosis and treatment scenarios, including emergency scenarios and chronic disease management scenarios.

[0025] Through virtual patient models and corresponding virtual diagnosis and treatment scenarios, users can perform standardized diagnosis and treatment exercises according to preset cases, or practice according to adaptively generated cases, and make operations and responses based on the patient's real-time adjusted symptoms and reactions; through simulation training in different virtual diagnosis and treatment scenarios, such as emergency, chronic disease management, health education, etc., users can develop communication skills and the ability to cope with different scenarios.

[0026] S2. The user uses voice and visual acquisition devices to interactively consult with the virtual patient model in the virtual diagnosis and treatment scenario and make corresponding treatment decisions; The virtual patient model is capable of generating natural and fluent human voices to simulate the speech styles of different ages, genders, and emotions. Furthermore, by simulating characteristic sound effects such as coughing and wheezing associated with illness, the realism of clinical scenarios is enhanced. This makes the speech more emotionally expressive in rhythm and intonation. The system also supports multi-round dialogue management, enabling natural language interaction based on context and the evolution of different disease conditions. The virtual patient model can simulate and adjust the patient's facial expressions and body language based on the context. Users can interact with the virtual patient using augmented reality (AR), virtual reality (VR), and mixed reality (MR). The constructed virtual patient can interact with real diagnostic objects, providing an immersive medical training environment.

[0027] S3. Acquire voice and visual information when the user interacts with the virtual patient model, and convert the voice information during the interaction into text information; the visual information includes facial expression information, body language information, and eye contact information; construct a medical knowledge database, and perform semantic analysis on the text information based on the medical knowledge database; and perform behavioral analysis on the user's emotional state, focus, and gestures based on the visual information; In a preferred embodiment, the process of obtaining voice information and visual information when the user interacts with the virtual patient model and converting the voice information during the interaction into text information specifically includes: By configuring a high-sensitivity, low-noise microphone array, the voice information during user interaction is captured, and digital signal processing algorithms (such as adaptive filters) are used to reduce noise and eliminate echoes on the collected voice information. The processed voice information is input into the automatic speech recognition system, and deep learning algorithms (such as recurrent neural networks and attention mechanisms) are used to convert the voice information into text data.

[0028] In a preferred embodiment, the construction of a medical knowledge database and the semantic analysis of the text information based on the medical knowledge database specifically include: Medical terms are stored in a structured manner, including disease names, symptoms, and treatment plans. Relationship graphs are used to connect medical terms and build a medical knowledge database. Combined with the medical knowledge database, the graph neural network model is used to identify and semantically parse the medical terms in the information in this article.

[0029] Using an end-to-end deep learning speech recognition model, combined with microphone array directional speech enhancement technology, it achieves high-precision speech acquisition and recognition while accurately analyzing the content of medical consultation conversations and medical terminology, and generates professional and natural diagnosis, treatment and evaluation dialogues.

[0030] Natural Language Processing (NLP) technology is used to deeply analyze user questions at a semantic level. A medical semantic understanding model based on Transformers (such as GPT / BERT) is combined with a medical knowledge database to accurately parse medical terminology, enabling the simulated patient to understand the professional consultation content and generate reasonable voice responses. Furthermore, prosody and emotion embedding technology is used to automatically adjust the tone and speed of speech, simulating the patient's emotional state and enhancing the authenticity of the interaction, enabling voice interaction between the user and the virtual patient model.

[0031] In a preferred embodiment, the behavioral analysis of the user's emotional state, focus, and gestures based on visual information specifically includes: By configuring a high-resolution camera, the system can capture the user's facial expressions, gestures, and eye direction in real time. The histogram equalization method is used to optimize the collected visual information and enhance the image quality; Use target detection algorithms to analyze and identify the user's emotional state, focus, and gestures in real time.

[0032] Based on the user's tone and questioning style, skeletal animation, facial driver technology (such as BlendShape and LiveLink), and emotion recognition models are used to simulate and adjust the virtual patient's facial expressions and body language in context. These expressions, such as frowning in pain, rapid breathing, clenched fists, and curled-up body movements, can be realistically displayed. Expression synthesis algorithms (such as GANs or 3D modeling) dynamically simulate the patient's emotional changes, such as anxiety and apathy, synchronizing expression and language for enhanced immersion. By integrating augmented reality (AR), virtual reality (VR) technologies (such as MetaQuest and HTC Vive), and mixed reality (MR) interactions, the virtual patient can interact with real objects (such as a stethoscope), providing an immersive medical training environment and enabling visual information exchange between the user and the virtual patient model.

[0033] S4, comparing the user's semantic parsing results and behavioral analysis results with standard diagnosis and treatment data to generate a personalized medical training and evaluation report for the user. Further, the step S4 specifically includes: Using machine learning algorithms, the system automatically analyzes the user's diagnosis, treatment, and assessment process based on the user's semantic analysis and behavioral analysis results. It evaluates the fluency and professionalism of the voice conversation, as well as the rationality of the diagnosis and treatment process and sequence, and provides personalized training recommendations. Combined with a deep learning model, the user's diagnosis, treatment and evaluation process is scored, mainly including the accuracy of the assessment, the rationality of treatment decisions, language fluency, tone and emotional management, the interactivity of facial expressions and body language, and the user's humanistic care ability. Ultimately, a personalized training report is generated, including the user's operation records, evaluation results and improvement suggestions.

[0034] By recording the user's questioning methods, diagnostic accuracy, treatment recommendations and other data, analyzing the user's diagnosis and evaluation path, providing scoring and improvement suggestions and generating personalized training reports, it can help users improve their diagnosis and treatment and nursing evaluation capabilities.

[0035] In addition, remote teaching support can also be provided. First, a cloud computing-based platform architecture is built to store data, including simulated patient cases, user operation records and other data; and a platform that supports multi-terminal access and multi-user online collaboration is established. For example, teachers create teaching groups, and each group is assigned different virtual cases for practice; the teacher's side provides a real-time monitoring interface to display the user's consultation process, including the content of the questions, examination operations, diagnosis results, etc., and directly sends feedback information to the user.

[0036] Furthermore, a medical training evaluation method based on an intelligent simulated patient also includes: Based on the visual information extracted during the user interaction process, the emotional change parameters of the virtual patient model are generated, and the expression synthesis model generates new expression images based on the emotional change parameters; an expression transition algorithm is designed and used to fuse the new expression images, and 3D modeling technology is combined to perform 3D modeling of the simulated patient's face, and update the stereo simulation module; Based on the consultation and treatment decisions made during the user interaction process, skeletal animation, facial drive technology, and emotion recognition model are used in combination with contextual simulation to adjust the symptom response of the virtual patient model in real time and update the case data module.

[0037] Another embodiment of the present invention provides a medical training evaluation system based on intelligent simulated patients, such as Figure 3 ,include: Virtual patient model construction module: used to generate a virtual patient model based on multiple preset disease cases or standard cases. The virtual patient model includes a case data module and a three-dimensional simulation module, and a virtual diagnosis and treatment scene is constructed based on the virtual patient model.

[0038] Multimodal perception module: It is used for users to conduct interactive consultations with virtual patient models in virtual diagnosis and treatment scenarios and make corresponding treatment decisions through voice and visual acquisition devices; the structural block diagram of the multimodal perception module is as follows Figure 4 shown.

[0039] Intelligent interaction module: used to obtain voice and visual information when the user interacts with the virtual patient model, and convert the voice information during the interaction into text information; the visual information includes facial expression information, body language information and eye information; build a medical knowledge database, and perform semantic analysis of the text information based on the medical knowledge database; perform behavioral analysis of the user's emotional state, focus and gestures based on the visual information.

[0040] Intelligent evaluation module: used to compare the user's semantic parsing results and behavioral analysis results with standard diagnosis and treatment data to generate personalized medical training and evaluation reports for the user.

[0041] Furthermore, it also includes: Real-time adjustment module for the virtual patient model: This module is used to generate emotion change parameters for the virtual patient model based on visual information extracted during user interaction. The expression synthesis model generates new expression images based on these emotion change parameters. The module also designs and utilizes an expression transition algorithm to fuse the new expression images, and combines 3D modeling technology to create a 3D model of the simulated patient's face, thereby updating the stereoscopic simulation module. Based on the consultation and treatment decisions made during the user interaction process, skeletal animation, facial drive technology, and emotion recognition model are used in combination with contextual simulation to adjust the symptom response of the virtual patient model in real time and update the case data module.

[0042] To enhance the intelligence of the system, the intelligent patient simulator needs to be continuously optimized and upgraded, mainly including: (1) Data accumulation and optimization: The system records all interactive data, including voice conversations, operational behaviors, diagnostic results, treatment recommendations, etc., to optimize voice recognition, emotion simulation, and disease deduction models; the distributed deep learning parameter update optimization system improves the training efficiency and accuracy of the model, and improves the accuracy of the system's understanding of medical terminology.

[0043] (2) User feedback and personalized adjustment: Allow users to adjust the interaction style of the simulated patient, the way the condition changes, etc. according to training needs. For example, provide serious, gentle or anxious interaction styles to meet the needs of different teaching scenarios; through medical literature crawlers and expert review mechanisms, continuously update the medical knowledge base to ensure the system's adaptability to the latest medical research.

[0044] The differences between the present invention and the prior art are mainly reflected in the following: 1. Integration and comprehensiveness of multimodal interaction: This invention integrates multiple multimodal interaction technologies, including speech recognition, visual recognition, natural language processing, and expression and emotion simulation, into a single system. It also constructs a 3D virtual patient image, creating a highly realistic and interactive intelligent simulated patient model that can simultaneously process multiple information, including speech, expression, and body movements, providing users with a more comprehensive and realistic clinical training experience. In contrast, existing technologies have a single interaction method, making it difficult to simulate real clinical scenarios and providing users with a more immersive experience.

[0045] 2. Personalization and customization capabilities: The present invention supports customizing information such as age, gender, race, medical history, and living habits through a virtual patient management system, generating different virtual patient images and different pathophysiological states to meet the user's personalized learning needs. In addition, the system also provides targeted training suggestions, which can adjust the training difficulty and content according to the user's performance to meet different training needs. However, existing simulation systems often use a unified teaching model, solidify the learning process, and it is difficult to provide targeted feedback on the user's learning effect and progress, and cannot meet the personalized needs of patients.

[0046] 3. Expansion of application scenarios and functions: This invention covers multiple aspects of medical education, including clinical skills training, communication skills training, and situational simulation training. The system can simulate physiological and psychological reactions at different stages of the disease and support full-process training from consultation and diagnosis to treatment recommendations. While training users in basic clinical skills, it can also promote the formation of their holistic clinical thinking. Existing multimodal technologies, on the other hand, are mostly focused on disease diagnosis assistance and survival prediction. The application scenarios are relatively simple and mostly auxiliary in nature, lacking a comprehensive simulation of the entire clinical diagnosis and treatment process.

[0047] 4. Emphasis on immersion and realism: This invention uses high-precision voice acquisition and recognition, and natural language processing to enable the system to mimic the natural and fluent human voice under different physiological and pathological conditions. Combined with the specific context, the system simulates and adjusts the patient's facial expressions and body language, and intuitively displays them through 3D virtual images. It also captures the user's facial and body information, provides real-time feedback, and provides users with natural interaction under different disease evolutions. However, most existing technologies only support simple voice questions and answers or touch interaction, which cannot truly simulate complex clinical scenarios and result in a poor user experience.

[0048] The foregoing description is merely a preferred embodiment of the present invention. It should be understood that the present invention is not limited to the form disclosed herein and should not be construed as excluding other embodiments. Rather, the present invention can be used in various other combinations, modifications, and environments and can be modified within the scope of the concept described herein through the above teachings or techniques or knowledge in the relevant field. Modifications and variations made by those skilled in the art that do not depart from the spirit and scope of the present invention are intended to be protected by the appended claims.

Claims

1. A medical training evaluation method based on intelligent simulated patients, characterized in that: The following steps are involved: S1. Generate a virtual patient model based on a plurality of preset disease cases or standard cases, wherein the virtual patient model includes a case data module and a three-dimensional simulation module, and construct a virtual diagnosis and treatment scene based on the virtual patient model; S2. The user uses voice and visual acquisition devices to interactively consult with the virtual patient model in the virtual diagnosis and treatment scenario and make corresponding treatment decisions; S3, obtaining voice information and visual information when the user interacts with the virtual patient model, and converting the voice information during the interaction into text information; the visual information includes facial expression information, body language information, and eye contact information; Construct a medical knowledge database and perform semantic analysis on the text information based on the medical knowledge database; perform behavioral analysis on the user's emotional state, focus, and gestures based on visual information; S4. Compare the user's semantic parsing results and behavioral analysis results with standard diagnosis and treatment data to generate a personalized medical training and evaluation report for the user.

2. The medical training and evaluation method based on intelligent simulated patients according to claim 1, characterized in that: The step S1 specifically includes: Based on Unity / Unreal Engine technology, 3D modeling and animation drive are carried out to build 3D virtual patient images through high-precision 3D character modeling; Extracting data on patients' basic conditions, symptom responses, and disease progression from multiple disease cases or standard cases, and using a generative adversarial network to construct a case data module that supports customizable age, gender, race, and pathological characteristics under different disease states; 3D modeling of the face of the 3D virtual patient image to construct a three-dimensional simulation module; Combining the case data module and the three-dimensional simulation module, different virtual diagnosis and treatment scenarios are constructed, including emergency scenarios and chronic disease management scenarios.

3. The medical training and evaluation method based on intelligent simulated patients according to claim 2, characterized in that: The 3D modeling of the face of the 3D virtual patient image and the construction of a three-dimensional simulation module specifically include: Adopting the generative adversarial network architecture, we build an expression synthesis model and generate corresponding expression images based on emotion parameters. The generated expression images are synthesized and combined with 3D modeling technology to perform 3D modeling of the simulated patient's face.

4. The medical training and evaluation method based on intelligent simulated patients according to claim 1, characterized in that: The process of obtaining voice information and visual information when the user interacts with the virtual patient model and converting the voice information during the interaction into text information specifically includes: By configuring a high-sensitivity, low-noise microphone array, the voice information during user interaction is captured, the collected voice information is processed for noise reduction and echo elimination, the processed voice information is input into the automatic speech recognition system, and the deep learning algorithm is used to convert the voice information into text data.

5. The medical training and evaluation method based on intelligent simulated patients according to claim 1, characterized in that: The step of constructing a medical knowledge database and performing semantic analysis on the information in the article based on the medical knowledge database specifically includes: Medical terms are stored in a structured manner, including disease names, symptoms, and treatment plans. Relationship graphs are used to connect medical terms and build a medical knowledge database. Combined with the medical knowledge database, the graph neural network model is used to identify and semantically parse the medical terms in the information in this article.

6. The medical training and evaluation method based on intelligent simulated patients according to claim 1, characterized in that: The behavioral analysis of the user's emotional state, focus, and gestures based on visual information specifically includes: By configuring a high-resolution camera, the system can capture the user's facial expressions, gestures, and eye direction in real time. The histogram equalization method is used to optimize the collected visual information and enhance the image quality; Use target detection algorithms to analyze and identify the user's emotional state, focus, and gestures in real time.

7. The medical training evaluation method based on intelligent simulated patients according to claim 3, characterized in that: Also includes: Based on the visual information extracted during the user interaction process, the emotional change parameters of the virtual patient model are generated, and the expression synthesis model generates new expression images based on the emotional change parameters; an expression transition algorithm is designed and used to fuse the new expression images, and 3D modeling technology is combined to perform 3D modeling of the simulated patient's face, and update the stereo simulation module; Based on the consultation and treatment decisions made during the user interaction process, skeletal animation, facial drive technology, and emotion recognition model are used in combination with contextual simulation to adjust the symptom response of the virtual patient model in real time and update the case data module.

8. The medical training and evaluation method based on intelligent simulated patients according to claim 1, characterized in that: The step S4 specifically includes: Using machine learning algorithms, the system automatically analyzes the user's diagnosis, treatment, and assessment process based on the user's semantic analysis and behavioral analysis results. It evaluates the fluency and professionalism of the voice conversation, as well as the rationality of the diagnosis and treatment process and sequence, and provides personalized training recommendations. Combined with a deep learning model, the user's diagnosis, treatment and evaluation process is scored, mainly including the accuracy of the assessment, the rationality of treatment decisions, language fluency, tone and emotional management, the interactivity of facial expressions and body language, and the user's humanistic care ability, and finally a personalized training report is generated.

9. A medical training and evaluation system based on an intelligent simulated patient, characterized by: include: Virtual patient model construction module: used to generate a virtual patient model based on multiple preset disease cases or standard cases. The virtual patient model includes a case data module and a three-dimensional simulation module, and a virtual diagnosis and treatment scene is constructed based on the virtual patient model; Multimodal perception module: This module allows users to conduct interactive consultations with virtual patient models in virtual diagnosis and treatment scenarios and make corresponding treatment decisions through voice and visual acquisition devices. Intelligent interaction module: used to obtain voice and visual information when the user interacts with the virtual patient model, and convert the voice information during the interaction into text information; the visual information includes facial expression information, body language information and eye contact information; Construct a medical knowledge database and perform semantic analysis on the text information based on the medical knowledge database; perform behavioral analysis on the user's emotional state, focus, and gestures based on visual information; Intelligent evaluation module: used to compare the user's semantic parsing results and behavioral analysis results with standard diagnosis and treatment data to generate personalized medical training and evaluation reports for the user.

10. The medical training and evaluation system based on intelligent simulated patients according to claim 9, characterized in that: Also includes: Real-time adjustment module for the virtual patient model: This module is used to generate emotion change parameters for the virtual patient model based on visual information extracted during user interaction. The expression synthesis model generates new expression images based on these emotion change parameters. The module also designs and utilizes an expression transition algorithm to fuse the new expression images, and combines 3D modeling technology to create a 3D model of the simulated patient's face, thereby updating the stereoscopic simulation module. Based on the consultation and treatment decisions made during the user interaction process, skeletal animation, facial drive technology, and emotion recognition model are used in combination with contextual simulation to adjust the symptom response of the virtual patient model in real time and update the case data module.

Citation Information

Patent Citations

  • Semantic graph network-based medical prediction method and system

    CN113035362A

  • General practice doctor-patient communication training examination method based on virtual reality technology

    CN117854341A

  • Semantic recognition method and device based on graph neural network, medium and electronic equipment

    CN118586401A

  • Automatic image segmentation and identification system for radiology department

    CN119991690A

Cited By

  • Physiological state multi-modal simulation method, device and equipment based on deep learning

    CN120954739A

  • Virtual standardized patient image generation and dialogue method and system

    CN121545653A

  • Clinical experiment environment simulation interaction method based on knowledge graph

    CN122314431A