Intelligent medical communication and interactive training system based on virtual characters and method thereof

TW202636459AActive Publication Date: 2026-09-01TAIPEI MEDICAL UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
TW114106119
Authority / Receiving Office
TW · TW
Patent Type
Applications
Current Assignee / Owner
Filing Date
2025-02-19
Publication Date
2026-09-01
Estimated Expiration
2045-02-18

AI Technical Summary

Technical Problem

Current medical education lacks effective training in empathy and medical communication, leading to subjective assessments and inefficiencies in patient interaction.

Method used

An intelligent medical communication and interaction training system using virtual characters, employing speech-to-text, voice recognition, facial expression analysis, and emotion analysis technologies to provide comprehensive training scores and responses.

Benefits of technology

Enhances empathy training for healthcare professionals through objective assessment and improved interaction skills using artificial intelligence.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure TWG2TA001073741_001
    Figure TWG2TA001073741_001
  • Figure TWG2TA001073741_002
    Figure TWG2TA001073741_002
  • Figure TWG2TA001073741_003
    Figure TWG2TA001073741_003
Patent Text Reader

Abstract

An intelligent medical communication interaction training system based on virtual characters and method thereof are disclosed. Trainee device uses identification data to log into communication interaction training and analysis server. Training scenario and virtual character are selected through training user interface by trainee device. Trainee message or trainee voice and trainee image are provided to communication interaction training and analysis server from trainee device. Features and types of trainee message or trainee voice and trainee image are analyzed according to expressions, emotions, and conversations of multimodal data by communication interaction training and analysis server using speech-to-text technology, voice recognition technology, natural language processing technology, expression analysis technology, and emotion analysis technology to obtain feature analysis results and type analysis results. Feature analysis results and type analysis results are analyze with interactive mode by communication interaction training and analysis server to obtain communication and interaction comprehensive score and comprehensive feature. Text response, voice response and emoji response are generated for virtual human character information separately based on communication and interaction comprehensive score and comprehensive feature by communication interaction training and analysis server. Text response, voice response and emoji response are feedback to trainee device to present and playback. Medical communication and interaction training at selected training scenario for trainee are realized. Therefore, the efficiency of intelligent medical communication interaction training based on virtual characters may be achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to a medical communication and interaction training system and method, particularly an intelligent medical communication and interaction training system and method based on virtual roles. Prior Technology

[0002] As Taiwan enters an aging society, the average life expectancy of its citizens continues to increase. Long-term care and mortality issues affect domestic medical care, the economy, and caregiver manpower. Therefore, respecting patients' medical autonomy and avoiding suffering caused by ineffective medical treatment and waste of medical resources are crucial. Communication between medical staff and patients should be based on respect for autonomy. Maintaining professionalism and empathy in assisting patients is essential. Medical staff also need to learn empathy through professionalism and the experience of their teachers during the teaching and learning process.

[0003] Because human emotions and reactions are highly complex and variable, they often carry hidden meanings beyond their surface meaning. For example, when different people say "I'm fine," they may express completely different emotions and reactions—it could be genuine calm or a hidden sadness. Furthermore, each person's experience of emotions differs; some are more easily empathetic than others. Current medical education lacks empathy training and is not convenient for direct patient interaction. Medical communication and interaction training presents certain difficulties and tends towards subjective assessment. Therefore, exploring the possibility of using artificial intelligence to assist in empathy training could help improve the empathy abilities of healthcare professionals.

[0004] In conclusion, it is clear that the existing teaching methods for medical staff have long suffered from a lack of training in assisting medical communication and interaction. Therefore, it is necessary to propose improved technical means to solve this problem. Summary of the Invention

[0005] In view of the fact that prior art cannot solve the problem of the lack of auxiliary medical communication and interaction training in the current teaching methods for medical staff, the present invention discloses a system and method thereof, wherein:

[0006] This invention discloses an intelligent medical communication and interaction training system and method based on virtual characters.

[0007] First, this invention discloses an intelligent medical communication and interaction training system based on virtual characters. This system includes: a trainee device and a communication and interaction training and analysis server. The trainee device further includes: a first non-transitory computer-readable storage medium and a first processor; the communication and interaction training and analysis server further includes: a second non-transitory computer-readable storage medium and a second processor.

[0008] A first non-transitory computer-readable storage medium stores a plurality of first computer-readable instructions; and a first processor is electrically connected to the first non-transitory computer-readable storage medium, executing the first computer-readable instructions to cause the trainee device to perform the following steps:

[0009] Log in to the Communication Interaction Training and Analysis Server using identity verification data to display the training user interface; select one of multiple training scenarios in the training user interface and provide the corresponding training scenario identification information to the Communication Interaction Training and Analysis Server; select at least one virtual character in the training user interface and provide the corresponding trainee virtual character identification information to the Communication Interaction Training and Analysis Server; receive trainee messages or trainee voice and trainee images corresponding to the selected virtual character and provide them to the Communication Interaction Training and Analysis Server; and receive, present, and play text responses, voice responses, and emoticon responses corresponding to unselected virtual characters from the Communication Interaction Training and Analysis Server.

[0010] A second non-transitory computer-readable storage medium stores a plurality of second computer-readable instructions and a plurality of virtual scene data, each virtual scene data containing at least one virtual human character data; and a second processor is electrically connected to the second non-transitory computer-readable storage medium and executes the second computer-readable instructions to cause the trainee device to perform the following steps:

[0011] After the trainee's device logs in using identification data, a training user interface is provided to the trainee's device; training scene recognition information and trainee virtual character recognition information are obtained from the trainee's device; at least one corresponding virtual human character data is retrieved based on the training scene recognition information, and the virtual human character data corresponding to the trainee's virtual character recognition information is excluded from the at least one virtual human character data; trainee messages or trainee voice and trainee images are obtained from the trainee's device; multimodal data facial expressions, emotions, and message features are analyzed using speech-to-text technology, voice recognition technology, natural language processing technology, facial expression analysis technology, and emotion analysis technology to obtain feature analysis results and type analysis results for the trainee's messages or trainee voice and trainee images; interaction pattern analysis is performed on the feature analysis results and type analysis results to obtain a comprehensive communication interaction score and comprehensive communication interaction features; based on the comprehensive communication interaction score and comprehensive communication interaction features, general artificial intelligence is used to generate corresponding text responses, voice responses, and facial expressions for at least one virtual human character data; and the text responses, voice responses, and facial expressions are provided to the trainee's device.

[0012] In addition, this invention discloses a training method for intelligent medical communication and interaction based on virtual characters, which includes the following steps:

[0013] First, the trainee's device logs into the Communication Interaction Training and Analysis Server using identification information. The server provides a training user interface to the trainee's device and displays it. Next, the trainee's device selects one of several training scenarios from the user interface and provides the corresponding training scenario identification information to the server. Then, the trainee's device selects at least one virtual character from the user interface and provides the corresponding virtual character identification information to the server. Next, the server stores multiple virtual scenario data, each containing at least one virtual character data. Then, the server retrieves the corresponding at least one virtual character data based on the training scenario identification information, excluding the virtual character data corresponding to the trainee's virtual character identification information. Finally, the trainee's device receives trainee messages or trainee voice and trainee image corresponding to the selected virtual character and provides... The training and analysis server then uses speech-to-text technology, voice recognition technology, natural language processing technology, facial expression analysis technology, and emotion analysis technology to perform multimodal data feature analysis on trainee messages, trainee voice, and trainee images to obtain feature analysis results and type analysis results. Next, the server performs interaction pattern analysis on the feature analysis results and type analysis results to obtain a comprehensive communication interaction score and comprehensive communication interaction features. Then, based on the comprehensive communication interaction score and comprehensive communication interaction features, the server uses general artificial intelligence to generate corresponding text responses, voice responses, and facial expressions for at least one virtual human character. The server then provides the text responses, voice responses, and facial expressions to the trainee's device. Finally, the trainee's device displays and plays the text responses, voice responses, and facial expressions corresponding to the unselected virtual character.

[0014] The system and method disclosed in this invention are as described above. The trainee device logs into the communication interaction training and analysis server using identification data. Through the training user interface, the trainee selects a training scenario and virtual role. The trainee device provides trainee messages, trainee voice, and trainee images to the communication interaction training and analysis server. The communication interaction training and analysis server uses speech-to-text technology, voice recognition technology, natural language processing technology, facial expression analysis technology, and emotion analysis technology to perform multimodal data feature analysis on the trainee messages, trainee voice, and trainee images to obtain feature analysis results and type analysis results. Then, interaction pattern analysis is performed on the feature analysis results and type analysis results to obtain a comprehensive communication interaction score and comprehensive communication interaction features. Based on the comprehensive communication interaction score and comprehensive communication interaction features, general artificial intelligence is used to generate corresponding text responses, voice responses, and facial responses for the virtual human role data and feed them back to the trainee device for presentation and playback, so as to realize medical communication interaction training for the trainee in the selected training scenario.

[0015] Through the aforementioned technical means, this invention can achieve the technical effect of intelligent medical communication and interaction training based on virtual characters. Simple Explanation of the Diagram

[0016] Figure 1 is a system block diagram of the intelligent medical communication and interaction training system based on virtual characters according to the present invention. Figure 2 illustrates the user interface for training intelligent medical communication and interaction training based on virtual characters according to the present invention. Figure 3 illustrates the training scenario and virtual character display diagram of the intelligent medical communication and interaction training based on virtual characters according to the present invention. Figures 4A and 4B illustrate the flowcharts of the intelligent medical communication and interaction training method based on virtual characters according to the present invention. Implementation

[0017] The following will describe the implementation of the present invention in detail with reference to the drawings and embodiments, so that the process of how the present invention uses technical means to solve technical problems and achieve technical effects can be fully understood and implemented accordingly.

[0018] The following will first describe the intelligent medical communication and interaction training system based on virtual characters disclosed in this invention, and please refer to "Figure 1", which is a system block diagram of the intelligent medical communication and interaction training system based on virtual characters of this invention.

[0019] First, this invention discloses an intelligent medical communication and interaction training system based on virtual characters. This system includes: a trainee device 10 and a communication and interaction training and analysis server 20. The trainee device 10 further includes: a first non-transitory computer-readable storage medium 11 and a first processor 12; the communication and interaction training and analysis server 20 further includes: a second non-transitory computer-readable storage medium 21 and a second processor 22. The trainee device 10 is, for example, a general computer, a laptop computer, a tablet computer, a smartphone, etc., which is only an example and is not intended to limit the application scope of this invention.

[0020] The first non-transitory computer-readable storage medium 11 stores a plurality of first computer-readable instructions; and the first processor 12 is electrically connected to the first non-transitory computer-readable storage medium 11, executing the first computer-readable instructions to cause the trainee device 10 to perform the following steps:

[0021] Log in to the communication and interaction training and analysis server 20 using identity verification data to display the training user interface 30. The identity verification data may be, for example, a combination of account and password, digital identity certificate, etc. This is only an example and is not intended to limit the application scope of the present invention. For a schematic diagram of the training user interface 30, please refer to "Figure 2". "Figure 2" is a schematic diagram of the training user interface of the intelligent medical communication and interaction training based on virtual roles of the present invention.

[0022] The trainee can operate the trainee device 10 to select one of the multiple training scenarios provided in the training scenario selection block 31 in the training user interface 30 by mouse clicking, touch clicking, etc. (this is only an example and does not limit the scope of application of the present invention). Then, the trainee device 10 provides the training scenario identification information corresponding to the selected training scenario to the communication and interaction training and analysis server 20. The aforementioned training scenarios are, for example, clinic scenarios, medical teaching scenarios, nursing scenarios, etc., which are only examples and do not limit the scope of application of the present invention.

[0023] The trainee then operates the trainee device 10 by clicking with a mouse, touching to select, etc., one of the virtual roles provided in the virtual role selection block 32 of the training user interface 30. Then, the trainee device 10 provides the trainee virtual role identification information corresponding to the selected virtual role to the communication and interaction training and analysis server 20. Specifically, if the selected training scenario is a clinic scenario, the training user interface 30 can provide doctor virtual roles, nurse virtual roles, and patient virtual roles for the trainee to choose from; if the selected training scenario is a medical teaching scenario, the training user interface 30 can provide medical student virtual roles and professor virtual roles for the trainee to choose from. This is only an example and does not limit the scope of application of the present invention.

[0024] Next, the trainee device 10 can provide the selected training scene in the training scene display block 33 on the training user interface 30, and provide the selected virtual character 341 and the unselected virtual character 342 in the selected training scene. For a schematic diagram of the training scene 33, the selected virtual character 341 and the unselected virtual character 342, please refer to "Figure 3". "Figure 3" is a schematic diagram of the training scene and virtual character display of the intelligent medical communication interaction training based on virtual characters of the present invention.

[0025] The trainee then operates the trainee device 10 using a keyboard (which can be a physical keyboard or a virtual keyboard, this is only for illustrative purposes and does not limit the scope of application of the present invention) to input text-mode trainee messages in the text input area 351. Alternatively, the trainee can input voice-mode trainee messages through a microphone embedded or externally connected to the trainee device 10. The trainee can also capture image-mode trainee images through a network camera external to the trainee device 10 or a camera embedded in a smart device (i.e., the trainee device 10). The text input area 351 is displayed next to the selected virtual character 341. The trainee device 10 can then provide trainee messages, trainee voice, and trainee images to the communication and interactive training and analysis server 20. The trainee messages, trainee voice, and trainee images correspond to the selected virtual character 341. The selected virtual character 341 in the training user interface 30 can then instantly display the trainee's trainee messages and trainee images.

[0026] Next, the trainee device 10 receives text responses, voice responses, and facial expressions corresponding to the unselected virtual character from the communication interaction training and analysis server 20. The trainee device 10 can then display the corresponding text response in the text response area 352 corresponding to the unselected virtual character in the training user interface 30. The trainee device 10 displays the corresponding facial expression of the virtual character in the unselected virtual character 342 in the training user interface 30. The trainee device 10 plays the voice response corresponding to the unselected virtual character in the training user interface 30. In this way, the trainee can use the trainee device 10 to provide feedback on trainee messages, trainee voice, and trainee images based on the text responses, voice responses, and facial expressions corresponding to the virtual character during the training period, thereby realizing medical communication interaction training for the trainee in the selected training scenario.

[0027] It is worth noting that, based on the rapid development of virtual character technology, a simple virtual character can be projected onto the virtual character by selecting the corresponding human facial expression through artificial intelligence calculations. Alternatively, deepfake technology can be used to modify and adjust the trainee's image or an authorized real face image based on the virtual character's corresponding facial expression and then overlay it onto the virtual character, so that the virtual character presents the corresponding facial expression. This is only an example for illustration. The presentation technology related to virtual characters can be referred to the existing technology, and this invention will not elaborate further.

[0028] The second non-transitory computer-readable storage medium 21 stores a plurality of second computer-readable instructions and a plurality of virtual scene data, each virtual scene data containing at least one virtual human character data; and the second processor 22 is electrically connected to the second non-transitory computer-readable storage medium 21 and executes the second computer-readable instructions to cause the trainee device 10 to perform the following steps:

[0029] After the trainee device 10 logs into the communication and interaction training and analysis server 20 using identification information, the communication and interaction training and analysis server 20 provides the training user interface 30 to the trainee device 10. After obtaining the training scene identification information and the trainee virtual role identification information from the trainee device 10, the communication and interaction training and analysis server 20 queries the corresponding at least one virtual human role data from the second non-transitory computer-readable storage medium 21 based on the training scene identification information, and the at least one virtual human role data excludes the virtual human role data corresponding to the trainee virtual role identification information.

[0030] Specifically, assuming the training scene recognition information is the identification number C101 corresponding to the clinic scene, and the trainee virtual role recognition information is the identification data DR corresponding to the doctor virtual role, the virtual human role data corresponding to the training scene recognition information C101 are doctor, nurse, and patient. The communication interaction training and analysis server 20 retrieves the corresponding virtual human role data of doctor, nurse, and patient from the second non-transitory computer-readable storage medium 21 based on the training scene recognition information C101. After excluding the doctor corresponding to the trainee virtual role recognition information DR, the retrieved virtual human role data are nurse and patient. This is only an example and does not limit the application scope of the present invention. It is worth noting that the training scene recognition information, in addition to C101 mentioned above, can also be Clinic_205, Examination Room, etc., and the trainee virtual role recognition information, in addition to DR mentioned above, can also be RN, PT, etc. This is only an example and does not limit the application scope of the present invention.

[0031] The communication interaction training and analysis server 20 obtains trainee messages, trainee voice, and trainee images from the trainee device 10. The communication interaction training and analysis server 20 uses speech-to-text technology, voice recognition technology, natural language processing technology, facial expression analysis technology, and emotion analysis technology to perform multimodal data facial expression, emotion, and message feature analysis on the trainee messages, trainee voice, and trainee images to obtain feature analysis results and type analysis results.

[0032] The Communication Interaction Training and Analysis Server 20 uses speech-to-text technology and voice recognition technology to convert the trainee's speech in the speech modality into trainee messages in the text modality. Speech-to-text technology uses acoustic models and language models to convert speech signals into text. The trainee's speech is converted into phonemes (the smallest phonemic units in a language) by the acoustic model, and then the language model uses context and grammatical rules to predict the word order corresponding to the phonemes to improve the accuracy of the converted text. Voice recognition technology analyzes the trainee's speech to obtain the trainee's voiceprint features. The aforementioned voiceprint features include, for example, pitch, intensity, speech rate, and formants, etc. These are only examples and are not intended to limit the application scope of this invention. By using speech-to-text technology and voice recognition technology, the trainee's speech in the speech modality can be converted into trainee messages in the text modality.

[0033] The Communication Interaction Training and Analysis Server 20 uses Facial Expression Recognition (FER) technology to analyze the trainee's facial expressions in the training video. This technology is based on Artificial Intelligence (AI) and computer vision. The Server 20 first uses facial detection algorithms (e.g., Haar feature classifier, deep learning model, etc.) to locate the trainee's face in the conversation video. Then, it marks key areas of the trainee's face (e.g., corners of the eyes, corners of the mouth, eyebrows, etc.) and extracts geometric or appearance features representing facial expressions from these key areas. Finally, based on the trainee's face area, the marked key areas, and the geometric or appearance features of the expressions, it uses machine learning or deep learning models to map these expressions to predefined expression types (e.g., happiness, sadness, surprise, anger, disgust, fear, etc.).

[0034] The Communication Interaction Training and Analysis Server 20 uses emotion recognition (or sentiment analysis) technology to analyze trainee messages, voice recordings, and images to determine trainee emotions. Emotion analysis is an artificial intelligence technology that identifies and analyzes individual emotions by processing data such as voice, text, and facial expressions to infer emotional responses. The Communication Interaction Training and Analysis Server 20 uses Natural Language Processing (NLP) to process trainee messages. NLP technology extracts emotionally charged words and phrases based on vocabulary and grammar to infer the trainee's predicted emotions. The communication interaction training and analysis server 20 infers the trainee's predicted emotions based on the voiceprint features of the trainee's speech. For example, a high-pitched voiceprint feature infers excitement, while a low-pitched voiceprint feature infers sadness. This is only an example and does not limit the scope of application of the present invention. The communication interaction training and analysis server 20 calls up the obtained trainee facial expressions and combines them with the inferred trainee emotions for comprehensive analysis to obtain the trainee's emotions.

[0035] The Communication Interaction Training and Analysis Server 20 uses natural language processing technology to perform semantic and word analysis on trainee messages to obtain message analysis results. The Communication Interaction Training and Analysis Server 20 uses the speech-to-text technology, voice recognition technology, natural language processing technology, facial expression analysis technology and emotion analysis technology described above to perform multimodal data facial expression, emotion and message feature analysis on trainee messages or trainee voice and trainee images to obtain feature analysis results and type analysis results.

[0036] Next, the communication interaction training and analysis server 20 performs interaction pattern analysis on the feature analysis results and type analysis results to obtain the comprehensive communication interaction score and comprehensive communication interaction features. It is worth noting that the communication interaction training and analysis server 20 uses an empathy analysis model to perform interaction pattern analysis on the feature analysis results and type analysis results to obtain the comprehensive communication interaction score and comprehensive communication interaction features. The empathy analysis model uses natural language processing technology, deep neural networks (ANNs), long short-term memory networks (LSTM), and Transformer models, etc., as examples only, and does not limit the application scope of the present invention. This enables the empathy analysis model to perform functions such as emotion analysis, semantic understanding, emotion simulation, and understanding empathic responses, etc., to analyze the interaction patterns of the training and then obtain the comprehensive communication interaction score and comprehensive communication interaction features based on the interaction pattern analysis.

[0037] The Communication Interaction Training and Analysis Server 20 uses general artificial intelligence to generate corresponding text, voice, and facial responses for each virtual human character based on a comprehensive communication interaction score and characteristics. Specifically, in a clinic setting, the server generates the following text responses for the nurse based on a communication interaction score of "40" and a communication interaction characteristic of "indifferent": a text modal response "The response is too indifferent, please adjust the response and train again" and a voice modal response "The response is too indifferent, please adjust the response and train again". The nurse's response to the virtual human character data is "I think the doctor's reply is a bit cold," and the corresponding facial expression is "shocked." Based on the overall communication interaction score of "40" and the overall communication interaction feature of "indifference," the communication interaction training and analysis server 20 uses general artificial intelligence to generate corresponding text responses for the nurse: a text modal "I think the doctor's reply is a bit cold," a voice modal "I think the doctor's reply is a bit cold," and the corresponding facial expression "surprised." This is only an example and does not limit the scope of application of the present invention. Then, the communication interaction training and analysis server 20 provides text responses, voice responses, and facial expression responses to the trainee device 10.

[0038] It is particularly noted that, in practical implementation, the invention may be implemented on the basis of partly or completely on hardware, for example, one or more components in the system may be implemented through an integrated circuit chip, System on Chip (SoC), Complex Programmable Logic Device (CPLD), Field Programmable Gate Array (FPGA) and other hardware processors (FPGAs). The non-transient computer-readable storage media of the invention, which uploads computer-readable instructions (or referred to as computer program instructions) for enabling the processor to implement various aspects of the invention, the non-transient computer-readable storage media may be a tangible device that can hold and store instructions used by the command execution equipment. Non-transient computer-readable storage media may be, but are not limited to, electrical storage equipment, magnetic storage equipment, optical storage equipment, electromagnetic storage equipment, semiconductor storage equipment, or any suitable combination of the above. More specific examples of computer-readable storage media (a nonexhaustive list) include: hard drives, random access memory, read-only memory, flash memory, optical discs, floppy disks, and any suitable combination of the above. The non-transient computer-readable storage media used here are not interpreted as instantaneous signals themselves, such as radio waves or other freely propagating electromagnetic waves, electromagnetic waves propagating through waveguides or other transmission media (e.g., optical signals via fiber optic cables), or electrical signals transmitted through wires. In addition, the computer-readable commands described here can be downloaded to various computing / processing equipment from non-transient computer-readable storage media, or to external computer equipment or external storage equipment via a network, such as: an Internet, LAN, WAN, and / or wireless network. The network may include copper transmission cables, fiber optic transmissions, wireless transmissions, routers, firewalls, switches, hubs and / or gateways. The network card or network interface in each of the computing / processing equipment receives computer-readable instructions from the network and forwards these computer-readable instructions for the non-transient computer-readable storage media stored in the respective computing / processing equipment. A computer-readable instruction performing the operations of the present invention may be a combination language instruction, an instruction set architecture instruction, a machine instruction, a machine-related instruction, a microinstruction, a firmware instruction, or a source or object code (Object Code) written in any combination of one or more programming languages, including object-oriented programming languages, such as Common Lisp, Python, C++, Objective-C, Smalltalk, Delphi, Java, Swift, C#, Perl, Ruby and PHP, etc., as well as conventional procedural (Procedural) programming languages, such as C language or similar programming languages.

[0039] Next, the operation method of the present invention will be described below, and please refer to Figures 4A and 4B, which are flowcharts of the intelligent medical communication and interaction training method based on virtual characters of the present invention.

[0040] First, the trainee's device logs into the Communication Interaction Training and Analysis Server using identification information. The server provides a training user interface to the trainee's device and displays it (step 401). Next, the trainee's device selects one of multiple training scenarios from the user interface and provides the corresponding training scenario identification information to the server (step 402). Then, the trainee's device selects at least one virtual character from the user interface and provides the corresponding virtual character identification information to the server (step 403). Next, the server stores multiple virtual scenario data, each containing at least one virtual character data (step 404). Next, the server retrieves the corresponding at least one virtual character data based on the training scenario identification information, excluding the virtual character data corresponding to the trainee's virtual character identification information (step 405). Finally, the trainee's device receives the trainee's message or voice and image corresponding to the selected virtual character and provides it to the communication server. Interactive training and analysis server (step 406); Next, the communication interaction training and analysis server uses speech-to-text technology, voice recognition technology, natural language processing technology, facial expression analysis technology, and emotion analysis technology to perform multimodal data facial expression, emotion, and message feature analysis on trainee messages or trainee voice and trainee images to obtain feature analysis results and type analysis results (step 407); Next, the communication interaction training and analysis server performs interaction pattern analysis on the feature analysis results and type analysis results to obtain a comprehensive communication interaction score and comprehensive communication interaction features (step 408); Next, the communication interaction training and analysis server uses general artificial intelligence to generate corresponding text responses, voice responses, and facial expression responses for at least one virtual human character data based on the comprehensive communication interaction score and comprehensive communication interaction features (step 409); Next, the communication interaction training and analysis server provides text responses, voice responses, and facial expression responses to the trainee device (step 410); And, the trainee device presents and plays the text responses, voice responses, and facial expression responses corresponding to the unselected virtual character (step 411).

[0041] In summary, the trainee device logs into the communication interaction training and analysis server using identification data. Through the training user interface, the trainee selects a training scenario and virtual role. The trainee device provides trainee messages, voice recordings, and images to the communication interaction training and analysis server. The server uses speech-to-text technology, voice recognition technology, natural language processing technology, facial expression analysis technology, and emotion analysis technology to perform multimodal data feature analysis on the trainee's messages, voice recordings, and images, obtaining feature analysis results and type analysis results. Then, interaction pattern analysis is performed on the feature analysis results and type analysis results to obtain a comprehensive communication interaction score and comprehensive communication interaction features. Based on the comprehensive communication interaction score and comprehensive communication interaction features, general artificial intelligence is used to generate corresponding text responses, voice responses, and facial expressions for the virtual role data, and these responses are fed back to the trainee device for presentation and playback, thereby enabling the trainee to conduct medical communication interaction training in the selected training scenario.

[0042] This technology can solve the problem of the lack of auxiliary medical communication and interaction training in the current teaching methods for medical staff, which is a problem of previous technologies, thereby improving the technical effectiveness of intelligent medical communication and interaction training based on virtual roles.

[0043] While the embodiments disclosed in this invention are as described above, the content is not intended to directly limit the scope of patent protection for this invention. Anyone skilled in the art to which this invention pertains may make minor modifications in form and detail without departing from the spirit and scope disclosed herein. The scope of patent protection for this invention shall still be determined by the appended claims.

[0044] Although the present invention has been disclosed above with reference to the foregoing embodiments, it is not intended to limit the present invention. Anyone skilled in the art can make some modifications and refinements without departing from the spirit and scope of the present invention. Therefore, the scope of patent protection of the present invention shall be determined by the scope of the patent application attached to this specification.

[0045] 10: Trainee Device 11: First non-transitory computer-readable storage medium 12: First Processor 20: Communication Interaction Training and Analysis Server 21 Second non-transitory computer-readable storage medium 22: Second Processor 30: Training the User Interface 31: Selecting a Block for the Training Scene 32: Selected Block for Virtual Character 33: Training Scene Display Block 341: Selected Virtual Character 342: No virtual character selected 351: Text Input Area 352: Text Response Area Step 401: The trainee's device logs into the Communication Interaction Training and Analysis Server using their identification information. The Communication Interaction Training and Analysis Server provides the training user interface to the trainee's device and displays it. Step 402: The trainee's device selects one of the multiple training scenarios in the training user interface and provides the corresponding training scenario recognition information to the communication and interactive training and analysis server. Step 403: The trainee device selects at least one virtual character in the training user interface and provides the corresponding trainee virtual character identification information to the communication and interaction training and analysis server. Step 404: The communication and interaction training and analysis server stores multiple virtual scene data, each containing data on at least one virtual human character. Step 405: The communication and interaction training and analysis server retrieves at least one corresponding virtual human character data based on the training scene recognition information, and excludes the virtual human character data corresponding to the trainee's virtual human character recognition information from the at least one virtual human character data. Step 406: The trainee's device receives trainee messages or trainee voice and trainee image corresponding to the selected virtual character and provides them to the communication and interactive training and analysis server. Step 407: The communication interaction training and analysis server uses speech-to-text technology, voice recognition technology, natural language processing technology, facial expression analysis technology, and emotion analysis technology to perform multimodal data feature analysis on trainees' messages, voices, and images, obtaining feature analysis results and type analysis results. Step 408: The communication interaction training and analysis server performs interaction pattern analysis on the feature analysis results and type analysis results to obtain the comprehensive communication interaction score and comprehensive communication interaction features. Step 409: The communication interaction training and analysis server uses general artificial intelligence to generate corresponding text responses, voice responses, and facial expressions for at least one virtual human character based on the comprehensive communication interaction score and comprehensive communication interaction features. Step 410: The communication interaction training and analysis server provides text responses, voice responses, and facial expressions to the trainee's device. Step 411: The trainee's device presents and plays the text response, voice response, and facial expression response corresponding to the unselected virtual character.

Claims

1. A smart medical communication and interaction training system based on virtual roles, the system comprising: a trainee device, the trainee device further comprising: a first non-transitory computer-readable storage medium storing a plurality of first computer-readable instructions; and a first processor electrically connected to the first non-transitory computer-readable storage medium, executing the first computer-readable instructions to cause the trainee device to perform the following steps: logging into a communication and interaction training and analysis server using identity information to display a training user interface; selecting one of a plurality of training scenarios in the training user interface and providing corresponding training scenario identification information to the communication and interaction training and analysis server; selecting at least one virtual role in the training user interface and providing corresponding trainee virtual role identification information to the communication and interaction training and analysis server; The system receives and provides a trainee message or trainee voice and trainee image corresponding to a selected virtual character to the communication interaction training and analysis server; and receives, presents, and plays a text response, a voice response, and an emoticon response corresponding to an unselected virtual character from the communication interaction training and analysis server; and the communication interaction training and analysis server further includes: a second non-transitory computer-readable storage medium storing a plurality of second computer-readable instructions and storing a plurality of virtual scene data, each virtual scene data including at least one virtual human character data; and a second processor electrically connected to the second non-transitory computer-readable storage medium, executing the second computer-readable instructions to cause the trainee device to perform the following steps: when the trainee device logs in using the identity information, providing the training user interface to the trainee device; obtaining the training scene identification information and the trainee virtual character identification information from the trainee device; querying the corresponding at least one virtual human character data based on the training scene identification information, and the at least one virtual human character data excluding the virtual human character data corresponding to the trainee virtual character identification information. The system acquires trainee messages, trainee voice, and trainee images from the trainee device; uses speech-to-text technology, voice recognition technology, natural language processing technology, facial expression analysis technology, and emotion analysis technology to perform multimodal data feature analysis on the trainee messages, trainee voice, and trainee images to obtain a feature analysis result and a type analysis result; performs interaction pattern analysis on the feature analysis result and the type analysis result to obtain a comprehensive communication interaction score and a comprehensive communication interaction feature; based on the comprehensive communication interaction score and the comprehensive communication interaction feature, uses general artificial intelligence to generate corresponding text responses, voice responses, and facial expression responses for the at least one virtual human character data; and provides the text responses, voice responses, and facial expression responses to the trainee device.

2. The intelligent medical communication and interaction training system based on virtual characters as described in claim 1, wherein the communication and interaction training and analysis server provides a different number of the at least one virtual character based on the selected training scenario.

3. The intelligent medical communication and interaction training system based on virtual characters as described in claim 1, wherein the communication and interaction training and analysis server uses artificial intelligence to map the corresponding facial expressions of the virtual characters to the virtual characters.

4. The intelligent medical communication and interaction training system based on virtual characters as described in claim 1, wherein the communication and interaction training and analysis server is superimposed onto the virtual character after modifying and adjusting the trainee's image or an authorized actual human face image based on the corresponding facial expression response of the virtual character using deepfake technology.

5. A method for intelligent medical communication and interaction training based on virtual roles, comprising the following steps: A trainee device logs into a communication and interaction training and analysis server using identification information; the communication and interaction training and analysis server provides a training user interface to the trainee device and displays it; The trainee device selects one of multiple training scenarios in the training user interface and provides corresponding training scenario identification information to the communication and interaction training and analysis server; The trainee device selects at least one virtual role in the training user interface and provides corresponding trainee virtual role identification information to the communication and interaction training and analysis server; The communication and interaction training and analysis server stores multiple virtual scenario data, each virtual scenario data containing at least one virtual human role data; The communication and interaction training and analysis server queries the corresponding at least one virtual human role data based on the training scenario identification information, and the at least one virtual human role data excludes the virtual human role data corresponding to the trainee virtual role identification information. The trainee device receives a trainee message or trainee voice and trainee image corresponding to a selected virtual character and provides them to the communication interaction training and analysis server. The communication interaction training and analysis server uses speech-to-text technology, voice recognition technology, natural language processing technology, facial expression analysis technology, and emotion analysis technology to perform multimodal data feature analysis on the trainee message or trainee voice and trainee image, obtaining a feature analysis result and a type analysis result. The communication interaction training and analysis server performs interaction pattern analysis on the feature analysis result and the type analysis result to obtain a comprehensive communication interaction score and a comprehensive communication interaction feature. Based on the comprehensive communication interaction score and the comprehensive communication interaction feature, the communication interaction training and analysis server uses general artificial intelligence to generate corresponding text responses, voice responses, and facial expression responses for the at least one virtual human character data. The communication interaction training and analysis server provides the text response, the voice response, and the facial expression response to the trainee device; and the trainee device presents and plays the text response, the voice response, and the facial expression response corresponding to the unselected virtual character.

6. The intelligent medical communication and interaction training method based on virtual characters as described in claim 5, wherein the communication and interaction training and analysis server provides a different number of the at least one virtual character based on the selected training scenario.

7. The intelligent medical communication and interaction training method based on virtual characters as described in claim 5, wherein the communication and interaction training and analysis server uses artificial intelligence to map the corresponding facial expressions of the virtual characters to the virtual characters.

8. The intelligent medical communication and interaction training method based on virtual characters as described in claim 5, wherein the communication and interaction training and analysis server is superimposed onto the virtual character after modifying and adjusting the trainee's image or an authorized actual face image based on the facial expression response corresponding to the virtual character using deepfake technology.