A basic nursing skill training system based on a large language model and deep learning
By using an intelligent nursing skills training system based on large language models and deep learning, combined with image and voice interaction technologies, the system solves the problem of lack of context and feedback in traditional nursing skills training. It enables multi-dimensional nursing operation assessment and personalized guidance, thereby improving trainees' clinical and communication skills.
Patent Information
- Application Number
- CN202511462286.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-14
- Publication Date
- 2026-01-27
- Estimated Expiration
- 2045-10-14
AI Technical Summary
Traditional nursing basic skills training methods are monotonous, lack contextualization and personalization, cannot provide real-time feedback, and are difficult to improve trainees' clinical response capabilities and communication skills.
An intelligent nursing skills training system based on large language models and deep learning is adopted. Through image acquisition and voice interaction devices, combined with multimodal data processing and deep learning algorithms, it realizes multi-dimensional evaluation and real-time feedback of nursing operations and supports diverse interactions between trainees and virtual patients.
It enables automated, multi-dimensional assessment of the entire process of nursing procedure standardization, clinical reasoning, and communication effectiveness, thereby improving trainees' clinical adaptability and communication skills and providing personalized learning guidance.
Smart Images

Figure CN120931451B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of nursing skills training technology, and in particular to a basic nursing skills training system based on large language models and deep learning. Background Technology
[0002] Practical training in basic nursing skills is a core component of nursing education, and its quality is crucial for cultivating nursing professionals with solid practical abilities. Traditional teaching relies mainly on classroom lectures and simulated demonstrations; this singular teaching method cannot fully meet the diverse needs of students and fails to fully unleash their learning initiative. Furthermore, teachers pay insufficient attention to students' communication skills and emotional exchange in clinical nursing during practical training, leading to difficulties in meeting patients' psychological and emotional needs in actual work.
[0003] Currently, most universities use overly simplistic methods for basic skills training, lacking complete training and assessment processes and contextualization. Simple operational drills hinder personalized learning, lack automatic recording and analysis functions, and lack real-time feedback mechanisms to encourage students to summarize and review their work after completion. Furthermore, students working with mannequins can only engage in skills training, failing to enhance their ability to cope with clinical situations. Summary of the Invention
[0004] The purpose of this invention is to overcome the shortcomings of existing technologies and provide a basic nursing skills training system based on large language models and deep learning. This system establishes an intelligent nursing skills training system that supports human-computer dialogue, incorporating multiple basic nursing operation cases. Trainees can verbally communicate with virtual patients of different personalities within the same case, receiving different evaluation dimensions at different stages of the operation. Simultaneously, through deep learning technology, algorithmic models are constructed for each nursing operation process to automatically evaluate the correctness of the nursing techniques and provide real-time feedback. By allowing students to train in a near-realistic environment, trainees can master standardized operations of various basic nursing skills while integrating clinical nursing thinking, thereby improving their clinical nursing abilities.
[0005] The objective of this invention is achieved through the following technical solution: a basic nursing skills training system based on a large language model and deep learning, comprising:
[0006] The user terminal device has a training system application installed, which is used for user login, selection of training modules and display of operation content;
[0007] Image acquisition equipment is used to acquire video data of a simulated operating area in real time during user operation.
[0008] A voice interaction device used to collect user voice input and play back the voice feedback of a virtual patient;
[0009] The processing and control unit is used to run large language models and deep learning algorithms to achieve intelligent processing of visual recognition and voice interaction of operation processes.
[0010] The evaluation and feedback module is used to conduct multi-dimensional assessments of users' operational compliance and communication performance and generate real-time feedback reports.
[0011] During simulation training, users first authenticate themselves through the user authentication module. After successful authentication, they enter the skill module manager and select the training skill they want to train. Then, the simulation scenario generator generates different nursing operation scenarios and virtual patient data with different personality traits based on the selected training skill.
[0012] Next, the user undergoes operation training. The multimodal interaction engine receives and processes various data streams generated during the operation training process. Visual data is input into the image recognition pipeline to analyze the user's operation actions and obtain image recognition results. Audio data is input into the speech processing pipeline to analyze the speech communication content and obtain speech processing results. The image recognition results are input into the stage judgment logic module to determine the correctness of the user's operation and whether all operation stages have been completed, and the operation is scored through the operation evaluation matrix. The speech processing results are directly input into the communication evaluation matrix for communication scoring.
[0013] The operational and communication scores are input into a comprehensive scoring engine for integration, generating and outputting a training evaluation report.
[0014] Preferably, during user training, phase detection is performed, including the following steps:
[0015] Before operation, voice recording analysis and instrument preparation recognition are performed. The system records and analyzes the communication between the user and the virtual patient, and uses image acquisition equipment to identify whether the medical instruments prepared by the user are correct and complete.
[0016] During operation, the system captures user hand movements and verifies the disinfection process. It tracks the user's hand movements and operation trajectory to determine whether the operation is standardized and accurate, and monitors whether the user has performed the correct disinfection procedure before executing the required disinfection steps. At the same time, it performs real-time operation correction.
[0017] After the operation, medical waste disposal and patient feedback analysis are performed to check whether the user can correctly classify and dispose of used medical waste, and to analyze the user's instructions and care for the virtual patient after the operation, as well as whether the user inquired about the patient's feelings.
[0018] Preferably, the system continuously performs phase detection, monitors the user's operation steps, and updates the evaluation of the current operation in real time based on the detected operations; at the same time, the system will determine whether the user has completed all operation steps. If not, it will return to phase detection and continue to monitor and evaluate the next operation step, forming a loop; if all operation steps are completed, it will exit the loop.
[0019] Preferably, when receiving and processing multiple data streams generated during operation training through a multimodal interaction engine, the following steps are included:
[0020] Multimodal data acquisition stage: Visual data, voice data, and time-series data are acquired. The visual data is used to capture the user's operating techniques and the entire operation process; the voice data is used to record the content of nurse-patient communication and humanistic care between the user and the virtual patient; and the time-series data is used to record the execution order and time of the operation steps.
[0021] Data preprocessing and feature extraction stage: Visual data is subjected to image frame extraction and standardization, followed by joint detection and motion trajectory analysis to quantify the degree of completion of the operation; speech data is subjected to speech-to-text and semantic parsing to generate response content and output it in speech and text form; all data streams are synchronized through operation time sequence segmentation and alignment to ensure that the analysis is consistent in time, and finally feature data is obtained.
[0022] In the multimodal data joint reasoning stage, preprocessed feature data is input in parallel into a deep learning-based operational standardization assessment model, a clinical reasoning model, and a communication effectiveness analysis model to obtain multiple analysis results. The operational standardization assessment model is used to evaluate the correctness of techniques, adherence to aseptic principles, and standardization of instrument use; the clinical reasoning model is used to evaluate the rationality of operational procedures, the ability to handle abnormal situations, and the appropriateness of clinical decisions; and the communication effectiveness analysis model is used to evaluate the accuracy of professional terminology use, the degree of humanistic care, and the patient's emotional response ability.
[0023] Real-time evaluation feedback and report generation phase: Based on the analysis results of each model, provide real-time visual feedback, real-time prompts for operational errors, and personalized improvement suggestions, and summarize all real-time feedback data to generate a training evaluation report.
[0024] Preferably, the user terminal device is a tablet computer with a training system application installed.
[0025] Preferably, the image acquisition device includes multiple cameras positioned in the simulated operation area.
[0026] Preferably, the voice interaction device includes a microphone and a speaker.
[0027] Preferably, the training evaluation report is output in PDF or Excel format.
[0028] The beneficial effects of this invention are:
[0029] 1) This invention realizes a complete intelligent teaching closed loop of "collection-processing-analysis-feedback-reporting". Compared with the traditional evaluation methods that only focus on operation results and pure software solutions that only simulate dialogue practice, it can automatically evaluate the three core competencies of operation standardization, clinical thinking and communication effectiveness in a full-process, multi-dimensional and granular manner, thereby providing strong technical support for cultivating high-quality nursing talents.
[0030] 2) The system covers multiple core basic nursing operations and uses common clinical cases as the communication task guide. The traditional mannequin can be upgraded and communicate with trainees in a realistic and natural language.
[0031] 3) The system integrates AI voice perception and interaction technology to intelligently simulate the diverse personalities of patients and simulate corresponding emotional feedback based on the communication content, thereby improving trainees' clinical adaptability.
[0032] 4) The system tracks trainees’ performance in real time, analyzes key communication points, humanistic care, communication skills and other evaluation criteria from multiple dimensions, and judges the standardization of operation in real time through image recognition, and generates AI evaluations in an instant to achieve accurate and personalized guidance. Attached Figure Description
[0033] Figure 1 This is a schematic diagram of the hardware system, software system, and data flow of the present invention;
[0034] Figure 2 This is a flowchart of the phased testing process;
[0035] Figure 3 For software flowcharts;
[0036] Figure 4 This is a flowchart of a multimodal intelligent response process. Detailed Implementation
[0037] The technical solution of the present invention will be clearly and completely described below with reference to the embodiments. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0038] See Figures 1-4 This invention provides a technical solution: a basic nursing skills training system based on large language models and deep learning, comprising:
[0039] The user terminal device has a training system application installed, which is used for user login, selection of training modules and display of operation content;
[0040] Image acquisition equipment is used to acquire video data of a simulated operating area in real time during user operation.
[0041] A voice interaction device used to collect user voice input and play back the voice feedback of a virtual patient;
[0042] The processing and control unit is used to run large language models and deep learning algorithms to achieve intelligent processing of visual recognition and voice interaction of operation processes.
[0043] The evaluation and feedback module is used to conduct multi-dimensional assessments of users' operational compliance and communication performance and generate real-time feedback reports.
[0044] During simulation training, users first authenticate themselves through the user authentication module. After successful authentication, they enter the skill module manager and select the training skill they want to train. Then, the simulation scenario generator generates different nursing operation scenarios and virtual patient data with different personality traits based on the selected training skill.
[0045] Next, the user undergoes operation training. The multimodal interaction engine receives and processes various data streams generated during the operation training process. Visual data is input into the image recognition pipeline to analyze the user's operation actions and obtain image recognition results. Audio data is input into the speech processing pipeline to analyze the speech communication content and obtain speech processing results. The image recognition results are input into the stage judgment logic module to determine the correctness of the user's operation and whether all operation stages have been completed, and the operation is scored through the operation evaluation matrix. The speech processing results are directly input into the communication evaluation matrix for communication scoring.
[0046] The operational and communication scores are input into a comprehensive scoring engine for integration, generating and outputting a training evaluation report.
[0047] In this embodiment, as Figure 1 As shown, the entire system includes a hardware system, a software system, and a data flow.
[0048] First, the hardware system, as the physical foundation, includes: User terminal device: a tablet computer used to display content to the user; Image acquisition device: several cameras used to capture images and transmit them to the video processing system (scene and micro-manipulation image acquisition and analysis); Voice interaction device: audio equipment that serves as both input and output, used for both voice interaction and audio output. The data collected by these hardware devices is then processed by the software system.
[0049] The core process of the software system includes: the user first authenticates their identity through the user authentication module;
[0050] After successful verification, enter the skill module manager and select the specific training skill;
[0051] The simulation scenario generator generates different nursing operation scenarios, as well as virtual patient data with different personality traits;
[0052] The multimodal interaction engine, as the core processing unit, simultaneously receives and processes multiple data streams. It splits the data into two processing pipelines: visual data is processed through an image recognition pipeline to analyze user actions; audio data is processed through a speech processing pipeline to analyze the content of the spoken communication. The results from the image recognition pipeline are sent to the stage judgment logic module to determine the stage and correctness of the user's operation, and then passed to the operation evaluation matrix for scoring. The results from the speech processing pipeline are directly sent to the communication evaluation matrix to evaluate communication ability.
[0053] Finally, the data flow module summarizes all evaluation results: the comprehensive scoring engine integrates evaluations of both operational and communication aspects; the evaluation report generation module generates detailed reports based on the scoring results; and the final report is output in PDF or Excel format for easy saving and review. The entire system achieves a multi-dimensional comprehensive evaluation of medical operational skills through a complete workflow of hardware acquisition, software processing, and data output.
[0054] Software process such as Figure 3 As shown, the specific steps from startup to final report generation are as follows:
[0055] 1. System Startup: The process begins with the user starting the system.
[0056] 2. Hardware preparation: The first step after the system starts is to turn on external devices such as speakers and cameras to prepare for training.
[0057] 3. System Initialization: After the equipment is ready, the next step is to initialize the training system itself.
[0058] 4. User Authentication: After system initialization, users (students) need to authenticate their identity to log in to the system.
[0059] 5. Select training content: After successful certification, trainees need to select the nursing skills module they wish to undertake.
[0060] 6. Configure training scenarios: After selecting a training case, trainees also need to select a patient personality template, which will determine the virtual patient's reaction and behavior patterns.
[0061] 7. Start Training: After completing the above settings, you will officially enter the operation training phase. The training phase consists of a loop-based operation sub-process (Operation Training):
[0062] Phase detection: The system will continuously perform phase detection to monitor the trainees' operation steps.
[0063] Real-time evaluation: The system will update the evaluation of the current operation in real time based on the detected operation.
[0064] Completion check: The system will determine whether the trainee has completed all steps; if not, the process returns to the stage check to continue monitoring and evaluating the next operation, forming a large loop; if all are completed, the loop will be exited.
[0065] 8. Report Generation: Once all steps are completed, the system will generate a detailed evaluation report.
[0066] 9. Output: Finally, the system will output this evaluation report containing multiple dimensions, and the process will end.
[0067] After identity verification, trainees select specific nursing scenarios and patient types for practice. The system tracks their actions step-by-step, providing real-time feedback, forming a "monitor-assessment-loop" process until all steps are completed. Ultimately, the system generates a comprehensive assessment report that fully reflects the trainee's performance.
[0068] In some embodiments, when a user performs operational training, phase detection is performed, including the following steps:
[0069] Before operation, voice recording analysis and instrument preparation recognition are performed. The system records and analyzes the communication between the user and the virtual patient, and uses image acquisition equipment to identify whether the medical instruments prepared by the user are correct and complete.
[0070] During operation, the system captures user hand movements and verifies the disinfection process. It tracks the user's hand movements and operation trajectory to determine whether the operation is standardized and accurate, and monitors whether the user has performed the correct disinfection procedure before executing the required disinfection steps. At the same time, it performs real-time operation correction.
[0071] After the operation, medical waste disposal and patient feedback analysis are performed to check whether the user can correctly classify and dispose of used medical waste, and to analyze the user's instructions and care for the virtual patient after the operation, as well as whether the user inquired about the patient's feelings.
[0072] In some embodiments, the system continuously performs phase detection, monitors the user's operation steps, and updates the evaluation of the current operation in real time based on the detected operations. At the same time, the system will determine whether the user has completed all operation steps. If not, it will return to phase detection and continue to monitor and evaluate the next operation step, forming a loop. If all operation steps are completed, it will exit the loop.
[0073] In this embodiment, as Figure 2The diagram illustrates the internal evaluation process after a user (trainee) enters the operational training phase. It highlights how the system monitors and evaluates a complete operation from multiple dimensions and in stages. The core of the process is the "stage assessment" step, which breaks down an operational training session into three clear stages and conducts specific competency assessments for each stage.
[0074] 1. Pre-operation (communication, preparation, and assessment): This stage assesses the preparations made before the operation begins.
[0075] Voice recording analysis: The system records and analyzes the communication between the trainee and the "patient", such as whether the operation process was explained and verbal consent was obtained.
[0076] Equipment preparation identification: The camera identifies whether the medical equipment prepared by the trainee is correct and complete.
[0077] 2. In-operation (real-time operation correction): This stage assesses the core skills during the operation execution process.
[0078] Operation motion capture: The system tracks the trainee's hand movements and operation trajectory to determine whether the operation is standardized and accurate.
[0079] Disinfection process verification: Monitor whether trainees have performed the correct disinfection procedures before executing key steps.
[0080] 3. Post-operative assessment (post-operative communication): This stage assesses the follow-up and communication work after the operation is completed.
[0081] Medical waste disposal: Check whether trainees can correctly classify and dispose of used medical waste.
[0082] Patient feedback analysis: This section analyzes the communication content between the trainee and the "patient" after the training, including the trainee's instructions, care, and whether the trainee inquired about the patient's feelings.
[0083] Voice communication recording analysis is integrated throughout all operational stages. Data from the six tracking channels (voice recording, instrument preparation, motion capture, disinfection process, patient feedback, and medical waste disposal) is analyzed and evaluated in real time by the system according to predefined rules and algorithms. The evaluation results from each channel are integrated into a comprehensive score. Subsequently, the system performs a completion check to determine whether all steps of the current operation have been completed.
[0084] If the process is not completed, it will return to the stage detection phase to continue monitoring and evaluating the next step, forming a loop until all steps are completed. If all steps are completed, the loop ends, the user clicks the "End Training" button, and the system will generate a comprehensive evaluation report. Ultimately, the system outputs this detailed evaluation report, encompassing multiple dimensions including communication, preparation, operation, safety, and humanistic care.
[0085] In some embodiments, receiving and processing multiple data streams generated during operation training via a multimodal interaction engine includes the following steps:
[0086] Multimodal data acquisition stage: Visual data, voice data, and time-series data are acquired. The visual data is used to capture the user's operating techniques and the entire operation process; the voice data is used to record the content of nurse-patient communication and humanistic care between the user and the virtual patient; and the time-series data is used to record the execution order and time of the operation steps.
[0087] Data preprocessing and feature extraction stage: Visual data is subjected to image frame extraction and standardization, followed by joint detection and motion trajectory analysis to quantify the degree of completion of the operation; speech data is subjected to speech-to-text and semantic parsing to generate response content and output it in speech and text form; all data streams are synchronized through operation time sequence segmentation and alignment to ensure that the analysis is consistent in time, and finally feature data is obtained.
[0088] In the multimodal data joint reasoning stage, preprocessed feature data is input in parallel into a deep learning-based operational standardization assessment model, a clinical reasoning model, and a communication effectiveness analysis model to obtain multiple analysis results. The operational standardization assessment model is used to evaluate the correctness of techniques, adherence to aseptic principles, and standardization of instrument use; the clinical reasoning model is used to evaluate the rationality of operational procedures, the ability to handle abnormal situations, and the appropriateness of clinical decisions; and the communication effectiveness analysis model is used to evaluate the accuracy of professional terminology use, the degree of humanistic care, and the patient's emotional response ability.
[0089] Real-time evaluation feedback and report generation phase: Based on the analysis results of each model, provide real-time visual feedback, real-time prompts for operational errors, and personalized improvement suggestions, and summarize all real-time feedback data to generate a training evaluation report.
[0090] In this embodiment, as Figure 4 As shown, the entire process can be divided into the following core stages:
[0091] 1. Multimodal data acquisition
[0092] The process begins with multimodal data acquisition. The system simultaneously collects three types of data through sensors: visual data: used to capture the trainee's operating techniques and the entire operating process; voice data: used to record the nurse-patient communication content and humanistic care between the trainee and the simulated patient; and time-series data: used to accurately record the execution sequence and rhythm of the operating steps.
[0093] 2. Data Preprocessing and Feature Extraction
[0094] The collected raw data is sent to the data preprocessing and feature extraction module for processing: visual data undergoes image frame extraction and standardization, followed by keypoint detection and motion trajectory analysis to quantify the degree of completion of the operation. Voice data undergoes speech-to-text and semantic parsing to understand the communication content, generate response content, and output it simultaneously in both voice and text formats. All data streams are synchronized through operation timing segmentation and alignment to ensure temporal consistency in the analysis.
[0095] 3. Multimodal Data Joint Inference (Core Analysis)
[0096] The preprocessed feature data is fed in parallel into three core deep learning models for in-depth analysis: Operational compliance assessment model: focuses on assessing hard skills of operation, including correctness of technique, adherence to aseptic principles, and standardization of instrument use; Clinical reasoning model: focuses on assessing decision-making and cognitive abilities, including the rationality of operational procedures, ability to handle abnormal situations, and appropriateness of clinical decisions; Communication effectiveness analysis model: focuses on assessing soft skills of interpersonal communication, including accuracy of professional terminology use, degree of humanistic care, and patient emotional responsiveness.
[0097] 4. Multi-dimensional real-time evaluation and feedback
[0098] The analysis results of all models are aggregated in a multi-dimensional real-time evaluation center, and various forms of feedback are generated immediately:
[0099] Real-time visual feedback: The analysis results are presented to trainees in a visual and intuitive way.
[0100] Real-time error alerts: Errors are immediately pointed out during training, allowing trainees to correct them on the spot.
[0101] Personalized improvement suggestions: Provide customized learning suggestions based on the student's specific performance.
[0102] 5. Generate the final evaluation report
[0103] All real-time feedback data is ultimately aggregated to generate a comprehensive formative assessment report. This report is not only a summary of a single operation, but also focuses on process assessment, including: skill mastery analysis, clinical reasoning ability assessment, and comprehensive ability growth trajectory.
[0104] In some embodiments, the user terminal device is a tablet computer with a training system application installed.
[0105] In some embodiments, the image acquisition device includes a plurality of cameras disposed in the simulated operation area.
[0106] In some embodiments, the voice interaction device includes a microphone and a speaker.
[0107] In some embodiments, the training evaluation report is output in PDF or Excel format.
[0108] Taking venous blood collection training as an example, an example is given below:
[0109] The treatment cart and supplies are ready (and will remain in a fixed position in front of the bed throughout the process). The operator holds a tablet on the task page, clicks "Start Communication," and then proceeds to the following stage:
[0110]
[0111]
[0112]
[0113]
[0114] The above description is merely a preferred embodiment of the present invention. It should be understood that the present invention is not limited to the forms disclosed herein and should not be construed as excluding other embodiments. It can be used in various other combinations, modifications, and environments, and can be altered within the scope of the concept described herein through the above teachings or related technologies or knowledge. Modifications and variations made by those skilled in the art that do not depart from the spirit and scope of the present invention should be within the protection scope of the appended claims.
Claims
1. A basic nursing skills training system based on large language models and deep learning, characterized in that: include: The user terminal device has a training system application installed, which is used for user login, selection of training modules and display of operation content; Image acquisition equipment is used to acquire video data of the simulated operation area in real time during user operation; A voice interaction device used to collect user voice input and play back the voice feedback of a virtual patient; The processing and control unit is used to run large language models and deep learning algorithms to achieve intelligent processing of visual recognition and voice interaction of operation processes. The evaluation and feedback module is used to conduct multi-dimensional assessments of users' operational compliance and communication performance and generate real-time feedback reports. During simulation training, users first authenticate themselves through the user authentication module. After successful authentication, they enter the skill module manager and select the training skill they want to train. Then, the simulation scenario generator generates different nursing operation scenarios and virtual patient data with different personality traits based on the selected training skill. Next, the user performs operation training. The multimodal interaction engine receives and processes various data streams generated during the operation training process. Visual data is input into the image recognition pipeline to analyze the user's operation actions and obtain image recognition results. Audio data is input into the speech processing pipeline to analyze the speech communication content and obtain speech processing results. The image recognition results are input into the stage judgment logic module to determine the correctness of the user's operation and whether all operation stages have been completed, and the operation is scored through the operation evaluation matrix. The voice processing results are directly input into the communication evaluation matrix for communication scoring; The operational and communication scores are input into a comprehensive scoring engine, integrated, and a training evaluation report is generated and output. During user training, phase detection is performed, including the following steps: Before operation, voice recording analysis and instrument preparation recognition are performed. The system records and analyzes the communication between the user and the virtual patient, and uses image acquisition equipment to identify whether the medical instruments prepared by the user are correct and complete. During operation, the system captures user hand movements and verifies the disinfection process. It tracks the user's hand movements and operation trajectory to determine whether the operation is standardized and accurate, and monitors whether the user has performed the correct disinfection procedure before executing the required disinfection steps. At the same time, it performs real-time operation correction. After the operation, medical waste disposal and patient feedback analysis are performed to check whether the user can correctly classify and dispose of used medical waste, and to analyze the user's instructions and care for the virtual patient after the operation, as well as whether the user asked about the patient's feelings. The system continuously performs phase checks, monitors the user's operation steps, and updates the evaluation of the current operation in real time based on the detected operations. At the same time, the system will determine whether the user has completed all operation steps. If not, it will return to the phase check and continue to monitor and evaluate the next operation step, forming a loop. If all operation steps are completed, the loop will be broken.
2. The basic nursing skills training system based on large language models and deep learning according to claim 1, characterized in that: When receiving and processing various data streams generated during operation training through a multimodal interaction engine, the following steps are included: Multimodal data acquisition stage: Visual data, voice data, and time-series data are acquired. The visual data is used to capture the user's operating techniques and the entire operation process; the voice data is used to record the content of nurse-patient communication and humanistic care between the user and the virtual patient; and the time-series data is used to record the execution order and time of the operation steps. Data preprocessing and feature extraction stage: Visual data is subjected to image frame extraction and standardization, followed by joint detection and motion trajectory analysis to quantify the degree of completion of the operation; speech data is subjected to speech-to-text and semantic parsing to generate response content and output it in speech and text form; all data streams are synchronized through operation time sequence segmentation and alignment to ensure that the analysis is consistent in time, and finally feature data is obtained. In the multimodal data joint reasoning stage, preprocessed feature data is input in parallel into a deep learning-based operational standardization assessment model, a clinical reasoning model, and a communication effectiveness analysis model to obtain multiple analysis results. The operational standardization assessment model is used to evaluate the correctness of techniques, adherence to aseptic principles, and standardization of instrument use; the clinical reasoning model is used to evaluate the rationality of operational procedures, the ability to handle abnormal situations, and the appropriateness of clinical decisions; and the communication effectiveness analysis model is used to evaluate the accuracy of professional terminology use, the degree of humanistic care, and the patient's emotional response ability. Real-time evaluation feedback and report generation phase: Based on the analysis results of each model, provide real-time visual feedback, real-time prompts for operational errors, and personalized improvement suggestions, and summarize all real-time feedback data to generate a training evaluation report.
3. The basic nursing skills training system based on large language models and deep learning according to claim 1, characterized in that: The user terminal device is a tablet computer with the training system application installed.
4. The basic nursing skills training system based on large language models and deep learning according to claim 1, characterized in that: The image acquisition device includes multiple cameras positioned in the simulated operation area.
5. The basic nursing skills training system based on large language models and deep learning according to claim 1, characterized in that: The aforementioned voice interaction device includes a microphone and a speaker.
6. The basic nursing skills training system based on large language models and deep learning according to any one of claims 1-5, characterized in that: The training evaluation report is output in PDF or Excel format.
Citation Information
Patent Citations
Interactive nursing teaching system for nurses and patients
CN120319084A
PICC puncture virtual simulation training system and method based on XR construction
CN120636224A