Clinical skill practical training evaluation method and device based on multi-modal large model

Through the multimodal large model, the user's action and voice data are collected and analyzed in real time, and error correction videos and voice feedback are generated, which solves the problems of limited teacher resources and inaccurate manual evaluation in traditional clinical skills training, and efficient and objective evaluation and supervision are achieved.

CN120565130APending Publication Date: 2025-08-29INSPUR ENTERPRISE CLOUD TECHNOLOGY (SHANDONG) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510706000.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-29
Publication Date
2025-08-29

AI Technical Summary

Technical Problem

During the traditional clinical skills training, teachers have limited resources and are difficult to conduct comprehensive supervision. Manual evaluation is affected by subjective factors, resulting in insufficient evaluation accuracy and fairness.

Method used

The multimodal large model is used to collect user action and voice data in real time, and combined with preset standards to perform error correction analysis and evaluation, generate error correction video and voice, provide real-time feedback, and generate evaluation scores based on machine learning algorithms.

Benefits of technology

Real-time and comprehensive monitoring of user operations is achieved, the subjective factors of manual evaluation are reduced, the accuracy and supervision efficiency of evaluation are improved, and the learning effect and interactive experience are enhanced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120565130A_ABST
    Figure CN120565130A_ABST
Patent Text Reader

Abstract

The invention provides a clinical skill practical training evaluation method and device based on a multi-modal large model. The method comprises the following steps: acquiring action data and voice data of a user in a clinical skill practical training process in real time; if the type of the clinical skill practical training is a training type, inputting the action data and the voice data into a multi-modal large model in real time, performing error correction analysis based on standard action data and standard voice data, and performing evaluation analysis based on a scoring standard to obtain a real-time error correction video and error correction voice; and obtaining an evaluation score at the end of the clinical skill training process. And if the type of the clinical skill practical training is an examination type, inputting the action data and the voice data into a multi-modal large model, and performing evaluation analysis based on a scoring standard to obtain an evaluation score. According to the scheme, real-time and comprehensive monitoring of the user operation process is achieved through the multi-mode large model technology, the problem that teacher resources are limited is solved, and the influence of subjective factors of manual evaluation in the practical training evaluation process is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of large models, and in particular to a clinical skills training and evaluation method and device based on a multimodal large model. Background Art

[0002] Traditional clinical skills training relies primarily on on-site instructor guidance and evaluation, which presents numerous challenges. For one thing, limited instructor resources make it difficult to comprehensively and meticulously monitor each user's operation, leading to some users' operational errors not being promptly detected and corrected. Furthermore, manual evaluation is significantly influenced by instructors' subjective factors, and evaluation criteria can be inconsistent, impacting both accuracy and fairness. Summary of the Invention

[0003] In view of this, an embodiment of the present invention provides a clinical skills training evaluation method and device based on a multimodal large model to achieve the purpose of improving monitoring efficiency and reducing the influence of subjective factors in evaluation.

[0004] To achieve the above objectives, the embodiments of the present invention provide the following technical solutions:

[0005] A first aspect of an embodiment of the present invention discloses a clinical skills training and evaluation method based on a multimodal large model, the method comprising:

[0006] Determine the type of clinical skills training to be performed according to the start instruction input by the user;

[0007] Real-time collection of user's motion data and voice data during clinical skills training;

[0008] If the type of the clinical skills training is a training type, the action data and the voice data are input into a preset multimodal large model in real time, error correction analysis is performed based on the preset standard action data and the preset standard voice data, and evaluation analysis is performed based on a preset scoring standard, to obtain an error correction video and error correction voice output by the multimodal large model in real time, and an evaluation score output by the multimodal large model at the end of the clinical skills training process; the error correction video is obtained by adding error correction text to the correct action video corresponding to the user's incorrect action;

[0009] Display the error correction video in real time and play the error correction voice in real time;

[0010] If the type of the clinical skills training is an examination type, the action data and the voice data are input into a preset multimodal large model, and an evaluation analysis is performed based on a preset scoring standard to obtain an evaluation score output by the multimodal large model.

[0011] Preferably, the real-time collection of the user's motion data and voice data during clinical skills training includes:

[0012] Using one or more of the following: cameras installed at the training site, pressure sensors installed in the patient model, and motion capture sensors installed at the user's joints, the user's motion data during clinical skills training is collected in real time;

[0013] Use voice acquisition equipment to collect users' voice data in real time during clinical skills training.

[0014] Preferably, the method further comprises:

[0015] If the type of the clinical skills training is a training type, the action data, the voice data, the error correction video, the error correction voice and the evaluation score are packaged to obtain a training folder, and a mark representing the user identity information is added to the training folder, and the marked training folder is stored in a preset database;

[0016] If the type of the clinical skills training is an examination type, the action data, the voice data and the evaluation score are packaged to obtain an examination folder, and a mark representing the user identity information is added to the examination folder, and the marked examination folder is stored in a preset database.

[0017] Preferably, the method further comprises:

[0018] Get the identity information entered by the user;

[0019] extracting the training folder or the test folder from the preset database based on the identity information;

[0020] According to the instructions input by the user, the content in the training folder or the test folder is displayed.

[0021] Preferably, the method further comprises:

[0022] When the evaluation score is obtained, the evaluation score is announced in the form of voice.

[0023] A second aspect of an embodiment of the present invention discloses a clinical skills training and evaluation device based on a multimodal large model, the device comprising:

[0024] A determination unit, configured to determine the type of clinical skills training to be performed according to a start instruction input by a user;

[0025] The acquisition unit is used to collect the user's motion data and voice data in real time during clinical skills training;

[0026] A training analysis unit is configured to, if the type of the clinical skills training is a training type, input the action data and the voice data into a preset multimodal large model in real time, perform error correction analysis based on preset standard action data and preset standard voice data, and perform evaluation analysis based on a preset scoring standard, thereby obtaining an error correction video and error correction voice output by the multimodal large model in real time, and obtaining an evaluation score output by the multimodal large model at the end of the clinical skills training process; the error correction video is obtained by adding error correction text to the correct action video corresponding to the user's incorrect action;

[0027] A real-time display unit, used to display the error correction video in real time and play the error correction voice in real time;

[0028] The examination analysis unit is used to input the action data and the voice data into a preset multimodal large model if the type of the clinical skills training is an examination type, perform evaluation and analysis based on a preset scoring standard, and obtain an evaluation score output by the multimodal large model.

[0029] Preferably, the collection unit is specifically used to:

[0030] Using one or more of the following: cameras installed at the training site, pressure sensors installed in the patient model, and motion capture sensors installed at the user's joints, the user's motion data during clinical skills training is collected in real time;

[0031] Use voice acquisition equipment to collect users' voice data in real time during clinical skills training.

[0032] Preferably, the device further comprises:

[0033] a storage unit for, if the type of the clinical skills training is a training type, packaging the action data, the voice data, the error correction video, the error correction voice, and the evaluation score to obtain a training folder, adding a mark representing user identity information to the training folder, and storing the marked training folder in a preset database;

[0034] If the type of the clinical skills training is an examination type, the action data, the voice data and the evaluation score are packaged to obtain an examination folder, and a mark representing the user identity information is added to the examination folder, and the marked examination folder is stored in a preset database.

[0035] Preferably, the device further comprises:

[0036] A tracing unit, used to obtain identity information input by the user;

[0037] extracting the training folder or the test folder from the preset database based on the identity information;

[0038] According to the instructions input by the user, the content in the training folder or the test folder is displayed.

[0039] Preferably, the device further comprises:

[0040] The score reporting unit is used to report the evaluation score in the form of voice when the evaluation score is obtained.

[0041] Based on the above-mentioned embodiment of the present invention, a clinical skills training evaluation method and device based on a multimodal large model are provided. According to the start instruction input by the user, the type of clinical skills training to be performed is determined; the action data and voice data of the user during the clinical skills training process are collected in real time; if the type of the clinical skills training is a training type, the action data and the voice data are input into the preset multimodal large model in real time, error correction analysis is performed based on the preset standard action data and the preset standard voice data, and evaluation analysis is performed based on the preset scoring criteria to obtain the error correction video and error correction voice output by the multimodal large model in real time, and the evaluation score output by the multimodal large model is obtained at the end of the clinical skills training process; the error correction video is obtained by adding error correction text to the correct action video corresponding to the user's incorrect action; the error correction video is displayed in real time, and the error correction voice is played in real time; if the type of the clinical skills training is an examination type, the action data and the voice data are input into the preset multimodal large model, and evaluation analysis is performed based on the preset scoring criteria to obtain the evaluation score output by the multimodal large model. In this solution, multimodal large model technology is used to achieve real-time and comprehensive monitoring of the user operation process, which not only overcomes the problem of limited teacher resources, but also reduces the influence of subjective factors of manual evaluation during the practical training evaluation process. BRIEF DESCRIPTION OF THE DRAWINGS

[0042] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are merely embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on the provided drawings without paying any creative work.

[0043] Figure 1 A flowchart of a clinical skills training and evaluation method based on a multimodal large model disclosed in an embodiment of the present invention;

[0044] Figure 2 This is a structural diagram of a clinical skills training and evaluation device based on a multimodal large model disclosed in an embodiment of the present invention. DETAILED DESCRIPTION

[0045] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.

[0046] In this application, the terms "comprises," "comprising," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, article, or apparatus that includes a list of elements includes not only those elements but also other elements not explicitly listed, or elements inherent to such process, method, article, or apparatus. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not preclude the presence of additional identical elements in the process, method, article, or apparatus that includes the element.

[0047] As can be seen from the background technology, traditional clinical skills training relies primarily on on-site guidance and evaluation by teachers, which presents numerous problems. On the one hand, limited teacher resources make it difficult to comprehensively and meticulously monitor each user's operation process, resulting in some users' erroneous operations not being discovered and corrected in a timely manner. On the other hand, manual evaluation is significantly influenced by the teacher's subjective factors, and evaluation criteria may be inconsistent, affecting the accuracy and fairness of the evaluation.

[0048] Therefore, the embodiment of the present invention discloses a clinical skills training evaluation method and device based on a multimodal large model. In this solution, multimodal large model technology is used to achieve real-time and comprehensive monitoring of the user's operation process, which not only overcomes the problem of limited teacher resources, but also reduces the influence of subjective factors of manual evaluation during the training evaluation process.

[0049] like Figure 1 FIG. 1 is a flow chart of a clinical skills training evaluation method based on a multimodal large model disclosed in an embodiment of the present invention, comprising the following steps:

[0050] Step S101: Determine the type of clinical skills training to be performed according to the start instruction input by the user.

[0051] In step S101 , the start instruction may represent the type of clinical skill training that the user needs to perform, including but not limited to: training and examination.

[0052] It is understandable that if the type of clinical skills training is a training type, it is necessary to point out the user's practical errors in real time and correct them, and give an evaluation score after the training is completed; if the type of clinical skills training is an examination type, it is not necessary to point out the user's practical errors during the examination in real time, and only need to give an evaluation score after the examination is completed.

[0053] Step S102: Real-time collection of user's motion data and voice data during clinical skills training.

[0054] In the specific implementation process of step S102, one or more of a camera set up in the training site, a pressure sensor set up in the patient model, and a motion capture sensor set up at the user's joints are used to collect the user's motion data during the clinical skills training process; and a voice collection device is used to collect the user's voice data during the clinical skills training process.

[0055] In an embodiment of the present invention, various types of sensors, such as high-definition cameras, pressure sensors, and motion capture sensors, are deployed at the training site to collect data on the user's movements, posture, and force during operation from multiple dimensions. High-definition cameras are used to capture the user's overall operating movements and limb motion trajectories; pressure sensors can be installed on the training equipment (i.e., the patient manikin) to monitor the intensity and changes in force during operation; and motion capture sensors can more accurately capture the user's joint motion information. Furthermore, highly sensitive voice acquisition equipment is used to collect user voice information during operation, including operation descriptions and questions.

[0056] Step S103: If the type of clinical skills training is a training type, the motion data and voice data are input into the preset multimodal large model in real time, and error correction analysis is performed based on the preset standard motion data and the preset standard voice data, and evaluation analysis is performed based on the preset scoring criteria to obtain the error correction video and error correction voice output by the multimodal large model in real time, and the evaluation score output by the multimodal large model is obtained at the end of the clinical skills training process.

[0057] Among them, the Multimodal Large Model (MLLM) is a large-scale AI model that integrates information from multiple modalities, such as text, images, speech, and video. Compared to traditional single-modal models, MLLMs can more comprehensively understand and process complex real-world information, demonstrating greater generalization capabilities and application potential.

[0058] During the error correction analysis of the multimodal large model, AI image recognition, motion capture technology, and speech recognition technology are used to deeply process and analyze the collected images (i.e., motion data) and speech data.

[0059] Through image recognition technology, the video captured by the camera is analyzed frame by frame to accurately identify whether the user's operating actions meet the preset standard action data.

[0060] For example, during cardiopulmonary resuscitation training, it can accurately determine whether the user's chest compression position is at the midpoint of the line connecting the sternum and the two nipples, whether the compression depth reaches 5-6cm, and whether the compression frequency is within the range of 100-120 times / minute.

[0061] Using motion capture technology, we can obtain the motion data of each joint of the user's limbs and analyze the standardization and smoothness of the operation movements.

[0062] For voice data, it is converted into text through voice recognition technology, and then natural language processing technology is used to analyze whether the user's expression during the operation is accurate and standardized, and whether it meets the clinical operation requirements (that is, the preset standard voice data).

[0063] When a user's action is detected to be inconsistent with standard action data, a corresponding correction video (including images and audio) is generated and popped up based on the pre-set error type and the corresponding correct action video library, providing the user with an intuitive operation demonstration. The correction video is created by adding correction text to the correct action video corresponding to the user's incorrect action.

[0064] The error correction text includes but is not limited to the time when the error occurred, the specific error content, and the operating environment data at the time, such as operating steps, surrounding environmental conditions, etc., to facilitate prompting users and subsequent comprehensive analysis and summary.

[0065] In an embodiment of the present invention, the pop-up error correction video is not only accompanied by detailed voice explanations, but also has key operating points and precautions marked on the video screen to facilitate user understanding and learning.

[0066] When it is detected that the user's expression during the operation does not conform to the preset standard voice data, a correction voice is generated. The correction voice is used to prompt the user that there is an error in the expression during the operation and provide the correct voice expression.

[0067] Specifically, the multimodal large model supports users to explain operation content and ask questions through voice. It uses speech recognition technology, combined with deep neural network models and language models, to process the collected voice data, automatically convert the voice into text, and then compare it with standard voice data in text form.

[0068] Preferably, in order to improve the accuracy of voice conversion, the conversion result is also verified and corrected in real time to ensure that the text content can accurately reflect the user's voice information.

[0069] During the evaluation and analysis process of the multimodal large model, the multimodal large model is used to automatically score each preset scoring item based on the collected action data and voice data. Using a machine learning method, a large amount of preset standard action data and preset standard voice data are used for training to learn the relationship between different operation performances and scores. In the actual evaluation process, the algorithm will compare the user's action data and voice data with the training model, comprehensively consider the completion of each scoring item, and automatically generate an evaluation score. Preferably, a weighted calculation can be performed based on the importance of different scoring items, highlighting the weight of key operation links, so that the evaluation results are more scientific and reasonable.

[0070] Among them, we refer to authoritative medical education standards, clinical practice guidelines and opinions of industry experts in advance to determine comprehensive and detailed scoring items:

[0071] Scoring covers multiple aspects, including pre-operation preparation, assessment, various operation steps, and post-operation processing. For example, pre-operation preparation includes personal preparation and item preparation. Personal preparation requires a dignified appearance and professional attire. Item preparation details the specific items that should be on and off the treatment cart. For example, the cart must be equipped with gauze or a portable mask, a blood pressure monitor and stethoscope (or electronic blood pressure monitor), a flashlight (or pupil pen), a watch, a record sheet, quick-disinfecting hand disinfectant, and a curved tray. The cart must be equipped with medical waste bins and recyclable waste bins. If necessary, a defibrillator (or AED), oxygen inhalation accessories, chest compression boards, and foot pedals.

[0072] The process of pre-setting the scoring criteria is as follows:

[0073] Each scoring item is scored based on a clear, quantitative standard score. For example, in a high-quality external chest compression session, 10 points are awarded for the correct compression location, 10 points for a rate between 100-120 times per minute, 10 points for a compression depth of 5-6 cm, 10 points for crossing hands, keeping the upper arm upright, the elbows firmly in place, the base of the lower palm pressed against the chest wall, and the fingers of the lower hand clear of the chest wall, 10 points, 5 points for sufficient chest wall recoil after each compression, and 5 points for interruptions in compressions not exceeding 10 seconds. This quantitative scoring standard ensures the objectivity and accuracy of the evaluation results.

[0074] Preferably, in terms of outputting evaluation results, the evaluation scores are not only displayed in text form, but also informed to users through voice broadcast, so that users can obtain information in a timely manner without having to manually view text content, thereby improving the efficiency of information acquisition.

[0075] Step S104: Display the error correction video in real time and play the error correction voice in real time.

[0076] It is understandable that the training process requires users to discover operational errors in real time, so it is necessary to display error correction videos in real time and play error correction voices in real time.

[0077] It should be noted that the error correction voice is for correcting the user's expressions during the training process, and is not the voice in the error correction video. The voice in the error correction video is the voice explaining the correct actions in the correct action video.

[0078] Step S105: If the type of clinical skills training is an examination type, the motion data and voice data are input into a preset multimodal large model, and evaluation and analysis are performed based on the standard motion data and standard voice data to obtain an evaluation score output by the multimodal large model.

[0079] In step S105, it is necessary to wait until the end of the test to obtain the evaluation score output by the multimodal large model.

[0080] In a specific implementation, the motion data and voice data can be input into the preset multimodal large model in real time, or all the motion data and voice data during the test can be input into the preset multimodal large model together after the test.

[0081] It should be noted that the process of evaluating and analyzing the multimodal large model is the same as that in the above-mentioned embodiment of the present invention, and reference can be made to each other, and no further details will be given here.

[0082] In one embodiment, if the type of clinical skills training is a training type, the motion data, voice data, error correction video, error correction voice and evaluation score are packaged to obtain a training folder, and a mark representing the user identity information is added to the training folder, and the marked training folder is stored in a preset database; if the type of clinical skills training is an examination type, the motion data, voice data and evaluation score are packaged to obtain an examination folder, and a mark representing the user identity information is added to the examination folder, and the marked examination folder is stored in a preset database.

[0083] In an embodiment of the present invention, the video, voice, sensor data, and evaluation scores of the user's clinical skills training process are uniformly stored to establish an operation record database. To facilitate management and query, the stored data is classified and managed according to user identity information, training time, training project, etc. For example, using the student user's student ID as a unique identifier, the relevant data of each training session is stored in the corresponding folder, and each folder is sorted according to the training time, so that users can quickly find the operation records they need.

[0084] In one embodiment, identity information input by a user is obtained; a training folder or an examination folder is extracted from a preset database based on the identity information; and the content of the training folder or the examination folder is displayed according to instructions input by the user.

[0085] In this embodiment of the present invention, after completing training, users can select the training records they want to review through the system interface. The system will display the video and voice recordings of the operation process, along with the corresponding evaluation information, in chronological order. Users can pause and replay the operation process at any time to review their performance at each stage. During the playback process, the system will simultaneously display the corresponding evaluation scores. Users can use the evaluation results to analyze their own operational problems, such as which operation steps scored low and where errors occurred, so as to make targeted improvements.

[0086] The clinical skills training and evaluation method based on a multimodal large model disclosed in the above embodiment of the present invention has the following beneficial effects:

[0087] Improved Supervision Efficiency: Multimodal large-scale model technology enables real-time, comprehensive monitoring of user operations, unconstrained by time and space, overcoming the challenges of limited teacher resources. Regardless of the number of users, the system accurately monitors each user's operations, ensuring effective oversight of each user's operations, promptly identifying errors and providing prompts, significantly improving supervision efficiency.

[0088] Improved evaluation accuracy: Based on scientific scoring criteria and advanced multimodal large-scale algorithms, the evaluation system of this invention reduces the subjective influence of manual evaluation. Through quantitative scoring criteria and machine learning algorithms, user operations are evaluated objectively, impartially, and accurately, making the evaluation results more reflective of the user's actual operation level, providing users with valuable reference feedback, and helping them to improve their clinical skills in a targeted manner.

[0089] Enhanced Learning: The process review function provides users with a platform for self-reflection and learning. By reviewing their procedures and combining them with evaluation results, users can clearly understand their strengths and weaknesses, summarize experiences and lessons, and improve their procedures in a targeted manner. This process of independent learning and reflection can effectively enhance learning outcomes and improve clinical skills and practical application.

[0090] Optimized interactive experience: Voice interaction eliminates the need for manual input, making it more convenient and efficient, meeting the needs of actual operational scenarios. Furthermore, the multimodal output of video and voice conveys information to users in a more intuitive and vivid manner, enhancing their learning enthusiasm and engagement, and creating a better learning environment for them.

[0091] Corresponding to the clinical skills training evaluation method based on a multimodal large model disclosed in the above embodiment of the present invention, Figure 2 As shown, this is a structural diagram of a clinical skill training and evaluation device based on a multimodal large model disclosed in an embodiment of the present invention, including: a determination unit 201, a collection unit 202, a training analysis unit 203, a real-time display unit 204 and an examination analysis unit 205.

[0092] The determination unit 201 is used to determine the type of clinical skill training to be performed according to the start instruction input by the user.

[0093] The collection unit 202 is used to collect the user's action data and voice data in real time during the clinical skills training process.

[0094] In one embodiment, the acquisition unit 202 is specifically configured to:

[0095] Using one or more of the following: cameras installed at the training site, pressure sensors installed in the patient model, and motion capture sensors installed at the user's joints, the user's motion data during clinical skills training is collected in real time;

[0096] Use voice acquisition equipment to collect users' voice data in real time during clinical skills training.

[0097] The training analysis unit 203 is used to input the action data and the voice data into a preset multimodal large model in real time if the type of the clinical skills training is a training type, perform error correction analysis based on the preset standard action data and the preset standard voice data, and perform evaluation analysis based on a preset scoring standard, to obtain the error correction video and error correction voice output by the multimodal large model in real time, and to obtain the evaluation score output by the multimodal large model at the end of the clinical skills training process; the error correction video is obtained by adding error correction text to the correct action video corresponding to the user's incorrect action.

[0098] The real-time display unit 204 is used to display the error correction video in real time and play the error correction voice in real time.

[0099] The examination analysis unit 205 is used to input the action data and the voice data into a preset multimodal large model if the type of the clinical skills training is an examination type, perform evaluation analysis based on a preset scoring standard, and obtain an evaluation score output by the multimodal large model.

[0100] In one embodiment, the apparatus further comprises:

[0101] a storage unit for, if the type of the clinical skills training is a training type, packaging the action data, the voice data, the error correction video, the error correction voice, and the evaluation score to obtain a training folder, adding a mark representing user identity information to the training folder, and storing the marked training folder in a preset database;

[0102] If the type of the clinical skills training is an examination type, the action data, the voice data and the evaluation score are packaged to obtain an examination folder, and a mark representing the user identity information is added to the examination folder, and the marked examination folder is stored in a preset database.

[0103] In one embodiment, the apparatus further comprises:

[0104] A tracing unit, used to obtain identity information input by the user;

[0105] extracting the training folder or the test folder from the preset database based on the identity information;

[0106] According to the instructions input by the user, the content in the training folder or the test folder is displayed.

[0107] In one embodiment, the apparatus further comprises:

[0108] The score reporting unit is used to report the evaluation score in the form of voice when the evaluation score is obtained.

[0109] Based on the above-mentioned embodiment of the present invention, a clinical skills training and evaluation device based on a multimodal large model is disclosed. In this solution, multimodal large model technology is used to achieve real-time and comprehensive monitoring of the user's operation process, which not only overcomes the problem of limited teacher resources, but also reduces the influence of subjective factors in manual evaluation during the training evaluation process.

[0110] An embodiment of the present invention further discloses an electronic device, comprising: a memory and a processor;

[0111] The memory is used to store computer programs; the processor is used to execute computer programs, specifically to implement a clinical skills training and evaluation method based on a multimodal large model disclosed in the above-mentioned embodiment of the present invention.

[0112] An embodiment of the present invention also provides a computer-readable storage medium, which includes a stored computer program, wherein when the computer program is executed by a processor, the device where the computer-readable storage medium is located is controlled to execute a clinical skills training and evaluation method based on a multimodal large model disclosed in the above-mentioned embodiment of the present invention.

[0113] Each embodiment in this specification is described in a progressive manner. The same or similar parts between the embodiments can be referred to each other. Each embodiment focuses on the differences from other embodiments. In particular, for system or system embodiments, since they are basically similar to method embodiments, the description is relatively simple. For relevant parts, refer to the partial description of the method embodiment. The system and system embodiments described above are merely schematic, wherein the units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. A person of ordinary skill in the art can understand and implement it without expending creative work.

[0114] Professionals may further appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the above description has generally described the components and steps of each example according to their functions. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professionals and technicians may use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of the present invention.

[0115] The above description of the disclosed embodiments is intended to enable one skilled in the art to implement or use the present invention. Various modifications to these embodiments will be readily apparent to one skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention is not limited to the embodiments shown herein but is intended to conform to the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A clinical skills training and evaluation method based on a multimodal large model, characterized by: The method comprises: Determine the type of clinical skills training to be performed according to the start instruction input by the user; Real-time collection of user's motion data and voice data during clinical skills training; If the type of the clinical skills training is a training type, the action data and the voice data are input into a preset multimodal large model in real time, error correction analysis is performed based on the preset standard action data and the preset standard voice data, and evaluation analysis is performed based on a preset scoring standard, to obtain an error correction video and error correction voice output by the multimodal large model in real time, and an evaluation score output by the multimodal large model at the end of the clinical skills training process; the error correction video is obtained by adding error correction text to the correct action video corresponding to the user's incorrect action; Display the error correction video in real time and play the error correction voice in real time; If the type of the clinical skills training is an examination type, the action data and the voice data are input into a preset multimodal large model, and an evaluation analysis is performed based on a preset scoring standard to obtain an evaluation score output by the multimodal large model.

2. The method according to claim 1, characterized in that The real-time collection of user's motion data and voice data during clinical skills training includes: Using one or more of the following: cameras installed at the training site, pressure sensors installed in the patient model, and motion capture sensors installed at the user's joints, the user's motion data during clinical skills training is collected in real time; Use voice acquisition equipment to collect users' voice data in real time during clinical skills training.

3. The method according to claim 1, characterized in that The method further comprises: If the type of the clinical skills training is a training type, the action data, the voice data, the error correction video, the error correction voice and the evaluation score are packaged to obtain a training folder, and a mark representing the user identity information is added to the training folder, and the marked training folder is stored in a preset database; If the type of the clinical skills training is an examination type, the action data, the voice data and the evaluation score are packaged to obtain an examination folder, and a mark representing the user identity information is added to the examination folder, and the marked examination folder is stored in a preset database.

4. The method according to claim 3, characterized in that The method further comprises: Get the identity information entered by the user; extracting the training folder or the test folder from the preset database based on the identity information; According to the instructions input by the user, the content in the training folder or the test folder is displayed.

5. The method according to any one of claims 1 to 4, characterized in that: The method further comprises: When the evaluation score is obtained, the evaluation score is announced in the form of voice.

6. A clinical skills training and evaluation device based on a multimodal large model, characterized in that: The device comprises: A determination unit, configured to determine the type of clinical skills training to be performed according to a start instruction input by a user; The acquisition unit is used to collect the user's motion data and voice data in real time during clinical skills training; A training analysis unit is configured to, if the type of the clinical skills training is a training type, input the action data and the voice data into a preset multimodal large model in real time, perform error correction analysis based on preset standard action data and preset standard voice data, and perform evaluation analysis based on a preset scoring standard, thereby obtaining an error correction video and error correction voice output by the multimodal large model in real time, and obtaining an evaluation score output by the multimodal large model at the end of the clinical skills training process; the error correction video is obtained by adding error correction text to the correct action video corresponding to the user's incorrect action; A real-time display unit, used to display the error correction video in real time and play the error correction voice in real time; The examination analysis unit is used to input the action data and the voice data into a preset multimodal large model if the type of the clinical skills training is an examination type, perform evaluation and analysis based on a preset scoring standard, and obtain an evaluation score output by the multimodal large model.

7. The device according to claim 6, characterized in that The acquisition unit is specifically used for: Using one or more of the following: cameras installed at the training site, pressure sensors installed in the patient model, and motion capture sensors installed at the user's joints, the user's motion data during clinical skills training is collected in real time; Use voice acquisition equipment to collect users' voice data in real time during clinical skills training.

8. The device according to claim 6, characterized in that The device further comprises: a storage unit for, if the type of the clinical skills training is a training type, packaging the action data, the voice data, the error correction video, the error correction voice, and the evaluation score to obtain a training folder, adding a mark representing user identity information to the training folder, and storing the marked training folder in a preset database; If the type of the clinical skills training is an examination type, the action data, the voice data and the evaluation score are packaged to obtain an examination folder, and a mark representing the user identity information is added to the examination folder, and the marked examination folder is stored in a preset database.

9. The device according to claim 8, characterized in that The device further comprises: A tracing unit, used to obtain identity information input by the user; extracting the training folder or the test folder from the preset database based on the identity information; According to the instructions input by the user, the content in the training folder or the test folder is displayed.

10. The device according to any one of claims 6 to 9, characterized in that The device further comprises: The score reporting unit is used to report the evaluation score in the form of voice when the evaluation score is obtained.