Clinical skills assessment method, apparatus, storage medium and program product
By acquiring videos of clinical skills operations through video capture equipment and automatically evaluating them using multimodal and single-modal assessment models, the problems of high cost and low efficiency of traditional manual assessment are solved, and efficient, accurate assessment results and readable report generation are achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- BEIJING FRIENDSHIP HOSPITAL CAPITAL MEDICAL UNIV
- Filing Date
- 2026-01-14
- Publication Date
- 2026-06-09
AI Technical Summary
Traditional clinical skills assessments rely on manual observation, resulting in high assessment costs, low efficiency, and low accuracy.
Clinical skills operation videos are acquired using video acquisition equipment, and the results are automatically evaluated using multimodal and single-modal assessment models. The assessment results are then integrated to generate an assessment report.
It reduced assessment costs, shortened the time, improved assessment efficiency and accuracy, and enhanced the readability of assessment reports.
Smart Images

Figure CN122177375A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of data processing, and more particularly to a clinical skills assessment method, device, storage medium, and program product. Background Technology
[0002] Currently, clinical skills procedures (such as intravenous infusion and intramuscular injection) are indispensable technical aspects of the medical field. The complete and accurate execution of clinical skills is a core element in ensuring medical quality and patient safety.
[0003] Traditional methods typically employ manual assessment, where assessors observe the clinical skills performed by the individuals being assessed. This approach is costly, inefficient, and lacks accuracy. Summary of the Invention
[0004] This application provides a clinical skills assessment method, device, storage medium, and program product to solve the problems of high assessment cost, low efficiency, and low accuracy caused by the assessment of clinical skills performed by the assessor on-site in traditional methods, thereby improving assessment efficiency and accuracy.
[0005] Firstly, this application provides a clinical skills assessment method, including: In response to an assessment request, a clinical skills operation video is acquired; the clinical skills operation video is obtained by a video acquisition device capturing clinical skills actions performed by a target user; the clinical skills operation video includes multimodal data; Extract target modality data from the multimodal data; The multimodal data is evaluated using a first evaluation model according to evaluation standard data to obtain a first evaluation result; the evaluation standard data includes multiple standard evaluation items and standard scores corresponding to the multiple standard evaluation items respectively; the first evaluation result includes multiple first evaluation items and first action execution results corresponding to the multiple first evaluation items respectively. The target modal data is evaluated using a second evaluation model according to the evaluation criteria data to obtain a second evaluation result; the second evaluation result includes multiple second evaluation items and the second action execution results corresponding to the multiple second evaluation items respectively; The first evaluation result and the second evaluation result are fused together to obtain the target evaluation result; the target evaluation result includes multiple target evaluation items and the target action execution results corresponding to the multiple target evaluation items respectively; Match the plurality of target evaluation items with the plurality of standard evaluation items; Based on the execution results of the target actions corresponding to the multiple target evaluation items and the standard scores corresponding to the matching standard evaluation items, the target scores corresponding to the multiple target evaluation items are determined. Based on the target scores corresponding to the multiple target evaluation items, a total score is calculated, and based on the total score and the target action execution results corresponding to the multiple target evaluation items, a target evaluation report for the target user is generated.
[0006] Secondly, this application provides a computing device, including a processing component and a storage component; The storage component stores a computer program; the computer program is invoked and executed by the processing component to implement the clinical skills assessment method as described in the first aspect.
[0007] Thirdly, this application provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processing component, implements the clinical skills assessment method as described in the first aspect.
[0008] Fourthly, this application provides a computer program product, including a computer program or instructions that, when executed by a processing component, implement the clinical skills assessment method as described in the first aspect.
[0009] In this embodiment of the application, during the clinical skills assessment process, in response to an assessment request, a video of clinical skills operations performed by a target user, captured by a video acquisition device, can be obtained. This video is multimodal data, from which single target modality data can be extracted. Using a first assessment model according to assessment standard data, the multimodal data can be assessed to obtain a first assessment result. The assessment standard data may include multiple standard assessment items and their corresponding standard scores. The first assessment result may include multiple first assessment items and their corresponding first action execution results. Using a second assessment model according to the assessment standard data, the single target modality data can be assessed to obtain a second assessment result. The second assessment result may include multiple second assessment items and their corresponding second action execution results. The first evaluation result and the second evaluation result are merged to obtain the target evaluation result. The target evaluation result may include multiple target evaluation items and their corresponding target action execution results. Multiple target evaluation items can be matched with multiple standard evaluation items. Based on the target action execution results corresponding to the multiple target evaluation items and the standard score values corresponding to the matched standard evaluation items, the target score values corresponding to the multiple target evaluation items are determined. Based on the target score values corresponding to the multiple target evaluation items, the total score value can be calculated. Based on the total score value and the target action execution results corresponding to the multiple target evaluation items, the target evaluation report corresponding to the target user can be generated.
[0010] By utilizing a first assessment model to automatically evaluate clinical skills operation videos from a multimodal data perspective, obtaining a first assessment result, and by utilizing a second assessment model to automatically evaluate clinical skills operation videos from a single modal data perspective, obtaining a second assessment result, and then fusing the first and second assessment results to obtain the target assessment result, this embodiment of the application significantly reduces the cost of clinical skills assessment, shortens the assessment time, and improves assessment efficiency and accuracy compared to traditional methods that require on-site observation and manual assessment by assessors. Furthermore, by comprehensively evaluating from both multimodal and single modal data perspectives, it further improves assessment accuracy compared to a single assessment method that relies solely on visual information recognition from the video.
[0011] Furthermore, by combining the execution results of the target actions corresponding to each of the multiple target evaluation items, as well as the standard scores corresponding to the matching standard evaluation items, the evaluation scores of each target evaluation item are calculated, and the total score corresponding to the target user is obtained. The target user is then evaluated intuitively and accurately based on the total score. Moreover, the readability of the evaluation report is improved by generating a target evaluation report for the target user based on the total score and the execution results of the target actions corresponding to each of the multiple target evaluation items.
[0012] These or other aspects of this application will become more apparent in the following description of the embodiments. Attached Figure Description
[0013] The accompanying drawings, which are included to provide a further understanding of this application and form part of this application, illustrate exemplary embodiments and are used to explain this application, but do not constitute an undue limitation of this application. In the drawings: Figure 1 A flowchart of one embodiment of a clinical skills assessment method provided in this application is shown; Figure 2 This diagram illustrates the architecture of a system in a practical application. Figure 3 This invention provides a schematic diagram of the structure of one embodiment of a clinical skills assessment device. Figure 4 A schematic diagram of a computing device in a practical application is shown. Detailed Implementation
[0014] To make the objectives, technical solutions, and advantages of this application clearer, the technical solutions of this application will be clearly and completely described below in conjunction with specific embodiments and corresponding drawings. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of them. Based on the embodiments in this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0015] It should be noted that, in the cases involving user information in the embodiments of this application, the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in the embodiments of this application are all information and data authorized by the user or fully authorized by all parties. Furthermore, the collection, use, and processing of related data must comply with the relevant laws, regulations, and standards of the relevant countries and regions, and corresponding operation entry points are provided for users to choose to authorize or refuse. In addition, the various models involved in this application (including but not limited to language models or large models) comply with relevant laws and standards.
[0016] Additionally, it should be noted that when user interaction operations or triggering operations are involved in the embodiments of this application, these operations include, but are not limited to, various interaction methods such as touch operations, gesture operations, voice operations, head movement operations, and eye movement operations. Touch operations include, but are not limited to, click operations, double-click operations, long-press operations, swipe operations, pinch operations, or mouse hover operations. Swipe operations include, but are not limited to, straight-line swipes and curved-line swipes.
[0017] As described in the background section, in order to solve the technical problem, the inventors have proposed the technical solution of this application, comprising: in response to an assessment request, acquiring a clinical skills operation video; the clinical skills operation video is acquired by a video acquisition device capturing clinical skills actions performed by a target user; the clinical skills operation video includes multimodal data; extracting target modality data from the multimodal data; evaluating the multimodal data using a first assessment model according to assessment standard data to obtain a first assessment result; the assessment standard data includes multiple standard assessment items and standard scores corresponding to the multiple standard assessment items respectively, the first assessment result includes multiple first assessment items and first action execution results corresponding to the multiple first assessment items respectively; evaluating the target modality data using a second assessment model according to the assessment standard data to obtain a second assessment result; The second evaluation result includes multiple second evaluation items and the execution results of the second actions corresponding to each of the multiple second evaluation items; the first evaluation result and the second evaluation result are fused to obtain a target evaluation result; the target evaluation result includes multiple target evaluation items and the execution results of the target actions corresponding to each of the multiple target evaluation items; the multiple target evaluation items are matched with the multiple standard evaluation items; based on the execution results of the target actions corresponding to each of the multiple target evaluation items and the standard scores corresponding to the matched standard evaluation items, the target scores corresponding to each of the multiple target evaluation items are determined; based on the target scores corresponding to each of the multiple target evaluation items, a total score is calculated; and based on the total score and the execution results of the target actions corresponding to each of the multiple target evaluation items, a target evaluation report for the target user is generated.
[0018] By utilizing a first assessment model to automatically evaluate clinical skills operation videos from a multimodal data perspective, obtaining a first assessment result, and by utilizing a second assessment model to automatically evaluate clinical skills operation videos from a single modal data perspective, obtaining a second assessment result, and then fusing the first and second assessment results to obtain the target assessment result, this embodiment of the application significantly reduces the cost of clinical skills assessment, shortens the assessment time, and improves assessment efficiency and accuracy compared to traditional methods that require manual assessment by assessors. Furthermore, by comprehensively evaluating from both multimodal and single modal data perspectives, it further improves assessment accuracy compared to a single assessment method that relies solely on visual information recognition from the video.
[0019] Furthermore, by combining the execution results of the target actions corresponding to each of the multiple target evaluation items, as well as the standard scores corresponding to the matching standard evaluation items, the evaluation scores of each target evaluation item are calculated, and the total score corresponding to the target user is obtained. The target user is then evaluated intuitively and accurately based on the total score. Moreover, the readability of the evaluation report is improved by generating a target evaluation report for the target user based on the total score and the execution results of the target actions corresponding to each of the multiple target evaluation items.
[0020] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0021] Figure 1 This is a flowchart of one embodiment of a clinical skills assessment method provided in this application. The technical solution of this embodiment can be applied to the server.
[0022] In a practical application, the system architecture to which the technical solution of this application applies may include a server and a client. The client and server can establish a connection via a network, which may include various connection types, such as wired, wireless communication links, or fiber optic cables. The client can interact with the server via the network to send evaluation requests or receive evaluation results or reports from the server.
[0023] The client can be geared towards evaluators, allowing them to trigger evaluation actions or view evaluation reports, etc.
[0024] The client can be a browser, an app (application), a web application such as an H5 (HyperText Markup Language 5) application, a mini-program (also known as a lightweight application), or a cloud application. The client can be deployed on electronic devices and depends on the device to run or on certain apps within the device. Electronic devices can have displays and support information browsing, such as personal mobile terminals like smartphones, tablets, personal computers, desktop computers, smart speakers, smartwatches, etc.
[0025] The server side can include servers that provide various services, such as servers that provide instant messaging, or servers that provide a first evaluation model and a second evaluation model.
[0026] It should be noted that the server can be implemented as a distributed server cluster consisting of multiple servers, or as a single server. The server can also be a server in a distributed system, or a server combined with blockchain. The server can also be a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery networks (CDN), and big data and artificial intelligence platforms, or an intelligent cloud computing server or intelligent cloud host with artificial intelligence technology.
[0027] It should be noted that the technical solutions of this application embodiment are applicable to a virtual network environment. The described users generally refer to "virtual users." Real users can register user accounts on the server through registration to obtain user identities in the network environment. The interaction between the server and the user can be implemented based on the user account. The corresponding data received or sent by the server to the user is also based on the user account. In fact, the user terminal corresponding to the user account receives or sends corresponding data to the generating server.
[0028] Figure 1 The clinical skills assessment method shown may include the following steps: 101: In response to an assessment request, obtain videos of clinical skills demonstrations.
[0029] This embodiment of the application is applicable to clinical skills operation assessment scenarios. The clinical skills operation video can be obtained by a video capture device capturing the clinical skills actions performed by the target user. The clinical skills operation video includes multimodal data, such as audio data and image data.
[0030] Video capture equipment may include cameras, video cameras, etc. Target users can refer to the personnel being assessed, such as medical staff participating in clinical skills assessments. Clinical skills actions can refer to the technical actions performed by medical staff, which may include physical actions such as connecting IV sets, applying tourniquets, and disinfecting puncture sites, as well as verbal communication actions such as informing about precautions, inquiring about health conditions, and explaining drug instructions.
[0031] Optionally, assessment criteria data can also be obtained. Assessment criteria data can include assessment criteria for multiple clinical skill actions, such as whether the action was performed, whether the method of performance was correct, and whether the order of performance was correct. For example, assessment criteria data could include whether the area of disinfection of the puncture site was correct, whether the method of applying the tourniquet was correct, and whether the puncture site was disinfected after applying the tourniquet.
[0032] Evaluation criteria data can be obtained from general databases or provided by relevant personnel.
[0033] 102: Extract target modal data from multimodal data.
[0034] The target modal data is single-modal data, such as audio data, image data, etc.
[0035] To improve assessment accuracy, this embodiment evaluates verbal communication actions in clinical skills demonstration videos from the perspective of speech information. In a practical application, the target modality data can be audio data. Optionally, to facilitate audio data processing, the above method may further include: Convert audio data into text data.
[0036] This involves using pre-defined video processing tools to extract audio data from clinical skills operation videos and then using a pre-defined audio processing model to convert the audio data into text data. In a practical application, the video processing tool could be MoviePy (a video processing library), which demultiplexes the clinical skills operation videos to separate the video and audio streams, thereby extracting the audio data from the video. The audio processing model could be the SenseVoiceSmall model (a multilingual speech recognition model), which decodes and analyzes the audio data, performing a series of processing operations such as acoustic feature extraction, acoustic model matching, and language model decoding to convert the audio data into text data. The specific process is consistent with traditional implementations and will not be elaborated here.
[0037] In a practical application, the server can establish a communication connection with the audio processing server corresponding to the audio processing model, and send requests to the audio processing server through the interface provided by the audio processing server to request the audio processing server to perform text conversion operations. Based on this, optionally, after extracting the audio data, the above method may further include: The extracted audio data is standardized and compressed. The compressed audio data is then encoded to obtain an audio string.
[0038] At this point, converting audio data into text data can include: A text conversion request is generated based on the audio string and sent to the audio processing server to obtain the text data returned by the audio processing server.
[0039] In a practical application, standardization can be achieved by converting the audio data to a 16kHz sampling rate, 16-bit depth mono PCM (Pulse Code Modulation) format and saving it as a WAV file (WaveForm). Compression can be performed using audio processing tools such as PyDub (an audio processing library) and the Opus encoding format (an open-source audio encoding format) to encode the standardized WAV file into a highly compressed audio format, maintaining excellent speech clarity even at low bitrates such as 32kbps. Encoding can be performed using Base64 encoding tools to encode the compressed audio data into an audio string.
[0040] By standardizing the extracted audio data, the consistency of subsequent audio data processing is ensured, and the accuracy of speech recognition is improved. By compressing the standardized audio data, the volume of audio data is reduced, which can improve transmission efficiency and text conversion efficiency while ensuring the speech recognition effect.
[0041] In a practical application, the text data returned by the audio processing server can be in JSON format. The server can parse it to obtain a text string that completely records all the language information in the original video.
[0042] 103: Using the first evaluation model and the evaluation standard data, evaluate the multimodal data to obtain the first evaluation result. The first evaluation result includes multiple first evaluation items and the execution results of the first actions corresponding to each of the multiple first evaluation items.
[0043] 104: Using the second evaluation model, the target modal data is evaluated according to the evaluation standard data to obtain the second evaluation result. The second evaluation result includes multiple second evaluation items and the execution results of the second actions corresponding to each of the multiple second evaluation items.
[0044] Taking audio data as an example, the first evaluation model can be used to evaluate the clinical skills operation video from the perspective of multimodal data, mainly visual information, and the second evaluation model can be used to evaluate the clinical skills operation video from the perspective of single modal data, mainly speech information, to obtain the corresponding evaluation results.
[0045] The first evaluation model can be a multimodal model (MM) based on artificial intelligence. This application does not limit the number of model parameters supported by the model, aiming to meet actual needs. The first evaluation model can be a deep learning model used to process and generate multimodal data, implemented based on a neural network architecture, and pre-trained on large amounts of data.
[0046] Optionally, the first evaluation model may include an encoder, a decoder, a self-attention layer, and a feed-forward neural network. The encoder primarily converts input data (usually in sequence form) into vector representations, capturing the semantic features of the input data. The decoder transforms the intermediate representations generated by the encoder into output data (usually in sequence form). The self-attention layer is a mechanism that allows the model to focus on other positions in the sequence to better encode information at the current position. The feed-forward neural network can perform nonlinear transformations on the output of the self-attention layer to enhance the model's expressive power. These components work together to enable the model built upon them to perform well in various complex processing tasks, such as natural language processing, computer vision, speech recognition, machine translation, text summarization, and intelligent question answering. In a practical application, the first evaluation model could be the Qwen2.5-VL-32B-Instruct model (Tongyi Qianwen Open Source Visual Language Model).
[0047] Specifically, evaluating clinical skills performance videos using the first assessment model can be achieved based on prompts. A prompt is a form of input used to guide or instruct the model to produce the expected output, indicating what action the model should take or what output to generate when performing a specific task. Prompts are a form of natural language input, letting the model know what it needs to do.
[0048] Optionally, using the first evaluation model to evaluate the multimodal data according to the evaluation criteria data, the first evaluation result may include: A first prompt instruction is generated based on the evaluation standard data and multimodal data, and then input into the first evaluation model to guide the first evaluation model to evaluate the multimodal data according to the evaluation standard data and output the first evaluation result.
[0049] For example, the initial prompt could be: You are a professional clinical skills assessment expert, possessing the ability to recognize and assess operations based on multimodal data, especially visual data. You will carefully observe the given clinical skills operation video, evaluate the video according to the assessment criteria data, and output a first assessment result that conforms to the format constraints.
[0050] In a practical application, the server can establish a communication connection with the first evaluation server corresponding to the first evaluation model, and send requests to the first evaluation server through the interface provided by the first evaluation server, so that the first evaluation server can perform video evaluation operations. Based on this, the above method may optionally further include: The clinical skills demonstration video was encoded into a video string, and the first text prompt string was constructed.
[0051] At this point, the multimodal data is evaluated using the first evaluation model according to the evaluation criteria data, and the first evaluation result may include: A video evaluation request is generated based on the evaluation standard data, the video string, and the first text prompt string. The video evaluation request is then sent to the first evaluation server to obtain the first evaluation result from the first evaluation server.
[0052] The video assessment request can trigger the first assessment server to generate a first prompt instruction based on the assessment standard data, the video string, and the first text prompt string, and input the first prompt instruction into the first assessment model to guide the first assessment model to assess the clinical skills operation video according to the assessment standard data and output the first assessment result.
[0053] In a practical application, a tool such as Base64 encoding can be used to encode clinical skills operation videos into video strings. The first text prompt string can include role settings, task context, core task instructions, and output format constraints. Role settings could be, for example, a professional clinical skills operation assessment expert. Task context could be, for example, providing the model with complete assessment criteria data as the basis for assessment. Core task instructions could be, for example, explicitly instructing the model to carefully observe the given clinical skills operation video and assess each action. Output format constraints could be, for example, providing an output format, such as xx action (completed) xx action (incomplete), prohibiting the addition of irrelevant characters.
[0054] By constructing a first text prompt string, the first evaluation model is accurately guided to complete the evaluation task, thereby improving the accuracy of the first evaluation results. By setting output format constraints, the standardization and parsability of the output results are ensured.
[0055] Optionally, before sending the video to the first assessment server, the clinical skills operation video can be processed by frame extraction according to a preset frame extraction rate to reduce the size of the video. The processed clinical skills operation video can then continue to be executed according to the steps encoded into a video string, thereby improving the transmission efficiency of the video assessment request and thus improving the video assessment efficiency.
[0056] It should be noted that the first prompt instruction and the first text prompt string described herein, as well as other prompt instructions or strings that may be mentioned below, are merely illustrative examples for ease of understanding. This application is not limited to these examples and can be specifically set according to actual circumstances.
[0057] The second evaluation model can be an artificial intelligence-based language model (LM). This application does not limit the number of model parameters supported by the model, aiming to meet practical needs. The second evaluation model can be a deep learning model used to process and generate natural language text, implemented based on a neural network architecture, and pre-trained on large amounts of data. In a practical application, the second evaluation model can be the DeepSeek-R1671B model (DeepSeek open-source large-scale language model).
[0058] Optionally, the target modal data can be evaluated using the second evaluation model according to the evaluation criteria data to obtain the second evaluation result, which may include: A second prompt instruction is generated based on the evaluation standard data and the target modal data, and then input into the second evaluation model to guide the second evaluation model to evaluate the target modal data according to the evaluation standard data and output the second evaluation result.
[0059] Taking the target modal data as audio data, and having already converted the audio data into text data as an example, a second prompt instruction can be generated based on the evaluation standard data and the text data. This second prompt instruction is then input into the second evaluation model to guide the second evaluation model to evaluate the text data according to the evaluation standard data and output the second evaluation result.
[0060] For example, the second prompt could be: You are a professional clinical skills assessment expert with text recognition and assessment capabilities. You will analyze and recognize the given text data in detail, evaluate the text according to the assessment criteria, and output the second assessment result.
[0061] In a practical application, the server can establish a communication connection with the second evaluation server corresponding to the second evaluation model, and send requests to it through the interface provided by the second evaluation server so that the second evaluation server can perform text evaluation operations. Based on this, the above method may optionally further include: Construct a second text prompt string.
[0062] At this point, based on the evaluation criteria data, the text data is evaluated using the second evaluation model, and the resulting second evaluation result may include: A text evaluation request is generated based on the evaluation standard data, text data, and a second text prompt string. The text evaluation request is then sent to the second evaluation server to obtain the second evaluation result from the second evaluation server.
[0063] The second evaluation request can trigger the second evaluation server to generate a second prompt instruction based on the evaluation standard data, text data, and second text prompt string, and input the second prompt instruction into the second evaluation model to guide the second evaluation model to evaluate the text data according to the evaluation standard data and output the second evaluation result.
[0064] In a practical application, the second text prompt string can include role setting, task context, core task instructions, and output format constraints. Role setting could be, for example, a professional clinical skills assessment expert. Task context could be, for example, providing the model with both assessment criteria data and text data, requiring the model to evaluate the latter based on the former. Core task instructions could be, for example, instructing the model to analyze and identify the given text data in detail, and to assess whether the text data can support the completion of each assessment item. Output format constraints could be, for example, providing an output format, such as xx action (completed) xx action (incomplete), prohibiting the addition of irrelevant characters.
[0065] 105: The first evaluation result and the second evaluation result are merged to obtain the target evaluation result.
[0066] In a practical application, the first evaluation result may include multiple first evaluation items and the execution results of the first actions corresponding to each of the multiple first evaluation items. The execution results of the first actions may include whether the action has been completed or not. The second evaluation result may include multiple second evaluation items and the execution results of the second actions corresponding to each of the multiple second evaluation items. The execution results of the second actions may include whether the action has been completed or not.
[0067] Based on this, the first evaluation result and the second evaluation result are merged to obtain the target evaluation result, which may include: The union of multiple first evaluation items and multiple second evaluation items is taken as multiple target evaluation items; In the case where any target evaluation item corresponds to the execution result of the first action or the execution result of the second action, the execution result of the first action or the execution result of the second action shall be taken as the execution result of the target action corresponding to the target evaluation item. In the case of any target evaluation item corresponding to the execution result of the first action and the execution result of the second action, if the execution result of the first action and / or the execution result of the second action are both considered to indicate that the action has been completed, the target action execution result of the target evaluation item is determined to be that the action has been completed; otherwise, the target action execution result of the target evaluation item is determined to be that the action has not been completed.
[0068] The assessment items may include applying a tourniquet, disinfecting the puncture site, informing the recipient of precautions, and explaining medication instructions. The first and second assessment items may be the same or different. The result of the first action may include whether the action was completed or not, and the result of the second action may also include whether the action was completed or not.
[0069] Specifically, multiple first assessment items and multiple second assessment items can be matched. For successfully matched first and second assessment items, either the first or second assessment item is retained. For unmatched first and second assessment items, both the first and second assessment items are retained. This determines multiple target assessment items, which are the union of multiple first and multiple second assessment items. For example, the first assessment result includes three first assessment items: applying a tourniquet, disinfecting the puncture site, and informing the recipient of precautions. The second assessment result includes two second assessment items: informing the recipient of precautions and explaining the medication instructions. Through matching, the first assessment item "informing the recipient of precautions" and the second assessment item "informing the recipient of precautions" are successfully matched, and either one can be retained. The remaining first and second assessment items are all unmatched and are all retained. The final target assessment items obtained by fusion include four items: applying a tourniquet, disinfecting the puncture site, informing the recipient of precautions, and explaining the medication instructions.
[0070] For any target evaluation item, if it corresponds only to the first action execution result or the second action execution result, the first action execution result or the second action execution result shall be taken as the target action execution result; if it corresponds to both the first action execution result and the second action execution result, if either action execution result indicates that the action execution has been completed, the target action execution result of the target evaluation item can be determined to be that the action execution has been completed; otherwise, it is determined to be that the action execution has not been completed.
[0071] Optionally, it can also be set that if both the first action execution result and the second action execution result are that the action execution has been completed, the target action execution result of the target evaluation item is determined to be that the action execution has been completed. Different action execution result determination rules can also be set for different evaluation items. These rules can be set according to actual needs, and this application embodiment does not limit them.
[0072] To improve the readability of the evaluation results, an evaluation score can also be calculated. The evaluation criteria data can include multiple standard evaluation items and their corresponding standard scores; the target evaluation results can include multiple target evaluation items and their corresponding target action execution results.
[0073] 106: Match multiple target evaluation items with multiple standard evaluation items.
[0074] 107: Based on the execution results of the target actions corresponding to multiple target evaluation items and the standard score values corresponding to the matching standard evaluation items, determine the target score values corresponding to the multiple target evaluation items.
[0075] 108: Calculate the total score based on the target scores corresponding to multiple target evaluation items, and generate a target evaluation report for the target user based on the total score and the target action execution results corresponding to the multiple target evaluation items.
[0076] In this embodiment, since the target evaluation items are generated by the first evaluation model and the second evaluation model, there may be inconsistencies with the standard evaluation items. Therefore, the two can be matched to determine the standard evaluation items that match each target evaluation item.
[0077] In a practical application, each standard assessment item corresponds to a standard score. When a standard assessment item is completed, the corresponding standard score is obtained. For example, 3 points are awarded for disinfecting the puncture site, and 2 points are awarded for applying a tourniquet. The result of the target action execution includes whether the action has been completed or not.
[0078] Based on this, and considering the execution results of the target actions corresponding to multiple target evaluation items, as well as the standard scores corresponding to the matching standard evaluation items, the target scores corresponding to the multiple target evaluation items can be determined as follows: For any target evaluation item, if the target action of the target evaluation item is completed and there is a standard evaluation item that matches the target evaluation item, the score of the target evaluation item is determined to be the standard score value corresponding to the matching standard evaluation item. Otherwise, the score for the target evaluation item is set to 0.
[0079] Calculating the total score based on the target scores corresponding to multiple evaluation items can include: The scores for each of the multiple target evaluation items are summed to obtain the total score.
[0080] For example, for the objective assessment item of disinfecting the puncture site, a score of 3 can be obtained when the objective action of this objective assessment item is completed and a matching standard assessment item exists. As another example, for the objective assessment item of applying a tourniquet, a score of 0 can be obtained when the objective action of this objective assessment item is incomplete (e.g., incorrect application of the tourniquet, incorrect placement, incorrect timing), indicating that the target user has not successfully performed this skill operation. Alternatively, a score of 0 can be obtained when the objective action of this objective assessment item is completed but no matching standard assessment item exists.
[0081] In this embodiment, during the clinical skills assessment process, in response to an assessment request, a video of the clinical skills operation performed by the target user, captured by a video acquisition device, can be obtained. This video is multimodal data, from which single target modality data can be extracted. Using a first assessment model according to assessment standard data, the multimodal data can be assessed to obtain a first assessment result. The assessment standard data may include multiple standard assessment items and their corresponding standard scores. The first assessment result may include multiple first assessment items and their corresponding first action execution results. Using a second assessment model according to the assessment standard data, the single target modality data can be assessed to obtain a second assessment result. The second assessment result may include multiple second assessment items and their corresponding second action execution results. The first evaluation result and the second evaluation result are merged to obtain the target evaluation result. The target evaluation result may include multiple target evaluation items and their corresponding target action execution results. Multiple target evaluation items can be matched with multiple standard evaluation items. Based on the target action execution results corresponding to the multiple target evaluation items and the standard score values corresponding to the matched standard evaluation items, the target score values corresponding to the multiple target evaluation items are determined. Based on the target score values corresponding to the multiple target evaluation items, the total score value can be calculated. Based on the total score value and the target action execution results corresponding to the multiple target evaluation items, the target evaluation report corresponding to the target user can be generated.
[0082] By utilizing a first evaluation model to automatically evaluate clinical skills operation videos from multimodal data perspectives, such as visual information, a first evaluation result is obtained. Conversely, by utilizing a second evaluation model to automatically evaluate clinical skills operation videos from single-modal data perspectives, such as speech information, a second evaluation result is obtained. The first and second evaluation results are then fused to obtain the target evaluation result. Compared to traditional methods requiring manual evaluation by assessors, the solution in this application significantly reduces the cost of clinical skills assessment, shortens assessment time, and improves assessment efficiency and accuracy. Furthermore, by comprehensively evaluating from both multimodal and single-modal data perspectives, such as visual and speech information, the accuracy of the evaluation is further improved compared to a single evaluation method that relies solely on visual information recognition from the video.
[0083] Furthermore, by combining the execution results of the target actions corresponding to each of the multiple target evaluation items, as well as the standard scores corresponding to the matching standard evaluation items, the evaluation scores of each target evaluation item are calculated, and the total score corresponding to the target user is obtained. The target user is then evaluated intuitively and accurately based on the total score. Moreover, the readability of the evaluation report is improved by generating a target evaluation report for the target user based on the total score and the execution results of the target actions corresponding to each of the multiple target evaluation items.
[0084] To improve the accuracy of assessment score calculation, in a practical application, a standard assessment item can include multiple assessment points. Each assessment point corresponds to a standard score value. When an assessment point is completed, the corresponding standard score value is obtained; when all assessment points are completed, the total standard score value is obtained. For example, the standard assessment item of disinfecting the puncture site includes three assessment points: correct disinfection range, no contamination, and no gaps. Each assessment point corresponds to 1 point, and when all three assessment points are completed, 3 points are obtained. In this case, the result of the target action can include action completed, action partially completed, and action not completed.
[0085] Based on this, and considering the execution results of the target actions corresponding to multiple target evaluation items, as well as the standard scores corresponding to the matching standard evaluation items, the target scores corresponding to the multiple target evaluation items can be determined as follows: For any target evaluation item, if the target action execution result of the target evaluation item is that the action execution is completed, and there exists a standard evaluation item that matches the target evaluation item, the score value of the target evaluation item is determined to be the sum of the multiple standard score values corresponding to the matching standard evaluation item; if the target action execution result of the target evaluation item is that the action execution is partially completed, and there exists a standard evaluation item that matches the target evaluation item, the score value of the target evaluation item is determined to be the sum of the multiple standard score values corresponding to the partially completed evaluation points of the matching standard evaluation item; otherwise, the score value of the target evaluation item is determined to be 0.
[0086] For example, for the target assessment item of disinfecting the puncture site, 3 points can be obtained when the target action of the target assessment item is completed and there is a matching standard assessment item; 2 points can be obtained when the target action of the target assessment item is partially completed, such as the disinfection area is correct and there is no contamination, but there are gaps, and there is a matching standard assessment item; otherwise, 0 points can be obtained.
[0087] By subdividing the standard assessment items into multiple assessment points, the system can accurately assess the details of clinical skills operations, thereby improving the accuracy of assessment score calculation.
[0088] In this embodiment of the application, there are multiple ways to match the target evaluation items and the standard evaluation items. The process is described below.
[0089] In a practical application, both the standard evaluation item and the target evaluation item can be implemented as strings. Optionally, before matching, the target evaluation item can be preprocessed to eliminate inconsistencies caused by differences in model expression habits or standard definitions. Preprocessing operations can include at least one of the following: unifying full-width / half-width characters, removing guiding text and category text, clearing preset symbols and their corresponding text content, and converting character formats to a specified format. Removing guiding text can include, for example, removing guiding words such as "whether"; removing category text can include, for example, removing category prefixes such as "disinfect:"; clearing preset symbols and their corresponding text content can include, for example, clearing parentheses and their internal content; converting character formats to a specified format can include, for example, removing non-alphanumeric characters and uniformly converting them to lowercase, etc. After preprocessing, the target evaluation item can be implemented as a "normalized" string for exact matching.
[0090] As an optional matching implementation, for any target evaluation item, a standard evaluation item with the same string as the target evaluation item can be found from multiple standard evaluation items. That is, the normalized string of the target evaluation item is directly compared with the normalized string of the standard evaluation item to see if they are completely identical. The standard evaluation item with the same string is taken as the matching standard evaluation item.
[0091] As an alternative matching method, for any target evaluation item, a search can be conducted among multiple standard evaluation items to find a standard evaluation item that contains the same character fragment as the target evaluation item; that is, a substring matching method can be used for the search. For example, for the target evaluation item "disinfection puncture", the standard evaluation item "disinfection puncture site" can be used as a matching standard evaluation item.
[0092] As another alternative matching approach, for any target evaluation item, the feature similarity between the target evaluation item and multiple standard evaluation items can be calculated based on the feature data of the target evaluation item. Then, based on the feature similarity, standard evaluation items that meet the similarity requirements with the target evaluation item can be determined from among the multiple standard evaluation items. For example, similarity algorithms such as edit distance can be used to calculate the similarity, and standard evaluation items that meet similarity requirements, such as a similarity greater than 80%, can be selected.
[0093] As another alternative matching implementation, for any target evaluation item, a standard evaluation item with the same string as the target evaluation item can be found from multiple standard evaluation items; if it does not exist, a standard evaluation item containing the same character segment as the target evaluation item can be found from multiple standard evaluation items; if it does not exist, based on the feature data of the target evaluation item, the feature similarity between the target evaluation item and multiple standard evaluation items can be calculated, and based on the feature similarity, a standard evaluation item that meets the similarity requirement with the target evaluation item can be determined from multiple standard evaluation items.
[0094] By prioritizing precise matching based on string similarity, employing fuzzy matching based on character fragment similarity when precise matching fails, and using similarity matching when fuzzy matching fails, a hierarchical intelligent matching strategy is adopted to achieve successful matching between standard evaluation items and target evaluation items, thereby improving matching efficiency and accuracy.
[0095] It should be noted that the implementation method for matching multiple first evaluation items with multiple second evaluation items described above is consistent with the implementation described here, and will not be repeated here.
[0096] As described above, the evaluation criteria data can be provided by the user. Therefore, in some embodiments, the above method may further include: Obtain the assessment criteria document; the assessment criteria document should include at least the names of multiple clinical skills and their corresponding standard scores; Describe information according to a preset expression method, and convert multiple operation names into corresponding standard evaluation items; Based on multiple standard evaluation items and their corresponding standard scores, evaluation standard data is constructed and stored.
[0097] In a practical application, the evaluation criteria file can be in JSON format. This JSON file can contain a list (array), where each object can include at least two key fields: the operation name and the standard score. By parsing this JSON file, all objects can be loaded into memory to form a data collection, such as a list of dictionaries.
[0098] For each operation name, information can be described according to a preset expression, such as the preset question keyword "whether", which transforms the operation name into a question with a clear judgment orientation. For example, the operation name "apply tourniquet" can be transformed into the standard evaluation item "whether to apply tourniquet".
[0099] The converted standard evaluation items can be concatenated using preset separators such as commas or pauses to form a single, continuous long string, which serves as the evaluation standard data. Alternatively, the standard evaluation items and their corresponding standard scores can be stored separately.
[0100] By generating a more user-friendly query format for the intelligent evaluation model, the first and second evaluation models are guided to make explicit "yes / no" judgments for each standard evaluation item, i.e., to conduct "completed / incomplete" evaluations. This ensures that the first and second evaluation models can clearly and completely receive all content to be evaluated and analyze and provide feedback according to a unified format. Simultaneously, by preserving the correspondence between operation names and standard score values, and between standard evaluation items and standard score values, technical support is provided for calculating evaluation scores.
[0101] To improve the comprehensiveness of the assessment report, in some embodiments, the above method may further include: Determine the target time for the steps to extract target modality data from multimodal data up to the steps to obtain the total score.
[0102] This system can monitor the execution time of each step and sum the execution times of all steps to obtain the total target time. For example, for the step of extracting audio data from clinical skills operation videos, the start time can be determined by calling a preset video processing tool, and the end time can be determined by calling a preset audio processing model, thereby determining the execution time of this step.
[0103] At this point, based on the total score and the execution results of the target actions corresponding to the multiple target evaluation items, the target evaluation report for the target user can include: Based on the total score, the execution results of the target actions corresponding to multiple target evaluation items, and the target duration, a target evaluation report is generated for the target user.
[0104] The above methods may also include: Convert the target assessment report into target data format and store it.
[0105] By generating a target evaluation report for the target user based on the target duration, the execution results of target actions corresponding to multiple target evaluation items, and the total score, and converting and storing this report in a target data format such as JSON, an objective digital record is provided for retrospective analysis. Compared to the traditional method of real-time evaluation by evaluators, this achieves evaluation traceability and facilitates problem analysis and experience summarization for common issues among evaluators. Furthermore, by setting preset encoding parameters to ensure correct processing of non-ASCII characters such as Chinese characters, and by using indented formatted output, the readability of the target evaluation report is further enhanced.
[0106] It should be noted that some processes described in the above embodiments and accompanying drawings include multiple operations appearing in a specific order. However, it should be clearly understood that these operations may not be executed in the order they appear in this document, or they may be executed in parallel. The operation numbers, such as 101, 102, etc., are merely used to distinguish different operations and do not represent any execution order. Furthermore, these processes may include more or fewer operations, and these operations may be executed sequentially or in parallel. It should also be noted that the descriptions such as "first" and "second" in this document are used to distinguish different messages, devices, modules, etc., and do not represent a sequential order, nor do they limit "first" and "second" to different types.
[0107] To facilitate understanding, the following will be combined with... Figure 2 The system architecture diagram shown illustrates the technical solution of this application.
[0108] like Figure 2 As shown, the system architecture may include a server 201 and a client 202. The client 202 may be directed to the assessor, and in response to user operations, the client 202 may send corresponding assessment requests to the server 201, which will then execute the clinical skills assessment scheme of this embodiment.
[0109] The server 201 may include a data acquisition module 2011, an evaluation standard data processing module 2012, an audio data preprocessing module 2013, a multi-channel evaluation module 2014, an evaluation result and scoring module 2015, and an evaluation report generation module 2016.
[0110] The data acquisition module 2011 can respond to the assessment request sent by the client 202 and obtain the clinical skills operation video and assessment standard document to be assessed. The clinical skills operation video can be obtained by video acquisition equipment capturing the clinical skills actions performed by the target user. The assessment standard document can include at least the operation names of multiple clinical skills and their corresponding standard scores.
[0111] The evaluation standard data processing module 2012 can describe information according to preset expression methods, such as adding preset question keywords, converting multiple operation names into corresponding standard evaluation items, and constructing and storing evaluation standard data based on multiple standard evaluation items and their corresponding standard scores. By generating a more user-friendly query format for the intelligent evaluation model, it guides the first and second evaluation models to make clear "yes / no" judgments for each standard evaluation item, i.e., to perform "completed / incomplete" evaluations. This ensures that the first and second evaluation models can clearly and completely receive all content to be evaluated and analyze and provide feedback according to a unified format. At the same time, by retaining the correspondence between operation names and standard scores, and between standard evaluation items and standard scores, it provides technical support for the calculation of evaluation scores.
[0112] The audio data preprocessing module 2013 can utilize preset video processing tools such as MoviePy to extract audio data from clinical skills operation videos. The extracted audio data undergoes standardization and compression. The compressed audio data is then encoded to obtain an audio string. Based on this audio string, a text conversion request is generated and sent to the audio processing server. The server then uses an audio processing model such as the SenseVoiceSmall model to convert the audio data into text data. Standardization of the extracted audio data ensures consistency in subsequent audio data processing and improves the accuracy of speech recognition. Compression of the standardized audio data reduces its size, improving transmission and text conversion efficiency while maintaining speech recognition performance.
[0113] The multi-path evaluation module 2014 can utilize a first evaluation model, such as the Qwen2.5-VL-32B-Instruct model (Tongyi Qianwen open-source visual language model), to evaluate clinical skills operation videos and obtain a first evaluation result. It can also utilize a second evaluation model, such as the DeepSeek-R1671B model (DeepSeek open-source large-scale language model), to evaluate text data and obtain a second evaluation result. By using the first evaluation model to automatically evaluate clinical skills operation videos from a visual information perspective and obtaining a first evaluation result, and by using the second evaluation model to automatically evaluate clinical skills operation videos from a speech information perspective and obtain a second evaluation result, and by fusing the first and second evaluation results to obtain the target evaluation result, compared to the traditional method requiring manual evaluation by evaluators, the solution in this application embodiment significantly reduces the cost of clinical skills evaluation, shortens the evaluation time, and improves evaluation efficiency and accuracy. Furthermore, by comprehensively evaluating from both visual and speech information perspectives, compared to a single evaluation method that only relies on visual information recognition from the video, the evaluation accuracy is further improved.
[0114] The Evaluation Results and Scoring Module 2015 can merge the first evaluation result and the second evaluation result to obtain the target evaluation result, and match multiple standard evaluation items with multiple target evaluation items; based on the target action execution results corresponding to the multiple target evaluation items and the standard score values corresponding to the matched standard evaluation items, the target score values corresponding to the multiple target evaluation items are determined, and the total score value is calculated based on the target score values corresponding to the multiple target evaluation items.
[0115] The evaluation report generation module 2016 can generate and store the target evaluation report for the target user based on the total score, the execution results of the target actions corresponding to multiple target evaluation items, the execution time of each step, and the total execution time.
[0116] By combining the execution results of the target actions corresponding to multiple target evaluation items and the standard scores corresponding to the matching standard evaluation items, the evaluation scores of each target evaluation item are calculated, and the total score of the target user is obtained. The target user is then evaluated intuitively and accurately based on the total score. Furthermore, the target evaluation report for the target user is generated and stored based on the total score and the execution results of the target actions corresponding to multiple target evaluation items, which improves the readability and traceability of the evaluation report.
[0117] Figure 3 A schematic diagram of one embodiment of a clinical skills assessment device provided in this application is shown. The device may include the following units: The first acquisition unit 301 is used to acquire a clinical skills operation video in response to an assessment request; the clinical skills operation video is acquired by a video acquisition device from the clinical skills actions performed by the target user; the clinical skills operation video includes multimodal data; Extraction unit 302 is used to extract target modality data from multimodal data; Evaluation unit 303 is used to evaluate multimodal data using a first evaluation model according to evaluation standard data to obtain a first evaluation result; the evaluation standard data includes multiple standard evaluation items and standard scores corresponding to the multiple standard evaluation items, and the first evaluation result includes multiple first evaluation items and first action execution results corresponding to the multiple first evaluation items; and to evaluate target modal data using a second evaluation model according to the evaluation standard data to obtain a second evaluation result; the second evaluation result includes multiple second evaluation items and second action execution results corresponding to the multiple second evaluation items. The fusion unit 304 is used to fuse the first evaluation result and the second evaluation result to obtain the target evaluation result; the target evaluation result includes multiple target evaluation items and the target action execution results corresponding to the multiple target evaluation items respectively; Matching unit 305 is used to match multiple standard evaluation items with multiple target evaluation items; The first determining unit 306 is used to determine the target score value corresponding to each of the multiple target evaluation items based on the target action execution results corresponding to the multiple target evaluation items and the standard score value corresponding to the matching standard evaluation items. The generation unit 307 is used to calculate the total score based on the target scores corresponding to multiple target evaluation items, and to generate a target evaluation report for the target user based on the total score and the target action execution results corresponding to multiple target evaluation items.
[0118] In some embodiments, the standard evaluation item and the target evaluation item include strings; The matching unit 305 can be specifically used to search for a standard evaluation item with the same string as the target evaluation item from the plurality of standard evaluation items for any target evaluation item; If not found, search among the multiple standard evaluation items for a standard evaluation item that contains the same character fragment as the target evaluation item; If not, based on the feature data of the target evaluation item, calculate the feature similarity between the target evaluation item and the plurality of standard evaluation items, and based on the feature similarity, determine the standard evaluation item that meets the similarity requirement with the target evaluation item from the plurality of standard evaluation items.
[0119] In some embodiments, the device may further include: The preprocessing unit is used to perform preprocessing operations on the target evaluation items. The preprocessing operations include at least one of the following: unifying full-width / half-width symbols, removing guiding text and category text, clearing preset symbols and their corresponding text content, and converting character formats to a specified format.
[0120] In some embodiments, the result of the target action execution includes whether the action execution has been completed or not. The first determining unit 306 can be used to determine the score of any target evaluation item as the standard score value corresponding to the matching standard score item if the target action execution result of the target evaluation item is that the action execution has been completed and there is a standard evaluation item that matches the target evaluation item; otherwise, the score of the target evaluation item is determined to be 0.
[0121] In some embodiments, the fusion unit 304 may be specifically used to take the union of multiple first evaluation items and multiple second evaluation items as multiple target evaluation items; when any target evaluation item corresponds to a first action execution result or a second action execution result, take the first action execution result or the second action execution result as the target action execution result corresponding to the target evaluation item; when any target evaluation item corresponds to both the first action execution result and the second action execution result, if the first action execution result and / or the second action execution result indicate that the action execution has been completed, determine that the target action execution result of the target evaluation item is that the action execution has been completed; otherwise, determine that the target action execution result of the target evaluation item is that the action execution has not been completed.
[0122] In some embodiments, the device may further include: The second acquisition unit is used to acquire the assessment criteria document; the assessment criteria document includes at least the operation names of multiple clinical skills and their corresponding standard score values; The conversion unit is used to describe information according to a preset expression method and convert multiple operation names into corresponding standard evaluation items. The building unit is used to construct and store evaluation standard data based on multiple standard evaluation items and their corresponding standard scores.
[0123] In some embodiments, the device may further include: The second determining unit is used to determine the target duration of the steps for extracting target modal data from multimodal data until the target evaluation result is obtained; The generation unit can be used to generate a target evaluation report for the target user based on the total score, the target action execution results corresponding to multiple target evaluation items, and the target duration. The device may also include: The storage unit is used to convert the target assessment report into the target data format and store it.
[0124] Figure 3 The aforementioned clinical skills assessment device can perform Figure 1 The implementation principle and technical effects of the clinical skills assessment method described in the illustrated embodiments will not be repeated here. The specific methods by which each module and unit of the clinical skills assessment device in the above embodiments performs its operations have been described in detail in the embodiments related to this method, and will not be elaborated upon here.
[0125] Figure 4 This is a schematic diagram of the structure of one embodiment of a computing device provided in this application. Figure 4 As shown, in practical applications, the computing device may include a storage component 401 and a processing component 402.
[0126] Storage component 401 is used to store computer programs and can be configured to store various other data to support operation on a computing device. Examples of this data include instructions for any application or method used to operate on the computing device, data structures, contact data, phone book data, messages, pictures, videos, etc.
[0127] Processing component 402, coupled to storage component 401, is used to execute computer programs in storage component 401 for implementing, etc. Figure 1 The clinical skills assessment method shown.
[0128] Furthermore, such as Figure 4 As shown, the computing device may also include other components such as a communication component 403, a display component 404, a power supply component 405, and an audio component 406. Figure 4 The diagram only shows some components and does not mean that the device includes only these components. Figure 4 The components shown. Additionally... Figure 4 The components within the dashed box are optional, not mandatory, and their specific requirements depend on the product form of the computing device. The computing device in this embodiment can be a terminal device such as a desktop computer, laptop computer, smartphone, or IoT (Internet of Things) device, or a server-side device such as a conventional server, cloud server, or server array. If the computing device in this embodiment is implemented as a terminal device such as a desktop computer, laptop computer, or smartphone, it may include... Figure 4 The components within the dashed box; if the computing device in this embodiment is implemented as a conventional server, cloud server, or server array, etc., then it may not include... Figure 4 The component within the dashed box.
[0129] The processing component described above includes one or more processors to execute computer instructions to complete all or part of the steps in the method described above. Alternatively, the processing component may be implemented as one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field-programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components to perform the method described above.
[0130] The aforementioned storage components can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as Static Random-Access Memory (SRAM), Electrically Erasable Programmable Read Only Memory (EEPROM), Erasable Programmable Read Only Memory (EPROM), Programmable Read-Only Memory (PROM), Read-Only Memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk.
[0131] The aforementioned communication component is configured to facilitate wired or wireless communication between the device housing the communication component and other devices. The device housing the communication component can access wireless networks based on communication standards, such as mobile communication networks, or combinations thereof. In one exemplary embodiment, the communication component receives broadcast signals or broadcast-related information from an external broadcast management system via a broadcast channel.
[0132] The aforementioned display components may include a screen, which may include a liquid crystal display (LCD) and a touch panel (TP). If the screen includes a touch panel, the screen can be implemented as a touchscreen to receive input signals from the user. The touch panel includes one or more touch sensors to sense touches, swipes, and gestures on the touch panel. The touch sensors can sense not only the boundaries of touch or swipe actions but also the duration and pressure associated with the touch or swipe operation.
[0133] The aforementioned power supply components provide power to various components within the device in which they reside. These power supply components may include a power management system, one or more power sources, and other components associated with generating, managing, and distributing power to the device in which they reside.
[0134] The aforementioned audio component can be configured to output and / or input audio signals. For example, the audio component includes a microphone (MIC) configured to receive external audio signals when the device containing the audio component is in an operating mode, such as call mode, recording mode, or voice recognition mode. The received audio signals can be further stored in memory or transmitted via a communication component. In some embodiments, the audio component also includes a speaker for outputting audio signals.
[0135] Accordingly, embodiments of this application also provide a computer-readable storage medium storing a computer program, which, when executed by a processor, enables the processor to implement the steps in the above-described method embodiments. The computer-readable storage medium includes volatile or non-volatile components, or a combination thereof, and can be removable or non-removable. Examples of computer-readable storage media include, but are not limited to, phase-change random access memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random-access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), flash memory or other memory technologies, CD-ROM, Digital Video Disc (DVD) or other optical storage, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other non-transmission medium. Accordingly, this application also provides a computer program product, which includes a computer program or instructions that, when executed by a processor, cause the processor to implement the steps in the above method embodiments. It should be understood that each step or combination of steps in the above method flow can be implemented by the computer program or instructions. Furthermore, these computer programs or instructions can be applied to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device, enabling the processor of the general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing device to function as an apparatus for implementing the corresponding functions in the above method embodiments.
[0136] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.
[0137] It should also be noted that the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article, or apparatus. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element.
[0138] Finally, it should be noted that the above are merely embodiments of this application and are not intended to limit the scope of this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the scope of the claims of this application.
Claims
1. A clinical skills assessment method, characterized in that, include: In response to the assessment request, obtain clinical skills demonstration videos; The clinical skills operation videos are obtained by video acquisition devices capturing the clinical skills actions performed by the target user; The clinical skills demonstration videos include multimodal data; Extract target modality data from the multimodal data; The multimodal data is evaluated using a first evaluation model according to evaluation standard data to obtain a first evaluation result; the evaluation standard data includes multiple standard evaluation items and standard scores corresponding to the multiple standard evaluation items respectively; the first evaluation result includes multiple first evaluation items and first action execution results corresponding to the multiple first evaluation items respectively. The target modal data is evaluated using a second evaluation model according to the evaluation criteria data to obtain a second evaluation result; the second evaluation result includes multiple second evaluation items and the second action execution results corresponding to the multiple second evaluation items respectively; The first evaluation result and the second evaluation result are fused together to obtain the target evaluation result; The target evaluation results include multiple target evaluation items and the target action execution results corresponding to each of the multiple target evaluation items; Match the plurality of target evaluation items with the plurality of standard evaluation items; Based on the execution results of the target actions corresponding to the multiple target evaluation items and the standard scores corresponding to the matching standard evaluation items, the target scores corresponding to the multiple target evaluation items are determined. Based on the target scores corresponding to the multiple target evaluation items, a total score is calculated, and based on the total score and the target action execution results corresponding to the multiple target evaluation items, a target evaluation report for the target user is generated.
2. The method according to claim 1, characterized in that, The standard evaluation item and the target evaluation item include strings; The step of matching the plurality of target evaluation items with the plurality of standard evaluation items includes: For any target evaluation item, find the standard evaluation item with the same string as the target evaluation item from the plurality of standard evaluation items; If not found, search among the multiple standard evaluation items for a standard evaluation item that contains the same character fragment as the target evaluation item; If not, based on the feature data of the target evaluation item, calculate the feature similarity between the target evaluation item and the plurality of standard evaluation items, and based on the feature similarity, determine the standard evaluation item that meets the similarity requirement with the target evaluation item from the plurality of standard evaluation items.
3. The method according to claim 2, characterized in that, Before matching the plurality of target evaluation items with the plurality of standard evaluation items, the method further includes: The target evaluation item is preprocessed; the preprocessing operation includes at least one of the following: unifying full-width / half-width symbols, removing guiding text and category text, clearing preset symbols and corresponding text content, and converting character format to a specified format.
4. The method according to claim 1, characterized in that, The target action execution result includes whether the action execution is completed or not; the determination of the target score value corresponding to each of the multiple target evaluation items based on the target action execution results corresponding to the multiple target evaluation items and the standard score value corresponding to the matching standard evaluation items includes: For any target evaluation item, if the target action execution result of the target evaluation item is that the action execution has been completed, and there is a standard evaluation item that matches the target evaluation item, the score value of the target evaluation item is determined to be the standard score value corresponding to the matching standard score item. Otherwise, the score of the target evaluation item is determined to be 0.
5. The method according to claim 1, characterized in that, The step of fusing the first evaluation result with the second evaluation result to obtain the target evaluation result includes: The union of the plurality of first evaluation items and the plurality of second evaluation items is taken as a plurality of target evaluation items; In the case where any target evaluation item corresponds to the execution result of the first action or the execution result of the second action, the execution result of the first action or the execution result of the second action shall be taken as the execution result of the target action corresponding to the target evaluation item; In the case of any target evaluation item corresponding to the first action execution result and the second action execution result, if the first action execution result and / or the second action execution result indicate that the action execution has been completed, the target action execution result of the target evaluation item is determined to be that the action execution has been completed; otherwise, the target action execution result of the target evaluation item is determined to be that the action execution has not been completed.
6. The method according to claim 1, characterized in that, The method further includes: Obtain the assessment criteria document; the assessment criteria document shall include at least the operation names of multiple clinical skills and their corresponding standard score values; Describe information according to a preset expression method, and convert multiple operation names into corresponding standard evaluation items; Based on multiple standard evaluation items and their corresponding standard scores, evaluation standard data is constructed and stored.
7. The method according to claim 1, characterized in that, The target modal data includes audio data; The step of evaluating the target modal data using the second evaluation model according to the evaluation criteria data includes: The audio data is converted into text data, and the text data is evaluated using a second evaluation model according to the evaluation criteria data. The method further includes: The target duration for determining the steps of extracting target modality data from the multimodal data until obtaining the total score is determined; The process of generating a target evaluation report for the target user based on the total score and the execution results of the target actions corresponding to the multiple target evaluation items includes: Based on the total score, the execution results of the target actions corresponding to the multiple target evaluation items, and the target duration, a target evaluation report is generated for the target user. The method further includes: The target evaluation report is converted into target data format and stored.
8. A computing device, characterized in that, This includes processing components and storage components; The storage component stores a computer program; the computer program is invoked and executed by the processing component to implement the clinical skills assessment method as described in any one of claims 1 to 7.
9. A computer-readable storage medium, characterized in that, It stores a computer program that, when executed by a processing component, implements the clinical skills assessment method as described in any one of claims 1 to 7.
10. A computer program product, characterized in that, Includes a computer program or instructions that, when executed by a processing component, implement the clinical skills assessment method as described in any one of claims 1 to 7.