Method and system for evaluating comprehensive clinical competency of medical students

By collecting video and audio data of medical students interacting with virtual patients, and combining speech recognition and retrieval enhancement generation technology, the accuracy of medical students' knowledge and communication performance are assessed. This solves the problem of the separation between professional knowledge and communication skills in the existing assessment system, and realizes an objective quantitative assessment of medical students' comprehensive clinical competence, thereby improving the scientificity and credibility of the assessment.

CN121998495APending Publication Date: 2026-05-08ARMY MEDICAL UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
ARMY MEDICAL UNIV
Filing Date
2026-01-21
Publication Date
2026-05-08

AI Technical Summary

Technical Problem

The existing medical student assessment system suffers from a disconnect between professional knowledge assessment and clinical communication skills assessment, strong subjectivity in assessment, lack of objective quantitative indicators, and insufficient assessment of non-verbal communication, resulting in a discrepancy between assessment results and actual clinical competence.

Method used

By collecting video and audio data from interactions between medical students and virtual patients, and combining this with speech recognition technology to obtain response text, the accuracy of the knowledge is assessed by semantically aligning the data with an authoritative medical knowledge base using retrieval enhancement generation technology. Simultaneously, communication performance is evaluated based on multimodal data, including language clarity, empathy, attitude, eye contact rate, and facial expression activity, and a weighted fusion is performed to generate a comprehensive score.

Benefits of technology

It enables a comprehensive, objective, and quantitative assessment of medical students' overall clinical competence, improves the scientific rigor and credibility of the assessment, provides medical students with precise and actionable improvement guidance, and overcomes the shortcomings of traditional assessments, such as strong subjectivity, lack of objective quantitative indicators, and insufficient non-verbal communication assessment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121998495A_ABST
    Figure CN121998495A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of medical education, and discloses a method and a system for evaluating comprehensive clinical competency of medical students, video and audio data in an interaction process of the medical students and virtual patients are collected, a response text is obtained in combination with a voice recognition technology, and the comprehensive clinical competency of the medical students is evaluated through a retrieval enhancement generation technology. Performing semantic alignment on the response text of the medical student and the authoritative medical knowledge base; the method comprises the steps of evaluating language definition, common situation performance, attitude performance, eye contact rate and expression activeness of medical students on the basis of multi-modal data, finally, performing weighted fusion on various quantitative indexes to obtain a comprehensive score, generating a feedback report, and performing encrypted storage on original data; therefore, the problems of mutual separation of professional knowledge and clinical communication ability assessment, strong assessment subjectivity, lack of objective quantitative indexes and insufficient non-language communication assessment in medical student assessment in a current medical education system are solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of medical education technology, and more specifically, to a method and system for assessing the comprehensive clinical competence of medical students. Background Technology

[0002] In the current medical education system, the assessment of medical students generally suffers from a disconnect between professional knowledge and clinical communication skills. Professional knowledge is primarily assessed through traditional written and oral examinations, focusing on the memorization and understanding of "hard knowledge" such as medical theories, disease diagnosis, and treatment plans. However, communication skills are often assessed independently of professional knowledge, typically through specific stations in the Objective Structured Clinical Examination (OSCE). This separate assessment model leads to a fundamental flaw: it fails to accurately reflect the comprehensive abilities required of physicians in clinical practice. In real medical settings, every interaction between doctors and patients involves a high degree of integration of professional knowledge and communication skills. Existing separate assessment methods cannot effectively measure this comprehensive ability, resulting in a discrepancy between assessment results and students' actual clinical competence.

[0003] Furthermore, existing assessment methods for medical students, particularly in the area of ​​communication skills, suffer from significant subjectivity and a lack of objective, quantifiable indicators. In traditional assessment models, whether through faculty observation, standardized patient (SP) feedback, or peer evaluation, the results largely depend on the assessor's personal experience, professional level, and subjective judgment. Different assessors may offer significantly different evaluations of the same communication behavior, raising questions about the reliability and validity of the assessment results. This highly subjective approach not only fails to provide students with precise and actionable improvement suggestions but also hinders the establishment of unified and comparable teaching quality standards among faculty members and institutions.

[0004] Meanwhile, nonverbal communication, including facial expressions, eye contact, body posture, gestures, and tone of voice, plays a crucial role in doctor-patient communication. However, in the current medical student assessment system, the evaluation of nonverbal communication skills is often neglected or reduced to a mere formality. Traditional assessment methods rely primarily on the assessor's visual observation, which is not only inefficient but also fails to capture subtle, momentary nonverbal cues. Due to the lack of objective quantitative standards, assessors find it difficult to accurately describe and evaluate students' nonverbal communication performance, often providing only general feedback and failing to offer specific, actionable guidance for improvement.

[0005] To address the aforementioned issues, existing technologies urgently need improvement. Summary of the Invention

[0006] The purpose of this application is to provide a method and system for assessing the comprehensive clinical competence of medical students, aiming to solve the problems in the current medical education system regarding the assessment of medical students, such as the separation between the assessment of professional knowledge and clinical communication skills, the strong subjectivity of the assessment, the lack of objective quantitative indicators, and the insufficient assessment of non-verbal communication.

[0007] Firstly, this application provides a method for assessing the comprehensive clinical competence of medical students, the method comprising the following steps: A1. Collect video and audio data during the interaction between medical students and virtual patients, and perform speech recognition on the audio data to obtain the medical student's response text; A2. Based on retrieval enhancement generation technology, the medical student's response text is semantically aligned with an authoritative medical knowledge base, and the knowledge accuracy is calculated based on the alignment results; A3. Based on the video data, the audio data, and the medical student's response text, the medical student's communication performance is evaluated to obtain communication performance indicators; the communication performance indicators include language clarity, empathy index, attitude index, eye contact rate, and facial expression activity. A4. Weight and integrate the various quantitative indicators to obtain a comprehensive score; the quantitative indicators include the knowledge accuracy and the various communication performance indicators. A5. Based on the comprehensive score and the various quantitative indicators, generate a feedback report containing dimensional scores, temporal positioning, and learning suggestions, and encrypt and store the original data; the original data includes the video data and the audio data.

[0008] Secondly, this application provides a system for assessing the comprehensive clinical competence of medical students, including a student terminal, a teacher terminal, and a server; The student terminal is used to provide medical students with an interactive interface to interact with virtual patients, and to collect video and audio data during the interaction between medical students and virtual patients, and upload them to the server. The server is configured with: The speech recognition module is used to perform speech recognition on the audio data to obtain the medical student's response text; The knowledge accuracy assessment module is used to semantically align the medical student's response text with an authoritative medical knowledge base based on retrieval enhancement generation technology, and calculate the knowledge accuracy based on the alignment results. The communication performance evaluation module is used to evaluate the communication performance of the medical students based on the video data, the audio data, and the medical students' response text, and obtain communication performance indicators; the communication performance indicators include language clarity, empathy performance index, attitude performance index, eye contact rate, and facial expression activity. The comprehensive scoring module is used to weight and integrate various quantitative indicators to obtain a comprehensive score; the quantitative indicators include the knowledge accuracy and various communication performance indicators. The first feedback module is used to generate a feedback report containing dimensional scores, temporal positioning, and learning suggestions based on the comprehensive score and the various quantitative indicators, encrypt and store the original data, and send the feedback report to the student's terminal; the original data includes the video data and the audio data; The second feedback module is used to generate feedback information containing the medical student's weaknesses and teaching strategy suggestions based on the comprehensive score and the various quantitative indicators, and send it to the teacher's end.

[0009] Beneficial Effects: This application provides a method and system for assessing the comprehensive clinical competence of medical students. By collecting video and audio data during interactions between medical students and virtual patients, and combining this with speech recognition technology to obtain response text, it achieves comprehensive data capture of medical students' clinical performance. Based on this, this application innovatively introduces retrieval-enhanced generation technology to semantically align medical students' response text with authoritative medical knowledge bases, thereby objectively quantifying the accuracy of medical students' knowledge and effectively solving the problem of the disconnect between professional knowledge assessment and clinical communication ability assessment in existing evaluations. Simultaneously, this application conducts a detailed assessment of medical students' communication performance based on multimodal data, covering multiple dimensions such as language clarity, empathy, attitude, eye contact rate, and facial expression activity, overcoming the shortcomings of traditional assessments, such as strong subjectivity, lack of objective quantitative indicators, and insufficient non-verbal communication assessment. Finally, a comprehensive score is obtained by weighted fusion of various quantitative indicators, and a feedback report containing dimensional scores, temporal positioning, and learning suggestions is generated, providing medical students with precise and actionable improvement guidance. The original data is encrypted and stored, ensuring data security and privacy protection. In summary, the method proposed in this application can comprehensively, objectively, and quantitatively assess the overall clinical competence of medical students, significantly improving the scientific rigor, effectiveness, and credibility of the assessment, and providing strong support for improving the quality of medical education. Attached Figure Description

[0010] Figure 1 A flowchart of a method for assessing the comprehensive clinical competence of medical students provided in this application.

[0011] Figure 2 A schematic diagram of a system for assessing the comprehensive clinical competence of medical students, provided in this application.

[0012] Figure 3 This is a schematic diagram of the server.

[0013] Labeling Explanation: 1. Student End; 2. Teacher End; 3. Server; 301. Speech Recognition Module; 302. Knowledge Accuracy Assessment Module; 303. Communication Performance Assessment Module; 304. Comprehensive Scoring Module; 305. First Feedback Module; 306. Low Confidence Backoff Control Module; 307. Second Feedback Module. Detailed Implementation

[0014] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of the embodiments. The components of the embodiments of this application described and shown in the accompanying drawings can generally be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of this application provided in the accompanying drawings is not intended to limit the scope of the claimed application, but merely represents selected embodiments of this application. All other embodiments obtained by those skilled in the art based on the embodiments of this application without inventive effort are within the scope of protection of this application.

[0015] It should be noted that similar reference numerals and letters in the following figures indicate similar items; therefore, once an item is defined in one figure, it does not need to be further defined and explained in subsequent figures. Furthermore, in the description of this application, terms such as "first," "second," etc., are used only to distinguish descriptions and should not be construed as indicating or implying relative importance.

[0016] Please refer to Figure 1 One method for assessing the comprehensive clinical competence of medical students, as described in some embodiments of this application, includes the following steps: A1. Collect video and audio data during the interaction between medical students and virtual patients, and perform speech recognition on the audio data to obtain the medical students' response text; A2. Based on retrieval enhancement generation technology, the medical student's response text is semantically aligned with an authoritative medical knowledge base, and the knowledge accuracy is calculated based on the alignment results; A3. Based on video data, audio data, and medical students' response texts, the communication performance of medical students was evaluated to obtain communication performance indicators. The communication performance indicators include language clarity, empathy index, attitude index, eye contact rate, and facial expression activity. A4. Weighted and integrated scores are obtained from the various quantitative indicators; the quantitative indicators include knowledge accuracy and various communication performance indicators. A5. Based on the comprehensive score and various quantitative indicators, generate a feedback report that includes dimensional scores, temporal positioning, and learning suggestions, and encrypt and store the original data; the original data includes video data and audio data.

[0017] This application aims to provide a method for assessing the comprehensive clinical competence of medical students. By integrating the assessment of professional knowledge and communication skills and introducing objective quantitative indicators, it seeks to overcome the limitations of existing assessment systems. This method collects multimodal data from interactions between medical students and virtual patients and analyzes it using advanced artificial intelligence technology, thereby achieving a comprehensive and objective assessment of medical students' overall clinical competence.

[0018] "Virtual patients" refer to computer-simulated patient images with specific medical histories, symptoms, and emotional responses, which medical students can interact with to simulate real clinical scenarios. These virtual patients provide a standardized interactive environment, ensuring all medical students are assessed under the same conditions.

[0019] The "authoritative medical knowledge base" refers to a database containing professionally reviewed and verified information such as medical theories, disease diagnoses, treatment plans, and clinical guidelines. This knowledge base provides a reliable reference for assessing the accuracy of medical students' professional knowledge.

[0020] Among them, "retrieval-enhanced generation technology" is an artificial intelligence technology that combines information retrieval and text generation. It can retrieve relevant information from a large-scale knowledge base and generate coherent and accurate text based on this information. In this application, this technology is used to semantically align medical students' response text with authoritative medical knowledge, thereby assessing the accuracy of their knowledge.

[0021] Among them, the "communication performance indicators" are a series of metrics used to quantify medical students' performance in doctor-patient communication, including verbal clarity, empathy index, attitudinal performance index, eye contact rate, and facial expression activity. These indicators reflect medical students' communication abilities from different dimensions.

[0022] This application provides a method for assessing the comprehensive clinical competence of medical students. Its core lies in achieving a comprehensive quantitative assessment of the accuracy of medical students' knowledge and their communication performance through multimodal data collection and intelligent analysis.

[0023] Specifically, in step A1, video and audio data of the interaction between the medical student and the virtual patient need to be collected first. Video data can be recorded using a high-definition camera, capturing the medical student's nonverbal behaviors, such as facial expressions, eye contact, and body posture. Audio data can be recorded using a high-sensitivity microphone, capturing the medical student's speech information, including speech rate, tone, and volume. After collection, the audio data undergoes speech recognition to convert it into the medical student's response text. Speech recognition technology can employ deep learning-based acoustic and language models to convert continuous speech signals into discrete text sequences. For example, this can be achieved using open-source speech recognition toolkits or commercial speech recognition services.

[0024] In step A2, based on retrieval enhancement generation technology, the medical student's response text is semantically aligned with an authoritative medical knowledge base, and the knowledge accuracy is calculated based on the alignment results. This step aims to evaluate the accuracy and completeness of the medical knowledge expressed by the medical student during the interaction. For example, the medical student's response text can first be preprocessed, including word segmentation, part-of-speech tagging, and named entity recognition. Subsequently, retrieval enhancement generation technology is used to retrieve medical knowledge fragments related to the content of the medical student's response text from the authoritative medical knowledge base. For example, a vector similarity matching method can be used to compare the semantic vector of the medical student's response text with the semantic vector of medical knowledge entries in the knowledge base to find the most relevant knowledge points. Next, the medical student's response text is semantically aligned with the retrieved authoritative medical knowledge, for example, by comparing the overlap and logical relationships between key concepts in the medical student's response text and authoritative knowledge points, thereby calculating the knowledge accuracy.

[0025] In step A3, the communication performance of medical students is evaluated based on video data, audio data, and their response texts to obtain communication performance indicators. These indicators include verbal clarity, empathy index, attitudinal index, eye contact rate, and facial expression activity. For example, verbal clarity can be analyzed by examining vocabulary usage, syntactic structure, and terminology in the students' response texts, quantified by metrics such as terminology ratio, average sentence length, and clause nesting ratio. For the empathy and attitudinal indices, acoustic features such as fundamental frequency, energy, speech rate, and intonation variations can be extracted from the audio data, and machine learning models can be used to analyze these features to identify the students' emotional state and level of empathy. For eye contact rate and facial expression activity, nonverbal visual features can be extracted from the video data; for example, facial recognition technology can be used to detect key facial features, head posture, and gaze direction to calculate the duration and frequency of eye contact and the intensity of facial expression changes.

[0026] In step A4, the various quantitative indicators are weighted and fused to obtain a comprehensive score. The quantitative indicators include knowledge accuracy and various communication performance indicators. For example, a linear weighting method can be used, assigning a weight coefficient to each indicator, and then multiplying the scores of each indicator by their corresponding weight coefficients and summing the results to obtain the final comprehensive score. These weight coefficients can be trained and optimized based on clinical expert experience, teaching objectives, or through machine learning methods (such as regression analysis).

[0027] In step A5, based on the comprehensive score and various quantitative indicators, a feedback report is generated, including dimensional scores, temporal positioning, and learning suggestions. The raw data is then encrypted and stored. The raw data includes video and audio data. The feedback report can detail the medical student's scores in various dimensions such as knowledge accuracy, language clarity, and empathy, and indicate specific points in time when the student performed well or needed improvement during the interaction (temporal positioning). For example, the report may indicate that the medical student's knowledge answer was inaccurate on a specific question, or that the medical student failed to demonstrate sufficient empathy at a specific moment. Simultaneously, the report will provide targeted learning suggestions, such as recommending relevant medical knowledge learning resources or providing communication skills training methods. To protect the privacy and data security of medical students, the collected raw video and audio data are encrypted before storage. For example, differential privacy technology is used to add noise, and the data is encrypted using the AES encryption algorithm before being stored on a secure server.

[0028] The method proposed in this application for assessing the comprehensive clinical competence of medical students achieves a comprehensive, objective, and quantitative assessment of their clinical competence by integrating multimodal data acquisition, intelligent speech recognition, retrieval enhancement generation technology, multi-dimensional communication performance assessment, and a weighted fusion scoring mechanism.

[0029] Compared to existing technologies, the core innovation of this application lies in constructing a comprehensive assessment framework that can objectively and quantitatively evaluate medical students' professional knowledge and clinical communication skills simultaneously. Traditional methods often separate knowledge assessment from communication skills evaluation, resulting in assessment results that fail to accurately reflect medical students' comprehensive performance in actual clinical situations. This application, by introducing retrieval-enhanced generation technology, achieves semantic alignment between medical students' response texts and authoritative medical knowledge bases, thereby enabling precise calculation of knowledge accuracy and overcoming the limitations of traditional written or oral examinations in assessing knowledge application abilities. Furthermore, this application utilizes multimodal data (video and audio data) to conduct a detailed analysis of medical students' nonverbal communication performance, quantifying indicators such as eye contact rate and facial expression activity—indicators that are difficult to achieve and often overlooked in existing assessment systems. By weighted fusion of various quantitative indicators, this application can generate a comprehensive score and provide a feedback report including dimensional scores, temporal positioning, and learning suggestions. This not only improves the objectivity and accuracy of the assessment but also provides medical students with specific and actionable directions for improvement, significantly outperforming the shortcomings of traditional assessments, which are often subjective and provide vague feedback. Therefore, this application represents a significant technological advancement and innovation in the assessment of medical students' comprehensive clinical competence.

[0030] In some implementations, step A2 includes: A201. Obtain dialogue context information; dialogue context information includes the type of question from the virtual patient and the department type information of the medical student; A202. Perform named entity recognition and key point extraction on the medical student's response text to obtain the medical entities and key points in the medical student's response text; A203. Based on dialogue context information, medical entities, and key points, and using retrieval enhancement generation technology, retrieve the Top-K most relevant authoritative documents from an authoritative medical knowledge base, and extract the key knowledge points from the retrieved authoritative documents; Top-K is a preset positive integer; A204. Compare key points with knowledge points, calculate the coverage of knowledge points and the matching degree of key points in the medical student's response text, and obtain semantic coverage and F1 score; A205. Weighted fusion of semantic coverage and F1 score yields knowledge accuracy.

[0031] Specifically, in step A201, the dialogue context information refers to the background information related to the interaction between the medical student and the virtual patient, aiming to provide necessary context for subsequent knowledge retrieval and assessment. The virtual patient's question type information can be understood as the main questions raised by the virtual patient during the interaction or the main symptoms they exhibit, such as "chest pain," "fever," or "dizziness," which helps the system focus on the relevant medical knowledge area. The medical student's department type information refers to the medical student's specialty department, such as "cardiology," "respiratory medicine," or "neurology," aiming to adjust the focus and depth of knowledge assessment based on the medical student's professional background. In practical applications, this dialogue context information can be preset by the system before the interaction begins, or dynamically acquired during the interaction based on the virtual patient's settings and the medical student's choices.

[0032] In step A202, named entity recognition and key point extraction are performed on the medical student's response text. The purpose is to accurately identify core medical concepts and important arguments from the medical student's free text response. Named entity recognition refers to identifying entities with specific meanings in the text, such as disease names, drug names, symptoms, and examination items. For example, a pre-trained named entity recognition model can be used to identify named entities. This model can employ deep learning models such as BERT or RoBERTa. During the training process of the named entity recognition model, the text sequence labeled with real entity tags is usually first converted into word vectors or character vectors as input to the model. The model predicts the entity tag sequence corresponding to each word or character based on these vectors. Then, the model error is quantified by calculating the loss function between the predicted tags and the real tags, and the optimizer iteratively updates the model parameters using the backpropagation algorithm. This process is repeated until the model performance reaches the expected level or the training epoch ends. When named entity recognition is needed, the medical student's response text is converted into word vectors or character vectors and input into the trained named entity recognition model to obtain the recognition result. Key point extraction refers to identifying key sentences or phrases from medical students' response texts that express core viewpoints, diagnostic approaches, or treatment suggestions. These key points can represent the medical students' understanding and ability to handle specific medical issues. For example, text summarization techniques can be used to extract key points. Specifically, first, the medical student's response text is preprocessed (including word segmentation, stop word removal, and word form restoration) to facilitate subsequent analysis. Next, the importance of each sentence or phrase in the text is evaluated. For example, an importance score is calculated based on at least one of the following: word frequency (e.g., sentences or phrases containing frequently occurring keywords are more important), the position of the sentence or phrase in the text (e.g., opening and closing sentences often contain the theme or conclusion, thus being more important), and similarity to the title (higher similarity means higher importance). (For example, an importance value is assigned to each of the above features, and then the weighted sum of these importance values ​​is calculated to obtain the importance score). Finally, based on a preset importance threshold, several sentences with importance scores higher than the threshold are selected from the medical student's response text and combined to form a concise summary, thereby effectively extracting the core points of the original text as key points.

[0033] In step A203, based on dialogue context information, medical entities, and key points, retrieval-augmented generation (RAG) technology is used to retrieve the top-K most relevant authoritative documents from an authoritative medical knowledge base, and the key knowledge points from the retrieved authoritative documents are extracted. RAG technology combines information retrieval and text generation, aiming to enhance the accuracy and reliability of the model's generated answers using an external knowledge base. RAG technology is existing and will not be detailed here. Specifically, dialogue context information, medical entities, and key points are used as query conditions to guide the retrieval process, ensuring that the retrieved documents are highly relevant to the current clinical context (e.g., integrating dialogue context information, medical entities, and key points to construct a multi-dimensional, structured query condition; then using a knowledge graph to semantically expand the medical entities in the query condition, incorporating their synonyms, hypernyms, or related concepts to obtain expanded query conditions; using the expanded query conditions for matching in the authoritative medical knowledge base, and using the K most relevant authoritative documents as the retrieval results). Authoritative medical knowledge bases can include medical textbooks, clinical guidelines, professional journals, and medical databases (such as PubMed and UpToDate). Their content undergoes professional review and possesses high accuracy and authority. Top-K is a preset positive integer representing the number of documents with the highest relevance retrieved; for example, K can be set to 3, 5, or 10. From these authoritative documents, key knowledge points relevant to the medical student's response text are further extracted. These key knowledge points represent accurate and comprehensive explanations of specific medical issues or concepts from the authoritative knowledge base.

[0034] In step A204, key points and knowledge points are compared, and the coverage of knowledge points and the matching degree of key points in the medical student's response text are calculated to obtain semantic coverage and F1 score. Semantic coverage refers to the degree of coverage between the knowledge points in the medical student's response text and the authoritative knowledge points. Its purpose is to measure the breadth of the medical student's answer, i.e., whether it covers all necessary knowledge points. For example, it can be obtained by calculating the proportion of key points in the medical student's response text whose semantic similarity to authoritative knowledge points exceeds a preset similarity threshold to the total number of knowledge points. The F1 score refers to the matching degree between key points and knowledge points. Its purpose is to measure the depth and accuracy of the medical student's answer, i.e., the combined performance of the precision and recall of its key points and authoritative knowledge points. The F1 score is the harmonic mean of precision and recall, which comprehensively reflects the accuracy of the matching.

[0035] In step A205, the semantic coverage rate and F1 score are weighted and fused to obtain the knowledge accuracy. Weighted fusion refers to linearly combining the semantic coverage rate and F1 score according to preset weight coefficients. Its purpose is to comprehensively consider the breadth and depth of medical students' knowledge, thereby obtaining a more comprehensive, objective, and accurate knowledge accuracy. For example, the knowledge accuracy KAS can be expressed as KAS = λ * Coverage + (1 − λ) * F1(KeyPoints), where KAS is the knowledge accuracy, Coverage is the semantic coverage rate, F1(KeyPoints) is the F1 score, and λ is the weight coefficient, which typically ranges from 0 to 1 (and can be set according to actual needs), used to balance the contributions of semantic coverage rate and F1 score to the final knowledge accuracy.

[0036] This application's solution significantly improves the accuracy and reliability of medical student knowledge accuracy assessment by introducing dialogue context information, performing named entity recognition and key point extraction on medical student response text, and using these refined features to guide retrieval enhancement generation technology. Specifically, dialogue context information provides the necessary context for knowledge retrieval and assessment, enabling the system to focus on the knowledge most relevant to the current clinical context and the medical student's professional background, avoiding generalized assessment results. Named entity recognition and key point extraction ensure accurate grasp of the core content of the medical student's response, effectively filtering out irrelevant information, making subsequent semantic alignment more focused and efficient. The retrieval enhancement generation technology based on these refined features can efficiently and accurately locate the most relevant authoritative documents from massive amounts of medical knowledge and extract authoritative knowledge points, thus ensuring the authority and relevance of the assessment. Furthermore, through multi-dimensional quantification of semantic coverage and F1 score, the breadth and depth of medical student knowledge can be comprehensively measured, overcoming the limitations of single-indicator assessment.

[0037] Through the above technical solutions, this application can significantly improve the accuracy and reliability of medical students' knowledge accuracy assessment. Specifically, the introduction of dialogue context information makes knowledge retrieval and assessment more context-relevant, avoiding vague assessment results. Named entity recognition and key point extraction ensure accurate grasp of the core content of medical students' responses and effectively filter out irrelevant information. Based on these refined features, the retrieval enhancement generation technology can efficiently and accurately locate the most relevant authoritative knowledge from massive amounts of medical knowledge, thereby ensuring the authority and relevance of the assessment. In addition, through multi-dimensional quantification of semantic coverage and F1 score, the breadth and depth of medical students' knowledge can be comprehensively measured, overcoming the limitations of single-indicator assessment. As a result, the obtained knowledge accuracy assessment results are more objective and detailed, providing medical students with more instructive feedback, helping them to accurately identify knowledge weaknesses and conduct targeted learning.

[0038] Preferably, step A204 may include: Obtain the semantic vectors of key points and knowledge points; Calculate the semantic similarity between the semantic vectors of each key point and the semantic vectors of each knowledge point; For each key point, the matching status of the key point is determined based on the semantic similarity between the semantic vector of the key point and the semantic vector of each knowledge point, as well as the preset similarity threshold. Calculate the F1 score by considering the matching status of each key point; For each knowledge point, determine whether there is a key point in the medical student's response text with a semantic similarity of not less than a preset similarity threshold, and determine the coverage status of the knowledge point based on the judgment result. Calculate semantic coverage based on the coverage status of each knowledge point.

[0039] Specifically, the semantic vectors of key points and knowledge points can be obtained by encoding the text using pre-trained language models (such as BERT, Word2Vec, or GloVe). These semantic vectors can capture the deep semantic information of the text, making similar words or phrases closer together in the vector space.

[0040] The semantic similarity between the semantic vectors of each key point and the semantic vectors of each knowledge point can be obtained by calculating the cosine similarity between the vectors. Cosine similarity can effectively measure the degree of proximity of two vectors in a direction, thus reflecting their semantic relevance.

[0041] Furthermore, for each key point, its matching status is determined based on the semantic similarity between its semantic vector and the semantic vectors of all knowledge points. When the semantic similarity between the semantic vector of a key point and the semantic vector of any knowledge point is not less than a preset similarity threshold, the key point is determined to be a match. The preset similarity threshold can be adjusted according to the actual application scenario and evaluation accuracy requirements, for example, set to 0.7-0.8.

[0042] Therefore, the F1 score is calculated by comprehensively considering the matching status of all key points. The F1 score is the harmonic mean of precision and recall, which comprehensively measures the degree of matching between key points in the medical student's response text and authoritative medical knowledge points. Specifically, precision can be defined as the proportion of the number of correctly matched key points (here, the number of matched key points can be used as the number of correctly matched key points) to the total number of key points in the medical student's response text, while recall can be defined as the proportion of the number of correctly matched key points to the total number of knowledge points in the authoritative medical knowledge base.

[0043] Furthermore, for each knowledge point, the coverage status of that knowledge point can be determined by judging whether there are any key points in the medical student's response text with a semantic similarity not less than the preset similarity threshold. If such key points exist, the knowledge point is considered to be covered.

[0044] Ultimately, the semantic coverage rate is calculated based on the coverage status of all knowledge points, representing the overall degree of coverage of authoritative medical knowledge points by the medical student's response text. For example, the semantic coverage rate can be defined as the proportion of the number of covered knowledge points to the total number of knowledge points in the authoritative medical knowledge base.

[0045] This application's solution, by introducing semantic vectors and semantic similarity calculation, overcomes the limitations of traditional keyword-based matching methods, more accurately capturing the deep semantic connections between medical students' responses and authoritative medical knowledge bases. By transforming key points and knowledge points into high-dimensional semantic vectors and calculating their similarity, it can identify medical concepts with similar meanings even if their expressions differ, thus providing a more comprehensive and objective assessment of medical students' knowledge accuracy. Specifically, by determining the matching status of key points and the coverage status of knowledge points, and calculating F1 scores and semantic coverage based on these, it can quantitatively assess medical students' knowledge mastery from two dimensions: matching accuracy and knowledge completeness, providing more refined and reliable data support for subsequent comprehensive scoring.

[0046] The aforementioned technical solutions significantly improve the precision and accuracy of knowledge accuracy assessments for medical students. Traditional keyword matching methods often fail to identify synonyms, near-synonyms, or the same concepts expressed in different ways, leading to biased assessment results. This application, however, utilizes semantic vectors and similarity calculations to gain a deeper understanding of the semantic content of medical students' responses, more accurately determining whether they have mastered authoritative medical knowledge points, thus making the knowledge accuracy assessment more objective and comprehensive. This semantically understanding-based assessment method not only reflects the degree of medical students' mastery of knowledge points but also assesses, to some extent, the accuracy and professionalism of their language expression, providing more valuable quantitative indicators for assessing medical students' clinical competence.

[0047] In some implementations, step A3 includes: A301. Calculate the language clarity based on the medical student's response text; A302. Based on the audio data, extract acoustic features and calculate the empathy performance index and attitude performance index based on the acoustic features; A303. Based on the video data, extract non-verbal visual features, and calculate eye contact rate and facial expression activity based on the non-verbal visual features.

[0048] Step A301 aims to quantify the clarity and professionalism of medical students' language expression by analyzing the content and structure of their responses. Language clarity is an important indicator of a medical student's ability to accurately and concisely convey medical information.

[0049] Furthermore, step A302 involves in-depth analysis of the audio data generated during the interaction between medical students and virtual patients to extract acoustic features. These acoustic features, such as tone of voice, speech rate, and volume, can reflect the emotional investment and attitudinal tendencies of medical students in communication. Based on these acoustic features, empathy performance index and attitudinal performance index can be calculated, thereby objectively assessing whether medical students demonstrate empathy and a positive, professional attitude in communication.

[0050] Furthermore, step A303 focuses on utilizing video data from the interaction between medical students and virtual patients. By extracting nonverbal visual features from the video data, such as facial expressions, eye direction, and body language, the effectiveness of medical students' nonverbal communication can be assessed. Specifically, eye contact rate reflects the focus and respect of medical students when interacting with patients, while facial expression activity reflects the emotional expression and approachability of medical students.

[0051] The aforementioned technical solution enables a more refined and comprehensive assessment of medical students' communication performance. Communication performance indicators are broken down into multiple sub-dimensions, including verbal clarity, empathy index, attitudinal index, eye contact rate, and facial expression activity. These are calculated using medical students' text responses, audio data, and video data, respectively. This allows the assessment results to not only reflect medical students' performance in different communication modalities but also provide more targeted dimensional scores and learning suggestions for subsequent feedback reports. Consequently, it is possible to more accurately identify medical students' strengths and weaknesses in communication skills, thereby providing strong support for personalized teaching and training.

[0052] Preferably, step A301 may include: The medical students' response texts were subjected to terminology recognition, syntactic analysis, and syntactic structure analysis to obtain the terminology ratio, average sentence length, and clause nesting ratio. The term ratio, average sentence length, and clause nesting ratio are standardized to obtain the standardized term ratio, standardized average sentence length, and standardized clause nesting ratio. The language clarity is obtained by weighting and fusing the standardized terminology ratio, the standardized average sentence length, and the standardized clause nesting ratio. When the term ratio exceeds a preset ratio threshold or the average sentence length exceeds a preset length threshold, a language expression improvement suggestion is generated.

[0053] Specifically, terminology recognition in medical student responses involves using natural language processing (NLP) technology to identify and extract specialized medical terms from the text. The aim is to quantify the extent to which medical students use professional vocabulary in their communication. Syntactic analysis involves analyzing the grammatical structure of sentences in the text, such as identifying subjects, predicates, and objects, to understand the basic structure of sentences. Sentence structure analysis further analyzes sentence complexity, such as identifying nested clauses and phrases, to assess the complexity of sentence structure. From this, we can obtain the terminology ratio, average sentence length, and clause nesting ratio. The terminology ratio can be understood as the proportion of specialized medical terms in the total vocabulary, reflecting the professionalism and accessibility of medical students' language. The average sentence length refers to the average number of words per sentence in the medical student's response text, measuring the conciseness and redundancy of the language. The clause nesting ratio refers to the frequency and complexity of clause structures in the text, assessing the complexity and comprehensibility of sentence structure.

[0054] In practical applications, the terminology ratio, average sentence length, and clause nesting ratio mentioned above are standardized, for example, using Z-score standardization or Min-Max standardization. The purpose is to eliminate the potential impact of differences in units and numerical ranges between different indicators, ensuring fairness and comparability in subsequent weighted fusion. Subsequently, the standardized terminology ratio, average sentence length, and clause nesting ratio are weighted and fused, for example, by calculating the Language Proficiency Index (LCI) using the formula LCI=1−(w1*JR+w2*ASL+w3*CR, where JR is the terminology ratio, ASL is the average sentence length, CR is the clause nesting ratio, and w1-w3 are preset weighting factors (preferably, to better balance professionalism and accessibility, simplicity and redundancy, and complexity and comprehensibility, a balance interval can be set for JR, ASL, and CR respectively; when these indicators are within their corresponding balance intervals, their corresponding weighting factors are zero; otherwise, they are preset positive values). The purpose is to comprehensively consider the contribution of each indicator to language proficiency, obtaining a comprehensive quantitative evaluation result.

[0055] Furthermore, when the terminology ratio exceeds a preset ratio threshold or the average sentence length exceeds a preset length threshold, the system will generate language expression improvement prompts. The purpose is to specifically point out the specific problems in medical students' language expression and provide actionable improvement suggestions, such as prompting medical students to reduce the use of professional terminology or simplify sentence structure.

[0056] This application's solution overcomes the limitations of a single numerical assessment of language clarity by conducting multi-dimensional, fine-grained linguistic analysis of medical students' response texts. Specifically, through terminology recognition, syntactic analysis, and syntactic structure analysis, it quantifies the professionalism, conciseness, and complexity of medical students' language expression. The terminology ratio, average sentence length, and clause nesting ratio—three indicators—reflect key elements of language clarity from different perspectives. Standardizing these indicators ensures fairness and effectiveness in weighted fusion. Therefore, the resulting Language Clarity Index (LCI) more comprehensively and objectively reflects medical students' language expression abilities. More importantly, when the terminology ratio or average sentence length exceeds a preset threshold, the system can promptly generate language expression improvement suggestions. This transforms the assessment from simply providing a score to directly identifying potential problems in medical students' communication, such as excessive use of technical terms or overly complex sentence structures, thus providing specific and actionable directions for improvement.

[0057] Through the aforementioned technical solution, this application enables a more refined and diagnostic assessment of medical students' language clarity. Compared to traditional methods that only provide a comprehensive score, this solution not only quantifies language clarity but also analyzes its constituent elements in depth, thereby accurately identifying specific weaknesses in medical students' language expression. Furthermore, by generating language expression improvement prompts when specific indicators exceed thresholds, this solution provides medical students with immediate and targeted feedback, guiding them to improve their communication strategies, such as adjusting the frequency of use of professional terminology or optimizing sentence structure. This significantly enhances the effectiveness of medical students' communication with patients and patients' comprehension, ultimately promoting the comprehensive development of medical students' clinical communication competence.

[0058] Preferably, step A302 may include: Based on the audio data, extract acoustic features; acoustic features include at least one of the following: fundamental frequency, energy, speech rate, pause ratio, interruption rate, and response delay; Based on acoustic features, the empathy performance index and attitude performance index are calculated using pre-trained classification or regression models.

[0059] Specifically, acoustic features refer to the objective physical attributes extracted from audio data during interactions between medical students and virtual patients, reflecting the speech and emotional state of the medical students. These features may include, but are not limited to, fundamental frequency, energy, speech rate, pause ratio, interruption rate, and response delay. Fundamental frequency reflects the pitch variation of speech and is related to the fluctuations in emotional expression; energy represents the loudness of speech and can reflect the intensity of the speaker's emotions; speech rate refers to the speed of pronunciation per unit of time and is often related to emotions such as tension and confidence; pause ratio refers to the proportion of pause time in the total speech duration and can reflect thinking, hesitation, or emphasis; interruption rate refers to the frequency with which medical students interrupt the virtual patient's speech, reflecting their listening habits and level of respect; and response delay refers to the time required for medical students to respond to questions from the virtual patient, reflecting their reaction speed and level of comprehension.

[0060] Furthermore, after extracting the aforementioned acoustic features, pre-trained classification or regression models can be used to calculate empathy and attitude performance indices. Pre-trained models are trained on large amounts of labeled data (e.g., speech samples containing different levels of empathy and attitude performance) and can identify and quantify emotions and intentions in speech. For example, a classification model can categorize medical students' speech performance into different levels of empathy or attitude (e.g., "high empathy," "medium empathy," "low empathy," each with a corresponding value), while a regression model can directly output continuous empathy or attitude performance index values. The aim is to objectively assess medical students' empathy and professional attitude in communication through quantitative methods.

[0061] This application's solution extracts acoustic features such as fundamental frequency, energy, speech rate, pause ratio, interruption rate, and response delay from audio data during interactions between medical students and virtual patients, enabling the capture of subtle vocal changes in medical students' communication. These acoustic features, as objective quantitative indicators, effectively reflect the emotional state, expression habits, and interaction patterns of medical students. Subsequently, by analyzing these acoustic features using pre-trained classification or regression models, the empathy and attitudinal expressions of medical students in communication can be accurately identified and quantified. For example, higher fundamental frequency and energy may be associated with emotional excitement, while longer response delays and higher pause ratios may reflect hesitation or contemplation. Through the comprehensive processing of these features by the model, complex vocal information can be transformed into assessable empathy and attitudinal indices, thereby overcoming the subjectivity and inconsistency of traditional manual assessments and providing technical support for the objective evaluation of medical students' communication performance.

[0062] The aforementioned technical solution enables a refined and objective assessment of medical students' empathy and attitude performance in communication. Specifically, by extracting multi-dimensional acoustic features and analyzing them in conjunction with a pre-trained intelligent model, the system can more comprehensively and accurately capture medical students' emotional expressions and professional attitudes during interactions, avoiding biases caused by single indicators or subjective judgments. Consequently, the resulting empathy and attitude performance indices possess higher reliability and discriminative power, providing a more solid data foundation for subsequent comprehensive scoring and feedback reports. This helps medical students better understand their strengths and weaknesses in communication, allowing for targeted improvements.

[0063] Preferably, step A303 may include: Non-verbal visual features were extracted from video data; these features included facial landmarks, head posture, and gaze direction of the medical student. Obtain camera calibration parameters; Based on camera calibration parameters, head posture, and gaze direction, it is determined whether the medical student and the virtual patient are in eye contact. The duration of eye contact between medical students and virtual patients was recorded. The total interaction time between medical students and virtual patients is obtained, and the ratio of the duration of interaction to the total interaction time is calculated to obtain the eye contact rate. Based on the movement amplitude of facial key points, the intensity of facial expression changes in medical students is calculated to obtain facial activity level.

[0064] Nonverbal visual features refer to visual information obtained by analyzing video data during interactions between medical students and virtual patients. While they do not directly contain verbal content, they reflect the students' emotions, attention, and interaction status. Specifically, facial landmarks can be understood as specific points on the student's face, such as the outlines of the eyes, nose, and mouth; these points can be used for subsequent expression analysis. Head posture refers to the spatial orientation of the student's head, usually represented by Euler angles (such as pitch, yaw, and roll). Gaze direction refers to the direction the student's eyes are looking, which can be estimated using eye-tracking technology or in conjunction with head posture.

[0065] Camera calibration parameters refer to the set of parameters used to describe the camera's internal geometry and external position and orientation. These include internal parameters such as focal length, principal point coordinates, and distortion coefficients, as well as external parameters such as the camera's rotation and translation matrices relative to the world coordinate system. The purpose of obtaining these parameters is to accurately transform visual information from the image coordinate system to the three-dimensional world coordinate system, thereby precisely calculating the medical student's head posture and gaze direction, and determining whether there is visual contact between the student and the virtual patient's preset position.

[0066] To determine whether a medical student and a virtual patient are in eye contact, the acquired camera calibration parameters, the student's head posture, and gaze direction can be used. For example, a virtual patient's position or area in three-dimensional space can be defined, and then it can be determined whether the student's gaze direction falls within that area, or whether the student's head posture is facing that area, and within a certain tolerable angle range.

[0067] The calculation of eye contact rate involves statistically analyzing the duration of eye contact. Specifically, throughout the interaction between the medical student and the virtual patient, the system continuously monitors and records the total duration of eye contact between the medical student and the virtual patient. This duration is then compared to the total interaction time, and the ratio between the two is used to quantify the medical student's eye contact rate.

[0068] Facial activity level is calculated by analyzing the range of motion of facial landmarks. The range of motion of these landmarks reflects the intensity of facial muscle activity in medical students, such as raising eyebrows, raising or lowering the corners of the mouth, etc. By quantifying these movements, the intensity of facial expression changes during interaction can be assessed, thus reflecting the richness and positivity of their emotional expression. For example, by weighting and synthesizing the range of motion of multiple landmarks, the system can ultimately derive a quantitative "facial activity level" index, with the weight of each landmark determined according to the organ region to which it belongs.

[0069] This application's solution, through refined extraction of nonverbal visual features and spatial positioning combined with camera calibration parameters, accurately determines the eye contact status between medical students and virtual patients and quantifies its duration, thereby objectively calculating the eye contact rate. Simultaneously, by analyzing the movement amplitude of facial key points, it captures subtle changes in the medical student's facial expressions, thus quantifying facial activity. These detailed steps ensure a comprehensive and accurate assessment of the medical student's nonverbal communication performance, providing reliable data support for subsequent comprehensive scoring.

[0070] The aforementioned technical solutions enable a refined assessment of medical students' nonverbal communication performance. Specifically, by extracting nonverbal visual features such as facial landmarks, head posture, and gaze direction, and combining this with precise spatial positioning using camera calibration parameters, the calculation of eye contact rate becomes more objective and accurate, avoiding potential misjudgments in traditional methods. Simultaneously, calculating facial expression activity based on the range of motion of facial landmarks allows for more sensitive capture of medical students' emotional changes and facial expressions, thus comprehensively reflecting their empathy and professional attitude in clinical interactions. The accurate acquisition of these quantitative indicators significantly improves the comprehensiveness and reliability of the assessment of medical students' overall clinical competence.

[0071] Preferably, before step A4, the following step is also included: A6. When the semantic coverage, speech recognition confidence, or raw data quality is lower than the corresponding preset threshold, low confidence assessment processing is triggered; low confidence assessment processing includes indicator downgrading or adjusting the virtual patient's speech; indicator downgrading is used to reduce the quantitative indicators used to calculate the comprehensive score.

[0072] Specifically, low-confidence assessment processing refers to a mechanism adopted by the system when the quality of certain key indicators or raw data fails to meet preset standards during the assessment process. Low semantic coverage (e.g., below a preset coverage threshold) indicates a low semantic match between the medical student's answer and the standard answer, or uncertainty in the semantic alignment process. Speech recognition confidence refers to the accuracy of converting medical student audio data into text. Low confidence (e.g., below a preset confidence threshold) leads to errors in the medical student's response text itself, thus affecting subsequent semantic alignment and communication performance assessment. Speech recognition confidence can be obtained from the speech recognition model (i.e., the deep learning-based acoustic and language models used in step A1, which output the confidence of the speech recognition results while performing speech recognition). The quality of raw data, such as video and audio data, may be compromised due to factors such as the acquisition environment and equipment performance; for example, blurry videos or noisy audio can directly affect the accuracy of feature extraction based on this data. For example, the quality of video data can be identified by calculating metrics such as peak signal-to-noise ratio (PSNR) and structural similarity index (SSIM); the quality of audio data can be identified by analyzing acoustic parameters such as signal-to-noise ratio (SNR) and total harmonic distortion (THD).

[0073] When any of the above conditions is determined to be below the corresponding preset threshold, the system will trigger low-confidence assessment processing. This processing can take two main forms: indicator downgrading or adjusting the virtual patient's dialogue. Indicator downgrading aims to reduce the impact of low-quality data or low-confidence results on the final score by reducing the quantitative indicators used to calculate the comprehensive score. For example, if the semantic coverage is too low, knowledge accuracy will not be considered when calculating the comprehensive score to avoid scoring bias caused by inaccurate semantic alignment. Similarly, when the confidence of speech recognition is too low, language clarity may be distorted due to calculations based on incorrect text, and can be removed. When the video data quality is low, eye contact rate and facial expression activity calculated based on video data may be inaccurate, and can therefore also be removed. As a preferred implementation, if indicator downgrading is triggered, the downgrading process will be clearly marked in the subsequently generated feedback report to inform the user of the special circumstances during the assessment process. Adjusting the virtual patient's dialogue refers to the system dynamically adjusting the virtual patient's questioning style or dialogue content based on the detected low-confidence conditions during the interaction. For example, if the confidence level of speech recognition remains low, the virtual patient can repeat the question, use simpler words, or ask the medical student to express themselves more clearly in an attempt to obtain higher quality speech input.

[0074] In practice, after triggering the low-confidence assessment, the virtual patient's script can be adjusted first. When the number of adjustments reaches a preset threshold (e.g., 2-3 times), the indicator is downgraded. This ensures that all quantitative indicators are considered as much as possible when performing the comprehensive scoring, improving the accuracy of the score. Furthermore, a correction coefficient can be generated based on the number of virtual patient script adjustments to correct the subsequent weighted fusion of the comprehensive score. For example, this correction coefficient is a value greater than 0 and less than or equal to 1. When the number of virtual patient script adjustments is 0, its value is 1; the more times the virtual patient script is adjusted, the smaller the correction coefficient.

[0075] This application's solution effectively addresses the issue of distorted assessment results that may arise from poor raw data quality or low confidence levels in intermediate processing steps during the comprehensive clinical competence assessment of medical students. Specifically, when the system detects that semantic coverage, speech recognition confidence, or raw data quality are below a preset threshold, it indicates that some inputs or intermediate results of the current assessment may be unreliable. In this case, by triggering low-confidence assessment processing, the system can take proactive measures. If indicator downgrading is selected, affected quantitative indicators are selectively removed, avoiding the inclusion of inaccurate or unreliable sub-item scores in the final comprehensive score, thus ensuring the overall robustness and fairness of the comprehensive score. For example, when speech recognition confidence is low, language clarity calculated based on erroneous text will no longer affect the final score. If adjusting the virtual patient's dialogue is selected, real-time intervention in the interaction process guides medical students to provide higher-quality input, improving data quality from the source and thus enhancing the accuracy of subsequent assessments. Both processing methods aim to reduce the negative impact of uncertainty on assessment results, enabling the final comprehensive score to more accurately reflect the clinical competence of medical students.

[0076] Through the aforementioned technical solution, this application significantly improves the robustness and reliability of the comprehensive clinical competence assessment method for medical students. Even when faced with suboptimal raw data quality or low confidence levels in intermediate processing results, the method effectively avoids assessment bias caused by data defects through intelligent low-confidence assessment processing. This not only ensures the accuracy and fairness of the comprehensive score but also increases the transparency of the assessment process by marking downgraded indicators in the feedback report. Furthermore, by adjusting the virtual patient's dialogue, this application can optimize the data collection process in real time, fundamentally improving assessment quality, thereby providing medical students with more accurate and instructive feedback and a more reliable basis for the development of teaching strategies.

[0077] In some preferred embodiments, step A4 uses a regression model trained based on gold standard data cross-validation to calculate a comprehensive score; The regression model is: TCS = a*KAS + b*LCI + c*EPI + d*API + e*ECR + f*FEI; Among them, TCS is the overall score, KAS is the knowledge accuracy, LCI is the language clarity, EPI is the empathy performance index, API is the attitude performance index, ECR is the eye contact rate, FEI is the facial expression activity, and a, b, c, d, e, and f are weighting factors.

[0078] Specifically, gold standard data refers to data generated by independent and authoritative assessments of medical students' clinical performance by experienced clinical experts or assessors, such as review data from the Objective Structured Clinical Examination (OSCE) or the mini-CEX. This data is considered the "gold standard" for measuring medical students' clinical competence, providing highly reliable and effective assessment results. Cross-validation is a statistical method used to evaluate the generalization ability of a model. It involves dividing the dataset into training and validation sets, repeatedly training and testing the model to ensure it performs well on unseen data. A regression model is a predictive model used to analyze the relationship between one or more independent variables (in this case, quantitative indicators) and a dependent variable (in this case, a composite score), and to predict the value of the dependent variable.

[0079] In this framework, TCS represents the final overall score, KAS represents the medical student's knowledge accuracy, LCI represents verbal clarity, EPI represents the empathy performance index, API represents the attitude performance index, ECR represents eye contact rate, and FEI represents facial expression activity. These indicators, as input variables to the regression model, collectively determine the overall score TCS. The weighting factors a, b, c, d, e, and f are the coefficients of the regression model. They are automatically learned through cross-validation training on the gold standard data. (Specifically, communication performance indicators, including KAS, LCI, EPI, API, ECR, and FEI, are assessed based on the clinical performance data of medical students in the gold standard data; the specific assessment method is described above. Then, using the communication performance indicators as samples and the review results in the gold standard data as the labels for each sample, a sample dataset is constructed. Next, the sample dataset is divided into a training set and a validation set. The regression model is trained using the training set to iteratively optimize the coefficients of the regression model, resulting in a trained regression model. During training, samples are input into the regression model, and the deviation between the comprehensive score output by the regression model and the label of the sample is calculated as the loss function. This loss function is converged by optimizing the coefficients of the regression model. Finally, the trained regression model is validated using the validation set.) These weighting factors reflect the relative contribution of each quantitative indicator to the comprehensive clinical competence of medical students, thus making the calculation of the comprehensive score more objective and data-driven.

[0080] This application's solution addresses the potential objectivity and accuracy deficiencies of traditional weighted fusion methods by introducing a regression model trained using gold-standard data and cross-validation. Specifically, the gold-standard data provides a reliable and authoritative evaluation benchmark, ensuring the effectiveness of model training. The cross-validation mechanism guarantees the generalization ability and stability of the regression model across different datasets, preventing overfitting and thus making the calculated comprehensive score more reliable. The regression model learns and quantifies the complex relationship between various quantitative indicators (such as Knowledge Accuracy (KAS), Language Clarity (LCI), Empathy Index (EPI), Attitude Index (API), Eye Contact Rate (ECR), and Facial Expression Index (FEI)) and medical students' comprehensive clinical competence, automatically determining the optimal weighting factors a, b, c, d, e, and f. Therefore, the calculation of the comprehensive score (TCS) no longer relies on subjective experience but is based on data-driven objective analysis, making the evaluation results more scientific and accurate.

[0081] Through the above technical solution, this application can significantly improve the objectivity, accuracy, and reliability of the comprehensive clinical competence assessment of medical students. Since the weighting factors are obtained through cross-validation training on gold standard data, the calculated comprehensive score can more realistically and accurately reflect the actual clinical performance of medical students, and is highly consistent with expert assessment results. This not only reduces subjective bias in the assessment process but also enhances the credibility of the assessment results, providing more precise guidance for the training of medical students and the improvement of their clinical abilities.

[0082] In some embodiments, the method for assessing the comprehensive clinical competence of medical students further includes the step of: A7. Based on the comprehensive score and various quantitative indicators, generate feedback information that includes information on the weaknesses of medical students and suggestions for teaching strategies.

[0083] Specifically, the weakness information refers to the identification of specific deficiencies in medical students' knowledge accuracy, language clarity, empathy index, attitude index, eye contact rate, and facial expression activity through in-depth analysis of their comprehensive scores and various quantitative indicators. For example, if a medical student's knowledge accuracy is low or their empathy index is unsatisfactory (determined by comparing it with a corresponding preset threshold) during interaction with a virtual patient, these aspects will be identified and marked as weaknesses by the system. The teaching strategy suggestions are generated based on the identified weakness information and aim to provide teachers with specific teaching guidance. These suggestions can be generated based on preset rules. For example, for the weakness of low knowledge accuracy, the system can suggest that teachers provide additional medical knowledge supplementary materials, arrange special discussions, or conduct case analyses; for the weakness of insufficient empathy, the system can suggest that teachers organize role-playing exercises, scenario simulation training, or provide communication skills guidance. These suggestions aim to help teachers more effectively improve medical students' clinical competence.

[0084] Through the aforementioned technical solution, this application can provide teachers with more targeted and actionable feedback, significantly improving teaching efficiency and quality. Teachers can quickly identify areas where medical students need improvement based on the weaknesses clearly pointed out in the feedback, and formulate personalized teaching plans and interventions according to the teaching strategy suggestions provided by the system. This not only helps medical students more effectively address their shortcomings and improve their overall clinical competence, but also greatly reduces the workload of teachers in assessment and teaching plan development, achieving a deep integration of assessment and teaching, thereby promoting the comprehensive development of medical students' clinical skills.

[0085] refer to Figure 2 , Figure 3 This application provides a system for assessing the comprehensive clinical competence of medical students, including a student terminal 1, a teacher terminal 2, and a server 3; Student terminal 1 is used to provide medical students with an interactive interface to interact with virtual patients, and to collect video and audio data during the interaction between medical students and virtual patients and upload them to server 3; Server 3 is configured with: The speech recognition module 301 is used to perform speech recognition on audio data to obtain the medical student's response text (for details, please refer to step A1 above). The knowledge accuracy assessment module 302 is used to semantically align the medical student's response text with an authoritative medical knowledge base based on retrieval enhancement generation technology, and calculate the knowledge accuracy based on the alignment results (for details, please refer to step A2 above). The communication performance assessment module 303 is used to evaluate the communication performance of medical students based on video data, audio data, and medical students' response texts to obtain communication performance indicators. The communication performance indicators include language clarity, empathy performance index, attitude performance index, eye contact rate, and facial expression activity (for details, please refer to step A3 above). The comprehensive scoring module 304 is used to weight and integrate various quantitative indicators to obtain a comprehensive score; the quantitative indicators include knowledge accuracy and various communication performance indicators (for details, please refer to step A4 above). The first feedback module 305 is used to generate a feedback report containing dimensional scores, time-series positioning, and learning suggestions based on the comprehensive score and various quantitative indicators, and to encrypt and store the raw data, as well as send the feedback report to the student's end; the raw data includes video data and audio data (for details, please refer to step A5 above). The second feedback module 307 is used to generate feedback information containing medical students' weaknesses and teaching strategy suggestions based on the comprehensive score and various quantitative indicators, and send it to the teacher's end 2 (the specific process can be referred to step A7 above).

[0086] Student terminal 1 is configured to provide medical students with an interactive interface for interacting with a virtual patient. This interface can be a graphical user interface (GUI), such as a tablet, personal computer, or dedicated simulation device. Student terminal 1 runs a corresponding application to provide the interactive interface. Through this interface, medical students can converse and interact with the virtual patient, simulating real clinical scenarios. Simultaneously, student terminal 1 is also used to collect video and audio data during the interaction between the medical student and the virtual patient. Video data can be recorded using an integrated camera or an external high-definition camera to capture the medical student's nonverbal behavior; audio data can be recorded using an integrated microphone or an external high-sensitivity microphone to capture the medical student's voice information. The collected data is then uploaded to server 3 for further processing. For example, data can be automatically uploaded via network transmission protocols (such as HTTP / HTTPS, FTP), or manually triggered by the user after the interaction ends.

[0087] Server 3 is configured to process and analyze data uploaded from student client 1. Server 3 can be a physical server cluster or a virtual server instance deployed on a cloud computing platform to provide scalable computing and storage capabilities. Server 3 is internally configured with multiple functional modules that work together to achieve comprehensive evaluation.

[0088] Teacher terminal 2 can be a tablet computer, personal computer, or dedicated management device, mainly used to receive and view students' assessment results, teaching suggestions, and conduct teaching management.

[0089] In some implementations, server 3 is also configured with: The low confidence fallback control module 306 is used to trigger low confidence assessment processing when the semantic coverage, confidence of speech recognition, or quality of raw data is lower than the corresponding preset threshold. The low confidence assessment processing includes indicator downgrading or adjusting the virtual patient's speech. The indicator downgrading is used to reduce the quantitative indicators used to calculate the comprehensive score (the specific process can be referred to step A6 above).

[0090] The above description is merely an embodiment of this application and is not intended to limit the scope of protection of this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the scope of protection of this application.

Claims

1. A method for assessing the comprehensive clinical competence of medical students, characterized in that, The method includes the following steps: A1. Collect video and audio data during the interaction between medical students and virtual patients, and perform speech recognition on the audio data to obtain the medical student's response text; A2. Based on retrieval enhancement generation technology, the medical student's response text is semantically aligned with an authoritative medical knowledge base, and the knowledge accuracy is calculated based on the alignment results; A3. Based on the video data, the audio data, and the medical student's response text, the medical student's communication performance is evaluated to obtain communication performance indicators; the communication performance indicators include language clarity, empathy index, attitude index, eye contact rate, and facial expression activity. A4. Weight and integrate the various quantitative indicators to obtain a comprehensive score; the quantitative indicators include the knowledge accuracy and the various communication performance indicators. A5. Based on the comprehensive score and the various quantitative indicators, generate a feedback report containing dimensional scores, temporal positioning, and learning suggestions, and encrypt and store the original data; the original data includes the video data and the audio data.

2. The method for assessing the comprehensive clinical competence of medical students according to claim 1, characterized in that, Step A2 includes: A201. Obtain dialogue context information; the dialogue context information includes the virtual patient's question type information and the medical student's department type information; A202. Perform named entity recognition and key point extraction on the medical student's response text to obtain the medical entities and key points in the medical student's response text; A203. Based on the dialogue context information, the medical entity, and the key points, using retrieval enhancement generation technology, retrieve the Top-K most relevant authoritative documents from the authoritative medical knowledge base, and extract the key knowledge points from the retrieved authoritative documents; Top-K is a preset positive integer; A204. Compare the key points with the knowledge points, and calculate the coverage of the medical student's response text with the knowledge points and the matching degree of the key points, respectively, to obtain the semantic coverage rate and F1 score; A205. The semantic coverage and the F1 score are weighted and fused to obtain the knowledge accuracy.

3. The method for assessing the comprehensive clinical competence of medical students according to claim 2, characterized in that, Step A204 includes: Obtain the semantic vectors of the key points and the knowledge points; Calculate the semantic similarity between the semantic vectors of each key point and the semantic vectors of each knowledge point; For each key point, the matching status of the key point is determined based on the semantic similarity between the semantic vector of the key point and the semantic vector of each knowledge point, as well as a preset similarity threshold. The F1 score is calculated by combining the matching status of each of the key points. For each knowledge point, determine whether there is a key point in the medical student's response text that has a semantic similarity of not less than the preset similarity threshold with the knowledge point, and determine the coverage status of the knowledge point based on the judgment result; The semantic coverage rate is calculated based on the coverage status of each knowledge point.

4. The method for assessing the comprehensive clinical competence of medical students according to claim 1, characterized in that, Step A3 includes: A301. Calculate the language clarity based on the medical student's response text; A302. Based on the audio data, extract acoustic features, and calculate the empathy performance index and attitude performance index based on the acoustic features; A303. Based on the video data, extract non-verbal visual features, and calculate eye contact rate and facial expression activity based on the non-verbal visual features.

5. The method for assessing the comprehensive clinical competence of medical students according to claim 4, characterized in that, Step A301 includes: The medical student's response text was subjected to terminology recognition, syntactic analysis, and syntactic structure analysis to obtain the terminology ratio, average sentence length, and clause nesting ratio. The terminology ratio, the average sentence length, and the clause nesting ratio are standardized to obtain standardized terminology ratio, standardized average sentence length, and standardized clause nesting ratio. The standardized terminology ratio, the standardized average sentence length, and the standardized clause nesting ratio are weighted and fused to obtain the language clarity. When the term ratio exceeds a preset ratio threshold or the average sentence length exceeds a preset length threshold, a language expression improvement suggestion is generated.

6. The method for assessing the comprehensive clinical competence of medical students according to claim 4, characterized in that, Step A302 includes: Based on the audio data, acoustic features are extracted; the acoustic features include at least one of fundamental frequency, energy, speech rate, pause ratio, interruption rate, and response delay; Based on the acoustic features, the empathy index and the attitude index are calculated using a pre-trained classification or regression model.

7. The method for assessing the comprehensive clinical competence of medical students according to claim 4, characterized in that, Step A303 includes: Non-verbal visual features are extracted from the video data; the non-verbal visual features include the medical student's facial landmarks, head posture, and gaze direction. Obtain camera calibration parameters; Based on the camera calibration parameters, the head posture, and the gaze direction, it is determined whether the medical student and the virtual patient are in eye contact. The duration of eye contact between the medical student and the virtual patient was recorded. The total interaction time between the medical student and the virtual patient is obtained, and the ratio of the duration to the total interaction time is calculated to obtain the eye contact rate. Based on the movement amplitude of the facial key points, the intensity of the medical student's facial expression changes is calculated to obtain the facial expression activity level.

8. The method for assessing the comprehensive clinical competence of medical students according to claim 2, characterized in that, Before step A4, the following steps are also included: A6. When the semantic coverage, the confidence level of speech recognition, or the quality of the raw data is lower than the corresponding preset threshold, a low confidence assessment process is triggered; the low confidence assessment process includes indicator downgrading or adjusting the virtual patient's speech. The indicator downgrading process is used to reduce the number of quantitative indicators used in calculating the overall score.

9. The method for assessing the comprehensive clinical competence of medical students according to claim 1, characterized in that, In step A4, the comprehensive score is calculated using a regression model trained based on gold standard data cross-validation. The regression model is: TCS = a*KAS + b*LCI + c*EPI + d*API + e*ECR + f*FEI; Wherein, TCS is the overall score, KAS is the knowledge accuracy, LCI is the language clarity, EPI is the empathy performance index, API is the attitude performance index, ECR is the eye contact rate, FEI is the facial expression activity, and a, b, c, d, e, and f are weighting factors.

10. A system for assessing the comprehensive clinical competence of medical students, characterized in that, Includes student client, teacher client, and server; The student terminal is used to provide medical students with an interactive interface to interact with virtual patients, and to collect video and audio data during the interaction between medical students and virtual patients, and upload them to the server. The server is configured with: The speech recognition module is used to perform speech recognition on the audio data to obtain the medical student's response text; The knowledge accuracy assessment module is used to semantically align the medical student's response text with an authoritative medical knowledge base based on retrieval enhancement generation technology, and calculate the knowledge accuracy based on the alignment results. The communication performance evaluation module is used to evaluate the communication performance of the medical students based on the video data, the audio data, and the medical students' response text, and obtain communication performance indicators; the communication performance indicators include language clarity, empathy performance index, attitude performance index, eye contact rate, and facial expression activity. The comprehensive scoring module is used to weight and integrate various quantitative indicators to obtain a comprehensive score; the quantitative indicators include the knowledge accuracy and various communication performance indicators. The first feedback module is used to generate a feedback report containing dimensional scores, temporal positioning, and learning suggestions based on the comprehensive score and the various quantitative indicators, encrypt and store the original data, and send the feedback report to the student's terminal; the original data includes the video data and the audio data; The second feedback module is used to generate feedback information containing the medical student's weaknesses and teaching strategy suggestions based on the comprehensive score and the various quantitative indicators, and send it to the teacher's end.