Multi-agent debate medical student ability assessment method and system
Through the multi-agent debate system, the role-giving examiner, thinking chain strategy and feedback mechanism are used to solve the problem of lack of interactivity and insufficient feedback mechanism of the existing medical student ability assessment methods, and a comprehensive and objective assessment of medical students' clinical thinking and decision-making abilities is achieved.
Patent Information
- Application Number
- CN202411867161.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-18
- Publication Date
- 2025-05-06
AI Technical Summary
The existing medical students' ability assessment methods lack interactivity and insufficient feedback mechanisms, so they cannot comprehensively evaluate medical students' clinical thinking and decision-making abilities.
The multi-agent debate method is adopted to give roles to the intelligent examiner, thinking chain strategy and feedback mechanism, and build a multi-agent system for medical students' ability assessment. The system includes a role-assignment module, an assessment module and a comprehensive evaluation module. Through interaction and feedback between multiple agents, the ability of medical students is evaluated.
It has achieved a comprehensive and objective assessment of medical students' clinical thinking and decision-making abilities, improved the fairness and effectiveness of the assessment, and can dynamically adjust the difficulty and type of questions, and evaluated medical students' logical reasoning and problem-solving abilities.
Smart Images

Figure CN119941006A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of intelligent assessment, and in particular to a method and system for assessing the ability of medical students in multi-agent debate. Background Art
[0002] In the field of medical education, the assessment of medical students' abilities is an important part of ensuring the quality of medical services. Traditional assessment methods mostly rely on standardized tests, clinical simulations, and tutor evaluations. Although these methods can assess the knowledge level and clinical skills of medical students to a certain extent, they have some limitations. For example, standardized tests are difficult to fully reflect the clinical thinking and decision-making ability of medical students; although clinical simulations can simulate real scenarios, they are costly and difficult to implement on a large scale; relying on tutor evaluations is greatly affected by their professional fields and personal preferences, and the comprehensiveness and fairness of the evaluation results are difficult to guarantee.
[0003] In recent years, with the development of artificial intelligence technology, medical student ability assessment systems based on natural language processing (NLP) and multi-agent systems have gradually become a research hotspot. These systems can more comprehensively and objectively evaluate the clinical thinking and decision-making abilities of medical students by simulating real medical school assessment scenarios and combining multi-agent interaction and feedback mechanisms. For example, a method for generating clinical case analysis questions based on multi-round prompts (CN202410224629.8).
[0004] The core of this multi-round prompt method for generating medical record problems is to generate clinical case analysis problems through multiple rounds of interaction with the conversational large language model. It fills the input text into the corresponding position in the prompt word template to obtain prompt words that can enable the conversational large language model to complete some tasks better, thereby mining the knowledge examination points in the medical records through different prompt word instructions. The entire system starts from inputting electronic medical records to generating the final clinical case analysis questions.
[0005] However, this method has the following problems and disadvantages:
[0006] 1. Lack of interactivity: This method mainly relies on a one-way generation process and lacks interaction with medical students, making it impossible to evaluate the actual abilities of medical students in real time.
[0007] 2. Insufficient feedback mechanism: The generated questions and answers lack a feedback mechanism and cannot be dynamically adjusted according to the medical students’ responses.
[0008] 3. Lack of thinking chain: Although the generated questions have a certain degree of complexity, they lack a comprehensive assessment of the medical students' thinking process and cannot evaluate their logical reasoning and problem-solving abilities. Summary of the invention
[0009] The purpose of the present invention is to provide a method and system for assessing the ability of medical students based on multi-agent debate, so as to solve the problems raised in the above-mentioned background technology.
[0010] To achieve the above-mentioned object of the invention, one aspect of the present invention provides a method for assessing the ability of medical students in multi-agent debate, comprising the following steps:
[0011] Step S1, formulate several roles for assessing medical students' abilities according to the assessment content and assign them to the intelligent examiners, and define the weights of different examiners;
[0012] Step S2: Generate several assessment dimensions and weights of each assessment dimension for evaluating the ability of medical students through the thought chain strategy, combine the thought chain with the feedback mechanism, and the intelligent examiner flexibly selects different communication strategies to conduct dialogue-based assessments with medical students;
[0013] In step S3, combined with the thinking chain in step S2, the decomposed assessment points are evaluated and scored by multi-agent examiners, and then the overall score of the medical students is obtained by comprehensive weights, and a competency assessment report of the medical students is generated.
[0014] Further, step S1 includes the following steps:
[0015] Step S101, according to the needs of medical student assessment, select the judging role, and select the appropriate large language model according to the judging role to form a debate agent;
[0016] Step S102: According to the emphasis of the medical student assessment, the user adjusts the weights of different examiner roles in the assessment process.
[0017] Further, step S2 includes the following steps:
[0018] Step S201, setting a communication strategy between the examinee and the intelligent examiner;
[0019] Step S202, combining the communication strategy and the thinking chain strategy to conduct multi-agent assessment;
[0020] Step S203, based on the dynamic feedback mechanism of the medical examinees' questions and answers, the feedback from the previous round of assessment is analyzed and the evaluation strategy of the intelligent agent is automatically adjusted.
[0021] Furthermore, the communication strategy in step S201 includes asking questions one by one, asking questions simultaneously, and asking questions simultaneously plus a summarizer, wherein:
[0022] Asking questions one by one means that in each round of assessment, the intelligent examiners generate questions in the set order. The intelligent examiners who speak later will include the questions of all previous examiners in their chat history.
[0023] Simultaneous questioning allows the agent examiner to asynchronously generate responses in each round of assessment to eliminate the influence of the order of questioning;
[0024] Asking questions at the same time plus a summarizer has one more agent as a summarizer than asking questions at the same time. After the debate, the summarizer will summarize the previous test content and the candidates' answers, and replace the chat history of all agents with this.
[0025] Further, step S3 includes the following steps:
[0026] Step S301, the multi-agent organizes and summarizes the thought chain in step S2 to ensure that the assessment process of each sub-topic and the examinee's answer are recorded in detail; according to the thought chain strategy of step S2, the multi-agent examiner evaluates and scores; according to the decomposed assessment points, the multi-agent formulates detailed evaluation standards and scoring rules, and calculates the overall score, where the overall score calculation formula is expressed as follows:
[0027]
[0028] Where N is the number of agent examiners, M is the number of evaluation dimensions, and w i is the weight of the i-th agent examiner, indicating the importance of the agent role in the overall evaluation, w j is the weight of the jth evaluation dimension, indicating the importance of this evaluation dimension in the overall evaluation; s ij Give the score of the ith agent on the jth evaluation dimension. To ensure the rationality of all weights and scores, the following conditions must be met:
[0029] The sum of the weights of all agent examiners is equal to 1, that is,
[0030] The sum of the weights of all evaluation dimensions is also equal to 1, that is,
[0031] The score of each evaluation dimension ij Between 0 and 10, or adjusted according to specific evaluation criteria;
[0032] Step S302, integrate and generate an evaluation report, integrating the discussion results.
[0033] Further, step S302 includes the following steps:
[0034] Step S321, the LLM agent in the role of "meeting minutes summarizer" summarizes the final evaluation and feedback of all agent examiners, including the identification, analysis and scoring of potential errors in the candidate's answer by each agent;
[0035] Step S322, classifying the identified errors, the error categories include medical diagnosis errors, medical knowledge errors, and attitude problems towards patients;
[0036] Step S323, providing a detailed description for each identified error, including the location of the error, why it is determined to be an error, and possible improvement suggestions;
[0037] Step S324: Count the scores of each agent in each assessment dimension according to the preset standard. ij , calculate the overall score of the medical examinee;
[0038] Step S325, using all the information to form a comprehensive evaluation report, the report content includes: the type and specific location of the deduction point, the error description and explanation, each agent examiner's sub-scores in each scoring dimension ij , the candidate's overall score.
[0039] Furthermore, the report formats include question-and-answer format and text report. The question-and-answer format is used to evaluate the relevance scenario, and the text report is used as a reference for candidates to improve.
[0040] Furthermore, the thinking chain sub-topics include medical knowledge proficiency, clinical event handling, and medical ethics, among which medical knowledge proficiency includes basic medical knowledge, clinical medical knowledge and the latest medical advances; clinical event handling includes the diagnosis and treatment of common diseases, first aid skills and patient communication skills; medical ethics includes doctor-patient relationship, medical ethics principles and laws and regulations.
[0041] Another aspect of the present invention provides a medical student ability assessment system for multi-agent debate, comprising a role assignment module, an assessment module, and a comprehensive evaluation module, wherein;
[0042] The role assignment module formulates several roles for assessing the medical students' abilities according to the assessment content and assigns them to the intelligent examiners, and defines the weights of different examiners;
[0043] The assessment module uses the thinking chain strategy to generate several assessment dimensions and the weight of each assessment dimension to evaluate the medical students' abilities. By combining the thinking chain with the feedback mechanism, the intelligent examiner can flexibly select different communication strategies to conduct dialogue-based assessments with medical students.
[0044] The comprehensive evaluation module combines the thinking chain in step S2, evaluates and scores the decomposed assessment points by multiple intelligent examiners, and then obtains the overall score of the medical students based on the comprehensive weights, and generates a medical student ability assessment report.
[0045] Compared with the prior art, the present system and method have the following advantages:
[0046] The present invention adopts a multi-agent system and can construct multiple agents, each of which is responsible for a different role, such as a medical school professor, a clinical teaching doctor, a patient, etc., and introduces a thinking chain mechanism to comprehensively evaluate its logical reasoning and problem-solving ability.
[0047] The present invention adopts a feedback mechanism to dynamically adjust the difficulty and type of questions according to the answers of medical students, thereby ensuring the fairness and effectiveness of the assessment.
[0048] The present invention adopts interactive debate, introduces interactive debate among multiple agents, simulates real clinical scenarios, and can evaluate the communication and teamwork abilities of medical students. BRIEF DESCRIPTION OF THE DRAWINGS
[0049] Figure 1 Flowchart of the method for assessing medical students' abilities based on multi-agent debate.
[0050] Figure 2 This is an assessment and evaluation diagram that combines the thinking chain and feedback mechanism.
[0051] Figure 3 Schematic diagram of the communication strategy principle. DETAILED DESCRIPTION
[0052] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0053] like Figure 1 The flowchart of the method of the present invention is shown. The embodiment of the present invention provides a method for assessing the ability of medical students based on multi-agent debates based on thought chains and feedback mechanisms. The specific steps are as follows:
[0054] Step S1: Set N examiners to evaluate the ability of medical students, that is, formulate N evaluation roles and assign them to the intelligent examiners, and define the weights w of different examiners. i .
[0055] Step S2: Generate M assessment dimensions and the weight w of each assessment dimension for evaluating medical students’ abilities through the thinking chain strategy. j Combining the thought chain with the feedback mechanism, N intelligent examiners flexibly select different communication strategies to conduct dialogue-style assessments with medical students.
[0056] Step S3: After the dialogue is over, combine the thought chain generated in step 2 to conduct multi-agent examiner evaluation and scoring on the decomposed assessment points to obtain s ij , calculate the total score according to the scoring formula Integrate and generate competency assessment reports for medical examinees.
[0057] Wherein, step 1 comprises the following steps:
[0058] Step S101, in the present invention, the debate agent is the core component, each agent is driven by an advanced artificial intelligence large language model (LLM), and is specifically responsible for generating responses to specific prompts. According to the needs of medical student assessment, users can select different judging roles, such as medical school professors, clinical teaching doctors, and patients. According to different roles, general large language models such as ChatGPT can be selected, or fine-tuned professional medical large models can be used.
[0059] Step S102, secondly, according to the emphasis of the medical student assessment, the user adjusts the weights of different examiner roles in the assessment process. For example, in the assessment of hospital regular trainees, the proportion of "medical school professor" agents and "clinical teaching doctor" agents is relatively high, and the proportion of "patients" is relatively low. The weights can be set to 0.4, 0.4, and 0.2.
[0060] Step 2 includes the following steps:
[0061] Step S201, setting a communication strategy. The different methods of maintaining and operating chat history records in this assessment system can be regarded as different ways of asking questions and assessments between candidates and AI agent examiners. Figure 3 As shown, this assessment system can use three different questioning strategies in the assessment:
[0062] a) Ask questions one by one: In each round of assessment, the intelligent examiner generates questions in the set order. The intelligent examiner who speaks later will include the questions of all previous examiners in his chat history.
[0063] b) Simultaneous questioning: Instead of asking questions one by one, this strategy allows the agent examiner to generate responses asynchronously in each round of assessment to eliminate the impact of the order of questions asked.
[0064] c) Simultaneous questioning plus summarizer: Similar to simultaneous questioning, but with an additional LLM as a summarizer. At the end of each round of debate, this summarizer will summarize the assessment content and candidates' answers so far, and replace the chat history of all agents with this summary.
[0065] For example, using the "simultaneous speaking plus summarizer" strategy, the agent will adjust its statement according to the prompts, and the summarizer will integrate the views of other agents. We can also combine chat strategies, such as conducting three rounds of assessments in sequence, and using the "simultaneous questioning plus summarizer" strategy in each round of assessment. This process helps each agent deepen its understanding of the problem and see the problem more comprehensively.
[0066] Step S202: Based on the communication strategy of step S201 and the thinking chain strategy, a multi-agent assessment is performed. Figure 2 As shown in the figure, the thinking chain strategy guides the agent to autonomously decompose the assessment process through prompts, and only focus on one sub-topic in each round of assessment and assign the assessment weight of the topic. For example, the assessment of medical students in regular training can be decomposed into sub-questions for specific fields, which may include medical knowledge proficiency, clinical event handling, medical ethics assessment, etc. After generating the thinking chain, the multi-agent examiner generates targeted scenarios and questions according to the sub-topics of each round of assessment.
[0067] In each round of multi-agent discussion, the agents will discuss and evaluate these sub-questions one by one. This method helps multi-agent examiners guide candidates to analyze the problem step by step, rather than considering all aspects at once, thereby improving the depth and accuracy of the assessment and achieving the purpose of medical education.
[0068] Step S203, using the feedback mechanism, according to the dynamic feedback mechanism of medical examinees' questions and answers, analyzes the feedback from the previous round of assessment, automatically adjusts the evaluation strategy of the intelligent agent, and achieves continuous optimization and adaptability of the evaluation of the medical question-and-answer model.
[0069] Step 3 includes the following steps:
[0070] Step S301: After the conversation is over, the multi-agent organizes and summarizes the thought chain recorded in the conversation process of step S2 to ensure that the assessment process of each sub-topic and the candidate's answer are recorded in detail. According to the thought chain strategy of step S2, the multi-agent examiner evaluates and scores. According to the decomposed assessment points, the multi-agent formulates detailed evaluation standards and scoring rules. The evaluation standards for each sub-topic should be specific and clear to ensure the fairness and objectivity of the evaluation.
[0071] The overall score calculation rules are as follows:
[0072]
[0073] Where: N is the number of agent examiners; M is the number of evaluation dimensions; w i is the weight of the i-th agent examiner, indicating the importance of the agent role in the overall evaluation; w j is the weight of the jth evaluation dimension, indicating the importance of this evaluation dimension in the overall evaluation; s ij is the score given by the i-th agent examiner on the j-th evaluation dimension.
[0074] To ensure that all weights and scores are reasonable, the following conditions need to be met:
[0075] The sum of the weights of all agents is equal to 1, that is
[0076] The sum of the weights of all evaluation dimensions is also equal to 1, that is,
[0077] The score of each evaluation dimension ij It is usually between 0 and 10, or adjusted according to specific evaluation criteria.
[0078] Step S302, integrating and generating an evaluation report, further comprising the following steps:
[0079] Step S321, integrate the discussion results. After multiple rounds of evaluation and feedback, a dedicated LLM agent with the role of "meeting minutes summarizer" is used to summarize the final evaluation and feedback of all agent examiners. This includes the identification, analysis and scoring of potential errors in the candidates' answers by each agent.
[0080] Step S322, error classification: Classify the identified errors, such as medical diagnosis errors, medical knowledge errors, attitude problems towards patients, etc.
[0081] Step S323: Detailed description: Provide a detailed description for each identified error, including the location of the error, why it is determined to be an error, and possible improvement suggestions.
[0082] Step S324: Count and calculate the scores. Count the scores of each agent in each assessment dimension according to the preset standard. ij . Calculate the overall score for the medical candidate.
[0083] Step S325: Generate a report. Use all the information to generate a comprehensive evaluation report. The report usually includes the following: the type and specific location of the deduction point; description and explanation of the error; the sub-scores of each agent examiner in each scoring dimension; ij The report can be in different formats to suit different application scenarios. For example, it can be a question and answer (Q&A) format to facilitate the evaluation of relevance; or a more detailed text report that can be issued by the medical school to candidates for self-improvement.
[0084] After these steps, the final evaluation report is output and provided to relevant medical school teachers or students for further analysis and evaluation.
[0085] In the specific implementation, after defining the assessment objectives, a thinking chain is generated through a large model multi-agent discussion. For example, the assessment objective is to evaluate the comprehensive ability of medical students in regular training. The sub-topics of the thinking chain include medical knowledge proficiency, clinical event handling, and medical ethics. Among them, medical knowledge proficiency includes basic medical knowledge, clinical medical knowledge, and the latest medical progress. Clinical event handling includes the diagnosis and treatment of common diseases, first aid skills, patient communication skills, etc. Medical ethics includes doctor-patient relationship, medical ethics principles, laws and regulations, etc.
[0086] Then, the multi-agent assessment is conducted by combining communication strategies, thought chains, and feedback mechanisms. The multi-agents use communication strategies to ask questions of students and conduct role-playing discussions to ensure that each agent can obtain the necessary information. The multi-agent examiner generates targeted scenarios and questions based on the sub-topics of each round of assessment. For example:
[0087] First round of assessment: medical knowledge proficiency
[0088] -Scenario: The patient has persistent high fever, cough and difficulty breathing, suspected pneumonia.
[0089] -Question: Please analyze the possible cause of this patient's illness and propose a preliminary diagnostic plan.
[0090] Second round of assessment: Clinical event handling:
[0091] -Scene: The patient suddenly fainted at home and his family called the emergency number.
[0092] -Question: Please describe the initial steps that should be taken at the first aid scene and explain the reasons for each step.
[0093] The third round of assessment: Medical ethics:
[0094] -Scenario: The patient refuses blood transfusion but is in critical condition.
[0095] -Question: Please explain how the doctor should handle this situation and give your reasons.
[0096] Based on the last round of questions and answers, the multi-agent examiner is given feedback through a mechanism, and the examiner dynamically adjusts questions and adds additional assessment points. For example:
[0097] In the first round of assessment, the candidates' analysis of the cause of the disease was relatively accurate, but the multi-agents collected feedback from this round, asked additional questions, and guided the candidates to supplement their analysis of laboratory test results in order to improve the accuracy of the diagnosis.
[0098] In the scoring process, assume that we have 2 agent examiners ((N=2)) and 3 assessment dimensions ((M=3)). After setting, the weights are as follows:
[0099] Weights of the agent examiner: (w_1=0.6), (w_2=0.4)
[0100] Weights of evaluation dimensions: (w_1=0.5), (w_2=0.3), (w_3=0.2)
[0101] The summarizer agent summarizes the thought chain recorded during the assessment process, ensures that the assessment process of each sub-topic and the candidate's answer are recorded in detail, and conducts multi-agent examiner evaluation and scoring. Based on the three assessment dimensions after decomposition, the two agent examiners discuss and formulate detailed evaluation standards and scoring rules, and give scores respectively:
[0102] The score given by the first agent on the first evaluation dimension is (s_{11}=8), and the reason is a;
[0103] The score given by the first agent on the second evaluation dimension is (s_{12}=7), and the reason is b;
[0104] The score given by the first agent on the third evaluation dimension is (s_{13}=6), and the reason is c;
[0105] The score given by the second agent on the first evaluation dimension is (s_{21}=7), and the reason is d;
[0106] The score given by the second agent on the second evaluation dimension is (s_{22}=8), and the reason is e;
[0107] The score given by the second agent on the third evaluation dimension is (s_{23}=9), and the reason is f;
[0108] Calculate the weighted total score:
[0109] Weighted total score = (0.6*0.5*8)+(0.6*0.3*7)+(0.6*0.2*6)+(0.4*
[0110] 0.5*7)+(0.4*0.3*8)+(0.4*0.2*9)=7.46
[0111] Therefore, the final score for this exam was 7.46.
[0112] Finally, a complete rating report is output by combining the scores and rating reasons.
[0113] Although embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions and variations may be made to the embodiments without departing from the principles and spirit of the present invention, and that the scope of the present invention is defined by the appended claims and their equivalents.
Claims
1. A method for assessing the ability of medical students based on multi-agent debate, characterized in that: The following steps are involved: Step S1, formulate several roles for assessing medical students' abilities according to the assessment content and assign them to the intelligent examiners, and define the weights of different examiners; Step S2: Generate several assessment dimensions and weights of each assessment dimension for evaluating the ability of medical students through the thought chain strategy, combine the thought chain with the feedback mechanism, and the intelligent examiner flexibly selects different communication strategies to conduct dialogue-based assessments with medical students; In step S3, combined with the thinking chain in step S2, the decomposed assessment points are evaluated and scored by multi-agent examiners, and then the overall score of the medical students is obtained by comprehensive weights, and a competency assessment report of the medical students is generated.
2. The method for assessing the ability of medical students through multi-agent debate according to claim 1, characterized in that: Step S1 includes the following steps: Step S101, according to the needs of medical student assessment, select the judging role, and select the appropriate large language model according to the judging role to form a debate agent; Step S102: According to the emphasis of the medical student assessment, the user adjusts the weights of different examiner roles in the assessment process.
3. The method for assessing the ability of medical students through multi-agent debate according to claim 1, characterized in that: Step S2 includes the following steps: Step S201, setting a communication strategy between the examinee and the intelligent examiner; Step S202, combining the communication strategy and the thinking chain strategy to conduct multi-agent assessment; Step S203, based on the dynamic feedback mechanism of the medical examinees' questions and answers, the feedback from the previous round of assessment is analyzed and the evaluation strategy of the intelligent agent is automatically adjusted.
4. The method for assessing the ability of medical students through multi-agent debate according to claim 3 is characterized in that: The communication strategies in step S201 include asking questions one by one, asking questions simultaneously, and asking questions simultaneously plus summarizing, wherein: Asking questions one by one means that in each round of assessment, the intelligent examiners generate questions in the set order. The intelligent examiners who speak later will include the questions of all previous examiners in their chat history. Simultaneous questioning allows the agent examiner to asynchronously generate responses in each round of assessment to eliminate the influence of the order of questioning; Asking questions at the same time plus a summarizer has one more agent as a summarizer than asking questions at the same time. After the debate, the summarizer will summarize the previous test content and the candidates' answers, and replace the chat history of all agents with this.
5. The method for assessing the ability of medical students through multi-agent debate according to claim 1, characterized in that: Step S3 includes the following steps: Step S301, the multi-agent organizes and summarizes the thought chain in step S2 to ensure that the assessment process of each sub-topic and the examinee's answer are recorded in detail; according to the thought chain strategy of step S2, the multi-agent examiner evaluates and scores; according to the decomposed assessment points, the multi-agent formulates detailed evaluation standards and scoring rules, and calculates the overall score, where the overall score calculation formula is expressed as follows: Where N is the number of agent examiners, M is the number of evaluation dimensions, and w i is the weight of the i-th agent examiner, indicating the importance of the agent role in the overall evaluation, w j is the weight of the jth evaluation dimension, indicating the importance of this evaluation dimension in the overall evaluation; s ij Give the score of the ith agent on the jth evaluation dimension. To ensure the rationality of all weights and scores, the following conditions must be met: The sum of the weights of all agent examiners is equal to 1, that is, The sum of the weights of all evaluation dimensions is also equal to 1, that is, The score of each evaluation dimension ij Between 0 and 10, or adjusted according to specific evaluation criteria; Step S302, integrate and generate an evaluation report, integrating the discussion results.
6. The method for assessing the ability of medical students through multi-agent debate according to claim 5, characterized in that: Step S302 includes the following steps: Step S321, the LLM agent in the role of "Meeting Minutes Summarizer" summarizes the final evaluation and feedback of all agent examiners, including the identification, analysis and scoring of potential errors in the candidates' answers by each agent; Step S322, classifying the identified errors, the error categories include medical diagnosis errors, medical knowledge errors, and attitude problems towards patients; Step S323, providing a detailed description for each identified error, including the location of the error, why it is determined to be an error, and possible improvement suggestions; Step S324: Count the scores of each agent in each assessment dimension according to the preset standard. ij , calculate the overall score of the medical examinee; Step S325, using all the information to form a comprehensive evaluation report, the report content includes: the type and specific location of the deduction point, the error description and explanation, each agent examiner's sub-scores in each scoring dimension ij , the candidate's overall score.
7. The method for assessing the ability of medical students through multi-agent debate according to claim 1, characterized in that: The report format includes question-and-answer format and text report. The question-and-answer format is used to evaluate relevant scenarios, and the text report is used as a reference for candidates to improve.
8. The method for assessing the ability of medical students through multi-agent debate according to claim 1, characterized in that: The thinking chain topics include medical knowledge proficiency, clinical event handling, and medical ethics. Medical knowledge proficiency includes basic medical knowledge, clinical medical knowledge and the latest medical advances; clinical event handling includes the diagnosis and treatment of common diseases, first aid skills and patient communication skills; medical ethics includes doctor-patient relationship, medical ethics principles and laws and regulations.
9. A multi-agent debate system for assessing the ability of medical students, characterized by: It includes role assignment module, assessment module and comprehensive evaluation module, among which; The role assignment module formulates several roles for assessing the medical students' abilities according to the assessment content and assigns them to the intelligent examiners, and defines the weights of different examiners; The assessment module uses the thinking chain strategy to generate several assessment dimensions and the weight of each assessment dimension to evaluate the medical students' abilities. By combining the thinking chain with the feedback mechanism, the intelligent examiner can flexibly select different communication strategies to conduct dialogue-based assessments with medical students. The comprehensive evaluation module combines the thinking chain in step S2, evaluates and scores the decomposed assessment points by multiple intelligent examiners, and then obtains the overall score of the medical students based on the comprehensive weights, and generates a competency assessment report for the medical students.
Citation Information
Patent Citations
Clinical case analysis problem generation method based on multiple rounds of prompts
CN118014083A
Cited By
Operation ticket auditing rule self-learning method, operation ticket auditing method, operation ticket auditing system and related equipment
CN120494469A
Hospital student comprehensive ability assessment method and system based on multi-source data
CN121095032A
Phishing website detection method, system and equipment based on multi-role agent debate
CN121125303A