Method, device and storage medium for evaluating disease diagnosis results
This method, which generates disease diagnosis results assessment by using a large language model to generate multiple diagnostic roles and dimensions, solves the problem of low accuracy in disease diagnosis results in existing technologies, achieves higher assessment accuracy and generalization, and is adaptable to complex medical scenarios.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- TSINGHUA UNIVERSITY
- Filing Date
- 2024-12-23
- Publication Date
- 2026-04-17
AI Technical Summary
The accuracy of disease diagnosis results in existing technologies is not high, and the generalization is poor. It is easy to miss information in medical records and it is difficult to adapt to complex and ever-changing medical scenarios.
The large language model is used to evaluate the diagnosis results of the disease. By generating different diagnostic roles and dimensions, and based on the diagnostic analysis report corresponding to the diagnostic role, the target evaluation result of the diagnosis result of the disease to be evaluated is determined. A multi-round voting mechanism and a comprehensive evaluation report are adopted to improve the accuracy and generalization of the evaluation results.
By mining rich medical knowledge through large language models, and avoiding the use of fixed review logic and rule templates, the accuracy and adaptability of disease diagnosis results are improved, making it suitable for complex and ever-changing medical scenarios.
Smart Images

Figure CN119864148B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of artificial intelligence technology, and in particular to a method, apparatus, device, and storage medium for evaluating the results of a medical diagnosis. Background Technology
[0002] In the modern medical system, diagnostic review is a key link in ensuring the quality and accuracy of medical care. During the clinical diagnostic review process, problems such as multiple diagnoses are often found in medical records. This not only affects the patient's treatment outcome, but also leads to unnecessary medical expenses.
[0003] Currently, to address the issue of multiple diagnostic tests, it is usually necessary to predefine some review logic and rule templates, extract key content from the patient's admission records and diagnostic texts, compare this key content with the review logic and rule templates to evaluate the diagnosis results, and then output the final evaluation results.
[0004] However, the above-mentioned method of evaluating disease diagnosis results based on review logic and rule templates has poor generalization and is prone to omitting information in medical records, resulting in low accuracy of evaluation results. Summary of the Invention
[0005] This invention provides a method, apparatus, device, and storage medium for evaluating disease diagnosis results, in order to address the shortcomings of low accuracy in the evaluation of disease diagnosis results in the prior art, and to achieve the goal of improving the accuracy of disease diagnosis results evaluation.
[0006] This invention provides a method for evaluating the results of a disease diagnosis, comprising:
[0007] Obtain at least one diagnosis of the condition to be evaluated based on the target medical record;
[0008] For each of the diagnostic results of the condition to be evaluated, the diagnostic results of the condition to be evaluated and the target medical record are input into the large language model to obtain at least two diagnostic roles and diagnostic dimensions of each diagnostic role output by the large language model;
[0009] For each diagnostic role, based on the diagnostic role, the diagnostic dimension corresponding to the diagnostic role, the diagnostic result of the condition to be evaluated, and the target medical record, a first prompting information is determined, and the first prompting information is input into the large language model to obtain the diagnostic analysis report corresponding to the diagnostic role output by the large language model;
[0010] Based on the diagnostic analysis reports corresponding to each of the diagnostic roles, a target evaluation result is determined for the diagnostic result of the condition to be evaluated. The target evaluation result is used to characterize whether the diagnostic result of the condition to be evaluated is correct.
[0011] According to a method for evaluating the diagnostic results of a disease according to the present invention, determining the target evaluation result of the diagnostic result to be evaluated based on the diagnostic analysis report corresponding to each diagnostic role includes:
[0012] Based on each of the diagnostic roles, the corresponding diagnostic analysis reports and diagnostic dimensions, the second prompt information is determined.
[0013] The second prompt information is input into the large language model to obtain the summary evaluation result output by the large language model. The summary evaluation result includes the diagnostic analysis report corresponding to all the diagnostic roles.
[0014] For each diagnostic role, based on the aggregated evaluation results, the voting information of the diagnostic role regarding the diagnostic result of the condition to be evaluated is determined, and the voting information includes the voting result of the diagnostic role regarding whether the diagnostic result of the condition to be evaluated is correct;
[0015] Based on the voting information of each diagnostic role, the target evaluation result of the diagnosis of the condition to be evaluated is determined.
[0016] According to a method for evaluating a disease diagnosis result provided by the present invention, determining the voting information of the diagnostic role regarding the disease diagnosis result to be evaluated based on the summarized evaluation results includes:
[0017] Based on the summarized assessment results and the diagnostic role, a third prompt message is determined;
[0018] The third prompt information is input into the large language model to obtain the voting information of the diagnostic role for the diagnosis result of the condition to be evaluated, which is output by the large language model.
[0019] According to the method for evaluating the results of a disease diagnosis provided by the present invention, the voting information also includes the reasons for voting;
[0020] The determination of the target assessment result for the diagnosis of the condition to be evaluated based on the voting information of each diagnostic role includes:
[0021] In the event of inconsistencies among all the voting results, a fourth prompt message is determined based on the voting results and reasons for each diagnostic role.
[0022] The fourth prompt information is input into the large language model to obtain the summary voting result output by the large language model. The summary voting result includes the voting information corresponding to all the diagnostic roles.
[0023] For each of the aforementioned diagnostic roles, a fifth prompt message is determined based on the aggregated voting results and the diagnostic role.
[0024] The fifth prompt information is input into the large language model to obtain the new voting information of the diagnostic role output by the large language model. The above steps are repeated until all new voting results are consistent, or the number of votes reaches the preset number.
[0025] Based on the final voting information, the target assessment result for the diagnosis of the condition to be evaluated is determined.
[0026] According to a method for evaluating the results of a medical condition diagnosis provided by the present invention, determining the target evaluation result of the medical condition diagnosis result to be evaluated based on the final voting information includes:
[0027] Based on the final voting information and the diagnosis results of the condition to be evaluated, the sixth prompt information is determined;
[0028] The sixth prompt information is input into the large language model to obtain a comprehensive evaluation report output by the large language model. The comprehensive evaluation report includes the target evaluation results and diagnostic opinions. The diagnostic opinions include the same diagnostic opinions and / or divergent diagnostic opinions of each of the diagnostic roles.
[0029] According to a method for evaluating the results of a disease diagnosis provided by the present invention, the method further includes:
[0030] If all the voting results are consistent, the voting result will be determined as the target evaluation result for the diagnosis of the condition to be evaluated.
[0031] According to the method for evaluating the results of a disease diagnosis provided by the present invention, when more than a first preset number of the disease diagnosis results to be evaluated belong to the same department, the at least two diagnostic roles include more than a second preset number of diagnostic roles belonging to the department.
[0032] The present invention also provides an assessment device for disease diagnosis results, comprising:
[0033] The acquisition module is used to acquire at least one diagnosis result of the condition to be evaluated based on the target medical record;
[0034] The input module is used to input the diagnosis results of the conditions to be evaluated and the target medical record into the large language model for each of the diagnosis results of the conditions to be evaluated, so as to obtain at least two diagnostic roles and diagnostic dimensions of each diagnostic role output by the large language model;
[0035] The determination module is used to determine the first prompt information for each of the diagnostic roles, based on the diagnostic role, the diagnostic dimension corresponding to the diagnostic role, the diagnostic result of the condition to be evaluated, and the target medical record;
[0036] The input module is also used to input the first prompt information into the large language model to obtain a diagnostic analysis report corresponding to the diagnostic role output by the large language model;
[0037] The determining module is further configured to determine the target evaluation result of the diagnosis result of the condition to be evaluated based on the diagnostic analysis report corresponding to each of the diagnostic roles, wherein the target evaluation result is used to characterize whether the diagnosis result of the condition to be evaluated is correct.
[0038] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the assessment method for the diagnostic results of the disease as described above.
[0039] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the method for evaluating the diagnostic results of a disease as described above.
[0040] The present invention also provides a computer program product, including a computer program that, when executed by a processor, implements an evaluation method for the diagnostic results of a disease as described above.
[0041] The present invention provides a method, apparatus, device, and storage medium for evaluating disease diagnosis results. This involves acquiring at least one disease diagnosis result to be evaluated based on a target medical record; inputting the disease diagnosis result and the target medical record into a large language model for each disease diagnosis result to be evaluated, thereby obtaining at least two diagnostic roles and diagnostic dimensions for each diagnostic role output by the large language model; determining a first prompt message for each diagnostic role based on the diagnostic role, the corresponding diagnostic dimension, the disease diagnosis result to be evaluated, and the target medical record, and inputting the first prompt message into the large language model to obtain a diagnostic analysis report corresponding to the diagnostic role output by the large language model; and determining a target evaluation result for the disease diagnosis result to be evaluated based on the diagnostic analysis report corresponding to each diagnostic role. This target evaluation result is used to characterize whether the disease diagnosis result to be evaluated is correct. Because large language models can be used to identify different diagnostic roles and their corresponding diagnostic dimensions, and then a diagnostic analysis report can be generated for each role based on the large language model and its corresponding diagnostic dimensions, the final target assessment result can be determined based on the diagnostic analysis reports of different roles. Large language models possess strong semantic understanding and processing capabilities, as well as the advantage of mining rich medical knowledge. Therefore, when using large language models for disease diagnosis review, fixed review logic and rule templates can be avoided, improving the generalization of the assessment method. Moreover, by fully mining useful information from the target case through large language models, it can adapt to complex and ever-changing medical scenarios, improving the accuracy of the assessment results. Attached Figure Description
[0042] To more clearly illustrate the technical solutions in this invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.
[0043] Figure 1 This is a flowchart illustrating the method for evaluating the diagnostic results of a disease as provided in an embodiment of the present invention.
[0044] Figure 2 A flowchart illustrating the method for evaluating disease diagnosis results provided in an embodiment of the present invention.
[0045] Figure 3 This is a schematic diagram of the structure of the device for evaluating the diagnosis results of a disease provided in an embodiment of the present invention.
[0046] Figure 4 This is a schematic diagram of the physical structure of an electronic device provided in an embodiment of the present invention. Detailed Implementation
[0047] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions of this invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this invention. All other embodiments obtained by those skilled in the art based on the embodiments of this invention without creative effort are within the scope of protection of this invention.
[0048] With the continuous advancement of medical technology, diagnostic techniques are playing an increasingly important role in improving diagnostic efficiency, accuracy, and patient experience. During clinical diagnostic review, it is frequently observed that medical records contain multiple diagnoses, which not only affects patient treatment outcomes but also leads to unnecessary medical expenses. Therefore, diagnostic review plays a crucial role in the healthcare system.
[0049] In existing technologies, when reviewing multiple medical records for diagnosis, medical experts typically need to pre-define review logic and rule templates. Medical record information, such as admission records and diagnostic texts, is then compared with these templates. After some yes / no logic judgments, an assessment result for the diagnosis is output. However, this method relies on specific review logic and rule templates, resulting in poor generalization. Furthermore, it depends on limited information from the cases, making it prone to overlooking details and difficult to adapt to complex and changing medical scenarios. Therefore, existing methods for assessing diagnosis results often lead to low accuracy.
[0050] To address the aforementioned issues, this invention provides a method for evaluating disease diagnosis results. This method utilizes a large language model to review disease diagnosis results. Because large language models possess strong semantic understanding and processing capabilities, as well as the advantage of mining rich medical knowledge, using them for disease diagnosis review avoids the use of fixed review logic and rule templates, improving the generalizability of the evaluation method. Furthermore, by fully mining useful information from cases using large language models, it can adapt to complex and ever-changing medical scenarios, improving the accuracy of the evaluation results.
[0051] The following is combined Figure 1 and Figure 2The method for evaluating disease diagnosis results provided in this embodiment of the invention is described below. The executing entity of this method can be an electronic device such as a terminal device, computer, or server, or a specially designed intelligent device for evaluating disease diagnosis results. It can also be a disease diagnosis result evaluation device installed in the electronic or intelligent device, which can be implemented through software, hardware, or a combination of both. This method can be applied in the medical field to scenarios where the correctness of disease diagnosis results is evaluated using a large language model, particularly in scenarios where a diagnosis result is considered a multiple diagnosis.
[0052] Figure 1 This is a flowchart illustrating the method for evaluating the diagnostic results of a disease provided in an embodiment of the present invention, as shown below. Figure 1 As shown, the method includes:
[0053] Step 101: Obtain at least one diagnosis of the condition to be evaluated based on the target medical record.
[0054] In this step, the target medical record may include the patient's personal information, examination results, and diagnostic information. Personal information includes the patient's age, gender, and past medical history; examination results include relevant laboratory results and imaging reports; and diagnostic information includes the disease name, symptom description, and clinical manifestations. By analyzing the target case, at least one diagnosis can be identified for evaluation, such as type 2 diabetes, hypertension, or chronic kidney disease.
[0055] Step 102: For each diagnosis result to be evaluated, input the diagnosis result and the target medical record into the large language model to obtain at least two diagnostic roles and the diagnostic dimensions of each diagnostic role output by the large language model.
[0056] In this step, each diagnostic result to be evaluated can be assessed according to the diagnostic result evaluation method provided in the embodiments of this invention to determine whether it is a duplicate diagnosis. In the task of reviewing duplicate diagnoses, since different diagnostic results may exist in different population groups and different departmental scenarios, and the conditions for confirming different diagnoses vary greatly, determining whether a diagnostic result is duplicated requires experts from different departments and specialties to make judgments and discussions on different review dimensions. Furthermore, determining the list of experts to be organized and the content to be discussed based on the dynamic diagnosis requires professional medical knowledge. Therefore, it is necessary to first determine the role of the moderator, and then the moderator analyzes the input personal information, examination results, and diagnostic information based on the key information in the target medical record, thereby dynamically generating at least two suitable diagnostic roles and the corresponding diagnostic dimensions for each role. The diagnostic roles can be virtual diagnostic experts.
[0057] Specifically, through cue word engineering, a large language model (LLM) can be guided to play the role of a medical conference host and identify key information, enabling it to have a professional understanding of medical conference organization. Examples of cue words are as follows: You are the main initiator and host of a medical conference on the theme of "Diagnosis Review and Approval." You understand the disease categories that multidisciplinary experts specialize in, and are proficient in disease diagnosis standards and medical record writing standards. For example, experts in endocrinology and nutrition are more proficient in diagnosing "diabetes," while experts in oncology, radiotherapy and chemotherapy, and hematology are more proficient in diagnosing "acute myeloid leukemia." For example, the diagnosis of "liver cysts" should be reviewed from the perspectives of past medical history and imaging examinations, and "hypokalemia" should be more concerned with "laboratory tests," etc.
[0058] Among them, keywords such as "medical conference", "disease diagnosis knowledge" and "multidisciplinary experts" in the prompt words can guide the large language model to fully activate its internal medical knowledge base, deeply analyze diagnosis and medical record related information, and make a general classification of the diagnosis.
[0059] In addition, key information can be extracted from target cases through preset rules or large language models, such as personal information, examination results, and diagnostic information.
[0060] For each diagnosis result to be evaluated, prompt words can be generated based on the diagnosis result and key information extracted from the target case. Through prompt word engineering, a large language model can be guided to generate diagnostic roles for different departments and specializing in different review dimensions. These diagnostic roles can be, for example, expert roles.
[0061] For example, assuming the diagnosis to be evaluated is type 2 diabetes with complications, the generated prompt could be: Given the following {diag}, medical record information {data}, to review whether the diagnosis has been overwritten, you need to generate a list of participating medical experts, such as specialists in internal medicine, surgery, and radiology, and determine the review dimensions for each expert, such as laboratory indicators, matching of imaging examinations, and medication regimens.
[0062] The input includes the patient's diagnosis (diag) and some medical record information (data), including:
[0063] Diagnosis pending evaluation: Type 2 diabetes mellitus with complications;
[0064] Partial medical record information includes: Other diagnostic information: hypertension, malnutrition, cholecystitis, mild pulmonary hypertension, chronic kidney disease; Department: Endocrinology; Age: 55 years old; Medical history: 10-year history of type 2 diabetes, 5-year history of coronary heart disease, no history of surgery; Laboratory test items: 2022.02.15 Biochemistry panel: Blood glucose: HbA1c 9.0%, Blood type: ABO blood type AB, Rh(D) blood type positive, stool routine + occult blood test showed no abnormalities.
[0065] The above information is input into the large language model. After analysis, the large language model can output the diagnostic role E and the diagnostic dimension AD for each diagnostic role. Examples of the diagnostic role E and the diagnostic dimension AD for each diagnostic role are as follows:
[0066] Diagnostic role set: {Endocrinologist 1, Endocrinologist 2, Nephrologist 1, ...};
[0067] Diagnostic dimensions include: {laboratory indicators, medication regimen, past medical history, ...}, among which endocrinologists should analyze from the perspective of reviewing laboratory test indicators, etc.
[0068] In summary, through prompt word engineering, large language models can be guided to generate expert roles from different departments and with expertise in different review dimensions, based on the diagnosis results of the disease to be evaluated and the information extracted from the target case.
[0069] Among them, the diagnostic roles and diagnostic dimensions of each diagnostic role can be dynamically generated based on the diagnosis results of the condition to be evaluated and the target medical record. This allows each diagnostic role to review the case from multiple diagnostic dimensions, ensuring that no information in the target case is missed and improving the comprehensiveness of the review.
[0070] Step 103: For each diagnostic role, based on the diagnostic role, the diagnostic dimension corresponding to the diagnostic role, the diagnostic result of the condition to be evaluated, and the target medical record, determine the first prompt information, and input the first prompt information into the large language model to obtain the diagnostic analysis report corresponding to the diagnostic role output by the large language model.
[0071] In this step, taking the diagnostic role as an expert role as an example, for each expert role, a first prompt message can be generated based on the diagnosis result of the condition to be evaluated, the target case, and the diagnostic dimension corresponding to that expert role. For example, the generated first prompt message could be: "You are an endocrinologist. The diagnosis result of the condition to be evaluated is 'Type 2 diabetes with complications.' Given the target medical record {data}, please review it based on the above information, paying particular attention to the diagnostic dimensions of laboratory indicators, medication regimen, and past medical history. Please provide your detailed review comments and information that needs further confirmation, and finally give a conclusion on whether it is overwritten or not." The target medical record data includes: "...Past medical history: ...10-year history of diabetes ...Daily course of disease: ...Laboratory tests: ...HbA1c 9.0% ...24-hour urine protein: 3g ...UAE: 300μg / min ...Medical orders: metformin, ...insulin, ...sacubitril / valsartan ...furosemide injection ...Billing details: ...insulin...".
[0072] After the generated initial prompt information is input into the large language model, the model analyzes the input information based on the expert's professional background, thereby generating a diagnostic analysis report for that expert role. This report may include the evaluation and analysis process, including analysis of each diagnostic dimension, further questions or clarifications regarding the analysis results, and the evaluation results themselves. An example of a diagnostic analysis report for an expert role is shown below:
[0073] Analysis Process: Laboratory Indicators: ① HbA1c reflects the average blood glucose level over the past 8 to 12 weeks and is the "gold standard" for assessing a patient's blood glucose control. The normal reference range for standard testing methods is 4% to 6%. The patient's HbA1c reached 9.0%, meeting the diagnostic criteria for diabetes. ② The patient's 24-hour urinary protein reached 3g, suggesting that the patient may have kidney disease.
[0074] Medication regimen: ① Insulin, a typical medication for diabetes, was found in both the doctor's orders and billing details. ② Metformin, a first-line drug for treating type 2 diabetes, was included in the patient's medical records, especially for overweight patients. The above analysis suggests that the patient has diabetes.
[0075] Past medical history: The patient's other diagnoses include chronic kidney disease, but it is uncertain whether it is a complication of diabetes.
[0076] Further questions or information requiring clarification: Through analysis of medical record information, the patient can be diagnosed with diabetes and kidney disease, but whether the kidney disease is a complication of diabetes requires comprehensive consideration of detailed past medical history, test results, etc.
[0077] Assessment result: There may be instances of extra or incorrect writing.
[0078] After obtaining the diagnostic analysis report corresponding to each expert role, each diagnostic analysis report will be stored in the system.
[0079] Step 104: Based on the diagnostic analysis reports corresponding to each diagnostic role, determine the target assessment result of the diagnosis of the condition to be evaluated. The target assessment result is used to characterize whether the diagnosis of the condition to be evaluated is correct.
[0080] In this step, after obtaining the diagnostic analysis report corresponding to each diagnostic role, the analysis process, further questions or information that needs clarification, and evaluation results in each diagnostic analysis report are analyzed. By comprehensively considering all diagnostic analysis reports, a target evaluation result for the diagnosis of the condition to be evaluated can be obtained to determine whether the diagnosis of the condition to be evaluated is correct, such as whether it is a case of over-written diagnosis.
[0081] In one possible implementation, the diagnostic analysis report corresponding to each diagnostic role can be input into a large language model. The large language model then analyzes each diagnostic analysis report to obtain the final target evaluation result. In another possible implementation, it can be determined from all diagnostic analysis reports corresponding to all diagnostic roles whether the number of identical evaluation results exceeds a preset number. If so, these identical evaluation results are used as the final target evaluation result.
[0082] The method for evaluating disease diagnosis results provided in this invention involves obtaining at least one disease diagnosis result to be evaluated based on a target medical record. For each disease diagnosis result to be evaluated, the disease diagnosis result to be evaluated and the target medical record are input into a large language model to obtain at least two diagnostic roles and diagnostic dimensions of each diagnostic role output by the large language model. For each diagnostic role, based on the diagnostic role, the diagnostic dimension corresponding to the diagnostic role, the disease diagnosis result to be evaluated, and the target medical record, a first prompt information is determined and input into the large language model to obtain a diagnostic analysis report corresponding to the diagnostic role output by the large language model. Based on the diagnostic analysis report corresponding to each diagnostic role, a target evaluation result for the disease diagnosis result to be evaluated is determined. This target evaluation result is used to characterize whether the disease diagnosis result to be evaluated is correct. Because large language models can be used to identify different diagnostic roles and their corresponding diagnostic dimensions, and then a diagnostic analysis report can be generated for each role based on the large language model and its corresponding diagnostic dimensions, the final target assessment result can be determined based on the diagnostic analysis reports of different roles. Large language models possess strong semantic understanding and processing capabilities, as well as the advantage of mining rich medical knowledge. Therefore, when using large language models for disease diagnosis review, fixed review logic and rule templates can be avoided, improving the generalization of the assessment method. Moreover, by fully mining useful information from the target case through large language models, it can adapt to complex and ever-changing medical scenarios, improving the accuracy of the assessment results.
[0083] For example, based on the above embodiments, when determining the target evaluation result of the diagnosis of the condition to be evaluated based on the diagnostic analysis reports corresponding to each diagnostic role, a second prompt information can be determined based on each diagnostic role, the diagnostic analysis reports corresponding to each diagnostic role, and the diagnostic dimensions. The second prompt information is then input into the large language model to obtain the summary evaluation result output by the large language model. The summary evaluation result includes the diagnostic analysis reports corresponding to all diagnostic roles. For each diagnostic role, the voting information of the diagnostic role regarding the diagnosis of the condition to be evaluated is determined based on the summary evaluation result. The voting information includes the voting results of the diagnostic role regarding whether the diagnosis of the condition to be evaluated is correct. Based on the voting information of each diagnostic role, the target evaluation result of the diagnosis of the condition to be evaluated is determined.
[0084] Specifically, after obtaining the diagnostic analysis report for each diagnostic role, all diagnostic analysis reports need to be summarized. Therefore, a second prompt message can be determined based on the diagnostic role, the diagnostic analysis report for each diagnostic role, and the diagnostic dimensions. This second prompt message could be, for example:
[0085] "You are the main initiator and moderator of a medical conference themed 'Review of Diagnostic Writing,' please summarize the review comments from the following expert roles: {report}"personal The process involves identifying points of agreement and disagreement among the expert opinions, generating preliminary summary assessment results based on diagnostic dimensions such as AD, and marking issues requiring further discussion.
[0086] Wherein, E represents at least one diagnostic role, i.e., an expert role, generated in the aforementioned embodiments; report personal This represents the diagnostic analysis report corresponding to each diagnostic role, and AD represents the diagnostic dimension corresponding to each diagnostic role.
[0087] After generating the second prompt information, it is input into the large language model. The large language model classifies and analyzes the diagnostic analysis reports based on similarity, summarizes the consensus, identifies areas of disagreement, and marks them as issues requiring further discussion. Based on the degree of influence and rationality of the opinions, the conflicts that need to be discussed in detail are ranked, thus forming a preliminary summary evaluation result. This summary evaluation result includes the diagnostic analysis reports corresponding to all diagnostic roles, that is, the main viewpoints and consensus of each diagnostic role, issues that need further clarification, and the overall conclusion regarding whether the diagnostic results for the condition to be evaluated are excessive.
[0088] The summary assessment results output by the large language model can be exemplified as follows: "Endocrinologist 1 believes it is uncertain whether there is overwriting, ... Laboratory indicators: ① HbA1c is the 'gold standard' for assessing a patient's blood glucose control, ... HbA1c reaching 9.0% confirms type 2 diabetes. ② The patient's 24-hour urinary protein reaches 3g, suggesting possible kidney disease. ③ UAE ... The patient's indicator reaches 300μg / min, confirming diabetic nephropathy. Medication regimen: ... Insulin and metformin are used, ... Furosemide injection is used for kidney disease, ... Past medical history: ... Issues requiring further clarification: ... Whether the patient's kidney disease is caused by diabetes is determined ... Overall conclusion: It is likely that there is no overwriting."
[0089] After generating the summary assessment results, a consensus-based voting mechanism can be used to form the final target assessment result. Each diagnostic role can determine its voting information based on the consensus reached among other diagnostic roles in the summary assessment results, the content with differing opinions from other diagnostic roles, and issues requiring further clarification. This voting information includes the diagnostic role's vote on whether the diagnosis of the condition to be assessed is correct.
[0090] For example, when determining the voting information of the diagnostic role for the diagnosis of the condition to be evaluated based on the summary evaluation results, a third prompt information can be determined based on the summary evaluation results and the diagnostic role, and the third prompt information can be input into the large language model to obtain the voting results of the large language model on whether the diagnostic role is correct for the diagnosis of the condition to be evaluated.
[0091] Specifically, for each diagnostic role, a third prompt message can be generated based on the summary assessment results and the diagnostic role. An example of this third prompt message is as follows: "Please refer to the summary assessment results {Report}..." interim Please vote on the given conclusions, taking into account the opinions of other experts, and choose to support or oppose them, explaining your reasons.
[0092] After generating the third prompt, it is input into the large language model. The large language model, by considering the opinions of this diagnostic role and those of other diagnostic roles, outputs the voting information of that diagnostic role regarding the diagnosis of the condition to be evaluated. An example of the voting information output by the large language model is as follows: "Endocrinologist 1: Support. Based on the summary assessment results, I ignored the UAE indicator, which is the gold standard for diabetic nephropathy… Endocrinologist 2: Dissent. Based on…, further clarification is needed. Nephrologist: Support. Based on…"
[0093] In the above method, the third prompt information generated based on the summary assessment results and diagnostic roles can be input into the large language model. The large language model can determine the voting information of each diagnostic role for the diagnosis result of the condition to be assessed. When voting, each diagnostic role will take into account the opinions and assessment results of other diagnostic roles, which can make the voting information more accurate.
[0094] After determining the voting information for each diagnostic role, the voting results in each voting information can be comprehensively considered to determine whether the target assessment result for the diagnosis result of the condition to be assessed should be written in excess.
[0095] In this embodiment, the diagnostic analysis reports corresponding to all diagnostic roles are comprehensively analyzed to obtain a summary evaluation result. Each diagnostic role can then vote based on this summary evaluation result, thereby determining the target evaluation result for the diagnosis of the condition to be evaluated based on the voting information of each diagnostic role. Because the voting information of all diagnostic roles can be comprehensively considered, the error caused by a single diagnostic role determining the evaluation result can be reduced, making the target evaluation result fairer and improving its accuracy.
[0096] For example, based on the above embodiments, the voting information also includes voting reasons. When determining the target evaluation result of the diagnosis of the condition to be evaluated based on the voting information of each diagnostic role, it can be achieved in the following way:
[0097] In the event of inconsistent voting results, a fourth prompt is determined based on the voting results and reasons for each diagnostic role. This fourth prompt is then input into the large language model to obtain the summarized voting results output by the large language model. The summarized voting results include the voting information for all diagnostic roles. For each diagnostic role, a fifth prompt is determined based on the summarized voting results and the diagnostic role itself. This fifth prompt is then input into the large language model to obtain the new voting information for the diagnostic role output by the large language model. The above steps are repeated until all new voting results are consistent, or the number of votes reaches a preset number. Based on the final voting information, the target assessment result for the diagnosis of the condition to be evaluated is determined.
[0098] Specifically, after each round of voting, if inconsistencies are found among the voting results, a summary voting result, or voting report, is generated based on the voting results and reasons for each diagnostic role. This report includes the statistical count of support and opposition votes, which allows analysis of the ratio of support to opposition. Additionally, it may include content on which consensus has been reached and content requiring further discussion. For example, the fourth prompt information determined based on the voting results and reasons for each diagnostic role can be input into the large language model to obtain the summary voting result output by the model. An example of the summary voting result is as follows:
[0099] Statistics of votes in favor and against:
[0100] Total votes: 3
[0101] Support votes: 2
[0102] 1 vote against
[0103] Summary of comments:
[0104] consensus:
[0105] Laboratory indicator dimensions: ...
[0106] Medication regimen dimensions: ...
[0107] Past medical history: ...
[0108] Disagreement: ...further clarification is needed.
[0109] Furthermore, for each diagnostic role, a fifth prompt can be generated based on the aggregated voting results and the specific diagnostic role. This fifth prompt is then input into the large language model, which analyzes the aggregated voting results to output new voting information for that diagnostic role. It should be understood that the new voting information output by the large language model for that diagnostic role is a revised version of its evaluation results, based on the consensus reached in the aggregated voting results and information from other diagnostic roles.
[0110] If inconsistent voting results still exist among all the new voting information, the above process can be repeated until all new voting results are consistent, or the number of votes reaches the preset number, so as to determine the target assessment result for the diagnosis of the disease to be evaluated based on the final voting information.
[0111] In this embodiment, when inconsistent voting results exist among all voting results, the final voting result can be determined through iterative voting. This involves continuously adjusting the voting information corresponding to each diagnostic role by referring to the voting information of other diagnostic roles. This ensures that each voting result incorporates and references the voting results of different diagnostic roles, improving the accuracy of the final voting information and further enhancing the precision of the target evaluation result. Furthermore, the aforementioned method for evaluating disease diagnosis results, which uses a large language model for multi-round voting by diagnostic roles to verify the evaluation results, can solve the error propagation problem caused by the numerous steps in the task itself when using deep learning for evaluation in existing technologies. This makes the final evaluation result controllable and improves the reliability of the target evaluation result.
[0112] For example, based on the above embodiments, when determining the target assessment result of the diagnosis result of the condition to be evaluated based on the final voting information, a sixth prompt information can be determined based on the final voting information and the diagnosis result of the condition to be evaluated, and the sixth prompt information can be input into the large language model to obtain a comprehensive assessment report output by the large language model. The comprehensive assessment report includes the target assessment result and the diagnostic opinion, and the diagnostic opinion includes the same diagnostic opinion and / or divergent diagnostic opinion of each diagnostic role.
[0113] Specifically, if inconsistent voting results still exist in the final voting information after a preset number of votes has been cast, a sixth prompt message can be generated based on the final voting information and the diagnosis result of the condition to be evaluated. This sixth prompt message can be as follows:
[0114] Based on the voting information {Votei}, please generate a comprehensive report on the multiple diagnostic reviews for this patient's {diag} diagnosis. The report should include the consensus reached on each diagnostic dimension and the remaining disagreements, and finally provide a recommended review conclusion.
[0115] Votei represents the final voting information, and diag represents the diagnosis result pending evaluation.
[0116] By inputting the aforementioned sixth prompt information into the large language model and guiding the large language model through prompt words, a comprehensive evaluation report can be generated. This comprehensive evaluation report includes the target evaluation results and diagnostic opinions, which include the same diagnostic opinions and / or divergent diagnostic opinions from various diagnostic roles.
[0117] In this embodiment, if the voting results are still inconsistent after the number of votes has reached a preset number, a comprehensive evaluation report can be output based on the large language model. This comprehensive evaluation report includes not only the target evaluation results but also diagnostic opinions, making the comprehensive evaluation report more comprehensive and thus providing a more comprehensive decision-making basis for subsequent diagnosis.
[0118] For example, if all voting results are consistent, the voting result is determined as the target assessment result for the diagnosis of the condition to be evaluated.
[0119] Specifically, when the voting results of all diagnostic roles are consistent, it means that all diagnostic roles have the same opinion on the assessment result of the diagnosis of the condition to be evaluated. Therefore, the final voting result can be directly determined as the target assessment result of the diagnosis of the condition to be evaluated, thereby ensuring the consistency of the final target assessment result.
[0120] For example, based on the above embodiments, if more than a first preset number of the diagnostic results of the conditions to be evaluated belong to the same department, then at least two diagnostic roles include more than a second preset number of diagnostic roles belonging to the same department.
[0121] Specifically, the diagnostic role can be understood as a virtual diagnostic expert. Diagnostic experts from different departments are skilled in diagnosing different diseases. Therefore, if more than a first preset number of the patient's diagnostic results belong to the same department, it means that most of the patient's symptoms belong to the same department. In this case, in order to fully explore the medical knowledge of the diagnostic experts in that department, more than a second preset number of diagnostic roles can be set to belong to this department, thereby improving the accuracy of the target evaluation results for the diagnostic results of the patient's condition.
[0122] Figure 2 A flowchart of the method for evaluating the diagnostic results of a disease provided in an embodiment of the present invention is shown below. Figure 2 As shown, in this embodiment of the invention, a large language model can be used in a multi-role-playing mode to diagnose or evaluate multiple writes. The method includes:
[0123] Step 1: In the meeting preparation stage, the large language model can be used through cue word engineering to play the role of a medical meeting host, receiving the diagnosis results of the disease to be evaluated and the target medical record in which it is located. The target medical record can be understood as the full case text. Based on the key information such as the diagnosis list, department, and personal information in the full case text, the diagnostic roles are generated, and the meeting outline is determined, that is, the diagnostic dimensions of each diagnostic role. The diagnostic role can be understood as an expert role.
[0124] Step 2: In the expert analysis phase, prompt word engineering can be used to make the large language model play the role of various experts. Each expert role conducts in-depth analysis based on the input diagnosis results of the disease to be evaluated and the target case, according to the corresponding diagnostic dimensions, and obtains a diagnostic analysis report containing key analysis points and conclusions of each diagnostic dimension.
[0125] Step 3: In the stage of summarizing the diagnostic analysis reports, multiple diagnostic analysis reports can be input into the large language model. The large language model can then analyze and extract the experts' analytical thinking and conclusions from the multiple diagnostic analysis reports to generate a summary evaluation result.
[0126] Step 4: In the multi-round voting phase, the summarized evaluation results can be input into the large language model. Through prompt word engineering, voting information corresponding to each expert role can be generated to provide a conclusion of agreement / disagreement and reasons. If the voting results of multiple expert roles are inconsistent, the voting information corresponding to each expert role can be input into the large language model. After the moderator role re-summarizes the results, a voting summary is generated, and the large language model outputs the new voting results for each expert role again. This process is repeated multiple times until the voting results of all expert roles are consistent or the prescribed number of votes is reached.
[0127] Step 5: Based on the final voting information, generate a comprehensive evaluation report using a large language model. This comprehensive evaluation report includes conclusions on overwriting / overwriting and analysis of multiple diagnostic dimensions.
[0128] The method for evaluating disease diagnosis results provided in this invention involves obtaining at least one disease diagnosis result to be evaluated based on a target medical record. For each disease diagnosis result to be evaluated, the disease diagnosis result to be evaluated and the target medical record are input into a large language model to obtain at least two diagnostic roles and diagnostic dimensions of each diagnostic role output by the large language model. For each diagnostic role, based on the diagnostic role, the diagnostic dimension corresponding to the diagnostic role, the disease diagnosis result to be evaluated, and the target medical record, a first prompt information is determined and input into the large language model to obtain a diagnostic analysis report corresponding to the diagnostic role output by the large language model. Based on the diagnostic analysis report corresponding to each diagnostic role, a target evaluation result for the disease diagnosis result to be evaluated is determined. This target evaluation result is used to characterize whether the disease diagnosis result to be evaluated is correct. Because large language models can be used to identify different diagnostic roles and their corresponding diagnostic dimensions, and then a diagnostic analysis report can be generated for each role based on the large language model and its corresponding diagnostic dimensions, the final target assessment result can be determined based on the diagnostic analysis reports of different roles. Large language models possess strong semantic understanding and processing capabilities, as well as the advantage of mining rich medical knowledge. Therefore, when using large language models for disease diagnosis review, fixed review logic and rule templates can be avoided, improving the generalization of the assessment method. Moreover, by fully mining useful information from the target case through large language models, it can adapt to complex and ever-changing medical scenarios, improving the accuracy of the assessment results.
[0129] The following describes the device for evaluating the diagnostic results of a disease provided by the present invention. The device for evaluating the diagnostic results of a disease described below can be referred to in correspondence with the method for evaluating the diagnostic results of a disease described above.
[0130] Figure 3 This is a schematic diagram of the structure of the disease diagnosis result assessment device provided in an embodiment of the present invention, with reference to... Figure 3 As shown, the assessment device 300 for the diagnosis results includes:
[0131] The acquisition module 11 is used to acquire at least one diagnosis result of the condition to be evaluated based on the target medical record;
[0132] Input module 12 is used to input the diagnosis results of the conditions to be evaluated and the target medical record into the large language model for each of the diagnosis results of the conditions to be evaluated, so as to obtain at least two diagnostic roles and diagnostic dimensions of each diagnostic role output by the large language model;
[0133] The determination module 13 is used to determine a first prompt information for each of the diagnostic roles, based on the diagnostic role, the diagnostic dimension corresponding to the diagnostic role, the diagnostic result of the condition to be evaluated, and the target medical record;
[0134] The input module 12 is also used to input the first prompt information into the large language model to obtain the diagnostic analysis report corresponding to the diagnostic role output by the large language model;
[0135] The determining module 13 is further configured to determine the target evaluation result of the diagnosis result of the condition to be evaluated based on the diagnostic analysis report corresponding to each of the diagnostic roles, wherein the target evaluation result is used to characterize whether the diagnosis result of the condition to be evaluated is correct.
[0136] In one example embodiment, the determining module 13 is specifically used for:
[0137] Based on each of the diagnostic roles, the corresponding diagnostic analysis reports and diagnostic dimensions, the second prompt information is determined.
[0138] The second prompt information is input into the large language model to obtain the summary evaluation result output by the large language model. The summary evaluation result includes the diagnostic analysis report corresponding to all the diagnostic roles.
[0139] For each diagnostic role, based on the aggregated evaluation results, the voting information of the diagnostic role regarding the diagnostic result of the condition to be evaluated is determined, and the voting information includes the voting result of the diagnostic role regarding whether the diagnostic result of the condition to be evaluated is correct;
[0140] Based on the voting information of each diagnostic role, the target evaluation result of the diagnosis of the condition to be evaluated is determined.
[0141] In one example embodiment, the determining module 13 is specifically used for:
[0142] Based on the summarized assessment results and the diagnostic role, a third prompt message is determined;
[0143] The third prompt information is input into the large language model to obtain the voting information of the diagnostic role for the diagnosis result of the condition to be evaluated, which is output by the large language model.
[0144] In one example embodiment, the voting information also includes the reasons for voting;
[0145] The determining module 13 is specifically used for:
[0146] In the event of inconsistencies among all the voting results, a fourth prompt message is determined based on the voting results and reasons for each diagnostic role.
[0147] The fourth prompt information is input into the large language model to obtain the summary voting result output by the large language model. The summary voting result includes the voting information corresponding to all the diagnostic roles.
[0148] For each of the aforementioned diagnostic roles, a fifth prompt message is determined based on the aggregated voting results and the diagnostic role.
[0149] The fifth prompt information is input into the large language model to obtain the new voting information of the diagnostic role output by the large language model. The above steps are repeated until all new voting results are consistent, or the number of votes reaches the preset number.
[0150] Based on the final voting information, the target assessment result for the diagnosis of the condition to be evaluated is determined.
[0151] In one example embodiment, the determining module 13 is specifically used for:
[0152] Based on the final voting information and the diagnosis results of the condition to be evaluated, the sixth prompt information is determined;
[0153] The sixth prompt information is input into the large language model to obtain a comprehensive evaluation report output by the large language model. The comprehensive evaluation report includes the target evaluation results and diagnostic opinions. The diagnostic opinions include the same diagnostic opinions and / or divergent diagnostic opinions of each of the diagnostic roles.
[0154] In one example embodiment, the determining module 13 is further configured to determine the voting result as the target evaluation result of the diagnosis result of the condition to be evaluated if all the voting results are consistent.
[0155] In one example embodiment, if more than a first preset number of the diagnostic results to be evaluated belong to diseases in the same department, the at least two diagnostic roles include more than a second preset number of diagnostic roles belonging to the department.
[0156] The device for evaluating the diagnosis results of the disease in this embodiment can be used to execute the method of any embodiment in the side embodiment of the method for evaluating the diagnosis results of the disease. Its specific implementation process and technical effects are similar to those in the side embodiment of the method for evaluating the diagnosis results of the disease. For details, please refer to the detailed description in the side embodiment of the method for evaluating the diagnosis results of the disease, which will not be repeated here.
[0157] Figure 4 This is a schematic diagram of the physical structure of an electronic device provided in an embodiment of the present invention, such as... Figure 4As shown, the electronic device may include: a processor 410, a communications interface 420, a memory 430, and a communications bus 440, wherein the processor 410, the communications interface 420, and the memory 430 communicate with each other through the communications bus 440. The processor 410 can call logical instructions in the memory 430 to execute a method for evaluating a disease diagnosis result. This method includes: acquiring at least one disease diagnosis result to be evaluated based on a target medical record; for each disease diagnosis result to be evaluated, inputting the disease diagnosis result to be evaluated and the target medical record into a large language model to obtain at least two diagnostic roles and diagnostic dimensions of each diagnostic role output by the large language model; for each diagnostic role, determining a first prompt based on the diagnostic role, the corresponding diagnostic dimension, the disease diagnosis result to be evaluated, and the target medical record, and inputting the first prompt based on the large language model to obtain a diagnostic analysis report corresponding to the diagnostic role output by the large language model; and based on the diagnostic analysis reports corresponding to each diagnostic role, determining a target evaluation result for the disease diagnosis result to be evaluated, wherein the target evaluation result is used to characterize whether the disease diagnosis result to be evaluated is correct.
[0158] Furthermore, the logical instructions in the aforementioned memory 430 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0159] On the other hand, the present invention also provides a computer program product, which includes a computer program that can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer is able to execute the evaluation method for the diagnosis results provided by the above methods. The method includes: acquiring at least one diagnosis result to be evaluated based on a target medical record; for each diagnosis result to be evaluated, inputting the diagnosis result to be evaluated and the target medical record into a large language model to obtain at least two diagnostic roles and diagnostic dimensions of each diagnostic role output by the large language model; for each diagnostic role, determining a first prompting information based on the diagnostic role, the diagnostic dimension corresponding to the diagnostic role, the diagnosis result to be evaluated, and the target medical record, and inputting the first prompting information into the large language model to obtain a diagnostic analysis report corresponding to the diagnostic role output by the large language model; and determining a target evaluation result for the diagnosis result to be evaluated based on the diagnostic analysis report corresponding to each diagnostic role, wherein the target evaluation result is used to characterize whether the diagnosis result to be evaluated is correct.
[0160] In another aspect, the present invention also provides a non-transitory computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements an evaluation method for the diagnostic results of the above-described methods. The method includes: acquiring at least one diagnostic result to be evaluated based on a target medical record; for each diagnostic result to be evaluated, inputting the diagnostic result to be evaluated and the target medical record into a large language model to obtain at least two diagnostic roles and diagnostic dimensions of each diagnostic role output by the large language model; for each diagnostic role, determining first prompt information based on the diagnostic role, the diagnostic dimension corresponding to the diagnostic role, the diagnostic result to be evaluated, and the target medical record, and inputting the first prompt information into the large language model to obtain a diagnostic analysis report corresponding to the diagnostic role output by the large language model; and determining a target evaluation result for the diagnostic result to be evaluated based on the diagnostic analysis report corresponding to each diagnostic role, wherein the target evaluation result is used to characterize whether the diagnostic result to be evaluated is correct.
[0161] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.
[0162] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.
[0163] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A method for evaluating the results of a disease diagnosis, characterized in that, include: Obtain at least one diagnosis of the condition to be evaluated based on the target medical record; For each of the diagnostic results of the condition to be evaluated, the diagnostic results of the condition to be evaluated and the target medical record are input into a large language model. The large language model acts as a medical conference host and obtains at least two diagnostic roles and diagnostic dimensions of each diagnostic role output by the large language model. For each diagnostic role, based on the diagnostic role, the diagnostic dimension corresponding to the diagnostic role, the diagnostic result of the condition to be evaluated, and the target medical record, a first prompting information is determined, and the first prompting information is input into the large language model to obtain the diagnostic analysis report corresponding to the diagnostic role output by the large language model; Based on the diagnostic analysis reports corresponding to each of the diagnostic roles, a target evaluation result is determined for the diagnostic result of the condition to be evaluated. The target evaluation result is used to characterize whether the diagnostic result of the condition to be evaluated is correct.
2. The method for evaluating the diagnostic results of a disease according to claim 1, characterized in that, The determination of the target assessment result for the diagnosis of the condition to be evaluated based on the diagnostic analysis report corresponding to each of the diagnostic roles includes: Based on each of the diagnostic roles, the corresponding diagnostic analysis reports and diagnostic dimensions, the second prompt information is determined. The second prompt information is input into the large language model to obtain the summary evaluation result output by the large language model. The summary evaluation result includes the diagnostic analysis report corresponding to all the diagnostic roles. For each diagnostic role, based on the aggregated evaluation results, the voting information of the diagnostic role regarding the diagnostic result of the condition to be evaluated is determined, and the voting information includes the voting result of the diagnostic role regarding whether the diagnostic result of the condition to be evaluated is correct; Based on the voting information of each diagnostic role, the target evaluation result of the diagnosis of the condition to be evaluated is determined.
3. The method for evaluating the diagnostic results of a disease according to claim 2, characterized in that, The step of determining the voting information of the diagnostic role regarding the diagnostic result of the condition to be evaluated based on the summarized evaluation results includes: Based on the summarized assessment results and the diagnostic role, a third prompt message is determined; The third prompt information is input into the large language model to obtain the voting information of the diagnostic role for the diagnosis result of the condition to be evaluated, which is output by the large language model.
4. The method for evaluating the diagnostic results of a disease according to claim 2, characterized in that, The voting information also includes the reasons for voting; The determination of the target assessment result for the diagnosis of the condition to be evaluated based on the voting information of each diagnostic role includes: In the event of inconsistencies among all the voting results, a fourth prompt message is determined based on the voting results and reasons for each diagnostic role. The fourth prompt information is input into the large language model to obtain the summary voting result output by the large language model. The summary voting result includes the voting information corresponding to all the diagnostic roles. For each of the aforementioned diagnostic roles, a fifth prompt message is determined based on the aggregated voting results and the diagnostic role. The fifth prompt information is input into the large language model to obtain the new voting information of the diagnostic role output by the large language model. The above steps are repeated until all new voting results are consistent, or the number of votes reaches the preset number. Based on the final voting information, the target assessment result for the diagnosis of the condition to be evaluated is determined.
5. The method for evaluating the diagnostic results of a disease according to claim 4, characterized in that, The determination of the target assessment result for the diagnosis of the condition to be evaluated based on the final voting information includes: Based on the final voting information and the diagnosis results of the condition to be evaluated, the sixth prompt information is determined; The sixth prompt information is input into the large language model to obtain a comprehensive evaluation report output by the large language model. The comprehensive evaluation report includes the target evaluation results and diagnostic opinions. The diagnostic opinions include the same diagnostic opinions and / or divergent diagnostic opinions of each of the diagnostic roles.
6. The method for evaluating the diagnostic results of a disease according to claim 4, characterized in that, The method further includes: If all the voting results are consistent, the voting result will be determined as the target evaluation result for the diagnosis of the condition to be evaluated.
7. The method for evaluating the diagnostic results of a disease condition according to any one of claims 1-6, characterized in that, If more than a first preset number of the diagnostic results of the conditions to be evaluated belong to the same department, then the at least two diagnostic roles include more than a second preset number of diagnostic roles belonging to the department.
8. A device for evaluating the results of a disease diagnosis, characterized in that, include: The acquisition module is used to acquire at least one diagnosis result of the condition to be evaluated based on the target medical record; The input module is used to input the diagnosis results of the conditions to be evaluated and the target medical record into the large language model for each of the diagnosis results of the conditions to be evaluated. The large language model acts as a medical conference host and obtains at least two diagnostic roles and diagnostic dimensions of each diagnostic role output by the large language model. The determination module is used to determine the first prompt information for each of the diagnostic roles, based on the diagnostic role, the diagnostic dimension corresponding to the diagnostic role, the diagnostic result of the condition to be evaluated, and the target medical record; The input module is also used to input the first prompt information into the large language model to obtain a diagnostic analysis report corresponding to the diagnostic role output by the large language model; The determining module is further configured to determine the target evaluation result of the diagnosis result of the condition to be evaluated based on the diagnostic analysis report corresponding to each of the diagnostic roles, wherein the target evaluation result is used to characterize whether the diagnosis result of the condition to be evaluated is correct.
9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that, When the processor executes the computer program, it implements the method for evaluating the diagnostic results of the disease as described in any one of claims 1 to 7.
10. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the method for evaluating the diagnostic results of the disease as described in any one of claims 1 to 7.
11. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, it implements the method for evaluating the diagnostic results of the disease as described in any one of claims 1 to 7.
Citation Information
Patent Citations
Medical diagnosis method, device and equipment based on large model and storage medium
CN118800442A