Automatic evaluation method and system for medical question-answering system

By combining large language model classification and medical entity graph evaluation methods, and dynamically adjusting weights for weighted summation, the problem of insufficient identification of logical connections of medical terms and question type characteristics in existing medical question answering system evaluations is solved, thus achieving a more efficient evaluation of medical question answering systems.

CN121072736APending Publication Date: 2025-12-05SHANGHAI SOUNDWISE TECHNOLOGY CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202511043778.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-28
Publication Date
2025-12-05

AI Technical Summary

Technical Problem

Existing automated evaluation methods in medical question-answering systems suffer from domain adaptability failure in the medical field. They are unable to parse the logical connections and structured relationships of medical terms, and cannot distinguish the characteristics of question types, resulting in a lack of professionalism and interpretability in the evaluation.

Method used

A large language model is used for question classification. The semantic similarity between the system's answer and the standard answer, as well as the entity similarity and logical relationship similarity of medical entities, are calculated. The evaluation is carried out by dynamic weighted summation and combined with medical entity graph and question type labels for accurate evaluation.

Benefits of technology

It achieves accurate analysis of the equivalence and logical relevance of medical terms, improves the professionalism and scenario adaptability of the assessment, and solves the problems of insufficient adaptability and assessment limitations of existing methods in the medical field.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121072736A_ABST
    Figure CN121072736A_ABST
Patent Text Reader

Abstract

The invention provides an automatic evaluation method and system for a medical question-answering system, and relates to the technical field of medical question-answering, and the method comprises the steps: classifying questions inputted into the medical question-answering system, and marking classification tags; calculating the semantic similarity between the system answer given by the medical question-answering system and the standard answer; calculating an entity overlapping degree based on the entity similarity and the logic relationship similarity between the system answer and the standard answer; selecting an evaluation weight based on the classification label, and then performing weighted summation according to the entity overlapping degree, the semantic similarity and the evaluation weight to obtain an evaluation score of the medical question and answer system. The method has the beneficial effects that based on the technical means of dynamically selecting problem classification labels and applying different evaluation weights to perform weighted summation, the technical effects of realizing adaptive adjustment according to the problem type characteristics and remarkably improving the scene adaptability are achieved; the problem that an existing method adopts a fixed weight assessment system, and medical problem type characteristics cannot be distinguished, so that assessment is lack of professionality and pertinence is solved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of medical question and answer, and particularly relates to an automatic evaluation method and system of a medical question and answer system. BACKGROUND

[0002] In the evaluation of question and answer systems (QAS) in the field of natural language processing, automatic evaluation methods have become the mainstream solution due to their efficiency and standardization advantages. Existing methods mainly include two categories: statistical-based methods (such as the BLEU index) and semantic similarity-based methods. However, in the medical question and answer scenario, these methods face the core bottleneck of failure in domain adaptation:

[0003] Existing statistical methods (such as BLEU) only evaluate quality through word surface alignment, and cannot analyze the logical correlation between medical terms. Such methods ignore the structured relationship of medical terms, resulting in serious distortion in the evaluation of abbreviations, synonymous terms and logical expressions.

[0004] Semantic similarity methods rely on general word embeddings and lack graph modeling of medical entity relationships. This is because existing methods only capture surface word meanings and do not construct structured relationships of diseases, symptoms and treatments.

[0005] Moreover, all existing methods use a fixed weight evaluation system and cannot distinguish the type characteristics of medical questions, which makes traditional methods unable to guarantee professional, explainability and scenario adaptability when evaluating medical question and answer systems, and there is an urgent need for an evaluation scheme that integrates medical knowledge structure analysis and dynamic weight allocation. SUMMARY

[0006] To solve the problems in the prior art, the present application provides an automatic evaluation method for a medical question and answer system, comprising:

[0007] Step S1, classifying and tagging the input medical question and answer system question with a classification label;

[0008] Step S2, calculating the semantic similarity between the system answer given by the medical question and answer system and the standard answer corresponding to the question;

[0009]

[0009] Step S3, calculating the entity overlap degree based on the entity similarity between the system answer and the standard answer and the logical relationship similarity between the structures of medical entities;

[0010] Step S4, selecting an evaluation weight based on the classification label, and then performing weighted summation according to the entity overlap degree, the semantic similarity and the evaluation weight to obtain an evaluation score of the medical question and answer system.

[0011] Preferably, the large language model is used for question classification in the step S1, and the large language model is trained for question classification before the step S1 is performed, and the training process comprises:

[0012] The question classification prompt words are input into the large language model, and the classification prompt words comprise diagnosis class prompt words and description class prompt words;

[0013] The large language model performs semantic extraction on the question to obtain question semantics, and gives the question a diagnosis class label when the question semantics conforms to the diagnosis class prompt words, and gives the question a description class label when the question semantics conforms to the description class prompt words.

[0014] Preferably, the step S2 comprises:

[0015] The step S21 converts the system answer and the standard answer into a system answer vector and a standard answer vector, respectively;

[0016] The step S22 calculates the cosine similarity between the system answer vector and the standard answer vector as the semantic similarity.

[0017] Preferably, the step S3 comprises:

[0018] The step S31 extracts medical entities in the system answer and the standard answer to obtain a system answer entity set and a standard answer entity set, respectively, and calculates the set similarity coefficient therebetween as the entity similarity;

[0019] The step S32 constructs a system answer graph and a standard answer graph according to the system answer entity set and the standard answer entity set, respectively, and calculates the graph structure similarity between the system answer graph and the standard answer graph as the logical relationship similarity;

[0020] The step S33 performs weighted summation processing on the entity similarity and the logical relationship similarity to obtain the entity overlap degree.

[0021] Preferably, the classification label comprises a diagnosis class label and a description class label; the diagnosis class label is associated with a diagnosis evaluation weight, and the description class label is associated with a description evaluation weight;

[0022] The step S4 comprises:

[0023] It is judged whether the classification label is a diagnosis class label or a description class label:

[0024] If it is a diagnosis class label, the entity overlap degree, the semantic similarity, and the diagnosis evaluation weight are weighted and summed to obtain an evaluation score of the medical question and answer system;

[0025] If the category label is a description label, the evaluation score of the medical question and answer system is obtained by weighted summation according to the entity overlap degree, the semantic similarity and the evaluation weight.

[0026] Preferably, the data preprocessing process further includes the following steps before step S2 is performed:

[0027] The system answer and the standard answer are subjected to word segmentation processing and stop word removal processing, and medical entities in the processed system answer and standard answer are identified.

[0028] The application further provides an automatic evaluation system of a medical question and answer system, which is applied to the automatic evaluation method and includes the following steps:

[0029] A question classification module is configured to classify an input question of the medical question and answer system and label the question with a classification label.

[0030] A semantic similarity calculation module is configured to calculate a semantic similarity between a system answer given by the medical question and answer system and a standard answer corresponding to the question.

[0031] An overlap degree calculation module is configured to calculate an entity overlap degree based on an entity similarity between medical entities in the system answer and the standard answer and a logical relationship similarity between structures of the medical entities.

[0032] An evaluation module is connected to the question classification module, the overlap degree calculation module and the semantic similarity calculation module, and is configured to select an evaluation weight based on the classification label, and then obtain an evaluation score of the medical question and answer system by weighted summation according to the entity overlap degree, the semantic similarity and the evaluation weight.

[0033] Preferably, the question classification module uses a large language model to classify the question, and further includes a training module configured to train the large language model to classify the question, and input a question classification prompt word to the large language model, wherein the classification prompt word includes a diagnosis class prompt word and a description class prompt word.

[0034] The large language model extracts a question semantic from the question, and labels the question with a diagnosis class label when the question semantic conforms to the diagnosis class prompt word, and labels the question with a description class label when the question semantic conforms to the description class prompt word.

[0035] Preferably, the semantic similarity calculation module includes:

[0036] A vector conversion unit is configured to convert the system answer and the standard answer into a system answer vector and a standard answer vector, respectively.

[0037] The similarity calculation unit is connected to the vector conversion unit and is configured to calculate a cosine similarity between the system answer vector and the standard answer vector as the semantic similarity.

[0038] Preferably, the overlap calculation module comprises:

[0039] The first calculation unit is configured to extract medical entities in the system answer and the standard answer respectively to obtain a system answer entity set and a standard answer entity set, and calculate a set similarity coefficient between the system answer entity set and the standard answer entity set as the entity similarity.

[0040] The second calculation unit is configured to construct a system answer graph and a standard answer graph according to the system answer entity set and the standard answer entity set respectively, and calculate a graph structure similarity between the system answer graph and the standard answer graph as the logical relationship similarity.

[0041] The overlap calculation unit is connected to the first calculation unit and the second calculation unit, and is configured to perform weighted summation processing on the entity similarity and the logical relationship similarity to obtain the entity overlap.

[0042] The above technical solution has the following advantages or beneficial effects:

[0043] 1. The automatic evaluation method of the present application realizes the technical effects of accurately analyzing the equivalence (such as abbreviations / full names) and logical correlation (such as disease-symptom-treatment chain, treatment step sequence) between medical terms by calculating the entity similarity and the logical relationship similarity, and solves the technical problem of evaluation distortion caused by the existing statistical method (such as BLEU) which only relies on surface alignment of words and ignores the structured relationship of medical terms.

[0044] 2. The automatic evaluation method of the present application realizes the technical effects of overcoming the lack of adaptability in the medical field and the complexity of deep semantic representation of medical terms by the existing single semantic similarity method, and solving the technical problem of evaluation limitation caused by the method based on semantic similarity which relies on general word embedding and lacks the ability of medical entity relationship graph modeling.

[0045] 3. The automatic evaluation method of the present application realizes the technical effects of self-adaptive adjustment of the evaluation system according to the characteristics of the problem type and significant improvement of scene adaptability by dynamically selecting and applying different evaluation weights for weighted summation based on the problem classification label, and solves the technical problem of lack of professional and targeted evaluation caused by the fixed weight evaluation system of all existing methods which cannot distinguish the characteristics of the medical problem type. BRIEF DESCRIPTION OF DRAWINGS

[0046] Figure 1 For the preferred embodiment of the present application, a flowchart of an automatic evaluation method of a medical question and answer system is shown in the figure;

[0047] Figure 2 For the preferred embodiment of the present application, a flowchart of a sub-process of step S2 is shown in the figure;

[0048] Figure 3 For the preferred embodiment of the present application, a flowchart of a sub-process of step S3 is shown in the figure;

[0049] Figure 4 For the preferred embodiment of the present application, a structure diagram of an automatic evaluation system of a medical question and answer system is shown in the figure. DETAILED DESCRIPTION

[0050] The present application will be described in detail below with reference to the accompanying drawings and specific embodiments. The present application is not limited to this embodiment, and other embodiments can also fall within the scope of the present application as long as they comply with the main idea of the present application.

[0051] In the preferred embodiment of the present application, based on the above-mentioned problems existing in the prior art, an automatic evaluation method of a medical question and answer system is provided, as shown in the figure, which comprises: Figure 1

[0052] Step S1, classifying the input medical question and answer system question and tagging with a classification label;

[0053] Step S2, calculating the semantic similarity between the system answer given by the medical question and answer system and the standard answer corresponding to the question;

[0054] Step S3, calculating the entity overlap degree based on the entity similarity between the system answer and the medical entity in the standard answer and the logical relationship similarity between the structures of the medical entity;

[0055] Step S4, selecting an evaluation weight based on the classification label, and then performing weighted summation according to the entity overlap degree, the semantic similarity and the evaluation weight to obtain an evaluation score of the medical question and answer system.

[0056] Specifically, in the prior art, the statistical method (such as BLEU index) exists the problem of lack of medical logical correlation, which is difficult to capture the semantic equivalence of medical terms (such as abbreviations and full names), the penalty mechanism for short texts is easy to lead to low score, and it relies too much on word order and ignores the flexibility of medical expression; for example, the statistical method such as BLEU cannot recognize the equivalence of terms such as "heart attack→myocardial infarction", and ignores the logical order of treatment steps (such as the equivalence of "anticoagulation before thrombolysis" and "anticoagulation before thrombolysis").

[0057] ​Further specifically, the prior art method based on semantic similarity is limited in evaluation effect because the general word embedding model cannot accurately represent complex medical terms, and lacks the modeling ability of disease context.

[0058] The automatic evaluation method of the present application realizes the technical effect of accurately analyzing the equivalence (such as abbreviation / full name) and logical correlation (such as disease-symptom-treatment chain, treatment step sequence) between medical terms through the technical means of calculating entity similarity and logical relationship similarity, and solves the technical problem that the existing statistical method (such as BLEU) only relies on word surface alignment and ignores the structured relationship of medical terms, resulting in distorted evaluation;

[0059] The automatic evaluation method of the present application realizes the technical effect of overcoming the lack of adaptability of existing single semantic similarity method in the medical field, and deeply representing the semantics and relationship of complex medical terms through the technical means of calculating the evaluation score by weighted summation of the semantic similarity and the entity overlap degree, and solves the technical problem that the method based on semantic similarity relies on general word embedding and lacks the ability of medical entity relationship graph modeling, resulting in limited evaluation;

[0060] The automatic evaluation method of the present application realizes the technical effect of self-adaptive adjustment of the evaluation system according to the characteristics of the problem type and significantly improves the scene adaptability through the technical means of dynamically selecting and applying different evaluation weights for weighted summation based on the problem classification label, and solves the technical problem that all existing methods use a fixed weight evaluation system and cannot distinguish the characteristics of the medical problem type, resulting in lack of professional and targeted evaluation.

[0061] In view of the technical problems in the prior art, the core improvement of the present application lies in the cooperative design of the following steps:

[0062] 1. Problem classification and labeling, step S1 first classifies and labels the classification label according to the problem type of the input problem, and the problem type includes but is not limited to:

[0063] Diagnosis type: problems related to disease diagnosis, symptom differentiation, etc. (such as "What diseases can cause chest pain?", "What disease may the patient have?") ;

[0064] Description type: problems related to medical knowledge explanation, treatment principle explanation, etc. (such as "How is coronary heart disease formed?", "What are the manifestations of thyroid nodules?").

[0065] The classification label provides context constraints for subsequent evaluation. For example, for diagnosis type problems, the entity logical chain related to disease characteristics and differential points should be matched first, and for description type problems, the semantic coherence and term accuracy should be evaluated.

[0066] In the preferred embodiment of the present application, a large language model is used for question classification in step S1, and the large language model is trained for question classification before step S1 is performed, and the training process includes:

[0067] The question classification prompt words are input into the large language model, and the classification prompt words include diagnosis class prompt words and description class prompt words.

[0068] The large language model extracts the semantics of the question to obtain question semantics, and gives the question a diagnosis class label when the question semantics conforms to the diagnosis class prompt words, and gives the question a description class label when the question semantics conforms to the description class prompt words.

[0069] Specifically, in the present embodiment, 100,000 pieces of medical question and answer data are collected from the Internal Medicine textbook, UpToDate clinical guidelines, and Dingxiangyuan forum.

[0070] The setting of the labeling rules includes:

[0071] Diagnosis class questions: label questions containing keywords such as "what disease is it possible" and "how to differential diagnosis" (such as "what emergency should chest pain with sweating be considered?").

[0072] Description class questions: label questions containing keywords such as "pathogenesis" and "clinical manifestations" (such as "What is the pathophysiological process of diabetic ketoacidosis?").

[0073] 1. Labeling process: independently labeled by 3 clinical physicians in a first-class hospital, cross-validated to retain samples with a consistency rate of > 90%, and finally 80,000 labeled data (40,000 for diagnosis class and 40,000 for description class) were formed.

[0074] 2. Semantic extraction: the base model uses the BioBERT model pre-trained on the PubMed corpus. Input question "patient with high fever for 3 days, accompanied by chills, what disease is it possible?" BioBERT generates a 768-dimensional semantic vector, focusing on the context embedding of keywords such as "high fever" and "chills".

[0075] 3. Prompt word matching: the diagnosis class prompt word library contains 200 templates such as "what disease is it possible" and "differential diagnosis". The model calculates the cosine similarity of the question semantics and the prompt word library, and the highest similarity of the diagnosis class reaches 0.82 (threshold > 0.7 is determined as diagnosis class).

[0076] 4. Label determination: the output diagnosis class probability is 0.88, which exceeds the threshold value of 0.5, and is finally marked as a diagnosis class question.

[0077] 5. Training effect verification:

[0078] Verification set performance: accuracy 92.3%, F1 value (diagnosis class) 91.8%.

[0079] Through the above embodiments, the complete training process of the large language model in the medical question classification task can be clearly shown, including data preparation, model tuning, semantic understanding, and label decision, and the like.

[0080] 2. Semantic similarity calculation (step S2), as shown in the preferred embodiment of the present application, Figure 2 as shown, step S2 includes:

[0081] Step S21, the system answer and the standard answer are respectively converted into a system answer vector and a standard answer vector;

[0082] Step S22, the cosine similarity between the system answer vector and the standard answer vector is calculated as the semantic similarity.

[0083] Specifically, in the present embodiment, the semantic similarity between the system answer and the standard answer is calculated by using a medical word embedding model combined with vector conversion, which accurately quantifies the semantic correlation degree between the system answer and the standard answer, and specifically includes the following sub-steps:

[0084] A domain-adapted medical word embedding model BioWordVec is selected, which is trained based on large-scale medical texts (such as PubMed literature and clinical guidelines) and can accurately represent the semantic relationship between complex medical terms such as "coronary atherosclerotic heart disease→coronary heart disease" and "heart attack→myocardial infarction" and abbreviations / full names, breaking through the recognition bottleneck of medical professional vocabulary for general models.

[0085] The words in the system answer and the standard answer are converted into high-dimensional vectors by BioWordVec, and each word is mapped into a fixed-dimensional vector (such as 300 dimensions), ensuring that different expressions of the same term (such as "high blood pressure" and "HTN") have highly similar coordinate distributions in the vector space.

[0086] Finally, the semantic similarity between the system answer and the standard answer is calculated using the cosine similarity formula.

[0087] A domain-adapted medical pre-training model (such as BioBERT) can also be used, which is fine-tuned on a million medical texts to enhance the representation ability of complex terms such as "coronary heart disease→coronary atherosclerotic heart disease", breaking through the recognition bottleneck of medical semantics for general models.

[0088] 3. Entity overlap calculation (step S3), as shown in the preferred embodiment of the present application, Figure 3 as shown, step S3 includes:

[0089] Step S31, respectively extract the medical entities in the system answer and the standard answer to obtain the system answer entity set and the standard answer entity set, and calculate the set similarity coefficient between them as the entity similarity;

[0090] Step S32, construct the system answer graph and the standard answer graph according to the system answer entity set and the standard answer entity set respectively, and calculate the graph structure similarity between the system answer graph and the standard answer graph as the logical relationship similarity;

[0091] Step S33, the entity similarity and the logical relationship similarity are weighted and summed to obtain the entity overlap degree.

[0092] Specifically, in the embodiment, the medical entity library and the knowledge graph are constructed, and the similarity of medical entities in the system answer and the standard answer and the logical relationship similarity are calculated.

[0093] The entity similarity calculation process: the medical entity sets extracted from the two groups of texts are respectively denoted as A and B;

[0094] Suppose the question is: How to deal with hypertension with dizziness in patients?

[0095] System answer: It is suggested to take antihypertensive drugs (such as amlodipine), and regularly monitor blood pressure.

[0096] Standard answer: It is recommended to use calcium channel blockers (such as amlodipine), take it every morning, and control sodium salt intake at the same time.

[0097] System answer entity set A: {hypertension, dizziness, antihypertensive drugs, amlodipine, blood pressure monitoring}

[0098] Standard answer entity set B: {hypertension, calcium channel blockers, amlodipine, sodium salt intake}

[0099] The Jaccard similarity coefficient formula is used to calculate the medical entity sets in the two groups of answers; the Jaccard similarity coefficient is used to measure the similarity of two sets, and the calculation formula is:

[0100]

[0101] Among them:

[0102] | A∩B| : The number of intersection elements of set A and B

[0103] | A∪B| : The number of union elements of set A and B

[0104] Step 1: Determine the entity set

[0105] System answer entity set A: {hypertension, dizziness, antihypertensive drugs, amlodipine, blood pressure monitoring} (5 elements in total);

[0106] Standard answer entity set B: {hypertension, calcium channel blockers, amlodipine, sodium salt intake} (4 elements in total);

[0107] Step 2: Calculate the intersection:

[0108] Intersection elements: {hypertension, amlodipine}, intersection size: |A∩B| = 2;

[0109] Step 3: Calculate the union

[0110] Union elements: {hypertension, dizziness, antihypertensive drugs, amlodipine, blood pressure monitoring, calcium channel blockers, sodium salt intake} (7 elements in total), union size: |A∪B| = 7;

[0111] Step 4: Substitute the formula, J(A,B) = 72 ≈ 0.2857

[0112] The Jaccard coefficient is 0.2857 (approximately 28.57%), indicating that the similarity between the system answer and the standard answer entity set is low.

[0113] The main differences are:

[0114] System answer contains "dizziness", "antihypertensive drugs", "blood pressure monitoring" and other practical entities;

[0115] The standard answer contains "calcium channel blockers", "sodium salt intake" and other mechanism entities.

[0116] This result reflects the difference in focus between the two (symptom treatment vs drug mechanism and lifestyle intervention).

[0117] Extended explanation: Threshold judgment: Jaccard coefficient > 0.5 is usually considered to have high similarity, and in this case, 0.2857 indicates that the entity coverage difference is significant.

[0118] Application scenario: In medical question and answer evaluation, low Jaccard coefficient may indicate that the system answer does not completely cover the key entities of the standard answer (such as drug classification, lifestyle intervention), and the answer strategy needs to be optimized.

[0119] Logical relationship similarity calculation process:

[0120] First, construct a graph based on the system answer entity set A and the standard answer entity set B, the construction process includes:

[0121] System answer graph:

[0122] Node: hypertension → dizziness (symptom association)

[0123] Node: Hypertension -> Antihypertensive (Therapeutic Means)

[0124] Node: Antihypertensive -> Amlodipine (Specific Drug)

[0125] Node: Blood Pressure Monitoring (Independent Measure)

[0126] Standard Answer Graph:

[0127] Node: Hypertension -> Calcium Channel Blockers (Drug Class)

[0128] Node: Calcium Channel Blockers -> Amlodipine (Specific Drug)

[0129] Node: Hypertension -> Sodium Salt Intake (Lifestyle Intervention)

[0130] Then proceed to graph structure similarity calculation:

[0131] Step 1: Node Matching

[0132] Matched Nodes: Hypertension, Amlodipine (Complete Match)

[0133] Unmatched Nodes: System Answer (Dizziness, Antihypertensive, Blood Pressure Monitoring) vs Standard Answer (Calcium Channel Blockers, Sodium Salt Intake)

[0134] Node Matching Rate = Average of (2 / 5 (System Graph Node Number) + 2 / 4 (Standard Graph Node Number)) = (0.4 + 0.5) / 2 = 0.45

[0135] Step 2: Edge Matching

[0136] Matched Edges: Hypertension -> Amlodipine (Both graphs exist)

[0137] Unmatched Edges: System Answer (Hypertension -> Dizziness, Antihypertensive -> Amlodipine, Blood Pressure Monitoring) vs Standard Answer (Hypertension -> Calcium Channel Blockers, Calcium Channel Blockers -> Amlodipine, Hypertension -> Sodium Salt Intake)

[0138] Edge Matching Rate = Average of (1 / 4 (System Graph Edge Number) + 1 / 3 (Standard Graph Edge Number)) = (0.25 + 0.33) / 2 ≈ 0.29

[0139] Step 3: Comprehensive Similarity

[0140] Use the weighted formula:

[0141] Logical Relationship Similarity = 0.6 × Node Matching Rate + 0.4 × Edge Matching Rate, where 0.6, 0.4 can be adjusted to other values;

[0142] Substitute the values: 0.6 x 0.45 + 0.4 x 0.29 ≈ 0.27 + 0.116 = 0.386

[0143] In this embodiment, by constructing the graph and calculating the structural similarity, the system answers the logical relationship similarity with the standard answer is 38.6%. The result reflects the difference between the two in the entity association method (such as the system answer emphasizes symptom monitoring, and the standard answer focuses on drug classification and life intervention), which provides a basis for subsequent weighted evaluation.

[0144] Finally, the entity overlap degree is obtained by weighted sum processing of the entity similarity and the logical relationship similarity.

[0145] 4. Dynamic weight weighted sum (step S4), in the preferred embodiment of the present application, the classification label includes a diagnosis class label and a description class label; the diagnosis class label is associated with a diagnosis evaluation weight, and the description class label is associated with a description evaluation weight;

[0146] Step S4 includes:

[0147] Determine whether the classification label is a diagnosis class label or a description class label:

[0148] If it is a diagnosis class label, the evaluation score of the medical question and answer system is obtained by weighted sum processing according to the entity overlap degree, the semantic similarity and the diagnosis evaluation weight;

[0149] If it is a description class label, the evaluation score of the medical question and answer system is obtained by weighted sum processing according to the entity overlap degree, the semantic similarity and the description evaluation weight.

[0150] Specifically, the weights of each index are adjusted adaptively according to the question type to realize dynamic weighted fusion. For example: the diagnosis class question focuses on entity and logical relationship matching, and the weight β(Qtype) of the entity overlap degree is increased; the description class question focuses on overall consistency, and the weight α(Qtype) of the semantic similarity is increased.

[0151] The final comprehensive score formula is as follows:

[0152] Score = α(Qtype) x semantic similarity + β(Qtype) x entity overlap degree;

[0153] Wherein α(Qtype), β(Qtype) are adaptively determined by the question type and α + β = 1; if the knowledge conflict detection is triggered, multiply by PenaltyFactor for weight reduction correction.

[0154] Further specifically, an embodiment in a specific scene is provided for further illustration:

[0155] 1. Scene setting

[0156] Diagnosis type question: Patient has chest pain for 2 hours, ECG shows ST segment elevation. What could it be?

[0157] Description type question: Explain the pathogenesis of coronary atherosclerosis.

[0158] 2. Classification label and weight configuration as shown in Table 1

[0159] Table 1: Label weight table configured in this example

[0160]

[0161] 3. Specific calculation process

[0162] Case 1: Diagnosis type question

[0163] Entity overlap: 0.7 (System answer contains core entities such as "heart attack", "ST segment elevation")

[0164] Semantic similarity: 0.8 (System answer is consistent with the pathological logic of the standard answer)

[0165] Evaluation score calculation:

[0166] 0.7 x 0.6 + 0.8 x 0.4 = 0.42 + 0.32 = 0.74

[0167] Case 2: Description type question

[0168] Entity overlap: 0.6 (System answer contains entities such as "lipid metabolism", "endothelial damage")

[0169] Semantic similarity: 0.9 (System answer is highly consistent with the pathogenesis description of the standard answer)

[0170] Evaluation score calculation:

[0171] 0.6 x 0.3 + 0.9 x 0.7 = 0.18 + 0.63 = 0.81

[0172] 4. Weight design basis

[0173] Diagnosis type question: Need to prioritize ensuring the accuracy of disease and symptoms, examination association (such as "ST segment elevation → heart attack"), so increase entity overlap weight.

[0174] Description type question: Need to focus on the completeness and consistency of the expression (such as "lipid metabolism abnormalities" whether corresponding to the standard answer "lipid infiltration"), so increase semantic similarity weight.

[0175] 5. Implementation effect

[0176] The diagnostic type problem score is more sensitive to entity omission (e.g., a 0.1 decrease in entity overlap degree results in a 0.06 decrease in total score);

[0177] The description type problem score is more sensitive to expression bias (e.g., a 0.1 decrease in semantic similarity results in a 0.07 decrease in total score);

[0178] The dynamic weight mechanism strongly associates the evaluation score with the characteristics of the problem type, thereby improving the evaluation pertinence.

[0179] In a preferred embodiment of the present application, the data preprocessing process is further included before the execution of step S2, and the data preprocessing process includes:

[0180] The system answer and the standard answer are subjected to word segmentation processing and stop word removal processing, and medical entities in the processed system answer and standard answer are identified.

[0181] The present application also provides an automatic evaluation system for a medical question and answer system, which applies the above automatic evaluation method, as shown in Figure 4 The automatic evaluation system comprises:

[0182] A question classification module 1 is configured to classify the input question of the medical question and answer system and label the classified question with a classification label;

[0183] A semantic similarity calculation module 2 is configured to calculate the semantic similarity between the system answer given by the medical question and answer system and the corresponding standard answer of the question;

[0184] An overlap degree calculation module 3 is configured to calculate the entity overlap degree based on the entity similarity between the system answer and the medical entities in the standard answer and the logical relationship similarity between the structures of the medical entities;

[0185] An evaluation module 4 is connected to the question classification module 1, the overlap degree calculation module 3 and the semantic similarity calculation module 2, and is configured to select an evaluation weight based on the classification label, and then perform weighted summation based on the entity overlap degree, the semantic similarity and the evaluation weight to obtain the evaluation score of the medical question and answer system.

[0186] In a preferred embodiment of the present application, a large language model is used in the question classification module for question classification, and a training module 5 for training the large language model for question classification is further included, which is configured to input a question classification prompt word to the large language model, and the classification prompt word includes a diagnostic type prompt word and a description type prompt word;

[0187] The large language model extracts the question semantics, and labels the question with a diagnostic type label when the question semantics conforms to the diagnostic type prompt word, and labels the question with a description type label when the question semantics conforms to the description type prompt word.

[0188] In a preferred embodiment of the present application, the semantic similarity calculation module 2 comprises:

[0189] a vector conversion unit 21 configured to convert the system answer and the standard answer into a system answer vector and a standard answer vector, respectively;

[0190] a similarity calculation unit 22 connected to the vector conversion unit 21 and configured to calculate a cosine similarity between the system answer vector and the standard answer vector as a semantic similarity.

[0191] In the preferred embodiment of the present application, the overlap calculation module 3 comprises:

[0192] a first calculation unit 31 configured to extract medical entities in the system answer and the standard answer to obtain a system answer entity set and a standard answer entity set, respectively, and to calculate a set similarity coefficient between the system answer entity set and the standard answer entity set as an entity similarity;

[0193] a second calculation unit 32 configured to construct a system answer graph and a standard answer graph according to the system answer entity set and the standard answer entity set, respectively, and to calculate a graph structure similarity between the system answer graph and the standard answer graph as a logical relationship similarity;

[0194] an overlap calculation unit 33 connected to the first calculation unit 31 and the second calculation unit 32 and configured to perform weighted summation processing on the entity similarity and the logical relationship similarity to obtain an entity overlap.

[0195] The above merely describes the preferred embodiments of the present application, and is not intended to limit the implementation manners and the protection scope of the present application. It should be realized by those skilled in the art that any equivalent replacement and obvious changes made according to the present application and the drawings should be included in the protection scope of the present application.

Claims

1. An automatic evaluation method of a medical question and answer system, characterized by, The method comprises the following steps: Step S1: classifying and tagging the input medical question and answer system question; Step S2: calculating the semantic similarity between the system answer given by the medical question and answer system and the standard answer corresponding to the question; Step S3: calculating the entity overlap degree based on the entity similarity between the system answer and the standard answer and the logical relationship similarity between the structures of medical entities; Step S4: selecting an evaluation weight based on the classification tag, and then performing weighted summation according to the entity overlap degree, the semantic similarity and the evaluation weight to obtain the evaluation score of the medical question and answer system.

2. The automated assessment method of claim 1, wherein, In step S1, a large language model is used for question classification. Before performing step S1, the large language model is trained for question classification, and the training process comprises: Inputting question classification prompt words into the large language model, wherein the classification prompt words include diagnosis class prompt words and description class prompt words; The large language model extracts the semantics of the question to obtain question semantics, and tags the question with a diagnosis class label when the question semantics conforms to the diagnosis class prompt words, and tags the question with a description class label when the question semantics conforms to the description class prompt words.

3. The automated assessment method of claim 1, wherein, Step S2 comprises: Step S21: converting the system answer and the standard answer into system answer vectors and standard answer vectors, respectively; Step S22: calculating the cosine similarity between the system answer vector and the standard answer vector as the semantic similarity.

4. The automated assessment method of claim 1, wherein, Step S3 comprises: Step S31: extracting medical entities in the system answer and the standard answer to obtain a system answer entity set and a standard answer entity set, respectively, and calculating the set similarity coefficient therebetween as the entity similarity; Step S32: constructing a system answer graph and a standard answer graph according to the system answer entity set and the standard answer entity set, respectively, and calculating the graph structure similarity between the system answer graph and the standard answer graph as the logical relationship similarity; Step S33: performing weighted summation processing on the entity similarity and the logical relationship similarity to obtain the entity overlap degree.

5. The automated assessment method of claim 1, wherein, The classification tags include diagnosis class tags and description class tags; the diagnosis class tags are associated with diagnosis evaluation weights, and the description class tags are associated with description evaluation weights; Step S4 comprises: Determine whether the classification tag is a diagnosis class tag or a description class tag: If it is a diagnosis class tag, perform weighted summation according to the entity overlap degree, the semantic similarity and the diagnosis evaluation weight to obtain the evaluation score of the medical question and answer system; If it is a description class tag, perform weighted summation according to the entity overlap degree, the semantic similarity and the description evaluation weight to obtain the evaluation score of the medical question and answer system.

6. The automated assessment method of claim 1, wherein, Before performing step S2, a data preprocessing process is further included, which comprises: Performing word segmentation and stop word removal processing on the system answer and the standard answer, and identifying the medical entities in the processed system answer and standard answer.

7. An automatic evaluation system of a medical question and answer system, characterized by, The application discloses an automatic evaluation method, comprising: a question classification module for classifying and tagging the input medical question and answer system question; a semantic similarity calculation module for calculating the semantic similarity between the system answer given by the medical question and answer system and the standard answer corresponding to the question; an overlap degree calculation module for calculating the entity overlap degree based on the entity similarity between the system answer and the medical entity in the standard answer and the logical relationship similarity between the structures of the medical entities; an evaluation module connected to the question classification module, the overlap degree calculation module and the semantic similarity calculation module, for selecting an evaluation weight based on the classification tag, and then performing weighted summation according to the entity overlap degree, the semantic similarity and the evaluation weight to obtain the evaluation score of the medical question and answer system.

8. The automated assessment system of claim 7, wherein, The question classification module uses a large language model for question classification, and further comprises a training module for training the large language model for question classification, for inputting question classification prompt words to the large language model, wherein the classification prompt words include diagnosis type prompt words and description type prompt words. The large language model extracts the question semantics from the question, and tags the question with a diagnosis type tag when the question semantics conforms to the diagnosis type prompt words, and tags the question with a description type tag when the question semantics conforms to the description type prompt words.

9. The automated assessment system of claim 7, wherein, The semantic similarity calculation module comprises: a vector conversion unit for converting the system answer and the standard answer into a system answer vector and a standard answer vector, respectively; a similarity calculation unit connected to the vector conversion unit, for calculating the cosine similarity between the system answer vector and the standard answer vector as the semantic similarity.

10. The automated assessment system of claim 7, wherein, The overlap degree calculation module comprises: a first calculation unit for extracting the medical entities in the system answer and the standard answer to obtain a system answer entity set and a standard answer entity set, respectively, and calculating the set similarity coefficient therebetween as the entity similarity; a second calculation unit for constructing a system answer graph and a standard answer graph according to the system answer entity set and the standard answer entity set, respectively, and calculating the graph structure similarity between the system answer graph and the standard answer graph as the logical relationship similarity; an overlap degree calculation unit connected to the first calculation unit and the second calculation unit, for performing weighted summation processing on the entity similarity and the logical relationship similarity to obtain the entity overlap degree.

Citation Information

Patent Citations

  • Medical question-answering system construction method based on language model and entity matching

    CN112667799A

  • Self-adaptive questionnaire question weight generation method and device, equipment and storage medium

    CN115952778A

  • Training scoring method based on natural language

    CN118467985A

  • Dialog system answering method based on sentence paraphrase recognition

    US20230069935A1