Graph matching medical text scoring method and device based on fact structure
By constructing triplets and knowledge graphs and combining with graph editing distance algorithms, the problem of insufficient accuracy of existing medical text evaluation methods is solved, and efficient and accurate evaluation of medical texts is achieved.
Patent Information
- Application Number
- CN202510257937.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-05
- Publication Date
- 2025-07-18
Smart Images

Figure CN120336865A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of artificial intelligence, and particularly to a graph matching medical text scoring method and device based on a fact structure. Background Art
[0002] With the rapid development of artificial intelligence and natural language processing technologies, text generation models are increasingly widely used in the medical field, especially showing great potential in medical text generation. These models can automatically generate structured medical texts based on information such as patients' symptoms, examination results, and medical histories.
[0003] However, medical texts are directly related to patients' health and life safety, and extremely high requirements are placed on their accuracy and quality. Traditional text similarity evaluation methods mainly focus on the similarity at the word and sentence levels, and often have difficulty capturing the semantic structure and key information in medical texts. These methods may ignore the accuracy of core contents such as diagnosis points and treatment suggestions, resulting in a deviation between the evaluation results and the actual medical value.
[0004] It can be seen that the medical text evaluation methods in the related technologies have the technical problem of low accuracy. Summary of the Invention
[0005] The present invention provides a graph matching medical text scoring method and device based on a fact structure, so as to solve the defect of low accuracy of the medical text evaluation method in the prior art, and realize providing a more accurate and comprehensive quality evaluation for medical texts.
[0006] The present invention provides a graph matching medical text scoring method based on a fact structure, including the following steps. Obtain a standard medical text and a test medical text; input the standard medical text and the test medical text into a pre-trained language model to obtain multiple standard keywords of the standard medical text and multiple test keywords of the test medical text output by the pre-trained language model; based on the multiple standard keywords and the multiple test keywords, respectively construct a standard triple of the standard medical text and a test triple of the test medical text; based on the standard triple and the test triple, respectively construct a standard knowledge graph of the standard medical text and a test knowledge graph of the test medical text; based on a preset graph edit distance algorithm, determine the graph edit distance between the standard knowledge graph and the test knowledge graph; normalize based on the graph edit distance to obtain an evaluation score of the test medical text.
[0007] A graph matching medical text scoring method based on a factual structure provided by the present invention. Based on the multiple standard keywords and the multiple test keywords, a standard triple of the standard medical text and a test triple of the test medical text are respectively constructed, including: constructing a standard triple based on each standard keyword in the multiple standard keywords and the weight of each standard keyword, where the standard triple includes: a standard head entity and a standard head entity weight, a standard relationship and a standard relationship weight, and a standard tail entity and a standard tail entity weight; constructing a test triple based on each test keyword in the multiple test keywords and the weight of each test keyword, where the test triple includes: a test head entity, a test relationship, and a test tail entity.
[0008] A graph matching medical text scoring method based on a factual structure provided by the present invention. Based on the standard triple and the test triple, a standard knowledge graph of the standard medical text and a test knowledge graph of the test medical text are respectively constructed, including: obtaining a preset knowledge graph structure; using the standard head entity and the standard tail entity in the standard triple as nodes and adding them to the preset knowledge graph structure to obtain a standard head entity node and a standard tail entity node; using the standard relationship in the standard triple as an edge and adding it to the preset knowledge graph structure to obtain a standard knowledge graph, where the edge is used to connect the standard head entity node and the standard tail entity node; wherein, the node weight in the standard knowledge graph corresponds to the weight of the standard keyword, and the edge weight in the standard knowledge graph corresponds to the sum of the node weights of the nodes associated with the edge.
[0009] A graph matching medical text scoring method based on a factual structure provided by the present invention. Based on the standard triple and the test triple, a standard knowledge graph of the standard medical text and a test knowledge graph of the test medical text are respectively constructed, including: obtaining a preset knowledge graph structure; using the test head entity and the test tail entity in the test triple as nodes and adding them to the preset knowledge graph structure to obtain a test head entity node and a test tail entity node; using the test relationship in the test triple as an edge and adding it to the preset knowledge graph structure to obtain a test knowledge graph, where the edge is used to connect the test head entity node and the test tail entity node; wherein, the node weight in the test knowledge graph corresponds to the weight of the test keyword, and the edge weight in the test knowledge graph corresponds to the sum of the node weights of the nodes associated with the edge.
[0010] A graph matching medical text scoring method based on a factual structure provided by the present invention, wherein the preset graph edit distance algorithm is a graph matching algorithm; determining the graph edit distance between the standard knowledge graph and the test knowledge graph based on the preset graph edit distance algorithm includes: performing weighted editing based on the standard knowledge graph and the test knowledge graph according to the graph matching algorithm to obtain a weighted graph edit distance, wherein the edit distance value of the weighted editing is the weight of the target editing content, and the target editing content includes at least one of the following: target nodes and target edges; determining the graph edit distance between the standard knowledge graph and the test knowledge graph based on the weighted graph edit distance, the number of nodes of the standard knowledge graph, and the number of nodes of the test knowledge graph.
[0011] A graph matching medical text scoring method based on a factual structure provided by the present invention, wherein determining the graph edit distance between the standard knowledge graph and the test knowledge graph based on the weighted graph edit distance, the number of nodes of the standard knowledge graph, and the number of nodes of the test knowledge graph includes: distance = Weighted GED / max(|G1|, |G2|); wherein, distance represents the graph edit distance, Weighted GED represents the weighted graph edit distance, G1 represents the number of nodes of the standard knowledge graph, and G2 represents the number of nodes of the test knowledge graph.
[0012] The present invention also provides a graph matching medical text scoring device based on a factual structure, including the following modules: an acquisition module for acquiring a standard medical text and a test medical text; a keyword module for inputting the standard medical text and the test medical text into a pre-trained language model to obtain a plurality of standard keywords of the standard medical text and a plurality of test keywords of the test medical text output by the pre-trained language model; a triple module for constructing a standard triple of the standard medical text and a test triple of the test medical text based on the plurality of standard keywords and the plurality of test keywords; a knowledge graph module for constructing a standard knowledge graph of the standard medical text and a test knowledge graph of the test medical text based on the standard triple and the test triple; an edit distance module for determining the graph edit distance between the standard knowledge graph and the test knowledge graph based on a preset graph edit distance algorithm; and an evaluation module for normalizing based on the graph edit distance to obtain an evaluation score of the test medical text.
[0013] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, the method for scoring medical texts based on fact structures as described in any one of the above is implemented.
[0014] The present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the method for scoring medical texts based on fact structures as described in any one of the above is implemented.
[0015] The present invention also provides a computer program product, including a computer program. When the computer program is executed by a processor, the method for scoring medical texts based on fact structures as described in any one of the above is implemented.
[0016] The method and device for scoring medical texts based on fact structures provided by the present invention automatically extract the keywords of standard medical texts and test medical texts through a pre-trained language model, avoiding the subjectivity and time-consuming nature of manual extraction; by constructing standard triples and test triples, the key information in medical texts is represented in a structured form, facilitating comparison and analysis; the construction of the standard knowledge graph and the test knowledge graph can capture the semantic relationships and context information in the corresponding keywords; using the graph edit distance algorithm to calculate the graph edit distance between the standard knowledge graph and the test knowledge graph can quantitatively evaluate the degree of difference between the test medical text and the standard medical text; through normalization processing, the graph edit distance is converted into an evaluation score, enabling texts of different lengths or medical contents of different complexities to be compared on the same scale. The normalized evaluation score is more intuitive and easy to understand, facilitating a quick judgment of the quality of the test medical text. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the accompanying drawings required for use in the description of the embodiments or the prior art will be briefly introduced one by one below. Obviously, the accompanying drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0018] Figure 1 is a schematic flowchart of the method for scoring medical texts based on fact structures provided by the present invention.
[0019] Figure 2 is the overall flowchart of the method for scoring medical texts based on fact structures provided by the present invention.
[0020] Figure 3 is a technical framework diagram of the method for scoring medical texts based on fact structures provided by the present invention.
[0021] Figure 4 It is a schematic structural diagram of a graph matching medical text scoring device based on a factual structure provided by the present invention.
[0022] Figure 5 It is a schematic physical structure diagram of an electronic device provided by the present invention. Detailed implementation manners
[0023] To make the objectives, technical solutions and advantages of the present invention clearer, the technical solutions in the present invention will be clearly and completely described below with reference to the accompanying drawings in the present invention. Obviously, the described embodiments are some but not all of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present invention without creative efforts shall fall within the protection scope of the present invention.
[0024] With the rapid development of artificial intelligence and natural language processing technologies, text generation models are increasingly widely used in the medical field, especially showing great potential in medical text generation. These models can automatically generate structured medical texts based on information such as patients' symptoms, examination results, and medical histories, and are expected to significantly improve medical efficiency and reduce the workload of medical staff.
[0025] However, medical texts are directly related to patients' health and life safety, and extremely high requirements are imposed on their accuracy and quality. Therefore, how to accurately evaluate the quality of medical texts has become an important challenge that needs to be solved urgently. Traditional text similarity evaluation methods, such as BLEU (Bilingual Evaluation Understudy), ROUGE (Recall-Oriented Understudy for Gisting Evaluation), etc., mainly focus on the similarity at the word and sentence levels, and often have difficulty capturing the semantic structure and key information in medical texts. These methods may ignore the accuracy of core contents such as diagnosis key points and treatment suggestions, resulting in a deviation between the evaluation results and the actual medical value.
[0026] In addition, medical texts have their particularities, including strong professionalism, containing a large number of medical terms and professional expressions; high structuralization degree, usually following specific formats and structures; large information density, where each description and diagnosis may contain key medical information; and high logical requirements, with strict logical relationships needed to be maintained among various parts. These characteristics pose numerous challenges when traditional text evaluation methods are applied to medical texts. For example, they may not be able to accurately identify and evaluate key elements such as causal relationships and diagnostic reasoning processes in medical texts. At the same time, with the extensive application of large language models in the medical field, some potential risks have emerged. For instance, the model may generate false or misleading medical information, which further highlights the importance of establishing an effective evaluation mechanism. Therefore, there is an urgent need for a method that can effectively evaluate the quality of medical texts and accurately capture the semantic structure and key information in medical texts. Based on the above background, the present invention proposes a graph matching evaluation method based on a factual structure, aiming to address the limitations of existing evaluation technologies and provide a more accurate and comprehensive quality evaluation for medical texts.
[0027] Optionally, the graph matching medical text scoring method based on a factual structure in the embodiments of the present application can be executed by a server, or by a terminal device, or jointly by a server and a terminal device. Taking the server to execute the graph matching medical text scoring method in this embodiment as an example.
[0028] Figure 1 is a schematic flowchart of the graph matching medical text scoring method based on a factual structure provided by the present invention, as Figure 1 shown, the method includes the following: Step 101, obtain a standard medical text and a test medical text.
[0029] In some embodiments, the standard medical text represents a text that is widely accepted and used in the medical field, with normativity, accuracy, and integrity. For example, the standard medical text can be a structured factual medical text in a preset format, including standard organs (or standard positions) and standard symptoms.
[0030] The test medical text includes description information and diagnostic information for the target symptom, but does not necessarily follow a preset format or structure. For example, the test medical text can come from unstructured or semi-structured data sources such as a patient's self-report, a doctor's preliminary diagnosis report, an emergency record, etc.
[0031] It should be noted that different data sources can be reasonably selected and utilized to obtain the standard medical text and the test medical text according to specific application scenarios and requirements to ensure the accuracy and effectiveness of the evaluation results.
[0032] Step 102: Input the standard medical text and the test medical text into the pre-trained language model to obtain multiple standard keywords of the standard medical text and multiple test keywords of the test medical text output by the pre-trained language model.
[0033] In the embodiments of the present invention, a pre-trained language model suitable for the medical field is selected, such as BioBERT, ClinicalBERT, or MedBERT, etc. These models have been pre-trained for medical texts and can better understand and process medical terms and concepts.
[0034] Preprocess the standard medical text and the test medical text, including removing irrelevant information, text cleaning (such as correcting spelling mistakes, unifying formats, etc.), and word segmentation.
[0035] Input the preprocessed standard medical text and test medical text into the selected pre-trained language model respectively; utilize the keyword extraction function of the pre-trained language model to extract multiple standard keywords from the standard medical text. Standard keywords usually represent the core medical concepts, disease names, symptom descriptions, etc. in the text. Similarly, extract multiple test keywords from the test medical text, and the test keywords reflect relevant content such as the target symptoms and diagnostic information described in the text.
[0036] In some embodiments, post-process the extracted keywords, such as removing duplicate keywords, merging near-synonyms or synonyms, etc., to improve the accuracy and practicality of the keywords.
[0037] Through the embodiments of the present invention, the pre-trained language model has mastered the professional terms and context information in the medical field through the training of a large number of medical texts, so it can more accurately extract medical-related keywords; using the pre-trained language model to extract the keywords of the standard medical text and the test medical text can improve the efficiency and quality of text processing.
[0038] Step 103: Based on the multiple standard keywords and the multiple test keywords, construct a standard triple of the standard medical text and a test triple of the test medical text respectively.
[0039] A triple usually consists of three parts: Entity, Relationship, and Attribute / Value. In medical text analysis, triples can be used to represent the association relationships between diseases, symptoms, treatments, etc.
[0040] The purpose of constructing the standard triple of the standard medical text and the test triple of the test medical text is to present the key information in the text and its mutual relationships in a formal way, facilitating subsequent information retrieval, comparison, and analysis.
[0041] In the embodiments of the present application, disease names, organ names, symptom descriptions, etc. are extracted from the keywords of standard medical texts as entities; the relationships between the entities are analyzed, such as "disease - symptom", "organ - disease", "treatment - disease", etc.; for each entity, its relevant attributes or values are extracted, such as the severity of the disease, the occurrence frequency of the symptom, etc.; the entities, relationships, and attributes / values are combined into the form of triples, such as (disease A, symptom, cough), (organ B, disease prone to, disease C), etc.
[0042] Similarly, disease names, symptom descriptions, etc. are extracted from the keywords of the test medical text as entities; the relationships between these entities in the text are analyzed, such as "patient symptom - diagnosed disease", "treatment attempt - effect", etc.; the identified entities, relationships, and attributes / values are combined into triples, such as (patient X, symptom, fever), (diagnosis, disease, pneumonia), etc.
[0043] Through the embodiments of the present invention, by constructing triples, the key information in the text can be presented in a structured form, improving the readability and analyzability of the information.
[0044] According to a graph matching medical text scoring method based on a fact structure provided by the present invention, based on multiple standard keywords and multiple test keywords, a standard triple of the standard medical text and a test triple of the test medical text are respectively constructed, including: Based on each standard keyword in the multiple standard keywords and the weight of each standard keyword, a standard triple is constructed, where the standard triple includes: a standard head entity and a standard head entity weight, a standard relationship and a standard relationship weight, and a standard tail entity and a standard tail entity weight; Based on each test keyword in the multiple test keywords and the weight of each test keyword, a test triple is constructed, where the test triple includes: a test head entity, a test relationship, and a test tail entity.
[0045] In the embodiments of the present invention, a pre - trained language model is used to extract the keywords and their weights of the standard medical text and the test medical text respectively, and triple information (head entity, relationship, tail entity) is extracted based on the keywords. For example, the weight of a node (entity) in the triple is obtained by the pre - trained language model; the edge weight is obtained by summing the weights of two adjacent nodes.
[0046] Among the extracted keywords (i.e., standard keywords or test keywords), identify the head entity (usually a noun or noun phrase that represents a main concept or object in the text) and the tail entity (another noun or noun phrase associated with the head entity). Determine the relationship between the head entity and the tail entity based on the contextual information in the text. This relationship can be "cause", "related to", "treatment", etc. Combine the head entity, relationship, and tail entity into a triple, such as (disease A, cause, symptom B).
[0047] Directly use the pre-trained language model to calculate the weights for the keywords (i.e., the head entity and the tail entity in the triplet) (according to the model's attention or importance score for the keyword, a weight is assigned to each keyword, and this weight reflects the relative importance of the keyword in the text). In order to reflect the strength of the relationship between the head entity and the tail entity in the triplet, the weights of the two adjacent nodes can be summed as the weight of the edge. That is, if the weight of the head entity is W1 and the weight of the tail entity is W2, the weight of the edge is W1+W2.
[0048] In some embodiments, one of the extracted standard keywords is selected as a head entity, which is usually a main concept or object in the text (such as a disease name, a drug name, etc.), and the weight corresponding to the head entity is used as the standard head entity weight.
[0049] Based on the contextual information in the text, determine the relationship between the head entity and other keywords. This relationship can be "lead to", "related to", "treatment", etc. The weight of the relationship can be set according to the actual situation, for example, it can be scored based on the relevance or importance between the head entity and the tail entity.
[0050] In some embodiments, if the relationship itself has clear weight hints in the text (such as "strongly related", "slightly related", etc.), these weights can be used directly. If there is no clear weight hint in the text, it can be inferred based on the weights of the head entity and the tail entity, such as taking the average of the weights of the two or performing weighted summation according to a preset rule.
[0051] A keyword associated with the head entity is selected from the extracted standard keywords as the tail entity, and the tail entity can be another concept, object, or symptom, etc. The weight corresponding to the tail entity is used as the standard tail entity weight.
[0052] The above three parts are combined into a standard triple, namely (standard head entity, standard relation, standard tail entity), and their respective weight information is attached.
[0053] The process of constructing the test triplet is similar to that in the above embodiment and will not be described again.
[0054] Through the embodiments of the present invention, by introducing keyword weights into structured triples, the core information and relationships in the text can be more accurately reflected, improving the accuracy of text analysis.
[0055] Step 104, based on the standard triples and test triples, construct the standard knowledge graph of the standard medical text and the test knowledge graph of the test medical text respectively.
[0056] In some embodiments, construct knowledge graphs for the standard medical text and the test medical text respectively. In the knowledge graph, entities are used as nodes and relationships are used as edges. The specific operations are as follows: create an empty graph structure; add the head entity and tail entity in each triple as nodes to the graph; use the relationship in the triple as an edge to connect the corresponding head entity and tail entity nodes; merge and calculate the weights for all nodes and edges.
[0057] In the embodiments of the present invention, select a graph database suitable for storing and querying the knowledge graph, such as Neo4j, OrientDB, etc.; these databases can efficiently store nodes (entities) and edges (relationships), and support complex graph queries and analyses.
[0058] Import the extracted standard triple data into the graph database. During the import process, it is necessary to ensure that the standard head entity, standard relationship, and standard tail entity of each standard triple can be correctly mapped to the nodes and edges in the database.
[0059] In the graph database, according to the standard triple data, the standard knowledge graph of the standard medical text includes creating standard nodes (representing entities), standard edges (representing relationships), and assigning attributes (such as weights, names, etc.) to the standard nodes and standard edges.
[0060] In the embodiments of the present invention, select a graph database suitable for storing and querying the knowledge graph, such as Neo4j, OrientDB, etc.; these databases can efficiently store nodes (entities) and edges (relationships), and support complex graph queries and analyses.
[0061] Import the extracted test triple data into the graph database. During the import process, it is necessary to ensure that the test head entity, test relationship, and test tail entity of each test triple can be correctly mapped to the nodes and edges in the database.
[0062] In the graph database, according to the test triple data, the test knowledge graph of the test medical text includes creating test nodes (representing entities), test edges (representing relationships), and assigning attributes (such as weights, names, etc.) to the test nodes and test edges.
[0063] A graph matching medical text scoring method based on a factual structure provided by the present invention constructs a standard knowledge graph of a standard medical text and a test knowledge graph of a test medical text based on a standard triple and a test triple, respectively, including: Obtain a preset knowledge graph structure; Use the standard head entity and the standard tail entity in the standard triple as nodes and add them to the preset knowledge graph structure to obtain a standard head entity node and a standard tail entity node; Use the standard relationship in the standard triple as an edge and add it to the preset knowledge graph structure to obtain a standard knowledge graph, where the edge is used to connect the standard head entity node and the standard tail entity node; Among them, the node weight in the standard knowledge graph corresponds to the weight of the standard keyword, and the edge weight in the standard knowledge graph corresponds to the sum of the node weights of the nodes associated with the edge.
[0064] In an embodiment of the present invention, all standard head entities and standard tail entities are extracted from the standard triple set; in the preset knowledge graph structure, a node is created for each unique standard head entity and standard tail entity; a unique node identifier is assigned to each node. According to the weight of the standard keyword, a corresponding weight attribute is set for each node; ensure that the node weight corresponds one-to-one with the weight of the standard keyword. All standard relationships are extracted from the standard triple set. In the preset knowledge graph structure, an edge is created between each pair of standard head entities and standard tail entities, and this edge represents the standard relationship between them; a unique edge identifier is assigned to each edge, and the type attribute of the edge is set to represent the type of the relationship; according to the sum of the weights of the two nodes associated with the edge, a corresponding weight attribute is set for each edge.
[0065] Through the embodiment of the present invention, when constructing the standard knowledge graph, not only the existence of entities (nodes) and relationships (edges) is considered, but also weight information is introduced. The node weight corresponds to the weight of the standard keyword, and the edge weight is calculated based on the weights of the associated nodes. The introduction of this weight information enables the knowledge graph to more accurately reflect the importance of entities and relationships.
[0066] A graph matching medical text scoring method based on a factual structure provided by the present invention constructs a standard knowledge graph of a standard medical text and a test knowledge graph of a test medical text based on a standard triple and a test triple, respectively, including: Obtain a preset knowledge graph structure; Use the test head entity and the test tail entity in the test triple as nodes and add them to the preset knowledge graph structure to obtain a test head entity node and a test tail entity node; Taking the test relationship in the test triple as an edge, add it to the preset knowledge graph structure to obtain a test knowledge graph, where the edge is used to connect the test head entity node and the test tail entity node; Among them, the node weight in the test knowledge graph corresponds to the weight of the test keyword, and the edge weight in the test knowledge graph corresponds to the sum of the node weights of the nodes associated with the edge.
[0067] In the embodiment of the present invention, all test head entities and test tail entities are extracted from the test triple set; in the preset knowledge graph structure, a node is created for each unique test head entity and test tail entity; a unique node identifier is assigned to each node. According to the weight of the test keyword, a corresponding weight attribute is set for each node; ensure that the node weight corresponds one-to-one with the weight of the test keyword. All test relationships are extracted from the test triple set. In the preset knowledge graph structure, an edge is created between each pair of test head entities and test tail entities, and this edge represents the test relationship between them; a unique edge identifier is assigned to each edge, and the type attribute of the edge is set to represent the type of the relationship; according to the sum of the weights of the two nodes associated with the edge, a corresponding weight attribute is set for each edge.
[0068] Through the embodiment of the present invention, the combination of the structured knowledge graph and the weight information enables the knowledge graph to support complex query and analysis tasks. For example, the path between specific entities can be found through the graph traversal algorithm, or sorting and filtering can be performed according to the weight information, so as to mine valuable information and patterns.
[0069] Step 105, based on the preset graph edit distance algorithm, determine the graph edit distance between the standard knowledge graph and the test knowledge graph.
[0070] In the embodiment of the present invention, the graph edit distance algorithm is used to calculate the distance between the standard knowledge graph and the test knowledge graph and obtain the weighted distance according to the node and edge weights. The specific operation is the graph matching algorithm, considering the operations of adding, deleting, and replacing nodes and adding, deleting, and replacing edges. The specific weighted calculation is that the cumulative distance value for each modification is not 1, but the weight size.
[0071] Here, the graph edit distance (GED) is a metric for measuring the similarity between two graphs. It measures the minimum number of edit operations required to transform one graph into another graph, and these operations include inserting, deleting, and modifying nodes and inserting, deleting, and modifying edges.
[0072] According to a graph matching medical text scoring method based on the fact structure provided by the present invention, the preset graph edit distance algorithm is the graph matching algorithm; Determine the graph edit distance between the standard knowledge graph and the test knowledge graph based on a preset graph edit distance algorithm, including: Based on the graph matching algorithm, perform weighted editing on the standard knowledge graph and the test knowledge graph to obtain the weighted graph edit distance, where the edit distance value of the weighted editing is the weight of the target editing content, and the target editing content includes at least one of the following: target nodes and target edges; Determine the graph edit distance between the standard knowledge graph and the test knowledge graph based on the weighted graph edit distance, the number of nodes in the standard knowledge graph, and the number of nodes in the test knowledge graph.
[0073] In the embodiments of the present invention, a graph matching algorithm (such as the Hungarian algorithm, the VF2 algorithm, etc.) is used to identify similar nodes and edges between the standard knowledge graph and the test knowledge graph. The goal of the graph matching algorithm is to find the best match between the two graphs, that is, to maximize the number of similar nodes and edges.
[0074] According to the result of the graph matching algorithm, determine the content to be edited, that is, the target editing content. This includes nodes (target nodes) and edges (target edges) that do not match or need to be modified. For each target node and target edge, calculate the cost of the editing operation according to its weight. The weight can be based on the importance, occurrence frequency, or other relevant attributes of the node or edge. The weighted graph edit distance is the sum of the weights of all target editing content, which reflects the minimum number of weighted editing operations required to convert the standard knowledge graph into the test knowledge graph (or vice versa).
[0075] Through the embodiments of the present invention, by calculating the weighted graph edit distance, the difference between the standard knowledge graph and the test knowledge graph can be accurately quantified. This quantification not only considers the presence or absence of nodes and edges, but also considers their weights, thus more comprehensively reflecting the similarity and difference degree between the graphs.
[0076] According to a graph matching medical text scoring method based on the fact structure provided by the present invention, determine the graph edit distance between the standard knowledge graph and the test knowledge graph based on the weighted graph edit distance, the number of nodes in the standard knowledge graph, and the number of nodes in the test knowledge graph, including: distance = Weighted GED / max(|G1|, |G2|); where distance represents the graph edit distance, Weighted GED represents the weighted graph edit distance, G1 represents the number of nodes in the standard knowledge graph, and G2 represents the number of nodes in the test knowledge graph.
[0077] In the embodiments of the present invention, the weighted graph edit distance (Weighted GED) is an important metric for measuring the difference between two graphs. It takes into account the weights of nodes and edges and can be calculated through a series of editing operations (such as insertion, deletion, replacement, etc.). The weighted graph edit distance usually better reflects the actual differences between graphs than the unweighted graph edit distance because it considers the importance of elements in the graphs.
[0078] The number of nodes is a simple and effective way to measure the scale of a graph. Here, |G1| and |G2| represent the number of nodes in the standard knowledge graph and the test knowledge graph respectively. Using max(|G1|, |G2|) as the denominator can ensure that the distance value does not fluctuate excessively due to different graph scales.
[0079] By dividing the weighted graph edit distance by max(|G1|, |G2|), a normalized graph edit distance is obtained. This distance value is between 0 and 1 (assuming that the Weighted GED is not greater than the maximum possible value of max(|G1|, |G2|)), and it is easier to interpret and compare. The closer the distance value is to 0, the more similar the two graphs are; the closer the distance value is to 1, the greater the difference between the two graphs.
[0080] Through the embodiments of the present invention, combining the weighted graph edit distance and the graph scale to obtain a normalized distance value helps to more accurately quantify the differences between graphs.
[0081] Step 106, normalize based on the graph edit distance to obtain the evaluation score of the test medical text.
[0082] In the embodiments of the present invention, the evaluation score of the test medical text can be determined by the following formula: Score = 100*(1 – distance / (max(distances)-min(distances))) Where score is the evaluation score of the final test medical text, distance is the graph edit distance between a single standard knowledge graph and the test knowledge graph, and distances is the set of all graph edit distances.
[0083] Through the embodiments of the present invention, converting the standard medical text (factual medical text) and the test medical text into graph structures and using the graph edit distance as a benchmark for scoring provides a new method for medical text quality assessment, overcomes the limitations of traditional text similarity assessment methods, can better capture the semantic structure and key information in medical texts, and helps to improve the accuracy of medical text assessment.
[0084] Through the above steps of the embodiments of the present invention, the keywords of the standard medical text and the test medical text are automatically extracted by a pre-trained language model, avoiding the subjectivity and time-consuming of manual extraction; by constructing the standard triples and test triples, the key information in the medical text is represented in a structured form, facilitating comparison and analysis; the construction of the standard knowledge graph and the test knowledge graph can capture the semantic relationships and context information in the corresponding keywords; by using the graph edit distance algorithm to calculate the graph edit distance between the standard knowledge graph and the test knowledge graph, the difference degree between the test medical text and the standard medical text can be quantitatively evaluated; through normalization processing, the graph edit distance is converted into an evaluation score, enabling texts of different lengths or medical contents of different complexities to be compared on the same scale, and the normalized evaluation score is more intuitive and easy to understand, facilitating the quick judgment of the quality of the test medical text.
[0085] Reference Figure 2 , Figure 2 is the overall flowchart of the graph matching medical text scoring method based on the fact structure provided by the present invention. It includes the following steps: Step 201, receive the factual medical text (i.e., the standard medical text) and the test medical text.
[0086] Step 202, use the language model to extract text keywords.
[0087] Step 203, extract triple information according to the text keywords.
[0088] Step 204, generate a knowledge graph according to the triple information.
[0089] Step 205, calculate the graph edit distance between the factual text and the test text.
[0090] Step 206, normalize the graph text distance to obtain a score.
[0091] Reference Figure 3 , Figure 3 is the technical framework diagram of the graph matching medical text scoring method based on the fact structure provided by the present invention. Specifically, it includes: 1. Keyword extraction, 2. Knowledge graph generation, and 3. Graph similarity.
[0092] Keyword extraction includes: inputting a preset text (A small amount of white viscous sputum can be seen in the tracheal lumen, without congestion, edema, no new growths, foreign bodies, or active bleeding) into a medical text model (a pre-trained language model) to obtain text keywords (A small amount (0.86) of white sputum (0.47) can be seen in the tracheal lumen. Among them, the numbers represent weights, and the yellow phrases represent keywords).
[0093] Knowledge graph generation includes: inputting preset text and text keywords into a large language model to obtain triple information, where the triple information includes (head entity, relationship, tail entity). For example, (mucosa, no, congestion), (mucosa, no, edema), (mucosa, no, erosion), (bronchus, no, tumor), (bronchus, no, foreign body), (bronchus, no, active bleeding).
[0094] Among them, the triple image includes: entities, entity-1, entity-2, entity-3, and the relationships between each entity.
[0095] Graph similarity includes: determining the graph similarity between a factual knowledge graph (standard knowledge graph) and a test knowledge graph based on entities and relationships. For example, the following formula can be used: Among them, represents the graph similarity between the factual knowledge graph (standard knowledge graph) and the test knowledge graph, represents the total number of nodes in the factual knowledge graph, represents the node index of the factual knowledge graph, represents the total number of nodes in the test knowledge graph, represents the node index of the test knowledge graph, represents the node of the factual knowledge graph and the node of the test knowledge graph the graph edit distance between them.
[0096] Taking the generation of reports by a large model in respiratory surgery as an example, the technical flow chart designed by this method is shown in Figure 2 . Specifically, the technical framework diagram is as shown in Figure 3 .
[0097] Step S0, use a pre-trained language model to extract the keywords of the standard respiratory report and the generated respiratory report respectively, and extract the triple information (head entity, relationship, tail entity) in the text based on the keywords.
[0098] Step S1, based on the triples extracted in step S1, construct graph structures for the standard report and the generated report respectively. In the graph, entities are used as nodes and relationships are used as edges. The specific operations are as follows: Create an empty graph structure; Add the head entity and tail entity in each triple as nodes to the graph; Use the relationship in the triple as an edge to connect the corresponding head entity and tail entity nodes; Merge and calculate the weights for all nodes and edges.
[0099] Step S2: Use the graph edit distance algorithm to calculate the distance between the standard report graph constructed in step S2 and the generated report graph, and obtain the weighted distance based on the node and edge weights. The specific operation is a graph matching algorithm, considering operations such as node addition, deletion, replacement, and edge addition, deletion, replacement.
[0100] The specific weighted calculation is that the cumulative distance value for each modification is not 1, but the weight size; Step S3: Normalize the graph edit distance calculated in step S3 to a score between 0 and 1.
[0101] distance = weightedGED / max(|G1|, |G2|); where weightedGED is the weighted graph edit distance, and |G1| and |G2| are the sizes (number of nodes) of the two graphs respectively.
[0102] Score = 100 * (1 – distance / (max(distances) - min(distances))); where score is the final score, distance is the distance of a single report, and distances is the set of all distances.
[0103] Step S4: Compare the normalized scores in step S4 to obtain the final score of the generated report. A threshold can be set to judge the report quality: 90 - 100 points: excellent; 80 - 89 points: good; 70 - 79 points: average; 60 - 69 points: pass; <60 points: fail.
[0104] The core of the technical solution of the present invention lies in using the graph edit distance to evaluate medical texts. By converting medical texts into a structured graph representation, the semantic information and entity relationships in the report can be better captured. Using the graph edit distance as an evaluation index not only considers the matching degree of facts and tests, but also can reflect the overall structural similarity of the report. This method is particularly suitable for medical texts with complex semantic structures such as medical reports and can provide more accurate and comprehensive evaluation results than traditional text similarity methods.
[0105] This method can be applied to various medical scenarios, such as the evaluation of medical AI-assisted diagnosis systems, report writing training in medical education, etc. Through objective and quantitative scoring, it can help improve the generation quality of medical reports and promote the application and development of artificial intelligence in the medical field.
[0106] The following describes the graph matching medical text scoring device based on the fact structure provided by the present invention. The graph matching medical text scoring device based on the fact structure described below can be correspondingly referred to the graph matching medical text scoring method based on the fact structure described above.
[0107] Reference Figure 4 , Figure 4 is a schematic structural diagram of the graph matching medical text scoring device based on the fact structure provided by the present invention.
[0108] An acquisition module 401, configured to acquire a standard medical text and a test medical text; A keyword module 402, configured to input the standard medical text and the test medical text into a pre-trained language model, and obtain a plurality of standard keywords of the standard medical text and a plurality of test keywords of the test medical text output by the pre-trained language model; A triple module 403, configured to construct a standard triple of the standard medical text and a test triple of the test medical text respectively based on the plurality of standard keywords and the plurality of test keywords; A knowledge graph module 404, configured to construct a standard knowledge graph of the standard medical text and a test knowledge graph of the test medical text respectively based on the standard triple and the test triple; An edit distance module 405, configured to determine the graph edit distance between the standard knowledge graph and the test knowledge graph based on a preset graph edit distance algorithm; An evaluation module 406, configured to normalize based on the graph edit distance to obtain an evaluation score of the test medical text.
[0109] Specifically, the above-mentioned graph matching medical text scoring device based on the fact structure provided by the present invention can implement all the method steps implemented by the above-mentioned graph matching medical text scoring method embodiment, and can achieve the same technical effects. The same parts and beneficial effects as those in the method embodiment will not be specifically described herein.
[0110] Figure 5 is a schematic physical structure diagram of the electronic device provided by the present invention, as Figure 5As shown in the figure, the electronic device may include: a processor 510, a communications interface 520, a memory 530, and a communication bus 540. Among them, the processor 510, the communications interface 520, and the memory 530 complete communication with each other through the communication bus 540. The processor 510 may call the logical instructions in the memory 530 to execute a graph matching medical text scoring method based on a factual structure. The method includes: obtaining a standard medical text and a test medical text; inputting the standard medical text and the test medical text into a pre-trained language model to obtain multiple standard keywords of the standard medical text and multiple test keywords of the test medical text output by the pre-trained language model; respectively constructing a standard triple of the standard medical text and a test triple of the test medical text based on the multiple standard keywords and the multiple test keywords; respectively constructing a standard knowledge graph of the standard medical text and a test knowledge graph of the test medical text based on the standard triple and the test triple; determining the graph edit distance between the standard knowledge graph and the test knowledge graph based on a preset graph edit distance algorithm; and normalizing based on the graph edit distance to obtain an evaluation score of the test medical text.
[0111] In addition, when the logical instructions in the above-mentioned memory 530 are implemented in the form of software functional units and sold or used as an independent product, they may be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of the technical solution, may be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The foregoing storage medium includes: various media such as a USB flash drive, a mobile hard disk, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), a magnetic disk, or an optical disc that can store program codes.
[0112] On the other hand, the present invention also provides a computer program product, which includes a computer program. The computer program can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the graph matching medical text scoring method based on a factual structure provided by each of the above methods. The method includes: obtaining a standard medical text and a test medical text; inputting the standard medical text and the test medical text into a pre-trained language model to obtain multiple standard keywords of the standard medical text and multiple test keywords of the test medical text output by the pre-trained language model; respectively constructing a standard triple of the standard medical text and a test triple of the test medical text based on the multiple standard keywords and the multiple test keywords; respectively constructing a standard knowledge graph of the standard medical text and a test knowledge graph of the test medical text based on the standard triple and the test triple; determining the graph edit distance between the standard knowledge graph and the test knowledge graph based on a preset graph edit distance algorithm; and normalizing based on the graph edit distance to obtain an evaluation score of the test medical text.
[0113] In another aspect, the present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it is configured to execute the graph matching medical text scoring method based on a factual structure provided by each of the above methods. The method includes: obtaining a standard medical text and a test medical text; inputting the standard medical text and the test medical text into a pre-trained language model to obtain multiple standard keywords of the standard medical text and multiple test keywords of the test medical text output by the pre-trained language model; respectively constructing a standard triple of the standard medical text and a test triple of the test medical text based on the multiple standard keywords and the multiple test keywords; respectively constructing a standard knowledge graph of the standard medical text and a test knowledge graph of the test medical text based on the standard triple and the test triple; determining the graph edit distance between the standard knowledge graph and the test knowledge graph based on a preset graph edit distance algorithm; and normalizing based on the graph edit distance to obtain an evaluation score of the test medical text.
[0114] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. A person of ordinary skill in the art can understand and implement it without creative efforts.
[0115] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on this understanding, the essence of the above technical solution or the part that contributes to the prior art can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.
[0116] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A graph matching medical text scoring method based on a factual structure, characterized in that, Including: Obtain a standard medical text and a test medical text; Input the standard medical text and the test medical text into a pre-trained language model to obtain multiple standard keywords of the standard medical text and multiple test keywords of the test medical text output by the pre-trained language model; Based on the multiple standard keywords and the multiple test keywords, construct a standard triple of the standard medical text and a test triple of the test medical text respectively; Based on the standard triple and the test triple, construct a standard knowledge graph of the standard medical text and a test knowledge graph of the test medical text respectively; Based on a preset graph edit distance algorithm, determine the graph edit distance between the standard knowledge graph and the test knowledge graph; Normalize based on the graph edit distance to obtain an evaluation score of the test medical text.
2. The graph matching medical text scoring method based on a factual structure according to claim 1, wherein, The constructing the standard triple of the standard medical text and the test triple of the test medical text respectively based on the multiple standard keywords and the multiple test keywords includes: Based on each standard keyword in the multiple standard keywords and the weight of each standard keyword, construct a standard triple, where the standard triple includes: a standard head entity and a standard head entity weight, a standard relation and a standard relation weight, and a standard tail entity and a standard tail entity weight; Based on each test keyword in the multiple test keywords and the weight of each test keyword, construct a test triple, where the test triple includes: a test head entity, a test relation, and a test tail entity.
3. The graph matching medical text scoring method based on a factual structure according to claim 2, wherein The constructing the standard knowledge graph of the standard medical text and the test knowledge graph of the test medical text respectively based on the standard triple and the test triple includes: Obtain a preset knowledge graph structure; Take the standard head entity and the standard tail entity in the standard triple as nodes and add them to the preset knowledge graph structure to obtain a standard head entity node and a standard tail entity node; Take the standard relation in the standard triple as an edge and add it to the preset knowledge graph structure to obtain a standard knowledge graph, where the edge is used to connect the standard head entity node and the standard tail entity node; Among them, the node weight in the standard knowledge graph corresponds to the weight of the standard keyword, and the edge weight in the standard knowledge graph corresponds to the sum of the node weights of the nodes associated with the edge.
4. The graph matching medical text scoring method based on a factual structure according to claim 2, wherein The constructing the standard knowledge graph of the standard medical text and the test knowledge graph of the test medical text respectively based on the standard triple and the test triple includes: Obtain a preset knowledge graph structure; Take the test head entity and the test tail entity in the test triple as nodes and add them to the preset knowledge graph structure to obtain a test head entity node and a test tail entity node; Take the test relation in the test triple as an edge and add it to the preset knowledge graph structure to obtain a test knowledge graph, where the edge is used to connect the test head entity node and the test tail entity node; Among them, the node weights in the test knowledge graph correspond to the weights of the test keywords, and the edge weights in the test knowledge graph correspond to the sum of the node weights of the nodes associated with the edge.
5. The method for scoring medical texts by graph matching based on a factual structure according to claim 1, wherein The preset graph edit distance algorithm is a graph matching algorithm; Based on the preset graph edit distance algorithm, determining the graph edit distance between the standard knowledge graph and the test knowledge graph includes: According to the graph matching algorithm, performing weighted editing based on the standard knowledge graph and the test knowledge graph to obtain a weighted graph edit distance, where the edit distance value of the weighted editing is the weight of the target edit content, and the target edit content includes at least one of the following: target nodes and target edges; Based on the weighted graph edit distance, the number of nodes in the standard knowledge graph, and the number of nodes in the test knowledge graph, determining the graph edit distance between the standard knowledge graph and the test knowledge graph.
6. The graph matching medical text scoring method based on a factual structure according to claim 5, wherein Based on the weighted graph edit distance, the number of nodes in the standard knowledge graph, and the number of nodes in the test knowledge graph, determining the graph edit distance between the standard knowledge graph and the test knowledge graph includes: distance = Weighted GED / max(|G1|, |G2|); Among them, distance represents the graph edit distance, Weighted GED represents the weighted graph edit distance, G1 represents the number of nodes in the standard knowledge graph, and G2 represents the number of nodes in the test knowledge graph.
7. A graph matching medical text scoring device based on a factual structure, characterized in that Including: An acquisition module for acquiring a standard medical text and a test medical text; A keyword module for inputting the standard medical text and the test medical text into a pre-trained language model to obtain multiple standard keywords of the standard medical text and multiple test keywords of the test medical text output by the pre-trained language model; A triple module for constructing a standard triple of the standard medical text and a test triple of the test medical text based on the multiple standard keywords and the multiple test keywords; A knowledge graph module for constructing a standard knowledge graph of the standard medical text and a test knowledge graph of the test medical text based on the standard triple and the test triple; An edit distance module for determining the graph edit distance between the standard knowledge graph and the test knowledge graph based on a preset graph edit distance algorithm; An evaluation module for normalizing based on the graph edit distance to obtain an evaluation score of the test medical text.
8. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and running on the processor, characterized in that, When the processor executes the computer program, it implements the fact-structure-based graph matching medical text scoring method according to any one of claims 1 to 6.
9. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the fact-structure-based graph matching medical text scoring method according to any one of claims 1 to 6.
10. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, it implements the fact-structure-based graph matching medical text scoring method according to any one of claims 1 to 6.
Citation Information
Cited By
LLM model-based text analysis method and device, medium and equipment
CN120929610A