Medical document quality detection method and system based on large model and knowledge graph

Through methods based on large models and knowledge graphs, combined with large language models to identify abnormal triple information, the problem of insufficient reliability of medical documents quality detection in the existing technology is solved, and more comprehensive and accurate quality detection is achieved.

CN120163144AActive Publication Date: 2025-06-17SINGULARITY INTELLIGENCE (BEIJING) TECH CO LTD

Patent Information

Application Number
CN202510648110.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-20
Publication Date
2025-06-17
Estimated Expiration
2045-05-20

AI Technical Summary

Technical Problem

The prior art is difficult to effectively detect complex medical knowledge and semantic scenarios in medical documents, resulting in insufficient reliability of quality testing.

Method used

Using a method based on large models and knowledge graphs, we can obtain the attribute labels of medical documents, match the quality detection standards in the standard database, and use the large language model to identify abnormal triple information to generate a quality detection report.

Benefits of technology

This method can fully cover the medical concepts, relationships and logic in medical documents, accurately identify various abnormal information, and greatly improve the reliability of quality testing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120163144A_ABST
    Figure CN120163144A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of data processing, in particular to a medical document quality detection method and system based on a large model and a knowledge graph, and the method comprises the steps: obtaining a corresponding target quality detection standard through matching a target attribute tag of a target medical document with a preset attribute tag of a preset quality detection standard in a standard database; in combination with deep understanding and analysis of a preset large language model on medical document contents and target quality detection standards, various medical concepts, relationships and logics involved in the medical document can be comprehensively covered, and various abnormal triple information such as logic contradictions, data errors, semantic errors, missing records, format errors and the like in the medical document can be accurately identified; and the reliability of quality detection is greatly improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of data processing, and in particular to a method and system for quality detection of medical documents based on large models and knowledge graphs. Background Art

[0002] As an important record carrier of medical activities, the quality of medical documents directly affects the reliability of clinical diagnosis, treatment decisions, and the handling of medical disputes. With the rapid development of medical informatization, the number of medical documents has increased explosively. Traditional manual review methods are difficult to meet the needs of large-scale medical document quality detection due to problems such as low efficiency, strong subjectivity, and being easily affected by differences in the professional levels of reviewers.

[0003] In the prior art, the quality of medical documents is automatically detected based on rule engines and simple machine learning algorithms. However, the method based on rule engines relies on a rule library written manually, which is difficult to cover complex and changing medical knowledge and semantic scenarios, and cannot effectively detect situations not covered by the rules; simple machine learning algorithms show deficiencies such as low accuracy and poor generalization ability when dealing with semantic understanding, logical reasoning, etc. in medical documents. In addition, medical knowledge has characteristics such as strong professionalism, rapid update, and complex relationships. Existing methods are difficult to make full use of the association relationships between medical knowledge and cannot deeply explore potential quality problems in medical documents, resulting in insufficient reliability of detection results.

[0004] Therefore, how to improve the reliability of medical document quality detection has become an urgent problem to be solved. Summary of the Invention

[0005] In view of the above technical problems, the technical solution adopted by the present invention is a method for quality detection of medical documents based on large models and knowledge graphs. The quality detection method includes: S10, obtaining a target medical document to be detected and target attribute tags corresponding to the target medical document.

[0006] S20, according to the target attribute tags, obtaining target quality detection standards corresponding to the target medical document from a standard database, wherein the standard database stores preset quality detection standards corresponding to each preset attribute tag, each preset quality detection standard includes preset quality evaluation content and preset quality scoring criteria, and all the preset quality detection standards are stored in the standard database in the form of a knowledge graph.

[0007] S30, inputting the target medical document and the target quality detection standards into a preset large language model to obtain abnormal triple information corresponding to the target medical document.

[0008] S40. Obtain the quality inspection report corresponding to the target medical document according to the target medical document and the abnormal triple information.

[0009] The present invention also provides a quality inspection system for medical documents based on a large model and a knowledge graph. The quality inspection system includes a processor and a memory. A standard database is stored in the memory. Among them, the standard database includes T standard medical texts, and the attribute labels and knowledge graph information corresponding to each standard medical document. T is a positive integer. The processor includes: A data acquisition module, configured to acquire the target medical document to be detected and the target attribute label corresponding to the target medical document.

[0010] A quality inspection standard acquisition module, configured to acquire the target quality inspection standard corresponding to the target medical document from the standard database according to the target attribute label. Among them, the standard database stores the preset quality inspection standards corresponding to each preset attribute label. Each preset quality inspection standard includes preset quality evaluation content and preset quality scoring criteria, and all the preset quality inspection standards are stored in the standard database in the form of a knowledge graph.

[0011] An abnormal information acquisition module, configured to input the target medical document and the target quality inspection standard into a preset large language model to obtain the abnormal triple information corresponding to the target medical document.

[0012] A detection report acquisition module, configured to obtain the quality inspection report corresponding to the target medical document according to the target medical document and the abnormal triple information.

[0013] The present invention has at least the following beneficial effects: By matching the target attribute label of the target medical document with the preset attribute label of the preset quality inspection standard in the standard database, the corresponding target quality inspection standard is obtained. Combining the in-depth understanding and analysis of the medical document content and the target quality inspection standard by the preset large language model, various medical concepts, relationships, and logics involved in the medical document can be comprehensively covered, and various abnormal triple information such as logical contradictions, data errors, semantic errors, omissions, and format errors in the medical document can be accurately identified, greatly improving the reliability of quality inspection. Description of the Drawings

[0014] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0015] Figure 1Flowchart of a medical document quality detection method based on a large model and a knowledge graph provided in Embodiment 1 of the present invention; Figure 2 Schematic diagram of the target processor in a medical document quality detection system based on a large model and a knowledge graph provided in Embodiment 2 of the present invention; Figure 3 Flowchart of a computer program executed by a medical document quality detection system provided in Embodiment 3 of the present invention. Detailed implementation manners

[0016] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative efforts fall within the protection scope of the present invention.

[0017] It should be noted that the terms "first", "second", etc. in the specification and claims of the present invention and the above-mentioned accompanying drawings are used to distinguish similar objects, and do not necessarily need to describe a specific order or sequence. It can be understood that, under appropriate circumstances, the above-mentioned terms for distinguishing similar objects can be interchanged so that the present invention can also implement other embodiments other than the illustrated embodiments or the described embodiments. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or server including a series of steps or units does not necessarily need to be limited to those clearly listed steps or units, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.

[0018] Embodiment 1 Embodiment 1 of the present invention provides a medical document quality detection method based on a large model and a knowledge graph. Refer to Figure 1 , and the quality detection method includes: S10. Obtain the target medical document to be detected and the target attribute label corresponding to the target medical document.

[0019] Among them, the target medical document refers to a specific medical document that needs to be subjected to quality detection, such as a medical record, a diagnosis report, a surgical record, etc., which record important medical information such as patient information, condition information, diagnosis process, treatment plan, etc.

[0020] The target attribute label can be an abstract description of the key features of the target medical document, and is used to identify attributes such as the type of the document, the department to which it belongs, the types of diseases involved, and the diagnosis and treatment stage. For example, the attribute label can be "Cardiology - Coronary Heart Disease - Inpatient Medical Record".

[0021] Obtain the target medical document from the medical information system through interface calls, file reading, etc. At the same time, according to the content and metadata of the document, automatically or manually label and generate target attribute tags, and establish the corresponding relationship between the medical document and the attribute tags.

[0022] As described above, the target attribute tags provide a basis for accurately retrieving the standard database and obtaining the corresponding quality inspection standards in the subsequent steps, making the inspection work more targeted, improving the inspection efficiency and reliability, avoiding blind retrieval and analysis, and reducing the waste of computing resources.

[0023] S20. According to the target attribute tags, obtain the target quality inspection standards corresponding to the target medical document from the standard database. Among them, the standard database stores the preset quality inspection standards corresponding to each preset attribute tag. Each preset quality inspection standard includes preset quality evaluation content and preset quality scoring criteria, and all the preset quality inspection standards are stored in the standard database in the form of a knowledge graph.

[0024] In a specific embodiment, S20 includes the following steps: Determine the preset quality inspection standard whose corresponding preset attribute tag is consistent with the target attribute tag as the target quality inspection standard.

[0025] Among them, the standard database is a database specifically used to store various preset quality inspection standards, which contains quality inspection standard information for a large number of different types of medical documents, providing a unified standard and basis for the quality inspection of medical documents. The preset attribute tags are a series of attribute tags predefined in the standard database, covering the attribute characteristics of various possible medical documents, used to identify the type of the document, the department it belongs to, the disease types involved, the diagnosis and treatment stage, etc. attributes, and can be matched with the target attribute tags of the target medical document to determine the corresponding quality inspection standards.

[0026] The preset quality inspection standards refer to the detailed quality inspection specifications formulated for each preset attribute tag, including preset quality evaluation content and preset quality scoring criteria, used to clarify for medical documents with specific attribute tags, from which aspects the quality should be evaluated and how to score.

[0027] The preset quality evaluation content is a part of the preset quality inspection standards, which specifically stipulates all aspects of the quality evaluation of medical documents, such as the integrity, accuracy, standardization, logic, etc. of the document content, and clarifies the specific items and indicators that need to be inspected and evaluated.

[0028] The preset quality scoring criteria are part of the preset quality detection criteria. According to the preset quality evaluation content, corresponding scoring rules and methods are formulated to quantitatively score the performance of medical documents in each evaluation item, so as to finally obtain a comprehensive quality score, thereby intuitively reflecting the quality level of medical documents.

[0029] A knowledge graph is a graph-based data structure composed of nodes and edges, which can clearly represent various nodes and the relationships between them. In this embodiment, each preset attribute label, preset quality evaluation content, and preset quality scoring criteria can be regarded as nodes in the knowledge graph, and the association relationships between the preset attribute labels, preset quality evaluation content, and preset quality scoring criteria are used as edges. Through the structure of the knowledge graph, complex quality detection standard information can be efficiently organized and managed, facilitating quick query and acquisition of the target quality detection criteria corresponding to the target attribute label.

[0030] By traversing the standard database, the target attribute label is matched with the predicted attribute label. If they are the same, the corresponding preset quality detection criteria are determined as the target quality detection criteria.

[0031] As described above, the preset quality detection criteria corresponding to the standard database and the knowledge graph provide a rich and accurate medical knowledge reference system for the target medical document, enabling the subsequent large language model to analyze the medical document based on the complete knowledge background and quality detection criteria, thereby more comprehensively and deeply discovering potential quality problems and improving the comprehensiveness and reliability of detection.

[0032] S30. Input the target medical document and the target quality detection criteria into the preset large language model to obtain the abnormal triple information corresponding to the target medical document.

[0033] In a specific embodiment, S30 includes the following steps: S301. Input the target medical document and the target quality detection criteria into the preset large language model to obtain a number of abnormal triples corresponding to the target medical document and the abnormal type corresponding to each abnormal triple.

[0034] S302. According to the abnormal type corresponding to each abnormal triple, obtain the quantity proportion corresponding to each abnormal type.

[0035] S303. Take the abnormal triples, quantity proportion, and preset severity weight value corresponding to each abnormal type as the abnormal triple information corresponding to the target medical document.

[0036] Among them, the pre-set large language model refers to a language model trained with a large amount of medical text data, which has the ability to understand and generate natural language, can analyze the semantics and logic in medical documents, compare the entities and relationships involved with the knowledge graph, and judge whether each triple is reasonable according to the pre-set rules and the knowledge learned through training, and identify the abnormal triple information with problems.

[0037] An abnormal triple refers to a combination of "entity-relationship-entity" that does not conform to the logic and norms of medical knowledge identified by the large language model in the target medical document. For example, "penicillin, treatment, patients allergic to penicillin" violates medical common sense and belongs to an abnormal triple.

[0038] The abnormal type refers to the classification of the nature of the error of the abnormal triple, including logical contradiction, which emphasizes the logical conflict between different contents in the medical document. For example, the diagnosis is "pneumonia", but the treatment plan is for "gastritis" medication; data error, which focuses on numerical and data-related errors, such as incorrect recording of the patient's age, incorrect entry of test index data and other data abnormalities; semantic error, which pays attention to problems at the semantic level such as medical terms and sentence expressions, such as miswriting "coronary atherosclerotic heart disease" as "coronary atherosclerosis heart disease", affecting the correct transmission of semantics; omission of records, which emphasizes the situation where medical information that should be recorded is missing, such as the patient's allergy history; format error, which emphasizes abnormal situations of the recording format.

[0039] The quantity proportion refers to the proportion of the number of abnormal triples corresponding to each abnormal type in the total number of all abnormal triples, which is used to measure the distribution of different abnormal types in the target medical document.

[0040] The pre-set severity weight refers to a value preset for quantifying the severity of the abnormality based on the differences in the impact of different abnormal types on the quality of medical documents and medical decisions. For example, the weight corresponding to the logical contradiction type involving the patient's life safety is relatively high, while the weight corresponding to the format error type is relatively low.

[0041] Integrate the specific abnormal triples corresponding to each abnormal type, the calculated quantity proportion, and the pre-set severity weight to form an abnormal triple information set containing multi-dimensional information, which comprehensively describes the specific content, distribution, and severity of the abnormal problems in the target medical document, providing rich and structured data support for the subsequent generation of quality inspection reports.

[0042] As described above, through the combination of the large language model and the target quality detection standard, the intelligent semantic understanding and logical reasoning of the content of medical documents are realized, and the errors and contradictions in the medical documents can be quickly and accurately located, and an abnormal triple information set containing multi-dimensional information can be obtained, which comprehensively describes the specific content, distribution and severity of the abnormal problems in the target medical document, and improves the reliability of subsequent abnormal detection of medical documents.

[0043] S40. According to the target medical document and the abnormal triple information, obtain the quality detection report corresponding to the target medical document.

[0044] Among them, the quality detection report refers to a document that records the quality status of the target medical document, comprehensively summarizes problems, analyzes reasons and gives improvement suggestions, and is an intuitive presentation of the detection results.

[0045] In a specific embodiment, S40 includes the following steps: S401. Input the target medical document and the abnormal triple information into a preset large language model to obtain the quality detection report corresponding to the target medical document.

[0046] Among them, the preset large language model, based on the medical knowledge, quality detection logic and report generation mode learned through previous training, receives the target medical document text and the abnormal triple information, analyzes and integrates the abnormal information, combines the overall content and characteristics of the medical document, and uses natural language generation technology to automatically generate a quality detection report covering problem description, quality scoring, reason analysis, improvement suggestions, etc.

[0047] In a specific embodiment, S40 further includes the following steps: S402. Obtain a preset quality detection report template.

[0048] S403. Input the target medical document, the abnormal triple information and the preset quality detection report template into a preset large language model to obtain the quality detection report corresponding to the target medical document.

[0049] Among them, the preset quality detection report template refers to a document template that is pre-designed and has a fixed format and content framework, including report title, basic information of the medical document, summary of abnormal situations, quality scoring, improvement suggestions, etc., and provides a unified and standardized format for the quality detection report.

[0050] Specifically, a quality inspection report template that meets the requirements can be obtained from a database or file system storing report templates through program instructions or interface calls. After receiving the target medical document, abnormal triple information, and the quality inspection report template, the pre-set large language model uses the quality inspection report template as a framework to fill in the corresponding sections of the template with various abnormal problems and their analysis results in the abnormal triple information. For example, the abnormal type, quantity ratio, and quality score are filled into the summary section of abnormal situations. Combining the specific content and characteristics of the target medical document, content generation is performed in the quality analysis conclusion and improvement suggestion sections to ensure that the generated report content is perfectly integrated with the template format, with logical coherence and accurate expression.

[0051] As described above, unifying the report format based on the quality inspection report template makes the quality inspection reports of different medical documents consistent and standardized, facilitating reading, comparison, and management. Moreover, the standardized template can guide the large language model to generate more comprehensive and well-organized report content, avoiding omission of important information, and enhancing the professionalism and readability of the quality inspection report.

[0052] As described above, by matching the target attribute tags of the target medical document with the preset attribute tags of the preset quality inspection standards in the standard database to obtain the corresponding target quality inspection standards, and combining the in-depth understanding and analysis of the medical document content and the target quality inspection standards by the pre-set large language model, various medical concepts, relationships, and logics involved in the medical document can be comprehensively covered, and various abnormal triple information such as logical contradictions, data errors, semantic errors, omission of records, and format errors in the medical document can be accurately identified, greatly improving the reliability of quality inspection, providing clear and intuitive medical document quality assessment results for medical staff and management personnel, helping to quickly understand the problems and severity of the document, so as to take targeted improvement measures in a timely manner, improve the quality of medical documents, and further ensure the reliability and dependability of medical information, providing strong support for medical decision-making, and also helping to improve the level of medical quality management.

[0053] Embodiment 2 Embodiment 2 provides a quality inspection system for medical documents based on a large model and a knowledge graph. The quality inspection system includes: a target processor and a target memory. The target memory stores a standard database, where the standard database includes T standard medical texts, and the attribute tags and knowledge graph information corresponding to each standard medical document. T is a positive integer, as Figure 2 shown, the target processor includes: A data acquisition module 21, configured to acquire a target medical document to be detected and the target attribute tags corresponding to the target medical document.

[0054] A quality inspection standard acquisition module 22, configured to obtain a target quality inspection standard corresponding to a target medical document from a standard database according to a target attribute label, where the standard database stores a preset quality inspection standard corresponding to each preset attribute label, each preset quality inspection standard includes a preset quality evaluation content and a preset quality scoring standard, and all the preset quality inspection standards are stored in the standard database in the form of a knowledge graph.

[0055] An exception information acquisition module 23, configured to input the target medical document and the target quality inspection standard into a preset large language model to obtain exception triple information corresponding to the target medical document.

[0056] A detection report acquisition module 24, configured to obtain a quality inspection report corresponding to the target medical document according to the target medical document and the exception triple information.

[0057] In a specific embodiment, the quality inspection standard acquisition module 22 includes: A target quality inspection standard acquisition sub-module, configured to determine a preset quality inspection standard whose corresponding preset attribute label is consistent with the target attribute label as the target quality inspection standard.

[0058] In a specific embodiment, the exception information acquisition module 23 includes: An exception type acquisition sub-module, configured to input the target medical document and the target quality inspection standard into a preset large language model to obtain a plurality of exception triples corresponding to the target medical document and an exception type corresponding to each exception triple.

[0059] A quantity proportion acquisition sub-module, configured to obtain a quantity proportion corresponding to each exception type according to the exception type corresponding to each exception triple.

[0060] An exception information acquisition sub-module, configured to use the exception triple, quantity proportion, and preset severity weight corresponding to each exception type as the exception triple information corresponding to the target medical document.

[0061] In a specific embodiment, the detection report acquisition module 24 includes: A first detection report acquisition sub-module, configured to input the target medical document and the exception triple information into a preset large language model to obtain a quality inspection report corresponding to the target medical document.

[0062] In a specific embodiment, the detection report acquisition module 24 includes: A report template acquisition sub-module, configured to obtain a preset quality inspection report template.

[0063] The first detection report acquisition sub-module is used to input the target medical document, abnormal triple information, and a preset quality detection report template into a preset large language model to obtain the quality detection report corresponding to the target medical document.

[0064] Embodiment 3 Embodiment 3 provides a quality detection system for medical documents. The detection system for medical documents includes a processor and a memory storing a computer program. The memory also stores a reference database. Among them, the reference database includes M reference triples and reference evaluation values respectively corresponding to the M reference triples. The reference triple includes a first reference entity, a second reference entity, and a reference relationship. The reference relationship includes N-level association relationships. Both M and N are positive integers. As Figure 3 shown, when the computer program is executed by the processor, the following steps are implemented: S1. Input the obtained target medical document and several preset attribute information into a preset large language model to obtain several target entities corresponding to the target medical document. Among them, the preset attribute information at least includes symptom information, drug information, and examination information.

[0065] S2. According to the association information of each preset attribute information, form K target triples from all target entities. Among them, the target triple includes a first target entity, a second target entity, and a target relationship. The target relationship includes N-level association relationships. K is a positive integer.

[0066] S3. For any target triple, match the current target triple with each reference triple to obtain the reference triple corresponding to the current target triple, and use the reference evaluation value of the reference triple corresponding to the current target triple as the target evaluation value of the current target triple.

[0067] S4. Determine the first comprehensive evaluation value according to each target triple and the target evaluation values respectively corresponding to each target triple.

[0068] S5. If the first comprehensive evaluation value does not meet the first preset condition, input all target triples and the target evaluation values corresponding to each target triple into a preset large language model to obtain the quality detection report corresponding to the target medical document.

[0069] Among them, the reference database can be constructed according to several reference texts, and the reference texts can refer to the obtained historical medical documents. For any reference text, extract several reference entities from the current reference text, determine the level of the association relationship between the current two reference entities as the reference relationship according to the preset attribute information to which any two reference entities respectively belong, and determine one reference entity among the current two reference entities as the first reference entity and the other as the second reference entity to form a reference triple.

[0070] It should be noted that when there are the same reference entities in a single reference text, only one reference entity is retained for constructing reference triples. When the constructed reference triples are the same, only one reference triple is retained.

[0071] In this embodiment, the preset attribute information may further include diagnostic information, treatment plan information, etc. There is a sequence for the preset attribute information, and the arrangement result of each preset attribute information in the sequence may be symptom information, diagnostic information, treatment plan information, drug information, and examination information.

[0072] In a specific implementation manner, the preset attribute information to which the reference entity belongs may also be name, feature, feature category, plan, project, object, etc., which can be determined by the implementer according to the actual application scenario.

[0073] The level of the association relationship can be determined according to the number of preset attribute information intervals between two preset attribute information in the sequence of preset attribute information arranged in sequence. When two preset attribute information are directly adjacent, the level of the association relationship is determined to be 1. When there is one preset attribute information interval between two preset attribute information, the level of the association relationship is determined to be 2, and so on. Continuing with the above example, the first target entity may correspond to symptom information, and the second target entity may correspond to examination information. There are a total of 3 preset attribute information intervals between them, namely diagnostic information, treatment plan information, and drug information. Then the level of the association relationship can be determined to be 4.

[0074] The reference evaluation value can represent the importance of the corresponding reference triple in the reference text.

[0075] When using a preset large language model for entity extraction in medical documents, the extraction of reference entities can also use this preset large language model.

[0076] The target evaluation value can represent the importance of the corresponding target triple in the target medical document, and the first comprehensive evaluation value can be used to evaluate the quality of the target medical document.

[0077] When the first comprehensive evaluation value does not meet the first preset condition, the preset large language model, based on the medical knowledge, quality detection logic, and report generation mode learned from previous training, receives all target triples and the target evaluation value corresponding to each target triple, analyzes and integrates the abnormal information, combines the overall content and characteristics of the medical document, and uses natural language generation technology to automatically generate a quality detection report covering problem description, cause analysis, improvement suggestions, etc.

[0078] In a specific implementation manner, when the computer program is executed by the processor, the following steps are also implemented: For any reference triple, determine the reference evaluation value corresponding to the current reference triple according to the reference relationship and the expert evaluation value in the current reference triple.

[0079] Among them, the expert evaluation value can be evaluated by experts in the field applied by the implementer on the importance of the current reference triple in the reference text, and the expert evaluation value can be a normalized value.

[0080] Specifically, the reference relationship can be normalized, that is, after subtracting the maximum value N - 1 of the reference relationship from the current reference relationship, and then dividing by the difference N - 2 between the maximum value N - 1 and the minimum value 1 of the reference relationship, the normalized result of the current reference relationship is obtained.

[0081] Multiply the normalized result of the current reference relationship by the expert evaluation value, and use the multiplication result as the reference evaluation value corresponding to the current reference triple.

[0082] In a specific embodiment, S2 includes the following steps: S21, for any two target entities, use one target entity as the current first target entity and the other target entity as the current second target entity.

[0083] S22, according to the association information between the preset attribute information corresponding to the current first target entity and the preset attribute information corresponding to the current second target entity, determine the level n of the association relationship corresponding to the current first target entity and the current second target entity, and form a target triple from the current first target entity, the current second target entity and the level n of the association relationship, where n is an integer within the range of [1, N].

[0084] Among them, in this embodiment, there is a sequence for the preset attribute information to facilitate the determination of the first target entity and the second target entity. The level of the association relationship can be determined according to the number of preset attribute information intervals between two preset attribute information in the sequence of preset attribute information arranged in sequence. For example, when two preset attribute information are directly adjacent, the level of the association relationship is determined to be 1, and when there is one preset attribute information interval between two preset attribute information, the level of the association relationship is determined to be 2, and so on.

[0085] In a specific embodiment, S3 includes the following steps: S31, for any target triple, match the current target triple with each reference triple. If there is a reference triple that is the same as the current target triple, use this reference triple as the reference triple corresponding to the current target triple, and use the reference evaluation value of the reference triple corresponding to the current target triple as the target evaluation value of the current target triple.

[0086] S32. If there is no reference triple that is the same as the current target triple, use the first preset value as the target evaluation value of the current target triple.

[0087] Among them, the matching can be implemented by calculating the similarity. In this embodiment, the similarity calculation can be implemented by the cosine similarity method. The value range of the cosine similarity is [-1, 1]. It should be noted that the matching in this embodiment is a complete match. That is, when using the cosine similarity, if the cosine similarity between the current target triple and a reference triple is 1, it is considered that the current target triple is the same as the reference triple; otherwise, it is considered that the current target triple is different from the reference triple.

[0088] Specifically, since there may be text errors resulting in no reference triple that is the same as the current target triple, in this case, use the first preset value as the target evaluation value of the current target triple. The first preset value can be set to 0.

[0089] In a specific implementation, S4 includes the following steps: S41. Form triple sets from the target triples with the same first target entity and target relationship, obtaining Q triple sets, where Q is a positive integer.

[0090] S42. For any triple set, calculate the average value of the target evaluation values corresponding to each target triple in the current triple set, and use the average calculation result as the comprehensive sub-evaluation value.

[0091] S43. Add up the comprehensive sub-evaluation values corresponding to all triple sets, and use the addition result as the first comprehensive evaluation value.

[0092] Among them, there may be multiple other target entities with the same target relationship for one target entity in the target medical document. For example, one target entity corresponds to symptom information, and this symptom information can correspond to multiple other target entities of drug information. In order to avoid the multiple target triples formed by one target entity and its multiple target entities with the same target relationship from affecting the normal detection of the target entity and other target entities with different target relationships, in this embodiment, the target triples with the same first target entity and target relationship are formed into triple sets, the average value of the target evaluation values corresponding to each target triple in a single triple set is calculated, and the average calculation result is used as the comprehensive sub-evaluation value of this triple set. Then, the comprehensive sub-evaluation values corresponding to all triple sets are added up, and the addition result is used as the first comprehensive evaluation value, so that multiple target triples with the same first target entity and target relationship can affect the first comprehensive evaluation value through the average value, avoiding the excessive influence of multiple target triples with the same first target entity and target relationship on the comprehensive evaluation.

[0093] In a specific embodiment, the first preset condition is that the first comprehensive evaluation value is greater than the first evaluation value threshold.

[0094] Among them, the first evaluation value threshold can be used to evaluate whether the first comprehensive evaluation value meets the first preset condition. When the first comprehensive evaluation value is greater than the first evaluation value threshold, it can be considered that the number of multiple target triples with the target entity and the target relationship in the target medical document meets the expectation at this time, and then it is determined that the first comprehensive evaluation value meets the first preset condition.

[0095] In a specific embodiment, step S4 further includes the following steps: S44, add up the target evaluation values corresponding to all the target triples respectively, and use the addition result as the second comprehensive evaluation value.

[0096] Correspondingly, step S5 includes: If the first comprehensive evaluation value does not meet the first preset condition, or the second comprehensive evaluation value meets the second preset condition, then input all the target triples and the target evaluation values corresponding to each target triple into the preset large language model to obtain the quality inspection report corresponding to the target medical document.

[0097] Among them, although the first comprehensive evaluation value can avoid the excessive influence of multiple target triples with the same first target entity and target relationship on the comprehensive evaluation, when facing the situation where multiple target triples with the same first target entity and target relationship need to be ensured to exist, especially when the target evaluation values of multiple target triples with the same first target entity and target relationship are relatively high, it is difficult to make an effective judgment. Therefore, in this embodiment, the target evaluation values corresponding to all the target triples are added up respectively, and the addition result is used as the second comprehensive evaluation value.

[0098] The second evaluation value threshold can be used to evaluate whether the second comprehensive evaluation value meets the second preset condition. When the second comprehensive evaluation value is greater than the second evaluation value threshold, it can be considered that the number of multiple target triples with the same target entity and target relationship and relatively high target evaluation values in the target medical document meets the expectation at this time, and then it is determined that the second comprehensive evaluation value meets the second preset condition.

[0099] Specifically, in this embodiment, through the collaborative judgment of the first preset condition and the second preset condition, when the first comprehensive evaluation value meets the first preset condition and the second comprehensive evaluation value meets the second preset condition, it is considered that the quality inspection of the target medical document passes. This not only avoids the influence of the completeness judgment of the target triples in the target medical document due to a higher target evaluation value and the existence of multiple identical target triples for the target entity and the target relationship, but also ensures the judgment requirements for the coexistence of a higher target evaluation value and the existence of multiple identical target triples for the target entity and the target relationship. When the preset conditions are not met, the abnormal information is analyzed and integrated through a preset large language model, and combined with the overall content and characteristics of the medical document, natural language generation technology is used to automatically generate a quality inspection report covering problem description, cause analysis, improvement suggestions, etc.

[0100] In a specific embodiment, when the computer program is executed by the processor, the following steps are further implemented: Obtain P sample texts of the same text type as the target medical document.

[0101] Determine a number of sample triples according to the P sample texts, where P is a positive integer.

[0102] For any reference triple, count the number of sample triples that are the same as the current reference triple, and use the normalized value of the statistical result as the statistical value corresponding to the current reference triple.

[0103] For any sample text, determine the first intermediate evaluation value and the second intermediate evaluation value according to the number of sample triples corresponding to the current sample, the number of reference triples whose corresponding statistical values are greater than the preset statistical threshold, and the reference evaluation values respectively corresponding to each reference triple.

[0104] Determine the first evaluation value threshold according to each first intermediate evaluation value.

[0105] Determine the second evaluation value threshold according to each second intermediate evaluation value.

[0106] Among them, the sample triple can include the first sample entity, the second sample entity, and the sample relationship. The sample text can refer to the historical text of the same text type as the target medical document. Here, the text type can be the disease type, that is, the sample text can be the text that historically stores the same disease as the target medical document records. Correspondingly, the sample text and the target medical document have high similarity, so that the first evaluation value threshold and the second evaluation value threshold can be obtained through the statistical results of the evaluation of P sample data, which are used as the data basis for subsequent evaluation of the target medical document.

[0107] The statistical value can characterize the occurrence frequency of the corresponding reference triple in each sample text. When determining the first intermediate evaluation value and the second intermediate evaluation value, it is determined by several reference triples whose statistical values are greater than the preset statistical threshold, which can effectively avoid the influence of reference triples with fewer occurrences on quality detection, improve the reliability of the first intermediate evaluation value and the second intermediate evaluation value, and further improve the reliability of quality detection.

[0108] Specifically, for any sample triple, match the current sample triple with several reference triples whose statistical values are greater than the preset statistical threshold to obtain the reference triple corresponding to the current sample triple, and use the reference evaluation value of the reference triple corresponding to the current sample triple as the sample evaluation value of the current sample triple. Similarly, if there is no corresponding reference triple for the current sample triple, determine that the sample evaluation value of the current sample triple is 0.

[0109] Sample triple sets are formed by sample triples with the same first sample entity and sample relationship. For any sample triple set, calculate the mean value of the sample evaluation values corresponding to each sample triple in the current sample triple set, and use the mean calculation result as the sample sub-evaluation value. Add up the sample sub-evaluation values corresponding to all sample triple sets, and use the addition result as the first intermediate evaluation value. Add up the sample evaluation values corresponding to all sample triples, and use the addition result as the second intermediate evaluation value.

[0110] Take the mean value of the first intermediate evaluation values of each sample text as the first evaluation value threshold, and take the mean value of the second intermediate evaluation values of each sample text as the second evaluation value threshold.

[0111] In a specific embodiment, since the methods for extracting abnormal information from the target medical document in Embodiment 1 and Embodiment 3 are different and the extracted differential information is different, the target medical document, abnormal triple information, all target triples, and the target evaluation value corresponding to each target triple can be input into a preset large language model to obtain a quality detection report corresponding to the target medical document, so as to more comprehensively extract and analyze the abnormal information in the target medical document, thereby improving the comprehensiveness and reliability of the quality detection report.

[0112] As described above, multiple target triples are formed from the target entities extracted from the target medical document. According to the matching result between the target triple and the reference triple in the reference database, the target evaluation value of the target triple is determined. Then, according to the target evaluation values of each target triple, the first evaluation value threshold of the target medical document is determined. According to the first evaluation value threshold, it is judged whether the target medical document passes the quality detection, which can more completely evaluate the target entity and the association relationship between the target entities, thereby improving the reliability of the quality detection of the target medical document.

[0113] Although some specific embodiments of the present invention have been described in detail by way of examples, those skilled in the art should understand that the above examples are for illustrative purposes only and not for limiting the scope of the present invention. Those skilled in the art should also understand that various modifications can be made to the embodiments without departing from the scope and spirit of the present invention. The scope of the present invention disclosed is defined by the appended claims.

Claims

1. A quality detection method for medical documents based on a large model and knowledge graph, characterized in that: The quality detection method comprises: S10, obtaining a target medical document to be detected and a target attribute label corresponding to the target medical document; S20, according to the target attribute tag, obtaining the target quality inspection standard corresponding to the target medical document from the standard database, wherein the standard database stores a preset quality inspection standard corresponding to each preset attribute tag, each preset quality inspection standard includes preset quality evaluation content and preset quality scoring standard, and all preset quality inspection standards are stored in the standard database in the form of a knowledge graph; S30, inputting the target medical document and the target quality detection standard into a preset large language model to obtain abnormal triplet information corresponding to the target medical document; S40: Acquire a quality inspection report corresponding to the target medical document according to the target medical document and the abnormal triplet information.

2. The quality detection method for medical documents based on a large model and a knowledge graph according to claim 1 is characterized in that: S20 includes the following steps: The preset quality detection standard whose corresponding preset attribute tag is consistent with the target attribute tag is determined as the target quality detection standard.

3. The quality detection method for medical documents based on a large model and knowledge graph according to claim 1 is characterized in that: S30 includes the following steps: S301, inputting the target medical document and the target quality detection standard into a preset large language model, and obtaining a plurality of abnormal triples corresponding to the target medical document and an abnormal type corresponding to each abnormal triple; S302, according to the exception type corresponding to each exception triplet, obtain the quantity ratio corresponding to each exception type; S303: The abnormal triplet, quantity ratio and preset severity weight corresponding to each abnormal type are used as the abnormal triplet information corresponding to the target medical document.

4. The quality detection method for medical documents based on a large model and knowledge graph according to claim 3 is characterized in that: S40 includes the following steps: S401, inputting the target medical document and the abnormal triplet information into the preset large language model to obtain a quality inspection report corresponding to the target medical document.

5. The quality detection method for medical documents based on a large model and knowledge graph according to claim 3 is characterized in that: S40 also includes the following steps: S402, obtaining a preset quality inspection report template; S403: Input the target medical document, the abnormal triplet information and the preset quality inspection report template into the preset large language model to obtain the quality inspection report corresponding to the target medical document.

6. A medical document quality inspection system based on a large model and knowledge graph, characterized in that: The quality inspection system includes a processor and a memory, wherein the memory stores a standard database, wherein the standard database includes T standard medical texts and attribute labels and knowledge graph information corresponding to each standard medical document, where T is a positive integer, and the processor includes: A data acquisition module, used to acquire a target medical document to be detected and a target attribute label corresponding to the target medical document; A quality inspection standard acquisition module, used to obtain the target quality inspection standard corresponding to the target medical document from a standard database according to the target attribute tag, wherein the standard database stores a preset quality inspection standard corresponding to each preset attribute tag, each preset quality inspection standard includes a preset quality evaluation content and a preset quality scoring standard, and all preset quality inspection standards are stored in the standard database in the form of a knowledge graph; An abnormal information acquisition module, used to input the target medical document and the target quality detection standard into a preset large language model to acquire abnormal triplet information corresponding to the target medical document; The test report acquisition module is used to acquire the quality test report corresponding to the target medical document according to the target medical document and the abnormal triplet information.

7. The quality inspection system for medical documents based on a large model and knowledge graph according to claim 6 is characterized in that: The quality inspection standard acquisition module includes: The target quality detection standard acquisition submodule is used to determine the preset quality detection standard whose corresponding preset attribute tag is consistent with the target attribute tag as the target quality detection standard.

8. The medical document quality inspection system based on a large model and knowledge graph according to claim 6 is characterized in that: The abnormal information acquisition module includes: An abnormality type acquisition submodule is used to input the target medical document and the target quality detection standard into a preset large language model to acquire a plurality of abnormal triples corresponding to the target medical document and an abnormality type corresponding to each abnormal triple; The quantity ratio acquisition submodule is used to obtain the quantity ratio corresponding to each abnormal type according to the abnormal type corresponding to each abnormal triplet; The abnormal information acquisition submodule is used to use the abnormal triplet, quantity ratio and preset severity weight corresponding to each abnormal type as the abnormal triplet information corresponding to the target medical document.

9. The medical document quality inspection system based on a large model and knowledge graph according to claim 6 is characterized in that: The test report acquisition module includes: The first test report acquisition submodule is used to input the target medical document and the abnormal triplet information into the preset large language model to obtain the quality test report corresponding to the target medical document.

10. The medical document quality inspection system based on a large model and knowledge graph according to claim 6, characterized in that: The test report acquisition module includes: The report template acquisition submodule is used to obtain the preset quality inspection report template; The first test report acquisition submodule is used to input the target medical document, the abnormal triplet information and the preset quality test report template into the preset large language model to obtain the quality test report corresponding to the target medical document.

Citation Information

Patent Citations

  • Medical record document review method and system based on electronic medical record management system

    CN118468886A

  • Physical examination report interpretation system and method based on rule base retrieval and large model training

    CN118471419A

  • Process industry safety knowledge graph error detection method and system based on large language model

    CN119740644A

  • Medical plan recommendation system and method based on knowledge graph representation learning

    WO2021189971A1

Cited By

  • Method and system for medical document report error correction and early warning

    CN122200702A

  • A method and system for error correction and early warning in medical document reports

    CN122200702B