Method and system for dynamic construction of chronic disease review knowledge graph under public health field

By constructing a knowledge graph for chronic disease verification, automated hierarchical extraction and correlation matching of mortality data were achieved, solving the problems of time-consuming and labor-intensive cause-of-death management and underreporting in existing technologies. This improved the recognition rate and real-time performance of chronic disease information verification, and supported the efficient management of public health events.

CN120727312BActive Publication Date: 2025-12-12ZHEJIANG CENT FOR DISEASE CONTROL & PREVENTION
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202511148816.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-08-18
Publication Date
2025-12-12
Estimated Expiration
2045-08-18

AI Technical Summary

Technical Problem

Existing methods for managing causes of death are outdated, time-consuming, and labor-intensive. They have a high rate of misidentification when faced with large amounts of cause of death data and chronic disease information screening, and they cannot detect unreported chronic disease medical records, resulting in poor management of public health events.

Method used

A dynamic construction method for chronic disease verification knowledge graphs in the field of public health is adopted, which includes collecting mortality data and medical record information, extracting cause-of-death chains hierarchically, and performing association matching and mapping transformation through a pre-set quality analysis rule base and a medical language model to construct a knowledge graph for dynamic association and updating of data.

Benefits of technology

It achieves highly automated hierarchical extraction of cause-of-death chains and disease information association matching, with a high recognition rate, can avoid missed reports, supports real-time monitoring and improves the accuracy of chronic disease information verification, and is suitable for the management of public health events.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120727312B_ABST
    Figure CN120727312B_ABST
Patent Text Reader

Abstract

The present application relates to medical data management technical field, specifically to the dynamic construction method and system of chronic disease verification knowledge graph in the field of public health, wherein the dynamic construction method of chronic disease verification knowledge graph in the field of public health comprises the following steps: S1: collecting death data and corresponding medical record information; S2: hierarchical extraction is carried out on the cause of death chain in the death data, at least one cause of death data is obtained, corresponding disease information is retrieved according to the cause of death data, association matching is carried out in the corresponding medical record information according to the disease information, whether there is relationship information related to the disease information is output, and the key entity in the death data is extracted and stored as the triple data of the death data; S3: the triple data of the death data is mapped and converted, and inserted into the knowledge graph of the preset graph database module. The present application can realize the dynamic construction of chronic disease verification knowledge graph in the field of public health, and has high automation degree.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of medical data management technology, specifically to a method and system for dynamically constructing a knowledge graph for chronic disease verification in the field of public health. Background Technology

[0002] With the rapid development of information technology and the increasing demand for public health management, a large amount of data has been accumulated in chronic disease management. This data is not only an important basis for formulating public health policies, but also the foundation for improving disease prevention and control capabilities and protecting public health. Among them, death report cards for individual cases, as one of the important sources of health information, contain key information such as the patient's cause of death and medical history. However, the current chronic disease information verification work largely relies on manual comparison of death report cards and chronic disease report cards. This method can only sample a portion of the data, which is time-consuming and labor-intensive, and it is difficult to identify potential underreported cases from massive amounts of data in a timely and accurate manner. Especially when faced with complex health records, manual analysis is difficult to fully identify the entire chain of causes of death in death report cards, which may lead to the omission of chronic disease cases that need to be reported, affecting the integrity of chronic disease monitoring data and the effectiveness of disease prevention and control work. In order to improve the management efficiency of death data and the monitoring of cause of death data, Chinese patents have disclosed a method for analyzing cause of death monitoring data based on Excel VBA (publication number: CN111063444A). In this patent technology, a data quality assessment model is constructed; a data analysis execution engine is constructed based on Excel. VBA uses a unified encoding rule and format to convert all variables, and all data is uniformly organized and imported into Excel, with ICD codes for different disease types embedded in the background. Using the imported Excel data as the execution object, data analysis and organization are performed in a horizontal or vertical manner. Horizontal analysis reflects population mortality and key indicators in a specific year, while vertical analysis reflects the changing trends of population mortality levels or key indicators across different years. This enables integrated analysis of mortality data and calculation of key indicators across the entire population and process. In contrast, the above-mentioned technical solution can only perform conversion work according to pre-defined specifications. In actual work, there are death report cards with different specifications and various medical records. For such complex organization work, the above-mentioned technical solution leads to a high misidentification rate, making real-time monitoring difficult. Furthermore, it cannot detect missed chronic disease medical records, resulting in poor management of public health events. Summary of the Invention

[0003] The technical problem to be solved by this invention is that existing methods for managing causes of death are relatively outdated, time-consuming and labor-intensive, have a high rate of misidentification when faced with large amounts of cause of death data and chronic disease information screening, and cannot detect missed chronic disease medical records, resulting in poor management of public health events.

[0004] To solve the above-mentioned technical problems, the first aspect of the present invention adopts the following technical solution: a method for dynamically constructing a knowledge graph for chronic disease verification in the field of public health, comprising the following steps:

[0005] S1: Collect mortality data and corresponding medical records;

[0006] S2: Perform hierarchical extraction of the cause of death chain in the death data to obtain at least one cause of death data. Retrieve the corresponding disease information based on the cause of death data. Perform association matching in the corresponding medical record information based on the disease information. Output whether there is a relationship information related to the disease information. Extract the key entities in the death data and store them together as the triple data of the death data.

[0007] S3: Map and transform the triple data of death data and insert it into the knowledge graph of the preset graph database module.

[0008] When this invention is in operation, it can perform a series of tasks such as hierarchical extraction of the cause-of-death chain from mortality data, correlation matching of disease information, and dynamic correlation and updating of data through knowledge graphs. It has a high degree of automation and a low proportion of human intervention. By extracting the underlying cause of death, direct cause of death, and indirect cause of death in a hierarchical manner, it can improve the screening of disease information and avoid underreporting. At the same time, the disease information can be correlated and matched with the corresponding medical record information to identify hidden information in the disease information. The recognition rate is high, making it convenient to be applied to the monitoring of public health events, and it has high real-time performance.

[0009] Preferably, step S1 further includes the following step: performing data cleaning and standardization on the collected death data and corresponding medical record information using a preset standard library.

[0010] Preferably, in step S2, when performing hierarchical extraction of the cause-of-death chain in the death data to obtain at least one layer of cause-of-death data, the following steps are adopted: according to the preset quality analysis rule base, hierarchical extraction of the cause-of-death chain in the death data is performed to identify and extract at least one layer of cause-of-death data among the fundamental cause of death, direct cause of death, and indirect cause of death.

[0011] Preferably, the quality analysis rule base is constructed using the following steps:

[0012] A1: Obtain historical mortality data, which should include at least one of the following: disease classification, cause of death classification, mortality data verification standards, cause of death verification rules, etiology association rules, and chronic disease underreporting detection rules;

[0013] A2: Define the logical relationship between the underlying cause of death, the direct cause of death, and the indirect cause of death, as well as the high-risk misreporting and underreporting patterns in chronic disease screening;

[0014] A3: Call the preset rule engine and / or ontology modeling method to build a structured rule set, which includes death chain hierarchical extraction rules and abnormal data identification rules;

[0015] A4: Optimize the rule set, validate rules based on historical mortality data, adjust the scope and confidence of the rule set to detect errors or omissions in chronic disease verification, and analyze the abnormal data patterns verified by preset machine learning methods to optimize the rule set;

[0016] A5: Use knowledge graph technology to structure and store rule sets into the quality analysis rule base.

[0017] When this invention is working, it performs hierarchical extraction of the cause-of-death chain in mortality data by constructing a quality analysis rule base. It can organize the logical relationship between the root cause of death, direct cause of death, and indirect cause of death. At the same time, it is convenient for operators to fine-tune the quality analysis rule base to adapt to various different needs and environments. It consumes less computing resources, has a fast extraction speed, and has a low cost of use. It is suitable for cause-of-death extraction work with moderate to low difficulty.

[0018] Preferably, in step S2, when performing hierarchical extraction of the cause-of-death chain in the death data to obtain at least one cause-of-death data, the following steps are adopted: by calling a preset extraction prompt word model through a pre-trained medical language model, the cause-of-death chain in the death data is inferred and extracted hierarchically to obtain at least one cause-of-death data.

[0019] Preferably, in step S2, when performing association matching based on the disease information in the corresponding medical record information and outputting whether there is relationship information related to the disease information, the following steps are adopted: using a pre-trained medical-specific language model to perform high-level semantic association on the medical record information using semantic matching technology, identifying potential disease information in the medical record information, and outputting whether there is relationship information matching, underreporting, or misreporting with the disease information retrieved from the cause of death data by association matching the disease information and potential disease information in the medical record information.

[0020] Preferably, the training of the medical-specific language model employs the following steps:

[0021] B1: Collect training data containing disease names, symptoms and test results. Preprocess the collected training data. Use natural language processing methods to segment, label and structure the training data. Standardize the disease names, symptoms and test results in the training data. Store the preprocessed training data as a training dataset.

[0022] B2: Obtain medical texts, construct a medical-specific language model using the Transformer architecture, and conduct large-scale unsupervised pre-training based on the medical texts and training objectives to enable the medical-specific language model to learn language features and contextual relationships related to chronic diseases, and establish the semantic understanding ability and cause-of-death chain reasoning ability of the medical-specific language model.

[0023] B3: Supervised learning of the medical language model using the training dataset to establish the model's ability to identify cause-of-death chain inference, chronic disease matching, and data quality assessment.

[0024] B4: Deploy a medical-specific language model and continuously learn to optimize its performance.

[0025] When this invention is working, it uses a pre-trained medical-specific language model to perform high-level semantic association, identify potential disease information in medical records, and is applicable to medical records recorded in different standards. It can avoid underreporting to a certain extent, and can also perform reasoning on the cause of death chain, which can further improve the recognition accuracy and help improve the accuracy and practicality of chronic disease information verification, thereby supporting chronic disease prevention and control and public health decision-making.

[0026] Preferably, in step S2, when extracting key entities from the death data and storing them together as triples of the death data, the following steps are adopted: key entities in the death data are extracted through a pre-trained medical language model. The key entities include at least one of the following: information about the deceased, information about the disease, and time of death.

[0027] Preferably, step S3 further includes the following steps: mapping and transforming the triple data of death data, inserting it into the knowledge graph of the graph database, and then performing graph visualization and graph analysis on the knowledge graph.

[0028] To solve the above-mentioned technical problems, the second aspect of the present invention adopts the following technical solution: a dynamic construction system for a knowledge graph of chronic disease verification in the field of public health, which applies the dynamic construction method for a knowledge graph of chronic disease verification in the field of public health as described above, including:

[0029] The file management module is used for real-time synchronous collection and preprocessing of death data and corresponding medical record information;

[0030] The quality analysis rules module is used to perform hierarchical extraction of the cause-of-death chain in mortality data and to perform correlation matching in the corresponding medical record information.

[0031] The triplet mapping graph conversion module is used to map and convert triplet data of death data;

[0032] The graph database module is used to store knowledge graphs;

[0033] The file management module is data-connected to the quality relationship rule module and transmits the collected death data and corresponding medical record information to the quality analysis rule module. The quality relationship module is data-connected to the triplet mapping graph conversion module and outputs the triplet data of the death data to the triplet mapping graph conversion module after extracting the cause of death data, relationship information and key entities. The triplet mapping graph conversion module is data-connected to the graph database module and stores the triplet data of the death data in the graph database module after performing mapping conversion on the triplet data of the death data.

[0034] The beneficial technical effects of this invention include:

[0035] 1. This invention enables hierarchical extraction of the cause-of-death chain from mortality data, correlation matching of disease information, and dynamic association and updating of data through knowledge graphs. It features a high degree of automation and a low proportion of human intervention. By hierarchically extracting the underlying cause of death, direct cause of death, and indirect cause of death, it can improve the screening of disease information and avoid underreporting. At the same time, the disease information can be correlated and matched with the corresponding medical record information to identify hidden information in the disease information. The recognition rate is high, making it convenient to apply to the monitoring of public health events, and it has high real-time performance.

[0036] 2. This invention enables hierarchical extraction of the cause-of-death chain in mortality data by constructing a quality analysis rule base. It can organize the logical relationship between the root cause of death, direct cause of death, and indirect cause of death. At the same time, it is convenient for operators to fine-tune the quality analysis rule base to suit various different needs. It consumes less computing resources, has a fast extraction speed, and has a low cost of use. It is suitable for extracting causes of death with moderate to low difficulty.

[0037] 3. This invention uses a pre-trained medical language model to perform high-level semantic association, identify potential disease information in medical records, and is applicable to medical records recorded in different standards. It can avoid underreporting to a certain extent and can also perform reasoning on the cause of death chain, which can further improve the recognition accuracy and help improve the accuracy and practicality of chronic disease information verification, thereby supporting chronic disease prevention and control and public health decision-making.

[0038] Other features and advantages of the present invention will be disclosed in detail in the following detailed description and accompanying drawings. Attached Figure Description

[0039] The invention will be further described below with reference to the accompanying drawings:

[0040] Figure 1 A flowchart illustrating the dynamic construction method of a knowledge graph for chronic disease verification in the public health field;

[0041] Figure 2 Workflow diagram for building a quality analysis rule base;

[0042] Figure 3 A flowchart for training a medical-specific language model;

[0043] Figure 4 A schematic diagram of the structure of a dynamic construction system for a knowledge graph of chronic disease verification in the field of public health. Detailed Implementation

[0044] The technical solutions of the embodiments of the present invention will be explained and described below with reference to the accompanying drawings. However, the following embodiments are only preferred embodiments of the present invention and not all of them. Other embodiments obtained by those skilled in the art based on the embodiments in the implementation methods without creative effort are all within the protection scope of the present invention.

[0045] In the following description, terms such as “inner,” “outer,” “upper,” “lower,” “left,” and “right” are used only to indicate orientation or positional relationship for the convenience of describing the embodiments and simplifying the description, and are not intended to indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation, and therefore should not be construed as a limitation of the present invention.

[0046] Example 1:

[0047] Please see Figure 1 This embodiment discloses a dynamic construction method for a knowledge graph of chronic disease verification in the field of public health, including the following steps:

[0048] S1: Collect mortality data and corresponding medical records;

[0049] S2: Perform hierarchical extraction of the cause of death chain in the death data to obtain at least one cause of death data. Retrieve the corresponding disease information based on the cause of death data. Perform association matching in the corresponding medical record information based on the disease information. Output whether there is a relationship information related to the disease information. Extract the key entities in the death data and store them together as the triple data of the death data.

[0050] S3: Map and transform the triplet data of death data and insert it into the knowledge graph of the preset graph database module 4.

[0051] In operation, this embodiment can perform a series of tasks, including hierarchical extraction of the cause-of-death chain from mortality data, correlation matching of disease information, and dynamic correlation and updating of data through knowledge graphs. It has a high degree of automation and a low proportion of human intervention. By extracting the underlying cause of death, direct cause of death, and indirect cause of death in a hierarchical manner, it can improve the screening of disease information and avoid underreporting. At the same time, the disease information can be correlated and matched with the corresponding medical record information to identify hidden information in the disease information. The recognition rate is high, which makes it convenient to be applied to the monitoring of public health events and has high real-time performance.

[0052] Preferably, step S1 also includes the following steps: using a preset standard library to perform data cleaning and standardization on the collected death data and corresponding medical record information, and to verify the required and unfilled items, data format, and value range format, which can improve data standardization, reduce the pressure of subsequent work, and improve recognition efficiency.

[0053] In specific implementation, in step S2, when performing hierarchical extraction of the cause-of-death chain in the death data to obtain at least one layer of cause-of-death data, the following steps are adopted: according to the preset quality analysis rule base, hierarchical extraction of the cause-of-death chain in the death data is performed to identify and extract at least one layer of cause-of-death data among the fundamental cause of death, direct cause of death, and indirect cause of death.

[0054] Please see Figure 2 In this embodiment, the construction of the quality analysis rule base adopts the following steps:

[0055] A1: Obtain historical mortality data. Historical mortality data should include at least one of the following: disease classification, cause of death classification, mortality data verification standards, cause of death verification rules, etiology association rules, and chronic disease underreporting detection rules. For example, the ICD-10 International Classification of Diseases, the Guidelines for Classification and Coding of Causes of Death, and related mortality data verification standards can all be obtained and selected.

[0056] A2: Define the logical relationship between the underlying cause of death, the direct cause of death, and the indirect cause of death, as well as the high-risk misreporting and underreporting patterns in chronic disease screening;

[0057] A3: Use the preset rule engine and / or ontology modeling method to construct a structured rule set. The rule set includes death chain hierarchical extraction rules and abnormal data identification rules. In specific implementation, quality analysis rules can be written using SQL query language, or any other suitable query method can be used to build the rule set. Death chain hierarchical extraction rules include related disease classifications. When extracting death cause data, only related disease information needs to be considered. Irrelevant information can be filtered out. For example, cerebral hemorrhage may be caused by hypertension, but hypertension cannot directly cause pneumonia, so pneumonia can be filtered out. Abnormal data identification includes mutually exclusive disease classifications. For example, if an individual under the age of 30 dies from Alzheimer's disease, there may be false positives.

[0058] A4: Optimize the rule set, validate rules based on historical mortality data, adjust the scope and confidence of the rule set to detect errors or omissions in chronic disease verification, and analyze the abnormal data patterns verified by preset machine learning methods to optimize the rule set;

[0059] A5: The rule set is structured and stored in the quality analysis rule base through knowledge graph technology. Ideally, new death classification standards or new chronic disease verification standards can be added in real time, and the quality analysis rule base can be dynamically adjusted to meet the needs of customized settings.

[0060] In this embodiment, a quality analysis rule base is constructed to perform hierarchical extraction of the cause-of-death chain in mortality data. It can organize the logical relationship between the underlying cause of death, the direct cause of death, and the indirect cause of death. At the same time, it is convenient for operators to fine-tune the quality analysis rule base to adapt to various different needs and environments. It consumes less computing resources, has a fast extraction speed, and has a low cost of use. It is suitable for cause-of-death extraction work with moderate to low difficulty.

[0061] In specific implementation, when extracting key entities from the death data and storing them together as triples of the death data in step S2, the following steps are adopted: key entities in the death data are extracted through a pre-trained medical language model. The key entities include at least one of the following: information about the deceased, information about the disease, and time of death.

[0062] As a further improvement to this embodiment, step S3 also includes the following steps: mapping and transforming the triple data of death data, inserting it into the knowledge graph of the graph database, and then performing graph visualization and graph analysis on the knowledge graph. This facilitates the establishment of a complete intelligent and automated workflow, thereby improving cross-departmental collaboration and data fusion capabilities, and supporting cross-organizational data sharing and dynamic verification functions.

[0063] Example 2:

[0064] Please see Figure 3This embodiment provides a dynamic construction method for a knowledge graph of chronic disease verification in the field of public health. The similarities with other embodiments will not be repeated here. The differences will be described in detail below.

[0065] In this embodiment, in step S2, when performing hierarchical extraction of the cause-of-death chain in the death data to obtain at least one cause-of-death data, the following steps are adopted: by calling a preset extraction prompt word model through a pre-trained medical language model, the cause-of-death chain in the death data is inferred and extracted hierarchically to obtain at least one cause-of-death data.

[0066] Preferably, in step S2, when performing association matching based on the disease information in the corresponding medical record information and outputting whether there is relationship information related to the disease information, the following steps are adopted: using a pre-trained medical-specific language model to perform high-level semantic association on the medical record information using semantic matching technology, identifying potential disease information in the medical record information, and outputting whether there is relationship information matching, missed reporting, or misreporting with the disease information retrieved from the cause of death data by association matching of the disease information and potential disease information in the medical record information.

[0067] In practical implementation, semantic matching technology is used to perform high-level semantic association of diagnostic information, chief complaint, medical orders, and test results in medical record texts, identifying records with different expressions but the same medical meaning. For example, type 2 diabetes is associated with non-insulin-dependent diabetes mellitus, and chronic obstructive pulmonary disease is associated with COPD. Ideally, knowledge graph reasoning can also be used to discover potentially related records. For example, the cause of death being chronic renal failure may be associated with diabetes not explicitly stated in the medical record. Of course, the semantic matching method can also be adjusted according to actual needs. For example, direct matching can be performed based on field values, or a match can be made based on identity information. If the identity information is completely consistent, it is directly determined as a match. Fuzzy matching association can also be used. That is, in the case of missing or inconsistent identity information, fuzzy matching can be performed using fields such as name, gender, date of birth, underlying cause of death, and date of death to calculate similarity and make preliminary association. Edit distance and Jaccard distance can also be used. Algorithms such as similarity can be used to correct and match misspelled names and variations of abbreviations. For date fields, a time window matching method can be used to allow for a certain error range and improve the fault tolerance of the matching. Ideally, multi-field cross-validation can also be performed. For example, cross-validation can be performed using multiple key fields such as identity information, name, date of birth, underlying cause of death, and date of death to ensure the accuracy of the final association.

[0068] As a further improvement to this embodiment, the training of the medical-specific language model adopts the following steps:

[0069] B1: Collect training data containing disease names, symptoms, and test results. For example, collect data from authoritative medical literature, electronic medical records, death case data, ICD-10 disease codes, chronic disease monitoring reports, public health guidelines, etc., and build a medical corpus. Preprocess the collected training data to remove irrelevant information, standardize the data format, handle missing values, and remove duplicates. Use natural language processing methods to segment, label, and structure the training data. Standardize the disease names, symptoms, and test results in the training data. Store the preprocessed training data as a training dataset.

[0070] B2: Obtain medical texts and construct a medical-specific language model using the Transformer architecture. Perform large-scale unsupervised pre-training based on the medical texts and training objectives to enable the medical-specific language model to learn language features and contextual relationships related to chronic diseases, and establish the semantic understanding and cause-of-death chain reasoning capabilities of the medical-specific language model. In specific implementation, the training objectives can be set according to needs, including vector representation optimization, medical terminology semantic relationship learning, and cause-of-death chain reasoning modeling. Ideally, Masked Language Model and Next Sentence Prediction tasks can also be used to improve the semantic understanding capabilities of the medical-specific model.

[0071] B3: Supervised learning of the medical language model is conducted using the training dataset to establish its ability to identify cause-of-death chain inference, chronic disease matching, and data quality assessment. Ideally, existing knowledge graphs can be integrated into the medical language model, and knowledge enhancement methods such as KG-BERT and KEPLER can be used to improve the inference ability of the medical language model. When the database is large, a quality analysis rule base can be combined to establish the model's logical consistency judgment of cause-of-death data by using rule learning and logical reasoning.

[0072] B4: Deploy a medical-specific language model and continuously learn to optimize its performance. In practice, different learning methods can be used, such as active learning, which combines expert feedback to highlight high-uncertainty samples predicted by the medical-specific language model and continuously optimize the training set. Alternatively, federated learning can be used, which supports collaborative training among different medical institutions, improving the generalization ability of the medical-specific language model while protecting data privacy. When there are significant regional differences, adaptive fine-tuning can also be performed to dynamically adjust the weights of each part of the medical-specific language model for different regions and disease spectra to improve applicability and versatility.

[0073] In this embodiment, a pre-trained medical language model is used to perform high-level semantic association to identify potential disease information in medical records. It is applicable to medical records with different standards and can avoid underreporting to a certain extent. At the same time, it can also perform reasoning on the cause of death chain, which can further improve the recognition accuracy and help improve the accuracy and practicality of chronic disease information verification, thereby supporting chronic disease prevention and control and public health decision-making.

[0074] Example 3:

[0075] This embodiment provides a dynamic construction method for a knowledge graph of chronic disease verification in the field of public health. The similarities with other embodiments will not be repeated here. The differences will be explained in detail below.

[0076] In this embodiment, to further improve the efficiency of knowledge graph reading and mining of potential related information, step S3 also includes the following steps:

[0077] C1: Obtain several nodes of the knowledge graph, obtain the occurrence frequency of each node by statistics, calculate the co-occurrence strength between nodes, and calculate the potential associations between nodes without direct connections;

[0078] C2: Call the preset weak association threshold. During operation, it can be selected according to actual needs. Its value range is set from 0.3 to 0.6. For example, in the event of a public health event, the weak association threshold can be increased to improve the mining of potential related information, thereby improving the response speed of corresponding disease information and enhancing the correlation of weak association nodes.

[0079] C3: Simplify the knowledge graph. In practice, the knowledge graph can be simplified by clustering and merging similar nodes. Of course, dynamic pruning can be performed according to actual needs, and the graph can be stored in the graph database module 4 for subsequent screening and management.

[0080] As a further improvement to this embodiment, in step C1, when counting the occurrence frequency of each node, calculating the co-occurrence strength between nodes, and calculating the potential associations of nodes without direct connections, the following formula is used:

[0081] ;

[0082] ;

[0083] in: For nodes with direct connections, the degree of association is... Let i be the number of times node i and node j appear together. Let i be the number of times node i appears. Let j be the number of times node j appears. Let G be the degree of association of nodes without direct connections, and let G be the set of common neighbors of nodes i and j.

[0084] Preferably, to further improve the enhancement effect, the correlation of weakly associated nodes can be enhanced by calculating disease level similarity and performing path propagation enhancement, thereby more efficiently identifying potential associations between nodes. In step C2, the following steps are adopted to enhance the correlation of weakly associated nodes:

[0085] D1: The following formula is used to calculate disease hierarchy similarity:

[0086] ;

[0087] ;

[0088] in: To enhance the relevance, To enhance the correlation coefficient, a value between 0.1 and 0.3 can be selected during operation. For example, in a public health event, if two disease types are located in unrelated chapters of ICD-10 but occur together at a high rate, a value of 0.3 can be used to maximize their correlation, thereby helping medical personnel to notice complications arising in public health events more quickly. Let be the hierarchical similarity between node i and node j;

[0089] D2: When strengthening indirect associations through path association propagation, the following formula is used:

[0090] ;

[0091] in: To further enhance the relevance, is the path attenuation coefficient, which can be selected between 0.5 and 1 during operation. By adjusting the attenuation rate of long path influence, the efficiency of potential information identification can be further enhanced. p is the connection path between node i and node j. Let be the length of the path connecting node i to node j.

[0092] Example 4:

[0093] Please see Figure 4 This embodiment provides a dynamic construction system for a knowledge graph of chronic disease verification in the public health field. The system applies the above-mentioned dynamic construction method for a knowledge graph of chronic disease verification in the public health field, including:

[0094] File management module 1 is used for real-time synchronous collection and preprocessing of death data and corresponding medical record information;

[0095] Quality analysis rule module 2 is used to perform hierarchical extraction of the cause-of-death chain in the mortality data and to perform correlation matching in the corresponding medical record information;

[0096] The triplet mapping graph conversion module 3 is used to map and convert triplet data of death data;

[0097] Graph database module 4 is used to store knowledge graphs;

[0098] The file management module 1 is connected to the quality relationship rule module and transmits the collected mortality data and corresponding medical record information to the quality analysis rule module 2. The quality relationship module is connected to the triple mapping graph conversion module 3 and outputs the triple data of the mortality data to the triple mapping graph conversion module 3 after extracting the cause of death data, relationship information and key entities. The triple mapping graph conversion module 3 is connected to the graph database module 4 and performs mapping conversion on the triple data of the mortality data before storing it in the graph database module 4. Through the intelligent workflow of collaboration between the file management module 1, the quality analysis rule module 2, the triple mapping graph conversion module 3 and the graph database module 4, seamless processing of mortality data can be achieved. It also supports multiple types of data input and dynamic data updates. Through automated processing throughout the entire process of cause of death chain analysis, ICD-10 information extraction, data comparison and anomaly marking, manual intervention is greatly reduced. It reduces the lag and error risk of manual verification, and improves the efficiency and accuracy of verification work. It is suitable for comparing large-scale death insurance card and chronic disease data. In specific implementation, the quality analysis rule base and medical-specific language model are both deployed in the quality analysis rule module 2.

[0099] Preferably, it also includes a knowledge graph management module 5 for graph visualization and graph analysis of the knowledge graph, which facilitates real-time verification and feedback, strengthens disease prevention and control early warning and intervention, and improves the rapid response capability for public health events.

[0100] The beneficial technical effects of this embodiment include: the present invention can realize a series of tasks such as hierarchical extraction of the cause-of-death chain in mortality data, association matching of disease information, and dynamic association and updating of data through knowledge graphs. It has a high degree of automation and a low proportion of human intervention. By hierarchically extracting the underlying cause of death, direct cause of death, and indirect cause of death, it can improve the screening of disease information and avoid underreporting. At the same time, the disease information can be associated and matched with the corresponding medical record information to identify hidden information in the disease information. The recognition rate is high, which is convenient for use in the monitoring of public health events and has high real-time performance.

[0101] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Those skilled in the art should understand that the present invention includes, but is not limited to, the contents described in the accompanying drawings and the specific embodiments above. Any modifications that do not depart from the functional and structural principles of the present invention will be included within the scope of the claims.

Claims

1. A dynamic construction method for a knowledge graph of chronic disease verification in the field of public health, characterized in that, Includes the following steps: S1: Collect mortality data and corresponding medical record information, and perform data cleaning and standardization processing on the collected mortality data and corresponding medical record information through a preset standard library; S2: Perform hierarchical extraction of the cause-of-death chain in the death data to obtain at least one cause-of-death data. Retrieve the corresponding disease information based on the cause-of-death data. Perform association matching in the corresponding medical record information based on the disease information. Output whether there is a relationship information related to the disease information. Extract key entities from the death data and store them together as triple data of the death data. When performing hierarchical extraction of the cause-of-death chain in the death data, the hierarchical extraction of the cause-of-death chain in the death data is performed by calling the preset extraction prompt word model through a preset quality analysis rule base or a pre-trained medical language model. When performing hierarchical extraction of the cause-of-death chain in mortality data using a pre-defined quality analysis rule base, at least one layer of cause-of-death data among the underlying cause of death, direct cause of death, and indirect cause of death is identified and extracted. The construction of the quality analysis rule base adopts the following steps: A1: Obtain historical mortality data, which should include at least one of the following: disease classification, cause of death classification, mortality data verification standards, cause of death verification rules, etiology association rules, and chronic disease underreporting detection rules; A2: Define the logical relationship between the underlying cause of death, the direct cause of death, and the indirect cause of death, as well as the high-risk misreporting and underreporting patterns in chronic disease screening; A3: Call the preset rule engine and / or ontology modeling method to build a structured rule set, which includes death chain hierarchical extraction rules and abnormal data identification rules; A4: Optimize the rule set, validate rules based on historical mortality data, adjust the scope and confidence of the rule set to detect errors or omissions in chronic disease verification, and analyze the abnormal data patterns verified by preset machine learning methods to optimize the rule set; A5: Use knowledge graph technology to structure and store rule sets into a quality analysis rule base; When a pre-trained medical language model calls a preset extraction prompt word model to perform hierarchical extraction of the cause of death chain in the death data, the cause of death chain in the death data is inferred and extracted hierarchically to obtain at least one cause of death data. When performing correlation matching based on the disease information in the corresponding medical record information, and outputting whether there is any relationship information related to the disease information, the following steps are adopted: using a pre-trained medical-specific language model to perform high-level semantic correlation on the medical record information using semantic matching technology, identifying potential disease information in the medical record information, and outputting whether there is any relationship information matching, underreporting, or misreporting with the disease information retrieved from the cause of death data by correlation matching of the disease information and potential disease information in the medical record information. S3: Map and transform the triple data of death data and insert it into the knowledge graph of the preset graph database module.

2. The method for dynamically constructing a knowledge graph for chronic disease verification in the field of public health according to claim 1, characterized in that: The training of the medical-specific language model employs the following steps: B1: Collect training data containing disease names, symptoms and test results. Preprocess the collected training data. Use natural language processing methods to segment, label and structure the training data. Standardize the disease names, symptoms and test results in the training data. Store the preprocessed training data as a training dataset. B2: Obtain medical texts, construct a medical-specific language model using the Transformer architecture, and conduct large-scale unsupervised pre-training based on the medical texts and training objectives to enable the medical-specific language model to learn language features and contextual relationships related to chronic diseases, and establish the semantic understanding ability and cause-of-death chain reasoning ability of the medical-specific language model. B3: Supervised learning of the medical language model using the training dataset to establish the model's ability to identify cause-of-death chain inference, chronic disease matching, and data quality assessment. B4: Deploy a medical-specific language model and continuously learn to optimize its performance.

3. The method for dynamically constructing a knowledge graph for chronic disease verification in the field of public health according to claim 1, characterized in that: In step S2, when extracting key entities from the death data and storing them together as triples of the death data, the following steps are adopted: key entities in the death data are extracted through a pre-trained medical language model. The key entities include at least one of the following: information about the deceased, information about the disease, and time of death.

4. The method for dynamically constructing a knowledge graph for chronic disease verification in the field of public health according to claim 1, characterized in that: Step S3 also includes the following steps: mapping and transforming the triple data of death data, inserting it into the knowledge graph of the graph database, and then performing graph visualization and graph analysis on the knowledge graph.

5. A dynamic construction system for a knowledge graph of chronic disease verification in the field of public health, employing the dynamic construction method for a knowledge graph of chronic disease verification in the field of public health as described in any one of claims 1 to 4, characterized in that, include: The file management module (1) is used to collect and preprocess death data and corresponding medical record information in real time. The quality analysis rule module (2) is used to perform hierarchical extraction of the cause-of-death chain in the mortality data and to perform correlation matching in the corresponding medical record information; The triplet mapping graph conversion module (3) is used to map and convert triplet data of death data; The graph database module (4) is used to store knowledge graphs; The file management module (1) is connected to the quality relationship rule module and transmits the collected death data and corresponding medical record information to the quality analysis rule module (2). The quality relationship module is connected to the triplet mapping graph conversion module (3) and outputs the triplet data of the death data to the triplet mapping graph conversion module (3) after completing the extraction of cause of death data, relationship information and key entities. The triplet mapping graph conversion module (3) is connected to the graph database module (4) and stores the triplet data of the death data in the graph database module (4) after completing the mapping conversion.

Citation Information

Patent Citations

  • Excel VBA-based death cause monitoring data analysis method

    CN111063444A

  • Multi-agent medical examination recommendation system and method

    CN120236740A

  • Dead cause chain auxiliary inference method and device

    CN120450052A