Medical data conjoint analysis system based on medical knowledge graph driving

By constructing a medical data joint analysis system based on medical knowledge graphs, integrating multi-source data and dynamically updating it, the problem of neglecting data connections in traditional methods is solved, and efficient and accurate analysis and intelligent decision support of medical data are achieved.

CN121191784APending Publication Date: 2025-12-23于瑶瑶
View PDF 0 Cites 8 Cited by

Patent Information

Application Number
CN202511296513.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-11
Publication Date
2025-12-23

AI Technical Summary

Technical Problem

Existing medical data analysis methods mainly rely on traditional data mining techniques and machine learning algorithms, neglecting the connections and knowledge background between various types of data. This results in the analysis results failing to fully and deeply reveal the internal mechanisms between medical data, and reducing the accuracy of medical knowledge path reasoning.

Method used

A medical data joint analysis system based on medical knowledge graph is adopted. By integrating medical clinical guidelines, biomedical literature, drug databases and historical medical data to construct a heterogeneous knowledge graph, feature extraction and knowledge path joint-driven analysis of multimodal medical data are carried out through delayed contradiction learning and update, combined with the path search mechanism of reinforcement learning.

Benefits of technology

It provides a comprehensive medical knowledge framework, enabling efficient retrieval and acquisition of relevant medical information, maintaining the timeliness and dynamism of the knowledge graph, improving the accuracy and efficiency of the technology, reducing the problems of information obsolescence and inconsistency, supporting the precision and reliability of clinical decision-making, and promoting the intelligent and refined development of medical services.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121191784A_ABST
    Figure CN121191784A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of medical information, in particular to a medical data conjoint analysis system based on medical knowledge graph driving. The system comprises a medical knowledge graph construction module, a medical data quality evaluation module, a medical feature engineering module and a medical knowledge driven analysis module, and heterogeneous knowledge graph modeling can be performed by integrating medical clinical guidelines, biomedical literatures, a drug database and historical medical data so as to generate a medical heterogeneous knowledge graph; obtaining new medical clinical test information and carrying out delayed contradictory learning update to generate a medical dynamic update knowledge graph; the method comprises the following steps: obtaining multi-modal medical data, carrying out quality verification evaluation and medical feature engineering analysis, carrying out knowledge path joint driving analysis at the same time, generating a decision support reasoning path corresponding to medical clinical knowledge, and outputting a corresponding medical knowledge path confidence coefficient. According to the method, fusion reasoning among cross-source data can be realized by constructing the multi-modal medical knowledge graph.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of medical information technology, and in particular to a medical data joint analysis system driven by a medical knowledge graph. BACKGROUND

[0002] With the rapid growth of medical field data, especially the wide application of various types of data such as electronic health records (EHR), genomic data, medical images, etc., the analysis and mining of medical data are facing great challenges. These data are usually derived from different medical devices, laboratory tests, clinical diagnoses and patient medical records, involving complex content and various formats. How to extract effective information from massive and heterogeneous medical data to help doctors conduct accurate analysis has become an important problem to be solved in the medical field. At present, the analysis method of medical data mainly relies on traditional data mining technology and machine learning algorithm. However, the traditional method mainly focuses on the analysis of single data source, ignoring the connection and knowledge background among various data, which leads to the inability of the analysis result to comprehensively and deeply reveal the internal mechanism among medical data, thereby reducing the accuracy of medical knowledge path reasoning. SUMMARY

[0003] Therefore, it is necessary to provide a medical data joint analysis system driven by a medical knowledge graph to solve at least one of the above technical problems.

[0004] To achieve the above-mentioned purpose, a medical data joint analysis system driven by a medical knowledge graph comprises the following modules: A medical knowledge graph construction module is configured to integrate medical clinical guidelines, biomedical literature, drug databases and historical medical data, and perform heterogeneous knowledge graph modeling according to the medical clinical guidelines, biomedical literature, drug databases and historical medical data, to generate a medical heterogeneous knowledge graph; obtain new medical clinical trial information, and perform delayed contradiction learning update on the medical heterogeneous knowledge graph based on the new medical clinical trial information, to generate a medical dynamic update knowledge graph; A medical data quality evaluation module is configured to access and protect the corresponding multi-modal medical data including electronic medical records, medical images and medical genomic data through federated query, and perform quality check and evaluation on the multi-modal medical data, to obtain multi-modal medical calibration data; A medical feature engineering module is configured to perform medical feature engineering analysis on the multi-modal medical calibration data, to generate multi-modal medical feature vectors including medical sub-features corresponding to electronic medical records, medical images and genomic data; The medical knowledge driven analysis module is configured to dynamically update a knowledge graph based on medical data, and to jointly drive analysis of a multi-modal medical feature vector based on a path search mechanism based on reinforcement learning, to generate a decision support reasoning path corresponding to medical clinical knowledge, and to output a corresponding medical knowledge path confidence.

[0005] Further, the medical knowledge graph construction module includes the following functions: By integrating medical clinical guidelines, biomedical literature, drug databases, and historical medical data; According to the medical clinical guidelines, biomedical literature, drug databases, and historical medical data, medical entity node extraction is performed to obtain 15 types of medical clinical entity node types, including disease, symptom / sign, drug, gene / biomarker, anatomical structure, laboratory examination index, imaging feature, surgical operation, patient population, clinical trial, medical device, literature evidence, medical institution, environmental factor, and time node corresponding to the medical entity node type; Based on the 15 types of medical clinical entity node types and combined with the medical clinical guidelines, biomedical literature, drug databases, and historical medical data, entity attribute edge relationship analysis is performed to obtain 30 types of medical entity attribute edge relationship types, including disease-complication, disease-symptom, disease-subtype, disease-risk factor, disease-genetic association, drug-indication, drug-contraindication, drug-adverse reaction, surgery-treatment disease, device-applicable scene, gene-target, pathway-disease, biomarker-prognosis, microorganism-disease, guideline-recommended solution, literature-supporting conclusion, clinical trial-result, medical data-therapeutic effect difference, disease-geographical high incidence, environment-disease risk, time-prognosis, hospital-specialty advantage, doctor-surgical proficiency, medical insurance-coverage, population-treatment response, comorbidity-treatment restriction, lifestyle-intervention, new evidence-overturn old knowledge, expert consensus-controversy, data source-credibility corresponding to the edge relationship type; According to the 15 types of medical clinical entity node types and the 30 types of medical entity attribute edge relationship types, and combined with the heterogeneous information network modeling technology, a heterogeneous knowledge graph is modeled to generate a medical heterogeneous knowledge graph; Obtain new medical clinical trial information, and update the medical heterogeneous knowledge graph based on the new medical clinical trial information to generate a dynamically updated medical knowledge graph.

[0006] Further, the obtaining of the new medical clinical trial information and the updating of the medical heterogeneous knowledge graph based on the new medical clinical trial information include: Obtain new medical clinical trial information; The new medical clinical trial information is analyzed by the newly added entity recognition analysis to obtain the medical clinical trial newly added entity, and the identification time delay of the medical clinical trial newly added entity is determined to be less than or equal to 1 hour; Based on the medical clinical trial newly added entity, the clinical knowledge relationship between the existing medical knowledge entities in the medical heterogeneous knowledge graph is inferred to generate the medical evidence knowledge relationship chain between the newly added entity and the existing medical knowledge entity; Based on the medical clinical trial newly added entity and the medical evidence knowledge relationship chain between the newly added entity and the existing medical knowledge entity, the medical heterogeneous knowledge graph is updated by contradiction detection learning to generate the medical dynamic update knowledge graph.

[0007] Further, the medical heterogeneous knowledge graph is updated by contradiction detection learning based on the medical clinical trial newly added entity and the medical evidence knowledge relationship chain between the newly added entity and the existing medical knowledge entity, comprising: Based on the medical evidence knowledge relationship chain between the newly added entity and the existing medical knowledge entity, the existing medical knowledge entity related to the newly added entity in the medical heterogeneous knowledge graph is inferred and searched to obtain the existing medical evidence knowledge entity related to the newly added entity; Based on the medical clinical trial newly added entity, the knowledge contradiction conflict difference analysis is performed on the existing medical evidence knowledge entity related to the newly added entity to obtain the contradiction conflict difference quantity between the newly added entity and the existing knowledge entity; Based on the contradiction conflict difference quantity between the newly added entity and the existing knowledge entity, the expert evidence level review is performed between the medical clinical trial newly added entity and the existing medical evidence knowledge entity related to the newly added entity to generate the medical evidence level comparison matrix between the newly added entity and the existing knowledge entity; By taking the contradiction conflict difference quantity between the newly added entity and the existing knowledge entity as the knowledge increment of the existing knowledge entity, and based on the knowledge increment of the existing knowledge entity and the medical evidence level comparison matrix between the newly added entity and the existing knowledge entity, the knowledge increment learning update is performed on the corresponding existing knowledge entity in the medical heterogeneous knowledge graph to generate the medical dynamic update knowledge graph.

[0008] Further, the expert evidence level review between the medical clinical trial newly added entity and the existing medical evidence knowledge entity related to the newly added entity based on the contradiction conflict difference quantity between the newly added entity and the existing knowledge entity comprises: Based on the contradiction conflict difference quantity between the newly added entity and the existing knowledge entity, the conflict potential impact assessment is performed between the medical clinical trial newly added entity and the existing medical evidence knowledge entity related to the newly added entity to obtain the medical knowledge conflict potential impact degree between the newly added entity and the existing knowledge entity; Based on the potential impact of medical knowledge conflicts between new entities and existing knowledge entities, the conflict priority of each existing medical evidence knowledge entity that is related to the new entity is ranked to generate a conflict priority sequence of existing knowledge entities that are related to the new entity. Based on the priority sequence of existing knowledge entities that have a relationship with the newly added entity, the expert evidence level is reviewed between the newly added entity in the medical clinical trial and each existing medical evidence knowledge entity that has a conflict, in order of sorting and intelligent push, so as to generate a medical evidence level comparison matrix between the newly added entity and the existing knowledge entities.

[0009] Furthermore, the medical data quality assessment module includes the following functions: By providing federated access to cross-agency medical databases and employing homomorphic encryption and differential privacy... The corresponding dual-protection access protection query includes multimodal medical data, including electronic medical records, medical images, and medical genomics data; Contrastive learning was used to perform cross-modal alignment of multimodal medical data, resulting in aligned multimodal medical field data. The quality of the aligned data for multimodal medical fields is evaluated to ensure that the quality score of the corresponding medical data is greater than 0.7 through consistency verification. The missing fields are then compensated by interpolation to obtain multimodal medical calibration data.

[0010] Furthermore, the response time for the access protection query is ≤500ms.

[0011] Furthermore, step S4 includes the following steps: Based on the knowledge entities within the dynamically updated medical knowledge graph, feature semantic association analysis is performed on the multimodal medical feature vectors to obtain the feature semantic association degree between the multimodal medical features and the knowledge entities within the knowledge graph. Based on the semantic correlation between multimodal medical features and various knowledge entities within the knowledge graph, and combined with a path search mechanism based on reinforcement learning, the corresponding reasoning hop count is used to perform joint knowledge path analysis between the multimodal medical feature vector and the knowledge entities within the dynamically updated medical knowledge graph whose semantic correlation is greater than or equal to a preset threshold, so as to generate decision support reasoning paths corresponding to medical clinical knowledge. The confidence score of the corresponding medical knowledge path is calculated and output based on the decision support reasoning path corresponding to the medical clinical knowledge.

[0012] Furthermore, the inference hop count corresponding to the path search mechanism is ≤5 layers.

[0013] Furthermore, the calculation of the confidence level of the medical knowledge path corresponding to the decision support reasoning path based on the medical clinical knowledge includes: By using the decision support reasoning path corresponding to medical clinical knowledge, we can obtain the medical clinical knowledge feature patterns corresponding to each medical knowledge entity on the reasoning path. The current medical pathological state of the patient is obtained, and the pathological knowledge similarity assessment of the current medical pathological state of the patient is performed based on the medical clinical knowledge feature patterns corresponding to each medical knowledge entity on the reasoning path, so as to obtain the similarity between the current pathological state and the medical knowledge features on the reasoning path. Based on the similarity between the current pathological state and the medical knowledge features along the reasoning path, a medical knowledge confidence calculation is performed on the decision support reasoning path corresponding to the clinical medical knowledge, so as to calculate and output the corresponding medical knowledge path confidence.

[0014] The beneficial effects of this invention are: The medical data joint analysis system based on medical knowledge graph proposed in this invention consists of a medical knowledge graph construction module, a medical data quality assessment module, a medical feature engineering module, and a medical knowledge-driven analysis module. Compared with the prior art, the beneficial effect of this application is that by integrating medical clinical guidelines, biomedical literature, drug databases, and historical medical data to model heterogeneous knowledge graphs, it can provide a comprehensive knowledge framework in the medical field, which includes various types of information such as diseases, symptoms, treatment plans, drug information, and patient historical data. Through this model, the system can help medical personnel efficiently retrieve and obtain relevant medical information, thereby improving the accuracy and efficiency of decision-making. Meanwhile, acquiring new medical clinical trial information and updating the medical knowledge graph using delayed contradiction learning based on this information can effectively maintain the timeliness and dynamism of the knowledge graph. Through delayed contradiction learning updates, not only can outdated or erroneous medical information be removed, but the existing knowledge graph can also be corrected based on the latest clinical trial results and research findings. This process effectively avoids the problems of information obsolescence and inconsistency, ensuring that the medical knowledge graph remains accurate and up-to-date. This can provide more accurate and reliable support for clinical decision-making, further improve the quality of medical services, and thus comprehensively and deeply reveal the intrinsic mechanisms between medical data. Secondly, by using federated access to protect and query multimodal medical data, including electronic medical records, medical images, and genomic data, the aim is to efficiently collect and use medical data while protecting patient privacy. The federated learning mechanism ensures distributed storage of data sources while avoiding the risks of centralized storage and leakage of patient data. For this multimodal medical data, quality verification and evaluation become crucial. By detecting the image quality of medical images, assessing the sequence quality of genomic data, and checking the consistency of data in electronic medical records, the accuracy and consistency of the data are ensured. This verification process can not only identify outliers, missing data, and noisy data, but also standardize the data through standardization methods, providing more accurate and complete data support for subsequent feature engineering and inference analysis, thereby enabling better analysis of the relationships and knowledge background between various types of data.Then, electronic medical records, medical images, and genomic data are processed and combined into medical feature vectors to extract key features from each data source. For example, information such as clinical symptoms, diagnoses, and treatment records in electronic medical records are converted into feature vectors using text mining technology. Medical images are processed using computer vision technology to extract features of lesion areas from the images. Genomic data is processed using gene sequence analysis technology to extract gene mutation information or other relevant biomarkers. Combining these features can comprehensively characterize a patient's health status, covering multiple dimensions such as clinical manifestations of diseases, imaging changes, and molecular biological characteristics. The greatest advantage of this process is that it can provide richer and more diverse input data for subsequent decision-making and reasoning, and deeply explore the patient's condition from different perspectives, thereby providing data support for subsequent processing. Finally, a joint knowledge path-driven analysis was conducted on multimodal medical feature vectors based on a dynamically updated medical knowledge graph and a reinforcement learning path search mechanism. The core of this step lies in automatically searching and optimizing knowledge paths within the medical knowledge graph through reinforcement learning to find the most suitable treatment plan. Reinforcement learning learns the optimal decision path through interaction with the environment, while the introduction of the knowledge graph provides a structured knowledge framework, enabling the decision-making process to not only rely on experience but also dynamically adjust based on real-time updated medical knowledge. The joint-driven analysis integrates features from different data modalities to form a multi-dimensional decision model. The output of medical knowledge path confidence helps clinicians assess the reliability of different paths, thereby selecting the optimal treatment strategy. Through this intelligent reasoning path, doctors can make quick and accurate decisions in complex clinical situations, greatly improving treatment efficiency and reducing human error. Simultaneously, it promotes the intelligent and refined development of medical services, thereby enhancing the accuracy of medical knowledge path reasoning within the medical knowledge graph. Attached Figure Description

[0015] Other features, objects, and advantages of the invention will become more apparent from the following detailed description of non-limiting embodiments with reference to the accompanying drawings: Figure 1 This is a schematic diagram of the modules of the medical data joint analysis system driven by medical knowledge graph of the present invention; Figure 2 for Figure 1 A functional flowchart of the TCM knowledge graph construction module; Figure 3 for Figure 1 A functional flowchart of the Traditional Chinese Medicine data quality assessment module. Detailed Implementation

[0016] The technical system of the present invention will now be clearly and completely described with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without inventive effort are within the scope of protection of the present invention.

[0017] Furthermore, the accompanying drawings are merely illustrative of the invention and are not necessarily drawn to scale. The same reference numerals in the drawings denote the same or similar parts, and therefore repeated descriptions of them will be omitted. Some block diagrams shown in the drawings are functional entities and do not necessarily correspond to physically or logically independent entities. These functional entities can be implemented in software, in one or more hardware modules or integrated circuits, or in different network and / or processor systems and / or microcontroller systems.

[0018] It should be understood that although the terms "first," "second," etc., may be used herein to describe various units, these units should not be limited by these terms. These terms are used merely to distinguish one unit from another. For example, without departing from the scope of the exemplary embodiments, a first unit may be referred to as a second unit, and similarly, a second unit may be referred to as a first unit. The term "and / or" as used herein includes any and all combinations of one or more of the associated listed items.

[0019] To achieve the above objectives, please refer to Figures 1 to 3 This invention provides a medical data joint analysis system driven by medical knowledge graphs, the system comprising the following modules: The medical knowledge graph construction module is used to integrate medical clinical guidelines, biomedical literature, drug databases, and historical medical data, and to model heterogeneous knowledge graphs based on these data to generate a medical heterogeneous knowledge graph; it also acquires new medical clinical trial information and updates the medical heterogeneous knowledge graph with delayed contradiction learning based on this new information to generate a dynamically updated medical knowledge graph. The medical data quality assessment module is used to query corresponding multimodal medical data through federated access protection, including electronic medical records, medical images and medical genomics data, and to perform quality verification and assessment on the multimodal medical data to obtain multimodal medical calibration data. The medical feature engineering module is used to perform medical feature engineering analysis on multimodal medical calibration data to generate multimodal medical feature vectors, including medical sub-features corresponding to electronic medical records, medical images, and genomes. The medical knowledge-driven analysis module is used to perform joint knowledge path analysis on multimodal medical feature vectors based on dynamically updated medical knowledge graphs and a path search mechanism based on reinforcement learning, so as to generate decision support reasoning paths corresponding to medical clinical knowledge and output the corresponding medical knowledge path confidence.

[0020] In the embodiments of this invention, please refer to Figure 1 The diagram shown illustrates the modules of the medical data joint analysis system driven by medical knowledge graphs according to the present invention. In this example, the medical data joint analysis system driven by medical knowledge graphs includes the following modules: S1: The medical knowledge graph construction module is used to generate a medical heterogeneous knowledge graph by integrating medical clinical guidelines, biomedical literature, drug databases, and historical medical data, and by modeling heterogeneous knowledge graphs based on these data; it also acquires new medical clinical trial information and updates the medical heterogeneous knowledge graph with delayed contradiction learning based on this information to generate a dynamically updated medical knowledge graph. In this embodiment of the invention, 5,000 biomedical articles are searched from platforms such as PubMed (US National Library of Medicine) and CNKI Medical Database; a dataset containing information on 1,000 drugs and 500 gene / biomarker data is obtained from the PharmGKB pharmacogenomics database and the OMIM Human Mendelian Inheritance Online Database via API interfaces; historical medical data of 10,000 patients are extracted from the HIS systems of 10 hospitals; and 20 authoritative international and domestic clinical guidelines are collected. Natural language processing technology is used to extract 15 categories of medical entity nodes, such as diseases and symptoms, using a Conditional Random Field (CRF) model. For example, entities such as "coronary heart disease" and "fever" are identified from the literature. A combination of rule-based and deep learning methods is used to determine the disease-complication. The system establishes 30 entity attribute relationships, such as "symptoms" and "drug-indication," for example, establishing a relationship of "aspirin-prevention of cardiovascular diseases." Entities and relationships are stored in the Neo4j graph database to construct a heterogeneous medical knowledge graph. When new medical clinical trial information is released, such as the results of a trial for a novel anticancer drug A, a delayed conflict learning update is adopted. The trial information is preprocessed, and a named entity recognition model is used to extract new entities such as "drug A" and "new target gene B." These are compared with existing knowledge in the graph. If a conflict is found between the indication of "drug A" and a contraindication of a drug in the graph, the credibility of the new evidence is calculated to be 0.8 and the credibility of the original knowledge is 0.6, based on factors such as the clinical trial sample size (set to 500 cases) and the research quality score (set to 8 points). The graph is then updated first, ultimately generating a dynamically updated medical knowledge graph.

[0021] S2: Medical data quality assessment module, used to query corresponding multimodal medical data through federated access protection, including electronic medical records, medical images and medical genomics data, and to perform quality verification and assessment on multimodal medical data to obtain multimodal medical calibration data; In this embodiment of the invention, a federated learning framework is used to connect the databases of three hospitals and one gene testing institution. The Paillier homomorphic encryption algorithm is used to encrypt query commands. Taking the query of diabetic patient data as an example, data such as age and blood glucose levels in the numerical query command are encrypted. For example, the age... Encryption Simultaneously, differential privacy technology is applied, and Laplace noise is added. Set a privacy budget =0.5, query function sensitivity =10, noise reduction was added to the returned 100 electronic medical records, 20 CT image files, and 15 gene data files to ensure a response time of 480ms, meeting the requirement of ≤500ms. Quality verification and evaluation were performed on the multimodal medical data. Regarding consistency verification, for diagnostic information and examination results in the electronic medical records, assuming the diagnostic information contains 10 keywords and the examination results contain 15 keywords, with 7 identical keywords, according to the formula S_consistency=2 The consistency score S_consistency = 2 × 7 / 10 + 15 = 0.56 is obtained from n_same / n_1 + n_2. The semantic matching score of medical image annotation and content is set to 0.8. The accuracy rate of gene data detection results is 0.9, and the comprehensive score S_total = 0.4 × 0.56 + 0.3 × 0.8 + 0.3 × 0.9 = 0.714 > 0.7. For missing medication records in electronic medical records, K-nearest neighbor interpolation (K=5) is used to supplement them based on similar patient records, and finally multimodal medical calibration data is obtained.

[0022] S3: Medical Feature Engineering Module, used to perform medical feature engineering analysis on multimodal medical calibration data to generate multimodal medical feature vectors, including medical sub-features corresponding to electronic medical records, medical images and genomes; In this embodiment of the invention, medical feature engineering analysis is performed on multimodal medical calibration data. For electronic medical record data, the TF-IDF algorithm is used to extract text features. Assuming there are 1000 diabetic patient medical records, and "polydipsia" appears in 200 of them, with a total of 10000 medical records, according to the formula... (in For words In the document word frequency in , Total number of documents For containing words (Number of documents), calculate the TF-IDF value of "excessive thirst", and normalize numerical test indicators such as blood glucose levels. The formula is: The electronic medical record medical sub-feature vector is obtained. For medical image data, ResNet50 convolutional neural network is used to extract CT image features. The image is input into the network, and after multiple convolution and pooling operations, a 2048-dimensional feature vector is output from the last fully connected layer. For genomic data, one-hot encoding is performed first, and then it is mapped to a 50-dimensional feature vector through a fully connected layer. At the same time, the medical sub-features corresponding to the electronic medical record, medical image and genome are concatenated to finally generate a multimodal medical feature vector.

[0023] S4: Medical Knowledge-Driven Analysis Module, which is used to perform joint knowledge path-driven analysis on multimodal medical feature vectors based on dynamically updated medical knowledge graphs and a path search mechanism based on reinforcement learning, in order to generate decision support reasoning paths corresponding to medical clinical knowledge and output the corresponding medical knowledge path confidence.

[0024] In this embodiment of the invention, taking the multimodal medical feature vector of diabetic patients as an example, a knowledge path joint-driven analysis is performed based on a dynamically updated medical knowledge graph. Cosine similarity is used to calculate the correlation between the feature vector and knowledge entities in the knowledge graph. For example, the similarity between the knowledge entity "diabetes" and electronic medical records, images, and gene-related features in the multimodal feature vector are 0.8, 0.7, and 0.6, respectively. Using the weighted summation formula S = 0.5 × 0.8 + 0.3 × 0.7 + 0.2 × 0.6 = 0.73, and setting a threshold of 0.7, knowledge entities with a correlation greater than or equal to the threshold are selected. Simultaneously, a path search mechanism based on reinforcement learning is used, employing a depth-first search algorithm, limiting the number of inference hops to 3. The system consists of 5 layers (≤5 layers). Starting with the entity "diabetes," the first layer identifies directly related entities such as "insulin" and "blood sugar," with edge weights set to 0.8 and 0.9 based on relevance and domain knowledge. The second layer identifies entities such as "insulin resistance" from "insulin," calculates new edge weights, and ultimately generates a decision support reasoning path for "diabetes-blood sugar-insulin resistance-hypoglycemic drugs." When calculating the confidence of the medical knowledge path, the system first extracts the feature patterns of each knowledge entity along the reasoning path, such as typical symptoms and weights of "diabetes" and detection indicators for "insulin resistance." Assuming the patient has symptoms of polydipsia and polyphagia, and a blood sugar level of 12 mmol / L, the patient's pathological state is transformed into a vector and compared with the feature pattern vectors of the knowledge entities. Cosine similarity is used to calculate the similarities of "diabetes," "blood sugar," "insulin resistance," and "hypoglycemic drugs," which are 0.8, 0.9, 0.7, and 0.6, respectively. The confidence of each knowledge entity in the graph is set to 0.9, 0.8, 0.7, and 0.8, respectively. The formula Confidence = ( 0.5×0.8+0.3×0.9+0.2×0.7)×0.6×0.9×0.8×0.7×0.8≈0.18, output the confidence score of the medical knowledge path for this path, and provide a basis for medical decision-making.

[0025] Furthermore, the medical knowledge graph construction module includes the following functions: By integrating medical clinical guidelines, biomedical literature, drug databases, and historical medical data; Based on medical clinical guidelines, biomedical literature, drug databases, and historical medical data, medical entity nodes were extracted to obtain 15 types of medical clinical entity nodes, including diseases, symptoms / signs, drugs, genes / biomarkers, anatomical structures, laboratory test indicators, imaging features, surgical procedures, patient groups, clinical trials, medical equipment, literature evidence, medical institutions, environmental factors, and medical entity node types corresponding to time points. Based on 15 types of medical and clinical entity nodes and combined with medical and clinical guidelines, biomedical literature, drug databases, and historical medical data, an analysis of entity attribute edge relationships was conducted to obtain 30 types of medical entity attribute edge relationships. These include disease-complications, disease-symptoms, disease-subtypes, disease-risk factors, disease-genetic associations, drugs-indications, drugs-contraindications, drugs-adverse reactions, surgery-disease treatment, equipment-applicable scenarios, genes-targets, pathways-diseases, biomarkers-prognosis, microorganisms-diseases, guidelines-recommended treatments, literature-supporting conclusions, clinical trials-results, medical data-efficacy differences, diseases-regional high incidence, environment-disease risk, time-prognosis, hospitals-specialty advantages, doctors-surgical proficiency, medical insurance-coverage, population-treatment response, comorbidities-treatment limitations, lifestyle-intervention, new evidence-overturning old knowledge, expert consensus-points of contention, and data source-credibility. Based on 15 types of medical clinical entity nodes and 30 types of medical entity attribute edge relationships, and combined with heterogeneous information network modeling technology, a heterogeneous knowledge graph model is constructed to generate a medical heterogeneous knowledge graph. Acquire new medical clinical trial information and update the medical heterogeneous knowledge graph with delayed contradiction learning based on the new medical clinical trial information to generate a dynamically updated medical knowledge graph.

[0026] As an embodiment of the present invention, reference Figure 2 As shown, Figure 1 A functional flowchart of the Traditional Chinese Medicine knowledge graph construction module. In this embodiment, the medical knowledge graph construction module includes the following functions: S11: By integrating medical clinical guidelines, biomedical literature, drug databases, and historical medical data; In this embodiment of the invention, biomedical literature was collected from multiple authoritative medical websites (such as PubMed from the U.S. National Library of Medicine and the medical database of CNKI). Search rules were set, and the literature was filtered and downloaded according to literature categories (such as clinical research, reviews, etc.) and keywords (such as disease names, drug names, etc.). A total of 5,000 relevant articles were collected. Drug information and gene / biomarker data were obtained from professional medical databases (such as the PharmGKB pharmacogenomics database and the OMIM online database of human Mendelian inheritance, etc.). Detailed information on 1,000 drugs was obtained through API interfaces or data download methods. A dataset containing detailed information (such as indications, contraindications, adverse reactions, etc.) and data related to 500 genes / biomarkers was created. Historical medical data, including patient medical records, examination reports, and treatment records, was extracted from the Hospital Information System (HIS). After data cleaning and organization, historical medical data covering 10,000 patients was obtained. At the same time, 20 authoritative international and domestic clinical guidelines (such as the American College of Cardiology's Cardiovascular Disease Guidelines and the Chinese Diabetes Prevention and Treatment Guidelines) were collected. These data were integrated to form a comprehensive medical data resource library, providing rich data support for subsequent knowledge extraction and atlas construction.

[0027] S12: Extract medical entity nodes based on medical clinical guidelines, biomedical literature, drug databases and historical medical data to obtain 15 types of medical clinical entity nodes, including diseases, symptoms / signs, drugs, genes / biomarkers, anatomical structures, laboratory test indicators, imaging features, surgical procedures, patient groups, clinical trials, medical equipment, literature evidence, medical institutions, environmental factors and medical entity node types corresponding to time nodes. In this embodiment of the invention, Named Entity Recognition (NER) is employed in Natural Language Processing (NLP) to extract medical entity nodes from data such as clinical guidelines and biomedical literature. For disease entity nodes, a Conditional Random Field (CRF)-based model is used. This model is trained on a large amount of annotated medical text to learn the grammatical and semantic features of disease names, thereby accurately identifying disease entities in the text, such as "coronary heart disease" and "diabetes." For symptom / sign entity nodes, a combination of dictionary matching and rule-based methods is used. First, a dictionary containing common symptoms / signs is constructed. Regular expressions were used to match relevant expressions in the text, such as "chest pain" and "fever." For drug entity nodes, deep learning-based methods, such as recurrent neural networks (RNNs) combined with long short-term memory networks (LSTMs), were employed to identify and classify information such as drug name, dosage form, and dosage. These methods extracted 15 types of medical clinical entity nodes from the data, specifically including diseases (diagnosis name, stage (e.g., TNM stage), phenotypic characteristics (e.g., HER2-positive breast cancer), symptoms / signs (e.g., fever, visual acuity score (VAS), pathological reflexes), and drugs (chemical drugs). The study covers pharmaceuticals, biological agents, and traditional Chinese medicine compound prescriptions (including dosage forms, specifications, and manufacturer information), gene / biomarkers (such as BRCA1 and PD-L1 expression levels, cfDNA mutation frequency), anatomical structures (organs, tissues, cell types (such as the left lobe of the liver, neurons)), laboratory test indicators (complete blood count, biochemical indicators (such as HbA1c), microbial culture results), imaging features (CT / MRI signs (such as "ground-glass opacity"), PET metabolic values ​​(SUVmax)), surgical procedures (surgical techniques (such as laparoscopic cholecystectomy), implant types), and patient populations (age stratification, ethnicity). The types of medical entity nodes corresponding to the following factors—comorbidity grouping (e.g., "diabetes + hypertension"), clinical trials (trial protocols for each phase, inclusion / exclusion criteria, primary endpoints), medical equipment (ventilator models, stent materials, AI-assisted diagnostic software), literature evidence (randomized controlled trials (RCTs), meta-analysis, expert consensus), medical institutions (hospital level, specialty advantages, differences in treatment pathways), environmental factors (air pollution index (PM2.5), occupational exposure history, epidemiological spectrum of the place of residence), and time nodes (duration of onset, treatment cycle, follow-up interval)—lay the foundation for constructing a knowledge graph.

[0028] S13: Based on 15 types of medical and clinical entity nodes and combined with medical and clinical guidelines, biomedical literature, drug databases, and historical medical data, entity attribute edge relationship analysis was conducted to obtain 30 types of medical entity attribute edge relationship types, including disease-complications, disease-symptoms, disease-subtypes, disease-high-risk factors, disease-genetic associations, drugs-indications, drugs-contraindications, drugs-adverse reactions, surgery-treatment of diseases, equipment-applicable scenarios, genes-targets, pathways-diseases, biomarkers-prognosis, microorganisms-diseases, guidelines-recommended treatments, literature-supporting conclusions, clinical trials-results, medical data-efficacy differences, diseases-regional high incidence, environment-disease risk, time-prognosis, hospitals-specialty advantages, doctors-surgical proficiency, medical insurance-coverage, population-treatment response, comorbidities-treatment limitations, lifestyle-intervention, new evidence-overturning old knowledge, expert consensus-points of contention, and data source-credibility. In this embodiment of the invention, entity attribute edge relationship analysis is performed on data such as medical clinical guidelines and biomedical literature based on 15 types of medical clinical entity nodes. For disease-complication relationships, the co-occurrence of disease names and related complications in the text is searched, and the causal relationship between them is determined by combining semantic analysis, such as "coronary heart disease-myocardial infarction". For drug-indication relationships, drug instructions and clinical research literature are analyzed to determine the diseases or symptoms for which the drug is applicable, such as "aspirin-prevention of cardiovascular diseases". For surgery-treatment of diseases relationships, the surgery name and the corresponding treatment disease are extracted from surgical records and clinical guidelines, such as "coronary artery bypass grafting-treatment of coronary heart disease". In this way, various relationships in the data are analyzed in detail to obtain 30 types of medical entity attribute edge relationship types, which specifically include disease-complication (e.g., "diabetes → diabetic nephropathy"), disease-symptom (e.g., "lung cancer → hemoptysis"), disease-type (e.g., "breast cancer → luminal ulcer"), and disease-subtype (e.g., "breast cancer → luminal ulcer"). Hepatitis B virus type (HBV), disease-risk factors ("liver cancer → hepatitis B virus infection"), disease-genetic association ("colorectal cancer → APC gene mutation"), drug-indications ("atorvastatin → hyperlipidemia"), drug-contraindications ("ACE inhibitor → pregnancy"), drug-adverse reactions ("cisplatin → nephrotoxicity"), surgery-disease treatment ("PCI → coronary artery disease"), device-applicable scenarios ("ECMO → ARDS"), gene-target ("EGFR → erlotinib"), pathway-disease ("Wnt / β-catenin → colorectal cancer"), Biomarkers - Prognosis ("PD-L1 positive → Immunotherapy response"), Microbes - Disease ("Helicobacter pylori → Gastric cancer"), Guidelines - Recommended regimens ("NCCN guidelines → Pembrolizumab first-line treatment"), Literature - Supporting conclusions ("PMID123456 → Confirms the efficacy of ibrutinib"), Clinical trials - Results ("NCT045356 → OS extended by 4.2 months"), Medical data - Efficacy differences ("ORR increased by 12% in Asian populations"), Disease - Geographical prevalence ("Esophageal cancer → North China"), Environment - Disease risk ("PM2.5")The relationships between different entity nodes—including 5-year survival rate (for stage III colon cancer), hospital-specialty advantages (hospital A has a 95% survival rate after heart transplantation), surgeon proficiency (surgeons perform more than 200 surgeries per year), medical insurance coverage (national medical insurance covers osimertinib), population-treatment response (African Americans may experience warfarin dosage reduction), comorbidities-treatment limitations (renal insufficiency may lead to metformin contraindication), lifestyle-interventions (smoking may require cessation of smoking 4 weeks prior to surgery), new evidence overturning old knowledge (2023 RCT negates hormone therapy), expert consensus-points of contention (controversy over breast-conserving surgery margin width), and data source-credibility (FDA black box warning may indicate level I evidence)—connect different entity nodes, forming the edges of the knowledge graph and enriching its semantic information.

[0029] S14: Based on 15 types of medical clinical entity nodes and 30 types of medical entity attribute edge relationships, and combined with heterogeneous information network modeling technology, heterogeneous knowledge graph modeling is performed to generate a medical heterogeneous knowledge graph. In this embodiment of the invention, a heterogeneous knowledge graph is modeled based on 15 types of medical clinical entity nodes and 30 types of medical entity attribute edge relationships, combined with heterogeneous information network modeling technology. First, each entity node is represented as a node object, containing information such as the entity's name, type, and attributes. For example, for the disease entity node "coronary heart disease," its node object contains the name "coronary heart disease," the type "disease," and attributes such as "incidence rate" and "symptoms." Then, connections are established between nodes based on the edge relationship types between entities. For example, "coronary heart disease" and "myocardial infarction" are connected through a "disease-complication" edge relationship. The knowledge graph is stored and managed using a graph database (such as Neo4j), storing node and edge information for efficient querying and reasoning. In this way, a medical heterogeneous knowledge graph containing rich medical knowledge is generated. This graph can intuitively display the relationships between medical entities, providing strong support for joint analysis of medical data.

[0030] S15: Acquire new medical clinical trial information and update the medical heterogeneous knowledge graph with delayed contradiction learning based on the new medical clinical trial information to generate a dynamically updated medical knowledge graph.

[0031] In this embodiment of the invention, by acquiring new medical clinical trial information, such as the results of a clinical trial on a novel anticancer drug, the new clinical trial information is first preprocessed, including data cleaning and standardization, to conform to the data format requirements of the knowledge graph. Then, based on the new medical clinical trial information, the medical heterogeneous knowledge graph is updated using delayed contradiction learning. For newly emerging entity nodes, such as the name of the novel anticancer drug and related gene targets, they are added to the knowledge graph, and corresponding edges are established according to their relationships with other entities. For existing entity nodes and edge relationships, if the new clinical trial results contradict or are inconsistent with the existing information in the knowledge graph, the correction is made by adjusting the edge weights or updating the node attributes. For example, if the new clinical trial shows that the adverse reactions of a certain drug are different from previous understandings, the "adverse reactions" attribute of the drug node is updated. In this way, the medical heterogeneous knowledge graph is continuously updated and improved, ultimately generating a dynamically updated medical knowledge graph that can reflect the latest medical research results and clinical practice experience in a timely manner.

[0032] Furthermore, the acquisition of new medical clinical trial information and the delayed contradiction learning update of the medical heterogeneous knowledge graph based on the new medical clinical trial information include: Obtain new medical clinical trial information; In this embodiment of the invention, an automated data search program is set up by real-time monitoring of authoritative domestic and international medical clinical trial registration platforms (such as ClinicalTrials.gov in the United States and the Chinese Clinical Trial Registry) and the official websites of top medical journals (The New England Journal of Medicine and The Lancet). Taking ClinicalTrials.gov as an example, the website is set to scan every 15 minutes. For newly published clinical trial information, it is filtered according to pre-set filtering rules (such as searching only for completed trials with publicly available results). When a journal's official website detects that a clinical trial result for a novel hypoglycemic drug A has been published, the program immediately starts the data collection process, extracting key information such as the trial name, participants, intervention measures, main results, and research conclusions from the webpage. At the same time, it captures links to original data documents related to the trial, downloads the complete PDF report, and performs data deduplication for information from multiple sources. Finally, complete and accurate new medical clinical trial information is obtained, providing basic data for subsequent analysis.

[0033] Preferably, new entity identification analysis is performed on the new medical clinical trial information to obtain the new entities in the medical clinical trial, and the identification time delay corresponding to the new entities in the medical clinical trial is determined to be ≤1 hour. In this embodiment of the invention, a named entity recognition model based on a combination of Bidirectional Long Short-Term Memory (BiLSTM) and Conditional Random Field (CRF) is used to perform new entity recognition analysis on newly acquired medical clinical trial information. The trial information text is segmented into sentences and input into the model. The BiLSTM network extracts features from the text in both forward and reverse directions, capturing the contextual semantic information of words. For example, in the sentence "A novel hypoglycemic drug A can significantly reduce the glycated hemoglobin level in patients with type 2 diabetes," the semantic features of potential entities such as "novel hypoglycemic drug A," "type 2 diabetes," and "glycated hemoglobin" can be effectively extracted. The CRF layer, based on the features extracted by BiLSTM, considers the sequence dependencies between words and labels each word with its entity category, determining which of the 15 categories of medical clinical entity nodes it belongs to, such as disease, drug, and examination indicators. To ensure that the recognition time delay is ≤1 hour, a high-performance computing cluster (equipped with a 32-core CPU, 256GB of memory, and 4 NVIDIA A100 processors) is used. The GPU can process multiple text data in parallel. When a clinical trial report containing 5,000 words is input, the model can complete the processing within 30 minutes, accurately identify new entities such as "drug A" and "postprandial blood glucose", and perform uniqueness verification on each new entity to avoid duplicate recording, thus ensuring the efficiency and accuracy of new entity recognition.

[0034] Preferably, clinical knowledge relationship reasoning is performed between existing medical knowledge entities in the medical heterogeneous knowledge graph based on newly added entities in medical clinical trials, so as to generate a medical evidence knowledge relationship chain between the newly added entities and existing medical knowledge entities; In this embodiment of the invention, based on newly identified entities from medical clinical trials, a clinical knowledge relationship reasoning method combining rules and graph neural networks (GNN) is first constructed to build a medical knowledge reasoning rule base. For example, "If a drug has a therapeutic effect on a specific disease, then there is a drug-treatment-disease relationship." Taking the newly added entity "Drug B" and the existing "hypertension" disease entity in the knowledge graph as an example, if the clinical trial information contains the statement "Drug B can effectively reduce the blood pressure level of hypertensive patients," it can be preliminarily determined that there is a "Drug B-treatment-hypertension" relationship according to the rules. Then, a graph neural network is used to model the knowledge graph, with entities as nodes and existing relationships as edges. New entities and related existing entities are combined to form a subgraph, which is then input into the GNN model. The GNN learns the potential relationships between entities through information transfer and aggregation between nodes. For example, by analyzing the structural similarity and mechanism of action of "Drug B" and "angiotensin receptor antagonist" (an existing drug entity), the relationship "Drug B - similar drugs - angiotensin receptor antagonist" is inferred. Ultimately, a complete chain of medical evidence knowledge relationships is generated, such as "Drug B - treatment - hypertension - complications - kidney failure - related examinations - serum creatinine detection," clarifying the association between new entities and existing medical knowledge entities.

[0035] Preferably, the medical heterogeneous knowledge graph is updated by performing contradiction detection, learning, and updating based on newly added entities in medical clinical trials and the medical evidence knowledge relationship chain between the newly added entities and existing medical knowledge entities, so as to generate a dynamically updated medical knowledge graph.

[0036] In this embodiment of the invention, based on newly added entities in medical clinical trials and the generated medical evidence knowledge relationship chain, a contradiction detection learning and updating method based on conflict detection algorithms and knowledge fusion strategies is adopted. For a newly added entity "Drug C", if its labeled indication overlaps with the contraindications of an existing drug "Drug D" in the knowledge graph (e.g., the indication for "Drug C" is "treatment of heart failure", and the contraindications for "Drug D" include "contraindicated in patients with heart failure"), a contradiction detection process is triggered. This process calculates the confidence level of the contradictory relationship (considering factors such as clinical trial sample size and research quality score). When the confidence level exceeds a threshold (set to 0.8), the new entity is preferentially adopted. For clinical trial evidence, the weight of the existing "Drug D - Contraindication - Heart Failure" relationship edge in the knowledge graph is reduced, and a new "Drug C - Indication - Heart Failure" relationship edge is added. At the same time, the source of evidence and update time are recorded in the node attributes. For the newly added relationship chain, such as "Drug E - Target - Gene F - Associated Disease - Tumor", it is fully integrated into the knowledge graph, and the neighbor node information and edge connection relationships of the relevant nodes are updated. After comprehensive contradiction detection and knowledge fusion operations, a dynamic updated medical knowledge graph containing the latest medical knowledge is finally generated, ensuring the accuracy and timeliness of the knowledge graph and providing a reliable basis for joint analysis of medical data.

[0037] Furthermore, the step of performing contradiction detection, learning, and updating of the heterogeneous medical knowledge graph based on newly added entities from medical clinical trials and the medical evidence knowledge relationship chain between these newly added entities and existing medical knowledge entities includes: Based on the medical evidence knowledge relationship chain between the new entity and the existing medical knowledge entity, the existing medical knowledge entities that are related to the new entity in the medical heterogeneous knowledge graph are inferred and searched out, so as to obtain the existing medical evidence knowledge entities that are related to the new entity. In this embodiment of the invention, taking the newly added entity "novel antidepressant A" as an example, a reasoning search is performed in the medical heterogeneous knowledge graph based on its medical evidence knowledge relationship chain. Assuming the relationship chain is "novel antidepressant A - mechanism of action - serotonin reuptake inhibition - associated neurotransmitter - serotonin - related diseases - depression", a depth-first search (DFS) algorithm is used. Starting from the "novel antidepressant A" node, the "serotonin reuptake inhibition" node is found along the "mechanism of action" relationship edge, then the "serotonin" node is found according to the "associated neurotransmitter" relationship edge, and finally the "depression" node is found through the "related diseases" relationship edge. These existing medical knowledge entities such as "serotonin reuptake inhibition", "serotonin", and "depression" connected by the relationship chain are identified as existing medical evidence knowledge entities that are related to the newly added entity. During the search process, the path information of each entity is recorded. If multiple paths connect the same entity, the relationship chain corresponding to the shortest path is selected to ensure the accuracy and efficiency of the search results.

[0038] Preferably, a knowledge conflict difference analysis is performed on the existing medical evidence knowledge entities that are related to the newly added entities in the medical clinical trial to obtain the conflict difference between the newly added entities and the existing knowledge entities. In this embodiment of the invention, a knowledge conflict analysis is performed on the existing medical evidence knowledge entity "depression" that is related to the newly added entity "novel antidepressant drug A". A quantitative comparison is made from three dimensions: treatment effect, side effects, and applicable population. Regarding treatment effect, the average remission rate of existing treatment regimens for depression is assumed to be... =65%, the average remission rate in clinical trials of "new antidepressant drug A" was 65%. =78%, then the difference in treatment effect =78%−65%=13%, Regarding side effects, let the incidence of side effects of the existing treatment regimen be 78%−65%=13%. =30%, the incidence of side effects of "new antidepressant drug A" is 30%. =20%, then the difference in side effects =30%−20%=10%, Regarding the applicable population, the statistical coverage of the applicable population for existing treatment plans is as follows: =80%, the coverage of the applicable population for "New Antidepressant Drug A" is 80%. =85%, then the difference in the applicable population is =85%−80%=5%. Taking into account the three dimensions, a weighted summation formula is used to calculate the difference in conflict levels. =0.5×13%+0.3×10%+0.2×5%=6.5%+3%+1%=10.5%, finally obtaining the difference in contradictions and conflicts between the newly added entities and the existing knowledge entities.

[0039] Preferably, based on the difference in the amount of contradictions and conflicts between the newly added entities and existing knowledge entities, an expert evidence level review is conducted between the newly added entities in medical clinical trials and the existing medical evidence knowledge entities that are related to the newly added entities, so as to generate a medical evidence level comparison matrix between the newly added entities and existing knowledge entities. In this embodiment of the invention, based on the discrepancy D=10.5% between the newly added entity "novel antidepressant drug A" and the existing knowledge entity "depression," an expert evidence level review is conducted. A review panel of five authoritative psychiatric experts is invited. Each expert scores the evidence level of "novel antidepressant drug A" and existing treatment options for "depression" based on three indicators: clinical trial sample size, research design rigor, and result reproducibility. The scoring range is 1-10 points. The scores for "New Antidepressant Drug A" in the three indicators of clinical trial sample size, rigor of study design, and reproducibility of results were as follows: , , The scores for existing treatment options for "depression" are as follows: , , The weighted average formula is used to calculate the level of evidence for each entity, with the following weights: =0.4、 =0.3、 If the evidence level is 0.3, then the level of evidence for "new antidepressant A" is... Level of evidence for existing treatments for depression Suppose one of the experts' calculations is =8.2 points, =7.5 points, thus combining the calculation results of the 5 experts to generate a medical evidence level comparison matrix, and finally generating a medical evidence level comparison matrix between the newly added entity and the existing knowledge entity.

[0040] Preferably, the knowledge increment corresponding to the existing knowledge entity is taken as the difference in contradiction and conflict between the newly added entity and the existing knowledge entity, and the knowledge increment learning and updating of the corresponding existing knowledge entity in the medical heterogeneous knowledge graph is performed based on the knowledge increment corresponding to the existing knowledge entity and combined with the medical evidence level comparison matrix between the newly added entity and the existing knowledge entity, so as to generate a medical dynamically updated knowledge graph.

[0041] In this embodiment of the invention, the difference in contradiction / conflict between the newly added entity "novel antidepressant drug A" and the existing knowledge entity "depression" is D=10.5% as the knowledge increment corresponding to the existing knowledge entity "depression". Combined with a medical evidence level comparison matrix, because... =8.2 points greater than =7.5 points, with significant knowledge increment. Knowledge increment learning and updating were performed on the "depression" entity in the medical heterogeneous knowledge graph. In the treatment plan section, information such as "new antidepressant A" and its treatment effects, applicable population, and side effects were added to the "depression" treatment plan node, and the recommendation priority of the treatment plan was updated. Let the original treatment plan recommendation priority be... Recommendation priority adjusted based on differences in evidence level ≈0.656. At the same time, side effect information of "new antidepressant drug A" is added to the side effect node, and the coverage data is updated in the applicable population node. Finally, a medical dynamic update knowledge graph containing the latest knowledge is generated, making the knowledge graph more in line with the actual progress of medicine.

[0042] Furthermore, the expert evidence level review based on the discrepancy between newly added entities and existing knowledge entities in medical clinical trials, as well as existing medical evidence knowledge entities related to the newly added entities, includes: Based on the difference in contradictions and conflicts between new entities and existing knowledge entities, the potential impact of conflict between new entities in medical clinical trials and existing medical evidence knowledge entities that are related to new entities is assessed to obtain the degree of potential impact of medical knowledge conflict between new entities and existing knowledge entities. In this embodiment of the invention, taking the newly added entity "novel antidepressant drug A" and the existing knowledge entity "depression" as examples, based on the calculated discrepancy D=10.5% between the two, the potential impact of the conflict is assessed from three dimensions: clinical application, academic research, and patient education. In the clinical application dimension, let the number of patients currently using the existing treatment plan be... =1000 people, if the new drug is used, the estimated number of patients affected is =300 people, influence coefficient =0.3; In terms of academic research, considering the citation count of relevant research papers, the citation count of the existing treatment plan in the past three years is: =500 times, with an estimated annual increase in citations for newly added drug-related research. =100 times, influence coefficient = 0.2; Patient education dimension, assuming the patient awareness rate of the existing treatment plan is 0.2. =70%, the expected increase in awareness of newly added drugs is =15%, Influence coefficient The potential impact of medical knowledge conflict was approximately 0.214, and a weighted summation formula was used to calculate the degree of impact. =0.4×0.3+0.3×0.2+0.3×0.214=0.12+0.06+0.0642=0.2442, thus obtaining the potential impact of the medical knowledge conflict between "new antidepressant drug A" and "depression".

[0043] Preferably, the conflict priority of each existing medical evidence knowledge entity that is related to the new entity is ranked based on the potential impact of medical knowledge conflict between the new entity and the existing knowledge entity, so as to generate a conflict priority sequence of existing knowledge entities that are related to the new entity. In this embodiment of the invention, the potential impact of medical knowledge conflict is calculated for existing medical evidence knowledge entities related to the newly added entity "novel antidepressant drug A," namely "depression," "serotonin," and "serotonin reuptake inhibition." The potential impact of "depression" has already been calculated. =0.2442; For "5-hydroxytryptamine", the evaluation is carried out from three sub-dimensions: theoretical system update, change in experimental research direction, and correlation with drug development. The influence coefficient of theoretical system update is set as follows: =0.1, the influence coefficient of changing the experimental research direction =0.15, drug development-related impact coefficient =0.08, using the weighted summation formula =0.5×0.1+0.3×0.15+0.2×0.08=0.05+0.045+0.016=0.111; For "serotonin reuptake inhibition", the same evaluation was conducted from three sub-dimensions: related mechanism research, drug classification adjustment, and correlation with clinical treatment regimens, resulting in... =0.18, sorted from largest to smallest impact, the conflict priority sequence is "depression", "serotonin reuptake inhibition" and "serotonin". Finally, the existing knowledge entity conflict priority sequence corresponding to the newly added entity "novel antidepressant drug A" is generated, which clarifies the order of subsequent processing.

[0044] Preferably, based on the priority sequence of conflicts between existing knowledge entities corresponding to the newly added entity, the expert evidence level is reviewed between the newly added entity in the medical clinical trial and each existing medical evidence knowledge entity that has a conflict, in order of arrangement, so as to generate a medical evidence level comparison matrix between the newly added entity and the existing knowledge entities.

[0045] In this embodiment of the invention, based on the generated priority sequence of existing knowledge entity conflicts corresponding to the newly added entity "novel antidepressant drug A", relevant information is intelligently pushed to a review panel composed of five authoritative psychiatric experts, targeting "depression" which is ranked first. The pushed information includes clinical trial data of "novel antidepressant drug A", the difference in conflict between it and existing treatment plans for "depression" (D=10.5%), and the potential impact of medical knowledge conflict. =0.2442, and detailed information on existing treatment plans. Experts scored "New Antidepressant A" and existing treatment plans for "Depression" based on three indicators: clinical trial sample size, research design rigor, and result reproducibility. Let expert i's score be... , , and , , The evidence level for each entity was calculated using a weighted average formula, with the weights being respectively... =0.4、 =0.3、 If the evidence level is 0.3, then the level of evidence for "new antidepressant A" is... Level of evidence for existing treatments for depression After completing the review of "depression", the same operation is performed on "serotonin reuptake inhibition" and "serotonin" in sequence, and finally a medical evidence level comparison matrix is ​​generated between the newly added entity and the existing knowledge entity.

[0046] Furthermore, the medical data quality assessment module includes the following functions: By providing federated access to cross-agency medical databases and employing homomorphic encryption and differential privacy... The corresponding dual-protection access protection query includes multimodal medical data, including electronic medical records, medical images, and medical genomics data; Contrastive learning was used to perform cross-modal alignment of multimodal medical data, resulting in aligned multimodal medical field data. The quality of the aligned data for multimodal medical fields is evaluated to ensure that the quality score of the corresponding medical data is greater than 0.7 through consistency verification. The missing fields are then compensated by interpolation to obtain multimodal medical calibration data.

[0047] As an embodiment of the present invention, reference Figure 3 As shown, Figure 1 A functional flowchart of the medical data quality assessment module in this embodiment is shown. The medical data quality assessment module includes the following functions: S21: Access cross-agency medical databases through federated access, and employ homomorphic encryption and differential privacy. The corresponding dual-protection access protection query includes multimodal medical data, including electronic medical records, medical images, and medical genomics data; In this embodiment of the invention, a federated learning framework is used to access a cross-institutional medical database, connecting the HIS systems of three different hospitals and the database of one gene testing institution. To ensure data security, the Paillier homomorphic encryption algorithm is used to encrypt the query command. This algorithm allows addition and multiplication operations to be performed on the ciphertext, for example, for numerical data in the query command. and After encryption, the result is and It can be calculated in encrypted state. and Simultaneously, differential privacy technology is applied by adding Laplace noise. (in For the sensitivity of the query function, For privacy budget, set =0.5) The query results are perturbed. Taking the query of electronic medical records, CT images and genetic data of diabetic patients as an example, the encrypted query command is sent to the database of each institution. After each institution performs the query operation in its local database, it adds noise to the results and returns them. By optimizing the network architecture and caching mechanism, it is ensured that the response time of access protection query is ≤500ms. If a query returns 100 electronic medical record data, 20 CT image files and 15 genetic data files, the actual time from sending the query command to receiving all the data is 480ms, which meets the response time requirement and successfully obtains multimodal medical data.

[0048] S22: Contrastive learning is used to perform cross-modal alignment of multimodal medical data to obtain multimodal medical field aligned data; In this embodiment of the invention, a cross-modal alignment method based on contrastive learning is used to process multimodal medical data. Electronic medical record text data is converted into vector representation using a word embedding model (e.g., Word2Vec, with a vector dimension of 300). Medical image data (e.g., CT images) has feature vectors extracted using a convolutional neural network (CNN, employing a ResNet50 architecture). Medical genomic data is converted into vectors through one-hot encoding and then mapped to feature vectors via a fully connected layer, in order to construct a contrastive learning loss function. ,in Let be the sample size. =500; For the sample eigenvectors, It is the feature vector of its positive samples (different modalities of data from the same patient). The feature vector of the negative samples (data from different patients); The cosine similarity function is used. Let temperature be the parameter. =0.1, by minimizing the loss function, the feature vectors of different modalities of the same patient are brought closer together in the feature space, while the feature vectors of different patients are moved further apart. For example, after training, the cosine similarity between the feature vector of a diabetic patient's electronic medical record and the corresponding CT image and gene data feature vectors is improved from 0.3 to 0.7, realizing cross-modal alignment of multimodal medical data, and finally obtaining multimodal medical field aligned data.

[0049] S23: Perform quality verification and evaluation on the multimodal medical field alignment data to ensure that the quality score of the corresponding medical data is greater than 0.7 through consistency verification, and compensate for the corresponding missing fields through interpolation to obtain multimodal medical calibration data.

[0050] In this embodiment of the invention, quality verification and evaluation are performed on the aligned data of multimodal medical fields. Regarding consistency verification, for diagnostic information and examination results in electronic medical records, a consistency score is calculated. Assuming the diagnostic information contains n_1 = 10 keywords and the examination results contain n_2 = 15 keywords, with n_same = 7 identical keywords, the consistency score S_consistency = 2 × n_same / n_1 + n_2 = 2×7 / 10+15=0.56; For the annotation information and actual image content of medical images, the consistency score is calculated using the image semantic matching algorithm, and the score is set as S_image=0.8; For the detection results and reference standards of gene data, the accuracy is calculated as the consistency score, set as S_gene=0.9. Then the comprehensive score S_total=0.4×S_consistency+0.3×S_image+0.3×S_gene=0.4×0.56+0.3×0.8+0.3×0.9=0.714>0.7, which meets the quality requirements. For data with missing fields, such as the missing medication record of a patient in the electronic medical record, the K-nearest neighbor interpolation method (K is set to 5) is used to interpolate and compensate based on the medication records of other similar patients (judged by the cosine similarity of the feature vectors of the medical records). Finally, complete and qualified multimodal medical calibration data is obtained, which provides a reliable data foundation for subsequent medical knowledge graph construction and joint analysis of medical data. Furthermore, the response time for the access protection query is ≤500ms.

[0051] Furthermore, step S4 includes the following steps: Based on the knowledge entities within the dynamically updated medical knowledge graph, feature semantic association analysis is performed on the multimodal medical feature vectors to obtain the feature semantic association degree between the multimodal medical features and the knowledge entities within the knowledge graph. In this embodiment of the invention, taking the multimodal medical feature vector of a diabetic patient as an example, this vector includes symptom descriptions and examination indicators from electronic medical records, features of CT images, and relevant information from gene data. In the dynamically updated medical knowledge graph, there are knowledge entities such as "diabetes," "insulin," "blood sugar," and "complications." For the "diabetes" knowledge entity, its semantic correlation with the multimodal medical feature vector is calculated. It is assumed that the semantic similarity between keywords such as "excessive thirst, excessive hunger, and excessive urination" and "elevated blood sugar" in the electronic medical record and the "diabetes" knowledge entity are 0.8 and 0.9, respectively. Features related to diabetes in CT images (such as changes in pancreatic morphology) are also considered. The correlation between the multimodal medical feature vector and the knowledge entity "diabetes" is 0.7; the correlation between the gene data and genes related to diabetes is 0.6. Therefore, a weighted summation formula is used to calculate the semantic correlation of the features. Let the weights be 0.4, 0.3, 0.2, and 0.1 respectively. Then, the semantic correlation between "diabetes" and this multimodal medical feature vector, S_diabetes = 0.4 × 0.8 + 0.3 × 0.9 + 0.2 × 0.7 + 0.1 × 0.6 = 0.32 + 0.27 + 0.14 + 0.06 = 0.79. Using the same method, the semantic correlation between the multimodal medical features and other knowledge entities within the knowledge graph is calculated.

[0052] Preferably, based on the semantic correlation between multimodal medical features and various knowledge entities within the knowledge graph, and combined with a path search mechanism based on reinforcement learning, the corresponding reasoning hop count is used to perform joint knowledge path driving analysis between the multimodal medical feature vector and the knowledge entities within the dynamically updated medical knowledge graph whose semantic correlation is greater than or equal to a preset threshold, so as to generate decision support reasoning paths corresponding to medical clinical knowledge. In this embodiment of the invention, based on previously obtained multimodal medical features and the semantic correlation between various knowledge entities within the knowledge graph, a joint knowledge path-driven analysis is performed using a reinforcement learning-based path search mechanism. A preset threshold of 0.7 is used. For knowledge entities with a semantic correlation greater than or equal to 0.7, such as "diabetes," "insulin," and "blood sugar," path search begins. The path search mechanism employs a depth-first search (DFS) algorithm, limiting the number of inference jumps to ≤5 layers. Starting with the "diabetes" knowledge entity, the first layer searches for directly related knowledge entities, such as "insulin," "blood sugar," and "complications," calculating the edge weights between them (based on the semantic correlation and...). (Domain knowledge setting) Assuming the edge weight between "diabetes" and "insulin" is 0.8, the edge weight between "diabetes" and "blood sugar" is 0.9, and the edge weight between "complications" is 0.7, the second-layer search starts from "insulin," "blood sugar," and "complications" and continues to search for related knowledge entities. For example, from "insulin," knowledge entities such as "insulin secretion" and "insulin resistance" can be found, and new edge weights are calculated. This process continues, and in each layer of search, the optimal path is selected based on the edge weights. Finally, a decision support reasoning path corresponding to medical clinical knowledge is generated, such as "diabetes-blood sugar-insulin-insulin secretion-blood sugar regulation," ensuring that the number of reasoning jumps corresponding to the path search mechanism is within the limit (i.e., ≤5 layers).

[0053] Preferably, the confidence level of the corresponding medical knowledge path is calculated and output based on the decision support reasoning path corresponding to the medical clinical knowledge.

[0054] In this embodiment of the invention, the confidence score of the corresponding medical knowledge path is calculated and output based on the decision support reasoning path "diabetes-blood glucose-insulin-insulin secretion-blood glucose regulation" corresponding to medical clinical knowledge. A Bayesian network model is then used, combining known medical knowledge and statistical data, to assign probability values ​​to each node and edge in the path. It is assumed that the probability of elevated blood glucose in diabetic patients is... The probability of abnormal insulin secretion when blood sugar is high is... The probability of abnormal insulin secretion leading to blood glucose regulation imbalance is... Calculate the confidence level of the path according to the chain rule of Bayesian networks. In this way, the confidence level is calculated for each decision support reasoning path corresponding to the generated medical clinical knowledge, providing a quantitative reference for medical decision-making and helping doctors to more accurately assess and formulate treatment plans.

[0055] Furthermore, the inference hop count corresponding to the path search mechanism is ≤5 layers.

[0056] Furthermore, the calculation of the confidence level of the medical knowledge path corresponding to the decision support reasoning path based on the medical clinical knowledge includes: By using the decision support reasoning path corresponding to medical clinical knowledge, we can obtain the medical clinical knowledge feature patterns corresponding to each medical knowledge entity on the reasoning path. In this embodiment of the invention, taking the decision support reasoning path corresponding to the medical clinical knowledge "pneumonia - elevated white blood cell count - antibiotic treatment - inflammation resolution" as an example, the corresponding medical clinical knowledge feature patterns are extracted for each medical knowledge entity on the path. For the "pneumonia" knowledge entity, its typical symptom feature patterns, such as "cough," "fever," and "dyspnea," are extracted from the medical dynamic update knowledge graph, and the weights of these symptoms in the diagnosis of pneumonia are set to 0.4, 0.3, and 0.3, respectively. The feature pattern of the "elevated white blood cell count" knowledge entity is a white blood cell count higher than the normal range (assuming the normal range is (4-10)×10). 9 / L, higher than 10×10 9 / L is considered elevated); the feature patterns of the "antibiotic treatment" knowledge entity include commonly used antibiotic types (such as penicillins and cephalosporins) and their corresponding dosages and courses of treatment; the feature patterns of the "inflammation resolution" knowledge entity are indicators such as symptom relief (such as body temperature returning to normal and cough lessening) and imaging examinations (such as lung CT showing a reduction in shadow area), and these features and their corresponding weights are organized into feature pattern vectors to form the medical clinical knowledge feature patterns corresponding to each medical knowledge entity on this reasoning path.

[0057] Preferably, the patient's current medical pathological state is obtained, and the pathological knowledge similarity assessment of the patient's current medical pathological state is performed based on the medical clinical knowledge feature patterns corresponding to each medical knowledge entity on the reasoning path, so as to obtain the similarity between the current pathological state and the medical knowledge features on the reasoning path. In this embodiment of the invention, it is assumed that the patient currently has symptoms of cough and fever, with a body temperature of 38.5°C and a white blood cell count of 12 × 10⁻⁶. 9 / L, a lung CT scan showed patchy shadows, which was taken as the patient's current medical pathological state. Based on the previously obtained medical clinical knowledge feature pattern on the reasoning path of "pneumonia-elevated white blood cell count-antibiotic treatment-inflammation resolution", the similarity of the patient's current pathological state was evaluated. The similarity was calculated using the cosine similarity formula. For the similarity between the "pneumonia" knowledge entity feature pattern and the patient's pathological state, the patient's cough and fever symptoms were matched with the corresponding symptoms in the "pneumonia" feature pattern. The symptoms were converted into vector form. Let the patient's symptom vector A = (1, 1, 0) (cough, fever, no dyspnea), and the "pneumonia" feature pattern vector B = (0.4, 0.3, 0.3). Then the cosine similarity between the two is calculated. ≈0.92, for the knowledge entity "elevated white blood cell count", the patient's white blood cell count is 12 × 10⁻⁶. 9The feature pattern of / L satisfies the higher-than-normal range, and the similarity is set to 1; for the knowledge entity "inflammation subsidence", since the patient is currently ill and the inflammation subsidence feature is not present, the similarity is set to 0. The similarity between the current pathological state and the medical knowledge features on the reasoning path is calculated comprehensively using a weighted average method. The weights of "pneumonia", "elevated white blood cell count", and "inflammation subsidence" are set to 0.5, 0.3, and 0.2, respectively. The total similarity is then calculated. =0.5×0.92+0.3×1+0.2×0=0.76, finally obtaining the similarity between the current pathological state and the medical knowledge features on the reasoning path.

[0058] Preferably, the decision support reasoning path corresponding to the clinical medical knowledge is evaluated based on the similarity between the current pathological state and the medical knowledge features on the reasoning path, so as to calculate and output the corresponding medical knowledge path confidence.

[0059] In this embodiment of the invention, the similarity between the previously obtained current pathological state and the medical knowledge features along the reasoning path is used. =0.76, perform medical knowledge confidence calculation on the decision support reasoning path corresponding to the medical clinical knowledge "pneumonia-elevated white blood cell count-antibiotic treatment-inflammation resolution". For example, the corresponding confidence adjustment formula can be used. (in The number of knowledge entities along the reasoning path, here =4; To determine the credibility of each knowledge entity in the medical knowledge graph, let the credibility of "pneumonia," "elevated white blood cell count," "antibiotic treatment," and "inflammation resolution" be 0.9, 0.8, 0.7, and 0.8, respectively. Then, the confidence level of this medical knowledge path is... =0.76×0.9×0.8×0.7×0.8≈0.31. In this way, the confidence level of the corresponding medical knowledge path is calculated for each decision support reasoning path of medical clinical knowledge, providing a quantitative reference for medical decision-making and helping doctors judge the applicability and reliability of the reasoning path under the current patient's condition.

[0060] Therefore, the embodiments should be considered exemplary and non-limiting in all respects, and the scope of the invention is defined by the appended claims rather than the foregoing description. Thus, all variations falling within the meaning and scope of the equivalents of the application are intended to be included within the invention.

[0061] The above description is merely a specific embodiment of the present invention, enabling those skilled in the art to understand or implement the invention. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the invention. Therefore, the present invention is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features of the invention herein.

Claims

1. A medical data joint analysis system driven by medical knowledge graph, characterized in that, Includes the following modules: The medical knowledge graph construction module is used to integrate medical clinical guidelines, biomedical literature, drug databases, and historical medical data, and to model heterogeneous knowledge graphs based on these data to generate a medical heterogeneous knowledge graph; it also acquires new medical clinical trial information and updates the medical heterogeneous knowledge graph with delayed contradiction learning based on this new information to generate a dynamically updated medical knowledge graph. The medical data quality assessment module is used to query corresponding multimodal medical data through federated access protection, including electronic medical records, medical images and medical genomics data, and to perform quality verification and assessment on the multimodal medical data to obtain multimodal medical calibration data. The medical feature engineering module is used to perform medical feature engineering analysis on multimodal medical calibration data to generate multimodal medical feature vectors, including medical sub-features corresponding to electronic medical records, medical images, and genomes. The medical knowledge-driven analysis module is used to perform joint knowledge path analysis on multimodal medical feature vectors based on dynamically updated medical knowledge graphs and a path search mechanism based on reinforcement learning, so as to generate decision support reasoning paths corresponding to medical clinical knowledge and output the corresponding medical knowledge path confidence.

2. The medical data joint analysis system based on medical knowledge graph driven by claim 1, characterized in that, The medical knowledge graph construction module includes the following functions: By integrating medical clinical guidelines, biomedical literature, drug databases, and historical medical data; Based on medical clinical guidelines, biomedical literature, drug databases, and historical medical data, medical entity nodes were extracted to obtain 15 types of medical clinical entity nodes, including diseases, symptoms / signs, drugs, genes / biomarkers, anatomical structures, laboratory test indicators, imaging features, surgical procedures, patient groups, clinical trials, medical equipment, literature evidence, medical institutions, environmental factors, and medical entity node types corresponding to time points. Based on 15 types of medical and clinical entity nodes and combined with medical and clinical guidelines, biomedical literature, drug databases, and historical medical data, an analysis of entity attribute edge relationships was conducted to obtain 30 types of medical entity attribute edge relationships. These include disease-complications, disease-symptoms, disease-subtypes, disease-risk factors, disease-genetic associations, drugs-indications, drugs-contraindications, drugs-adverse reactions, surgery-disease treatment, equipment-applicable scenarios, genes-targets, pathways-diseases, biomarkers-prognosis, microorganisms-diseases, guidelines-recommended treatments, literature-supporting conclusions, clinical trials-results, medical data-efficacy differences, diseases-regional high incidence, environment-disease risk, time-prognosis, hospitals-specialty advantages, doctors-surgical proficiency, medical insurance-coverage, population-treatment response, comorbidities-treatment limitations, lifestyle-intervention, new evidence-overturning old knowledge, expert consensus-points of contention, and data source-credibility. Based on 15 types of medical clinical entity nodes and 30 types of medical entity attribute edge relationships, and combined with heterogeneous information network modeling technology, a heterogeneous knowledge graph model is constructed to generate a medical heterogeneous knowledge graph. Acquire new medical clinical trial information and update the medical heterogeneous knowledge graph with delayed contradiction learning based on the new medical clinical trial information to generate a dynamically updated medical knowledge graph.

3. The medical data joint analysis system based on medical knowledge graph driven by claim 2, characterized in that, The process of acquiring new medical clinical trial information and updating the medical heterogeneous knowledge graph using delayed contradiction learning based on this information includes: Obtain new medical clinical trial information; New entity identification analysis is performed on new medical clinical trial information to obtain new entities in medical clinical trials, and the identification time delay corresponding to the new entity in medical clinical trials is determined to be ≤1 hour. Based on newly added entities in medical clinical trials, clinical knowledge relationship reasoning is performed between existing medical knowledge entities in the medical heterogeneous knowledge graph to generate medical evidence knowledge relationship chains between newly added entities and existing medical knowledge entities; Based on newly added entities in medical clinical trials and combined with the medical evidence knowledge relationship chain between the newly added entities and existing medical knowledge entities, the heterogeneous medical knowledge graph is subjected to contradiction detection, learning and updating to generate a dynamically updated medical knowledge graph.

4. The medical data joint analysis system based on medical knowledge graph driven by claim 3, characterized in that, The method of performing contradiction detection, learning, and updating of the heterogeneous medical knowledge graph based on newly added entities from medical clinical trials and the medical evidence knowledge relationship chain between these new entities and existing medical knowledge entities includes: Based on the medical evidence knowledge relationship chain between the new entity and the existing medical knowledge entity, the existing medical knowledge entities that are related to the new entity in the medical heterogeneous knowledge graph are inferred and searched out, so as to obtain the existing medical evidence knowledge entities that are related to the new entity. Based on the newly added entities in medical clinical trials, a knowledge conflict difference analysis is conducted on the existing medical evidence knowledge entities that are related to the newly added entities, in order to obtain the amount of conflict difference between the newly added entities and the existing knowledge entities. Based on the difference in contradictions and conflicts between newly added entities and existing knowledge entities, expert evidence level review is conducted between newly added entities in medical clinical trials and existing medical evidence knowledge entities that are related to the newly added entities, in order to generate a medical evidence level comparison matrix between newly added entities and existing knowledge entities. By taking the difference in contradictions and conflicts between newly added entities and existing knowledge entities as the knowledge increment corresponding to the existing knowledge entity, and based on the knowledge increment corresponding to the existing knowledge entity and combined with the medical evidence level comparison matrix between newly added entities and existing knowledge entities, the knowledge increment of the corresponding existing knowledge entities in the medical heterogeneous knowledge graph is learned and updated to generate a medical dynamically updated knowledge graph.

5. The medical data joint analysis system based on medical knowledge graph driven by claim 4, characterized in that, The expert evidence level review based on the discrepancy between newly added entities and existing knowledge entities in medical clinical trials, and the existing medical evidence knowledge entities related to the newly added entities, includes: Based on the difference in contradictions and conflicts between new entities and existing knowledge entities, the potential impact of conflict between new entities in medical clinical trials and existing medical evidence knowledge entities that are related to new entities is assessed to obtain the degree of potential impact of medical knowledge conflict between new entities and existing knowledge entities. Based on the potential impact of medical knowledge conflicts between new entities and existing knowledge entities, the conflict priority of each existing medical evidence knowledge entity that is related to the new entity is ranked to generate a conflict priority sequence of existing knowledge entities that are related to the new entity. Based on the priority sequence of existing knowledge entities that have a relationship with the newly added entity, the expert evidence level is reviewed between the newly added entity in the medical clinical trial and each corresponding existing medical evidence knowledge entity that has a conflict, in order of sorting and intelligent push, so as to generate a medical evidence level comparison matrix between the newly added entity and the existing knowledge entities.

6. The medical data joint analysis system based on medical knowledge graph driven by claim 1, characterized in that, The medical data quality assessment module includes the following functions: By providing federated access to cross-agency medical databases and employing homomorphic encryption and differential privacy... The corresponding dual-protection access protection query includes multimodal medical data, including electronic medical records, medical images, and medical genomics data; Contrastive learning was used to perform cross-modal alignment of multimodal medical data, resulting in aligned multimodal medical field data. The quality of the aligned data for multimodal medical fields is evaluated to ensure that the quality score of the corresponding medical data is greater than 0.7 through consistency verification. The missing fields are then compensated by interpolation to obtain multimodal medical calibration data.

7. The medical data joint analysis system based on medical knowledge graph driven by claim 6, characterized in that, The response time for the access protection query is ≤500ms.

8. The medical data joint analysis system based on medical knowledge graph driven by claim 1, characterized in that, Step S4 includes the following steps: Based on the knowledge entities within the dynamically updated medical knowledge graph, feature semantic association analysis is performed on the multimodal medical feature vectors to obtain the feature semantic association degree between the multimodal medical features and the knowledge entities within the knowledge graph. Based on the semantic correlation between multimodal medical features and various knowledge entities within the knowledge graph, and combined with a path search mechanism based on reinforcement learning, the corresponding reasoning hop count is used to perform joint knowledge path analysis between the multimodal medical feature vector and the knowledge entities within the dynamically updated medical knowledge graph whose semantic correlation is greater than or equal to a preset threshold, so as to generate decision support reasoning paths corresponding to medical clinical knowledge. The confidence score of the corresponding medical knowledge path is calculated and output based on the decision support reasoning path corresponding to the medical clinical knowledge.

9. The medical data joint analysis system based on medical knowledge graph driven by claim 8, characterized in that, The path search mechanism corresponds to a reasoning jump count of ≤5 layers.

10. The medical data joint analysis system based on medical knowledge graph driven by claim 8, characterized in that, The calculation and output of the corresponding medical knowledge path confidence based on the decision support reasoning path corresponding to medical clinical knowledge includes: By using the decision support reasoning path corresponding to medical clinical knowledge, we can obtain the medical clinical knowledge feature patterns corresponding to each medical knowledge entity on the reasoning path. The current medical pathological state of the patient is obtained, and the pathological knowledge similarity assessment of the current medical pathological state of the patient is performed based on the medical clinical knowledge feature patterns corresponding to each medical knowledge entity on the reasoning path, so as to obtain the similarity between the current pathological state and the medical knowledge features on the reasoning path. Based on the similarity between the current pathological state and the medical knowledge features along the reasoning path, a medical knowledge confidence calculation is performed on the decision support reasoning path corresponding to the clinical medical knowledge, so as to calculate and output the corresponding medical knowledge path confidence.

Citation Information

Cited By

  • Medical data processing system and processing method

    CN121393703A

  • A medical data processing system and processing method

    CN121393703B

  • Federal learning-based multi-modal remote diagnosis and treatment method and medium

    CN121416058A

  • Enhanced learning medical record data association mining method and system

    CN121460216A

  • Drug sales data analysis method

    CN121685016A