Construction method of biomedical knowledge graph
By combining differential privacy technology with the BERT-BiLSTM-CRF model, the cross-source alignment and privacy protection issues in biomedical data integration are solved, efficient integration and continuous updating of multimodal data are achieved, and the quality and credibility of the knowledge graph are improved.
Patent Information
- Application Number
- CN202510697709.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-28
- Publication Date
- 2025-09-16
AI Technical Summary
Existing technologies lack cross-source data alignment mechanisms in biomedical data integration, resulting in entity redundancy, relationship conflicts, and update lags in knowledge graphs, as well as a lack of joint processing capabilities and privacy protection mechanisms for multimodal data.
Differential privacy technology is used for data preprocessing, the BERT-BiLSTM-CRF model is used for entity recognition and relationship labeling, and a graph embedding algorithm is combined for cross-source alignment. A quality control system including a rule engine and expert feedback is built, and multi-institutional collaboration and privacy protection are achieved through blockchain.
It achieves efficient integration of multimodal data and accurate alignment of cross-source data, provides dynamic privacy protection and continuous updating capabilities, supports multi-institutional collaboration, and improves the quality and credibility of knowledge graphs.
Smart Images

Figure CN120653782A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of knowledge graph technology, and in particular relates to a method for constructing a biomedical knowledge graph. Background Art
[0002] With the explosive growth of biomedical data, efficiently integrating multi-source heterogeneous data and constructing high-quality knowledge graphs has become a major challenge in biomedical research. Traditional approaches typically rely on a single data source (such as a literature database or electronic medical record) and lack cross-source data alignment mechanisms, leading to problems in knowledge graphs such as entity redundancy, relationship conflicts, and delayed updates. Furthermore, biomedical data involves patient privacy and multimodal data such as imaging, leading to the following limitations in existing technologies: 1) Lack of fine-grained classification and privacy protection mechanisms in the data preprocessing stage; 2) The knowledge extraction framework lacks the ability to jointly process unstructured text and image data.
[0003] In recent years, although some studies have attempted to combine deep learning with knowledge graph technology, such as BERT-based entity recognition or relational reasoning based on graph neural networks, there are still technical gaps in scenarios such as structured information extraction from medical images and federated graph updates.
[0004] Therefore, there is an urgent need for a method to construct a biomedical knowledge graph that can integrate multimodal data, support privacy protection, and has dynamic evolution capabilities. Summary of the Invention
[0005] In order to solve the technical problems mentioned in the background technology, a method for constructing a biomedical knowledge graph is provided.
[0006] To achieve the above objectives, the specific technical solutions of the method for constructing a biomedical knowledge graph of the present invention are as follows: A method for constructing a biomedical knowledge graph comprises the following steps: S1. Data preprocessing: Classify and grade multi-source biomedical data from literature databases, electronic medical records, genomic data, and medical images, and desensitize sensitive information using differential privacy technology; S2. Knowledge Extraction and Annotation: Entities are extracted using a joint model based on BERT-BiLSTM-CRF, relationships between entities are identified using an attention mechanism, and semantic annotation is performed using an ontology framework. S3, cross-source alignment and fusion: Calculate entity similarity based on a graph embedding algorithm and integrate entities and relationships from different data sources through a probabilistic soft alignment model; S4. Quality control system: Build a three-level verification mechanism including rule engine, statistical analysis and expert feedback to dynamically monitor the integrity, consistency and timeliness of the knowledge graph.
[0007] Furthermore, in S1, it includes: Use named entity recognition-driven hierarchical tagging for text data; implement DICOM format standardization and anonymization for image data; and implement fine-grained data access control through attribute-based encryption.
[0008] Furthermore, in S2, it includes: Multi-scale 3D CNN is used to extract lesion area features from medical images, an image-text cross-modal alignment model is constructed, image features are mapped to the ontology concept space, and an active learning strategy is used to optimize the selection of labeled samples.
[0009] Furthermore, in S3, we include: The TransEdge algorithm is used for cross-source relationship alignment, the conflict resolution algorithm is used to handle entity attribute contradictions, and the alignment threshold is dynamically adjusted based on reinforcement learning.
[0010] Furthermore, in S4, it includes: Deploy a redundancy detection module based on subgraph isomorphism, deploy a timeliness decay model to automatically mark expired knowledge, and deploy visualization tools to support manual correction and version backtracking.
[0011] The method for constructing a biomedical knowledge graph of the present invention has the following advantages: This invention integrates multimodal data such as text, genome, and image, supports cross-modal knowledge association, and solves the limitation of single data source of traditional methods. It builds a multi-level privacy protection system through differential privacy, ABE, homomorphic encryption, and zero-knowledge proof to achieve safe and compliant multi-institutional collaboration. The time series graph model and three-level verification mechanism ensure that the knowledge graph is continuously updated and reliable and available. The combination of TransEdge algorithm, reinforcement learning threshold adjustment, subgraph isomorphism detection and other technologies improves the efficiency and accuracy of cross-source data integration. The blockchain federation architecture and cross-chain technology design support seamless access to new data sources and collaborative institutions in the future, and provide high-quality knowledge support for clinical decision-making, drug development and other scenarios through dynamic evolution, secure collaboration and precise quality control. BRIEF DESCRIPTION OF THE DRAWINGS
[0012] Figure 1 This is a flow chart of the method for constructing the biomedical knowledge graph of the present invention. DETAILED DESCRIPTION
[0013] To make the purpose, technical solutions, and advantages of the embodiments of the present invention more clear, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without making any creative efforts shall fall within the scope of protection of the present invention.
[0014] Those skilled in the art will appreciate that although some embodiments herein include certain features included in other embodiments but not other features, the combination of features from different embodiments is intended to be within the scope of the present invention and to form different embodiments. For example, in the claims, any one of the claimed embodiments may be used in any combination.
[0015] Please refer to the attached Figure 1 Describe the method for constructing the biomedical knowledge graph of the present invention.
[0016] A method for constructing a biomedical knowledge graph, such as Figure 1 As shown, the following steps are included: S1. Data Preprocessing: Classify and grade multi-source biomedical data from literature databases, electronic medical records, genomic data, and medical images, and desensitize sensitive information using differential privacy technology. Specifically, multi-source data integration covers four heterogeneous data sources: literature databases, electronic medical records, genomic data, and medical images, breaking through the limitations of traditional single data sources. At the same time, differential privacy technology is used to desensitize sensitive information and address the risk of biomedical data privacy leakage. Named entity recognition-driven hierarchical tagging is used for text data, and NER-based hierarchical tagging is used to achieve refined data management, thereby solving the problem of lack of fine-grained classification; DICOM format standardization and anonymization processing are implemented for image data, and DICOM format processing is compatible with multimodal image data input, expanding multimodal support capabilities; fine-grained data access control is achieved through attribute-based encryption, and attribute-based encryption (ABE) supports on-demand decryption to meet the differentiated permission requirements of medical institutions and strengthen the privacy protection mechanism.
[0017] S2. Knowledge Extraction and Annotation: A joint model based on BERT-BiLSTM-CRF is used to extract entities, combined with an attention mechanism to identify relationships between entities, and an ontology framework is used for semantic annotation. Specifically, the BERT-BiLSTM-CRF model improves entity extraction accuracy, and the attention mechanism enhances relationship recognition capabilities, thereby optimizing the deficiencies in knowledge extraction. A multi-scale 3D CNN is used to extract lesion area features from medical images, and an image-text cross-modal alignment model is constructed to address the limitations of insufficient joint processing of image data. Image features are mapped to the ontology concept space, and an active learning strategy is used to optimize the selection of annotation samples to improve knowledge extraction efficiency.
[0018] S3. Cross-source alignment and fusion: Calculate entity similarity based on the graph embedding algorithm, and integrate entities and relationships from different data sources through a probabilistic soft alignment model; use the TransEdge algorithm for cross-source relationship alignment, use the conflict resolution algorithm to handle entity attribute contradictions, and dynamically adjust the alignment threshold based on reinforcement learning. Specifically, in terms of relationship alignment, the TransEdge algorithm is used to capture the semantic differences of cross-source relationships and resolve relationship conflicts. In terms of dynamic adjustment, reinforcement learning is used to optimize the alignment threshold to adapt to quality fluctuations of different data sources and improve alignment robustness.
[0019] S4. Quality control system: Build a three-level verification mechanism including rule engine, statistical analysis and expert feedback to dynamically monitor the integrity, consistency and timeliness of the knowledge graph. Specifically, through the three-level verification mechanism rule engine, statistical analysis and expert feedback, ensure the continuous credibility of the knowledge graph, thereby solving the problem of update lag; deploy a redundancy detection module based on subgraph isomorphism, deploy a timeliness decay model to automatically mark expired knowledge, and deploy visualization tools to support manual correction and version backtracking. Specifically, the subgraph isomorphism algorithm identifies duplicate entities, reduces storage redundancy, and solves the entity redundancy problem. The decay model automatically marks expired knowledge to prevent outdated data from misleading applications, thereby solving the problem of update lag.
[0020] Preferably, another embodiment of the present invention further includes: S5. Federalized Graph Update: Implement secure collaborative updates across multiple institutions within a blockchain framework, managing data contribution and usage rights through smart contracts. The blockchain framework enables multi-institutional data collaboration, breaking the limitations of centralized architectures and addressing the challenges of secure multi-institutional collaboration. Use homomorphic encryption to perform cross-institutional entity alignment calculations; design a contribution assessment algorithm based on Shapley values; and verify the authenticity of data sources through zero-knowledge proofs. Specifically, homomorphic encryption ensures data privacy during the cross-institutional entity alignment process, meeting the compliance requirements of medical data sharing. The Shapley value algorithm quantifies institutional contributions, promotes data collaboration enthusiasm, and addresses the problem of insufficient motivation for multi-institutional collaboration. Zero-knowledge proofs verify data sources, prevent forged data from contaminating the graph, and enhance the system's anti-attack capabilities. The cross-institutional entity alignment calculation adopts the Paillier homomorphic encryption algorithm and performs the following operations in the ciphertext state: attribute matching between authorized institutions is achieved through proxy re-encryption technology; encrypted entity vector similarity calculation is accelerated based on SIMD batch processing; a trusted execution environment is deployed to decrypt and verify the alignment results; the Paillier algorithm and SIMD batch processing are used to accelerate the ciphertext similarity calculation to solve the efficiency bottleneck of homomorphic encryption. The trusted execution environment (TEE) is combined with proxy re-encryption to achieve controllable "calculation-decryption-verification" process. Through the collaboration of encryption technology and TEE, man-in-the-middle attacks and data leakage in cross-institutional collaboration are prevented.
[0021] S6. Dynamic evolution mechanism: Based on the time series graph model, the knowledge evolution path is tracked and historical version snapshots with confidence weights are generated. The time series graph model retains historical version snapshots, supports knowledge traceability and evolution analysis, and enhances the interpretability of the knowledge graph.
[0022] This invention integrates multimodal data such as text, genome, and image, supports cross-modal knowledge association, and solves the limitation of single data source of traditional methods. It builds a multi-level privacy protection system through differential privacy, ABE, homomorphic encryption, and zero-knowledge proof to achieve safe and compliant multi-institutional collaboration. The time series graph model and three-level verification mechanism ensure that the knowledge graph is continuously updated and reliable and available. The combination of TransEdge algorithm, reinforcement learning threshold adjustment, subgraph isomorphism detection and other technologies improves the efficiency and accuracy of cross-source data integration. The blockchain federation architecture and cross-chain technology design support seamless access to new data sources and collaborative institutions in the future, and provide high-quality knowledge support for clinical decision-making, drug development and other scenarios through dynamic evolution, secure collaboration and precise quality control.
[0023] This invention systematically addresses the three core challenges of data heterogeneity, privacy security, and dynamic evolution in the construction of biomedical knowledge graphs, providing a reliable infrastructure for precision medicine and translational medicine research.
[0024] Obviously, the above embodiments of the present invention are merely examples for the purpose of clearly illustrating the present invention, and are not intended to limit the embodiments of the present invention. Those skilled in the art will appreciate that other variations or modifications can be made based on the above description. It is not necessary and impossible to enumerate all embodiments here. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention shall be included within the scope of protection of the claims of the present invention.
Claims
1. A method for constructing a biomedical knowledge graph, characterized in that: The following steps are involved: S1. Data preprocessing: Classify and grade multi-source biomedical data from literature databases, electronic medical records, genomic data, and medical images, and desensitize sensitive information using differential privacy technology; S2. Knowledge Extraction and Annotation: Entities are extracted using a joint model based on BERT-BiLSTM-CRF, relationships between entities are identified using an attention mechanism, and semantic annotation is performed using an ontology framework. S3, cross-source alignment and fusion: Calculate entity similarity based on a graph embedding algorithm and integrate entities and relationships from different data sources through a probabilistic soft alignment model; S4. Quality control system: Build a three-level verification mechanism including rule engine, statistical analysis and expert feedback to dynamically monitor the integrity, consistency and timeliness of the knowledge graph.
2. The method for constructing a biomedical knowledge graph according to claim 1, characterized in that: In S1, including: Use named entity recognition-driven hierarchical tagging for text data; implement DICOM format standardization and anonymization for image data; and implement fine-grained data access control through attribute-based encryption.
3. The method for constructing a biomedical knowledge graph according to claim 2, characterized in that: In S2, including: Multi-scale 3D CNN is used to extract lesion area features from medical images, an image-text cross-modal alignment model is constructed, image features are mapped to the ontology concept space, and an active learning strategy is used to optimize the selection of labeled samples.
4. The method for constructing a biomedical knowledge graph according to claim 1, wherein: In S3, this includes: The TransEdge algorithm is used for cross-source relationship alignment, the conflict resolution algorithm is used to handle entity attribute contradictions, and the alignment threshold is dynamically adjusted based on reinforcement learning.
5. The method for constructing a biomedical knowledge graph according to claim 1, wherein: In S4, it includes: Deploy a redundancy detection module based on subgraph isomorphism, deploy a timeliness decay model to automatically mark expired knowledge, and deploy visualization tools to support manual correction and version backtracking.