Biomedical knowledge extraction, fusion and sharing method and system based on federal graph data

By using a biomedical knowledge extraction, fusion, and sharing method based on federated graph data, localized nodes and communication networks are built within medical institutions. This solves the problem of high privacy risks in centralized data processing, enables dynamic updates of the knowledge graph under privacy protection, improves the completeness and clinical applicability of the knowledge graph, adapts to the expansion of medical alliances, and ensures privacy and compliance.

CN120977608BActive Publication Date: 2026-02-03THE FIRST AFFILIATED HOSPITAL OF ARMY MEDICAL UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511507092.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-10-21
Publication Date
2026-02-03
Estimated Expiration
2045-10-21

AI Technical Summary

Technical Problem

Existing methods for updating biomedical knowledge graphs pose a high risk to patient medical data privacy, especially in centralized data processing models where data is easily leaked during transmission and storage. Furthermore, the relaxation of privacy protection procedures in pursuit of efficiency increases the risk of privacy information leakage.

Method used

A biomedical knowledge extraction, fusion, and sharing method using federated graph data is employed. Localized nodes are deployed within medical institutions, and a central knowledge graph is constructed through a communication network architecture. Only graph-level summary data is transmitted, and multimodal diagnosis and treatment data is extracted, aligned, and updated. Combined with professional scoring and geolocation mechanisms, cross-institutional cross-validation and personalized updates are achieved, ensuring dynamic updates of the knowledge graph while protecting privacy.

Benefits of technology

It enables dynamic updates of the knowledge graph while protecting privacy, improves the overall integrity and clinical applicability of the knowledge graph, shortens update delays, improves the efficiency and accuracy of knowledge sharing, adapts to the expansion needs of medical alliances, and ensures the compliance and accuracy of clinical decision-making.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120977608B_ABST
    Figure CN120977608B_ABST
Patent Text Reader

Abstract

The scheme belongs to the field of artificial intelligence and biomedical information technology, and specifically relates to a biomedical knowledge extraction fusion sharing method and system based on federated graph data. The biomedical knowledge extraction fusion sharing method based on federated graph data comprises the following steps: S10: deploying a localized institution node in a participating medical institution, and establishing a communication network architecture between each institution node and a center node; the institution nodes jointly construct a center knowledge graph according to stored biomedical knowledge, and the center node stores the center knowledge graph; S20: the institution node obtains the diagnosis and treatment data of local patients. Through the communication network architecture and multi-modal processing, the scheme solves the problem of high privacy risk in the existing method when processing patient diagnosis and treatment data, improves the accuracy in the expansion process of the center knowledge graph, retains the characteristics of special institution sub-graphs, shortens the update delay of the center knowledge graph, shortens the deployment time of new institution nodes, and adapts to the expansion of medical alliances.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present scheme belongs to the field of artificial intelligence and biomedical information technology, and particularly relates to a biomedical knowledge extraction fusion sharing method and system based on federated graph data. BACKGROUND

[0002] At present, most systems update the biomedical knowledge graph through multi-modal data fusion and knowledge extraction technology: natural language processing is used to analyze electronic medical record texts, computer vision is used to extract pathological image features, bioinformatics is used to analyze genomic sequences, multi-modal data is mapped to unified entities through medical standard terms, and the entities are aligned and the relationships are predicted to integrate into the graph, and through an automatic interface or batch import, new knowledge (such as new diagnosis and treatment guidelines, drug targets) is dynamically identified and updated in real time in the form of triples.

[0003] However, most current methods have significant privacy risks in patient diagnosis and treatment data processing. The existing methods obtain patient diagnosis and treatment data from medical institutions and rely on centralized data processing mode when processing patient diagnosis and treatment data. The multi-modal sensitive data such as electronic medical records, pathological images, and genomic sequences are transmitted across institutions to the centralized platform for unified integration. During the process, the data is separated from the original secure environment and exposed to the transmission link and storage node. Encryption or access control failure can easily lead to leakage. In addition, most current methods simplify the privacy protection process to ensure the efficiency of updating the knowledge graph: such as relaxing data access restrictions, simplifying encryption strategies, or omitting the sensitive information desensitization link. Although this reduces data processing delays and speeds up graph iteration, it allows sensitive information of some patients to be unprotected, which increases the risk of leakage of patient privacy information. SUMMARY

[0004] The purpose of the present scheme is to provide a biomedical knowledge extraction fusion sharing method and system based on federated graph data to solve the problem of high privacy risk in patient diagnosis and treatment data processing of existing methods.

[0005] In order to achieve the above purpose, the present scheme provides a biomedical knowledge extraction fusion sharing method based on federated graph data, comprising the following steps:

[0006] S10: deploying a local institution node in each participating medical institution, establishing a communication network architecture between each institution node and a center node; the institution nodes jointly construct a center knowledge graph according to the stored biomedical knowledge, and the center node stores the center knowledge graph;

[0007] S20: The institution node obtains the diagnosis and treatment data of the local patient, extracts entities and entity relationships in the same modality from the diagnosis and treatment data to construct graph data, aligns each diagnosis and treatment data according to the entity relationship in the graph data, and generates an institution sub-graph. When the entity or entity relationship extracted from the newly added diagnosis and treatment data cannot be matched in the institution sub-graph, the institution node adds the entity or entity relationship related to the new graph data to the new graph data, obtains the entity and entity relationship related to the new graph data in the center knowledge graph as the verification content from the center node, compares the verification content with the new graph data, modifies the institution sub-graph according to the comparison result, and generates a verification application according to the modification content and sends it to the center node;

[0008] S30: After receiving the verification application, the center node sends the verification application to other institution nodes. After receiving the verification application, the institution node compares the content of the verification application with the local institution sub-graph, generates verification feedback information according to the comparison result, and sends it to the center node. The center node summarizes the feedback information and modifies the center knowledge graph according to the summary result;

[0009] S40: The institution node monitors the content updated by the center knowledge graph according to the entity and entity relationship in the local institution sub-graph, and synchronously updates the local institution sub-graph according to the monitoring result.

[0010] The principle and technical effect of the scheme is that: first, the extraction, alignment and update of multi-modal diagnosis and treatment data are all completed locally in each medical institution, and only the highly abstract graph-level summary such as new graph data and verification application is uploaded, and the original electronic medical record, image or genome sequence is no longer transmitted. The center node only makes consensus and summary at the graph level, and does not touch the original sensitive data of the patient, so as to compress the data exposure surface from the cross-institution transmission link and centralized storage to the local institution, thereby directly eliminating the core privacy risk that the centralized processing causes the data to be out of the original safe environment, encryption or access control failure to be easily leaked, and there is no need to relax the desensitization or encryption strategy for efficiency, which realizes the purpose of dynamic updating of the knowledge graph under the premise of protecting privacy. At the same time, the verification process of the abstracted knowledge graph (i.e. the center knowledge graph) still retains the cross-validation of multiple institutions, which guarantees the knowledge quality.

[0011] Secondly, the scheme maps multi-modal diagnosis and treatment data to a unified semantic abstract sub-graph at the institution side, preserving regional characteristics, ethnic customs, local norms, and other distinctive regional imprints while keeping the original data within the domain. After receiving the new semantic abstract, the center node performs logical verification based on the universal knowledge graph and initiates cross-institution cross-validation. The returned results are layered according to their universality and specificity. The consensus without regional restrictions is upgraded to global knowledge and written into the center knowledge graph. The characteristic knowledge that is locally established and does not conflict with universal knowledge is retained in the local area with a regional tag. Conflicting content is returned to manual review. This way, the individual characteristics of each institution sub-graph are preserved while ensuring the objectivity and universality of the center knowledge graph. A federal biomedical knowledge network with universal and characteristic dual tracks is constructed, significantly improving the overall integrity and clinical applicability of the knowledge system.

[0012] Finally, after the center knowledge graph reaches a consensus through the communication network architecture, only the universal knowledge of the consensus is used to update the center knowledge graph, and the institution sub-graph updates locally using special knowledge. This way, doctors at the institution node automatically obtain the latest consensus knowledge when they access the institution sub-graph next time. Institutions that do not have special knowledge or rarely encounter such patients can also update the universal knowledge in a timely manner, eliminating data islands and knowledge lag in the updating process of biomedical knowledge. The original diagnosis and treatment data remains in the local institution throughout the process, and only semantic-level abstracts are transmitted, reducing the risk of privacy leakage. New medical institutions can quickly synchronize universal knowledge through the same communication architecture when they access, eliminating the need for full replication, improving knowledge sharing efficiency and accuracy, and meeting the flexible expansion needs of medical groups.

[0013] In summary, the scheme solves the problem of high privacy risk in patient diagnosis and treatment data processing through communication network architecture and multi-modal processing, improves the accuracy of center knowledge graph expansion, preserves the characteristics of special institution sub-graphs (or supports the expansion of special institution sub-graphs), and shortens the update delay of center knowledge graph, shortens the deployment time of new institution nodes, and adapts to the expansion of medical groups.

[0014] Further, the center node classifies entities according to biomedical knowledge, calculates professional scores of institution nodes in different fields based on entity types, entity quantities, and entity relationships in institution sub-graphs, and adjusts the weight of feedback information based on the professional scores of each institution node.

[0015] First, the entities are classified in the medical ontology library according to the professional field, and the professional score is calculated according to the entity accumulation depth, relationship complexity and historical verification performance of the institution in the specific field, so that the specialized medical institution has a higher decision weight in the knowledge verification of the corresponding field. For example, the tumor hospital has significantly higher speech power than the ordinary clinic in the anti-tumor drug rule update, ensuring compliance with the requirements of clinical diagnosis and treatment specifications. Second, institutions with insufficient professional scores automatically lose voting qualifications in related fields, effectively preventing cross-field malicious tampering, and only high-scoring institutions in the field are selected to participate in verification, significantly reducing the knowledge update cycle and network communication load. Finally, to ensure clinical decision compliance, the scoring weight mechanism requires that changes in knowledge in a specific medical field must be led by authoritative institutions in that field, significantly improving the ability to intercept incorrect knowledge, and allowing institutions to appeal to the score through encrypted credentials, building a medical knowledge governance system that takes into account professionalism and fairness in a distributed architecture.

[0016] Further, after the institution node joins the communication network, it sends a knowledge graph acquisition request containing the field classification label to the center node based on the locally stored institution subgraph, the center node filters and returns the corresponding center knowledge graph subset according to the field characteristics in the request; the institution node aligns the local subgraph with the center knowledge graph subset, identifies structural differences or missing content through entity matching and entity relationship comparison, generates a verification application with field annotation for the identified results and submits it to the center node; then, the center node evaluates the entity coverage, relationship complexity and field professionalism of the institution subgraph, expands the coverage of the center knowledge graph subset based on the evaluation results and sends it to the institution node, and the institution node modifies the local institution subgraph according to the received center knowledge graph subset.

[0017] When an institutional node receives a subset of the central knowledge graph and modifies its local subgraph by combining it with multi-source backup data, it can preload multi-source backup content of commonly used domains and directly use the entity relationship data already stored in the local cache for subsequent knowledge expansion. This reduces repeated query requests to the central node, significantly reduces the communication load of the distributed network (communication network), reduces the synchronization delay between the acquisition of the institutional subgraph and the central knowledge graph, and improves the efficiency of network resource utilization. Meanwhile, the extended recommendations pushed by the central node based on the evaluation results of the institutional sub-graph include domain-specific knowledge of higher-scoring institutions within the communication network (such as local endemic microbial toxin association data obtained by institutions in a certain region). By integrating this precisely matched professional knowledge, institutional nodes can quickly supplement the breadth of entity coverage in the local graph (such as adding 30% more regionally specific entities) and the complexity of relationships (such as increasing the depth of multi-hop relationships by 2 layers), thereby improving their own professional scores in the corresponding fields (such as increasing the score in the field of microbial poisoning from 0.7 to 0.9), forming a positive cycle of "data accumulation - recommendation optimization - score improvement". This continuously promotes the development of institutional nodes' professional capabilities in subdivided medical fields, enabling them to gradually grow from ordinary data participants into authoritative nodes in the field, and strengthening the knowledge specialization and clinical decision support capabilities of the entire communication network.

[0018] Furthermore, when an institutional node joins the communication network, the central node collects its geographical location information and establishes a mapping relationship by combining the institutional node's professional field and the professional score of that professional field. The central node obtains the mapping relationship related to the geographical location information of each institutional node, calculates the correlation between geographical location and professional field based on the professional score, and establishes a connection between geographical location and professional field based on the calculation result. When selecting a verification node, the central node obtains the associated geographical location based on the professional field included in the verification application, and expands the institutional nodes as verification nodes based on the obtained geographical location.

[0019] The central node prioritizes regionally relevant institutions when selecting verification nodes by associating them with the geographical location, professional fields, and professional scores of the associated institutional nodes. This ensures that the verification results are more aligned with local medical practices (e.g., verification of knowledge about fungal poisoning in a certain region is led by local institutions with high scores), significantly improving the matching degree between knowledge and regional clinical needs. The cross-geographical expansion mechanism for verification nodes (e.g., institutions in neighboring provinces serving as auxiliary verification nodes) avoids the limitations of single-region data, promotes cross-verification and complementarity of medical knowledge from different regions (e.g., medical institutions in Southwest China collaborate to improve the diagnosis and treatment rules for wild fungal poisoning), and enhances the comprehensiveness of knowledge verification in the distributed network. After receiving verification feedback that includes regional associations, institutional nodes can supplement local disease data (e.g., associations with specific diseases in ethnic minority areas), promoting the formation of regionally distinctive professional sub-graphies and accelerating the knowledge accumulation of grassroots institutional nodes in specific fields (e.g., county-level hospitals enhance their ability to diagnose and treat diabetic foot complications through local case data). Meanwhile, the deep integration of geographical location and professional fields enables the central node to quickly locate and activate surrounding high-scoring institutional nodes to participate in knowledge updates in response to public health emergencies (such as regional infectious disease outbreaks). This builds a more responsive regional medical knowledge collaboration network (for example, when a food poisoning outbreak occurs in a certain area, nearby poisoning treatment center nodes will be given priority to participate in relevant knowledge verification and push). Ultimately, this achieves a multi-dimensional effect of "precise verification of regional knowledge - enhanced cross-regional collaboration - cultivation of primary care specialist capabilities - improved emergency response efficiency," allowing the distributed medical knowledge network to maintain both professional authority and regional adaptability.

[0020] Furthermore, after the central node summarizes the feedback information from each institutional node, it extracts the summary result record with attached modification confirmation requests, identifies the institutional node that sent the modification confirmation request, and calculates and updates the contribution value of the institutional node based on the summary result; when the central node adjusts the weight of the feedback information, it adjusts the weight according to the difference in contribution value of each institutional node; when the central node expands the subset of the central knowledge graph, it adjusts the expansion range according to the contribution value, and the expansion range includes expansion in terms of the number of entities, types, and entity relationships.

[0021] This solution links the contribution value of institutional nodes to the richness of professional resources in their geographical location, the density of institutional collaboration, and professional scores. This makes institutions with high contribution values ​​more authoritative in their professional fields, thereby improving the accuracy and credibility of knowledge updates. At the same time, this mechanism incentivizes institutions to actively contribute high-quality data (e.g., institutions in scarce areas can achieve higher growth rates by improving data quality), and dynamically adjusts the collaboration density to avoid the monopoly of weight by resource-rich areas, achieving a balanced distribution of resources across regions. In addition, the contribution value calculation combined with geographical location factors makes the knowledge graph more aligned with the medical practice needs of different regions, enhancing regional adaptability and providing accurate data support for hierarchical medical treatment.

[0022] Furthermore, when the central node calculates the contribution value of the institutional nodes, it takes the professional fields involved in the aggregated feedback information as the evaluation field, the geographical location of the institutional node as the evaluation location, obtains the professional scores of other institutional nodes in the same region and evaluation field as the evaluation location, and adjusts the contribution value of the current institutional node by floating up or down based on the average of the scores.

[0023] This solution dynamically adjusts contribution values ​​by matching them with professional scores from institutions in the same field and region, strengthening the collaborative optimization of professional field and regional data. This allows institutions with strong geographical expertise to receive weighted contributions in their respective fields, accelerating knowledge breakthroughs. Simultaneously, this mechanism balances the differences in contribution weights across different regions, preventing institutions in scarce regions from being underestimated due to geographical limitations and promoting knowledge accumulation in long-tail areas such as rare diseases. Furthermore, it guides institutions to focus on their local strengths, promoting targeted cross-regional flow of professional knowledge (e.g., disseminating marine disease prevention knowledge from coastal areas to inland regions), optimizing data coverage in less popular fields within federated learning, and filling gaps in traditional methods.

[0024] Furthermore, the central node divides the domain grid based on geographical location and the professional fields associated with those locations. It then establishes a consensus communication group by acquiring the institutional nodes within the domain grid. These institutional nodes select a group leader node based on their contribution values ​​within the consensus communication group and jointly define the voting rules for data updates within the group. Consensus group members aggregate subgraphs related to their professional fields from their local knowledge graph to generate a domain consensus subgraph. Before submitting a modification confirmation request to the central node, each member in the consensus communication group sends the request to other members in the group for multiple rounds of voting. The group leader node collects feedback from all parties and coordinates conflicts until a consensus is reached. The group leader node then sends the consensus result to the central node, which modifies the central knowledge graph based on the consensus result.

[0025] This solution improves the accuracy and regional adaptability of knowledge updates by dividing the data into domain-regional grids and selecting high-weight institutions. At the same time, it reduces the load on central nodes and the pressure on data transmission by leveraging the local negotiation mechanism of consensus groups. By strengthening the knowledge leadership role of high-authority institutions within consensus groups, a virtuous cycle of "local consensus - global optimization" is formed. This not only ensures the privacy and compliance of medical data, but also improves the efficiency of cross-institutional knowledge collaboration through multi-round voting mechanisms. Ultimately, it achieves the subdivision and optimization of medical specialty fields, making the central knowledge graph updates more in line with the needs of regional medical practice, and ensuring the priority inclusion and efficient integration of cutting-edge knowledge.

[0026] Furthermore, the central node divides the consensus group into basic, advanced, and specialized layers. Each institutional node analyzes the coverage of entity types in its local data within its professional field, as well as the complexity of relationships between entities, to generate personalized knowledge requirements. When expanding the central knowledge graph, the central node prioritizes expanding the entity coverage of basic layer institutional nodes, expands advanced layer institutional nodes by balancing entity coverage with relationship complexity, and expands specialized layer institutional nodes by increasing the proportion of entity types in the professional field and increasing the complexity of relationships between entities.

[0027] This solution divides consensus groups into basic, advanced, and specialized levels and implements differentiated knowledge expansion strategies to ensure that institutions at different levels receive knowledge updates tailored to their diagnostic and treatment capabilities: Basic level institutions prioritize expanding entity coverage, efficiently acquiring universal knowledge of common diseases and enhancing the breadth of primary healthcare services; Advanced level institutions balance entity coverage with relationship complexity, addressing both advanced treatment of common diseases and basic knowledge of rare diseases, meeting comprehensive diagnostic and treatment needs; Specialized level institutions focus on increasing the proportion of entities treating difficult and complex diseases and deepening the understanding of complex diagnostic and treatment relationships, strengthening the ability to accurately diagnose and treat complex diseases. This solution generates personalized knowledge requirements by analyzing local data, avoiding a "one-size-fits-all" approach to knowledge supply and achieving precise matching of medical knowledge across different levels. Simultaneously, the differentiated expansion strategies for different levels help build a collaborative system characterized by "comprehensive coverage of primary hospitals, balanced capabilities of general hospitals, and in-depth specialization of specialized hospitals," improving the efficiency of regional medical resource allocation and knowledge application effectiveness. Ultimately, it provides patients with tiered medical support, ensuring "guaranteed treatment for basic diseases and more precise referrals for complex diseases," forming a virtuous cycle of "demand-driven, tiered supply, and dynamic optimization" for knowledge updates.

[0028] Furthermore, after receiving the verification information, the central node or institutional node parses the entity and entity relationship types to generate a structured request. The central node compares the verification content with the entity coverage and relationship complexity of the knowledge graph, and checks for logical conflicts and redundancy. The institutional node evaluates the adaptability based on local data. Multi-dimensional verification of logical, relational, and regional adaptability is performed. If logical conflicts or regional incompatibility exist, the central node or institutional node immediately returns the verification information with specific reasons. If the verification passes, the central node summarizes the verification results to form a consensus report and updates the node weights.

[0029] This solution parses verification information and generates structured requests, transforming complex knowledge verification requirements into quantifiable entity and relationship check objects, thus improving processing efficiency. Secondly, the central node performs logical conflict and redundancy checks based on the global knowledge graph, ensuring that the verification content conforms to the overall domain specifications and preventing disruptive errors from entering the central knowledge graph. Meanwhile, institutional nodes assess adaptability by combining local clinical data, filtering information that does not match their own clinical capabilities or regional disease characteristics (e.g., automatically filtering unreasonable modifications related to complex surgeries in primary hospitals), enhancing the practicality of knowledge application. Furthermore, multi-dimensional verification based on logic, relationships, and regional adaptability... The system constructs a defense mechanism from three levels: medical logical rigor, completeness of the knowledge system, and applicability to regional practices. This reduces the limitations of single-dimensional judgment (for example, preventing both deviations from evidence-based medicine recommendations and knowledge updates that are "unsuitable" for local contexts). Finally, clear hierarchical processing rules (immediate return of problematic information and updates to weights through information aggregation) form a virtuous cycle of "incentivizing high-quality contributions and filtering invalid information." This not only ensures the authority of the central knowledge graph but also incentivizes institutional nodes to actively participate in verification through weight adjustments. Ultimately, this improves the quality and efficiency of knowledge updates throughout the entire solution, ensuring that the included information is both scientifically rigorous and practically feasible. Attached Figure Description

[0030] Figure 1 This is a flowchart of a biomedical knowledge extraction, fusion, and sharing method based on federated graph data in an embodiment of the present invention. Detailed Implementation

[0031] The following will describe the concept and technical effects of the present invention clearly and completely with reference to embodiments, so as to fully understand the purpose, features and effects of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Other embodiments obtained by those skilled in the art based on the embodiments of the present invention without creative effort are all within the scope of protection of the present invention.

[0032] like Figure 1 As shown, the biomedical knowledge extraction, fusion, and sharing method based on federated graph data includes the following steps:

[0033] S10: Deploy localized institutional nodes within participating medical institutions and establish a communication network architecture between each institutional node and the central node; the institutional nodes jointly construct a central knowledge graph based on the stored biomedical knowledge, and the central node stores the central knowledge graph;

[0034] S20: The institution node obtains the local patient's diagnosis and treatment data, extracts entities and entity relationships of the same modality from the diagnosis and treatment data to construct graph data, and then aligns each diagnosis and treatment data according to the entity relationships in the graph data to generate an institution subgraph. When the entities or entity relationships extracted from the newly added diagnosis and treatment data cannot be matched in the institution subgraph, the institution node takes the entities or entity relationships that cannot be matched as new graph data, adds the entities or entity relationships that have structural relationships with the new graph data to the new graph data, obtains the entities and entity relationships related to the entities in the new graph data from the central knowledge graph as verification content from the central node, compares the verification content with the new graph data, modifies the institution subgraph according to the comparison results, and generates a verification application based on the modification content and sends it to the central node.

[0035] S30: After receiving the verification request, the central node sends the verification request to other institutional nodes; after receiving the verification request, the institutional nodes compare the content of the verification request with their local institutional subgraph, generate verification feedback information based on the comparison results, and send it to the central node; the central node summarizes the feedback information and modifies the central knowledge graph based on the summary results.

[0036] S40: The organization node updates its local organization subgraph synchronously based on the monitoring results, according to the updated content of the entity and entity relationship monitoring center knowledge graph in the local organization subgraph.

[0037] The process includes several key components: First, entity information extraction within the same modality: This involves extracting entity features and inter-entity relationships from large-scale unstructured same-modality data such as medical records and pathological images using natural language processing and image feature extraction to generate same-modality graph data, including protein-protein interactions. Second, cross-modality knowledge construction: This first transforms the graph data into a low-dimensional, dense vectorized representation. Then, leveraging the scalability of the knowledge graph that stores entity associations, it uses inference learning to extract unknown inter-entity relationships from multimodal data, thereby expanding the institutional subgraph, such as associations between genomes and tissue cells. Third, information aggregation of different relationship types: Based on the different relationships between heterogeneous multimodal data, a multi-head spatial self-attention mechanism is used to allocate different weight parameters to aggregate neighbor information of entities in multimodal data, updating the expert knowledge rules contained in the institutional subgraph. Fourth, an indexing mechanism based on the institutional subgraph: This efficiently enhances the in-depth understanding of medical data associations and finds alternative routes. Entities in specific data can be used to query related entities in heterogeneous data, enabling the structural subgraph to provide responses.

[0038] Most current methods can only extract explicit knowledge when constructing knowledge graphs, while medical data such as electronic medical records contain a large amount of implicit knowledge that is difficult to extract. This solution can effectively extract implicit medical knowledge through multimodal fusion. Conventional methods for entity relation extraction are used to obtain common sense knowledge sets from guidelines, textbooks, etc., as training data. A pre-trained model such as GPT-3 is used to fine-tune the production model. By learning the structure and relationships of the central knowledge graph, the hyperparameters of the corresponding analysis model are adjusted to deeply mine implicit knowledge and obtain semantic triples. New nodes and edges are continuously generated on the central knowledge graph, thereby generating an updatable and transferable disease-centered knowledge graph.

[0039] For graph data acquired from heterogeneous multimodal data with different node categories and edge relationships, the institutional nodes employ efficient graph neural network models with classification and regression capabilities to generate graph embeddings for medical rule reasoning. Graph Convolutional Neural Networks (PCNs): Generate embedding representations for small-scale biomedical graph data through Laplacian matrix eigenvalue decomposition; Graph Attention Networks: Process large-scale graph data with complex relationships using a graph attention network with an adjacency node weight allocation mechanism, rationally allocating attention weights to different adjacency nodes to generate graph embedding representations; Dynamic Graph Neural Networks: Based on combining graph neural networks for acquiring spatial information with recurrent neural networks for acquiring temporal information, dynamic graph spatiotemporal information can be obtained from continuous-time biomedical data to obtain dynamic graph embeddings, enabling disease progression prediction.

[0040] A federated learning framework for graph data is built based on a central node (central server), enabling multiple graph data holders to collaboratively train models in a distributed manner without sharing data. Global model parameter sharing: The central node initializes the global model and distributes model parameters to each holder (institutional node), selecting holders for model training based on the importance and type of graph data they possess. Parameter update aggregation: Selected holders construct graph neural network models according to different graph data types, initialize their local models using global model parameters, and then upload the updated local model parameters generated from local model training to the central node for aggregation and global model parameter updates. Privacy-preserving distributed model training: A trade-off is struck between global and local differential privacy based on the confidence level of the central node or data holder. Global differential privacy adds Laplace-distributed noise to the global model parameters and distributes it to holders for local model training. Local differential privacy adds perturbations to the node attributes or edge weights of the held graph data, enabling perturbed model training on the graph data and uploading local model parameters.

[0041] The central node classifies entities by domain based on biomedical knowledge, and calculates the professional scores of institutional nodes in different domains based on the types, number, and relationships between entities in the institutional sub-graph. When summarizing feedback information, the central node obtains the professional scores of institutional nodes in the corresponding domains of the verification applications, and adjusts the weight of feedback information based on the comparison of the professional scores of each institutional node.

[0042] Specifically, the subgraph of the central node defines the mechanism node i as follows: ,in For a set of entity nodes, Let d be the set of relation edges, corresponding to the specialized sub-fields of the central node after classifying entities based on biomedical knowledge. The central node is classified using a domain classification function. Entity Mapped to domain d, forming a subset of domain entities. .

[0043] Professional rating of institution node i in domain d The definition is shown in the following formula (1):

[0044] (1),

[0045] in To account for the diversity of entity types in the domain, Shannon entropy is used for calculation, as shown in the following formula (2):

[0046] (2),

[0047] in For the set of entity types in domain d, This represents the proportion of entity type t in domain d for institution i. This indicator reflects the balance of entity types (e.g., in a certain region, the entity types of institutions in the domain of "bacterial poisoning" cover toxins, bacterial species, symptoms, etc., and the higher the entropy value, the stronger the diversity).

[0048] Standardized value for the number of entities in the domain: ,in Let be the number of entities of institution i in domain d. Normalize by the maximum value to avoid scale bias (e.g., the normalized value of the amount of bacterial poisoning case data handled by an institution in a certain region).

[0049] The relation complexity index is determined by combining graph density and relation type entropy, as shown in the following formula (3):

[0050] (3),

[0051] in For graph density, it is used to reflect the degree of correlation between entities in the subgraph of the domain d of institution i; The denominator is the number of relation edges that actually exist for institution i in the subgraph corresponding to domain d; Under the assumption of "undirected graphs" (i.e., in biomedicine, entity relationships are mostly bidirectional and traceable, such as the semantic equivalence of "disease to symptom" and "symptom to disease"), The maximum number of relationship edges that an entity can theoretically form (i.e., the upper limit of the number of edges when all entities are paired). The entropy of the relation type for the i-th institution node in the d-th biomedical specialty: Let r be the single relation type that the i-th institutional node exists in the d-th domain. , It is the set of all relation types of the i-th institution node in the d-th domain; Let i be the proportion of relation type r of organization i in domain d. For a predefined set of relations in the domain d. For sets All relation types r corresponding to The terms are summed, and then the relation type entropy is obtained by adding them together and passing the negative sign. This indicator quantifies the diversity of relationship types (such as the richness of the types of relationships between toxins and targets, symptoms and diseases in the field of bacterial poisoning).

[0052] In formula (1) Must meet Ensure professional scoring The value range is controllable (usually mapped to the [0,1] interval) to avoid excessive weighting of a certain dimension leading to scoring distortion. Domain-adaptive dynamic weights are employed, and accuracy is optimized through historical validation: Loss Function In the loss function It is a response to The overall reference of the weight vector (which can be represented as) = , (represents transpose) Not independent of Instead of using new weights, the weights of these three dimensions are optimized as a whole set. Here, K is the number of historical validation samples, 1 ≤ k ≤ K. This is the true and authoritative verification result, corresponding to a single sample with index k in the historical verification sample set; This represents the weighted prediction result for a single sample of the k-th rank. For example, in the field of "bacterial poisoning," it can be trained using regional clinical data to... (Diversity) and (Relationship complexity) receives higher weight, highlighting the characteristics of regional knowledge.

[0053] Specifically, for the field d corresponding to the verification application, the professional scores of each institutional node need to be normalized to avoid absolute value bias: ,in To standardize the scoring of organization i in domain d, and ensure .

[0054] Feedback information from the central node to the d-th professional field involved by the i-th institutional node Assign weights And the final verification result of the d-th professional field is generated through weighted aggregation. , As shown in formula (4) below:

[0055] (4),

[0056] in The sharpness parameter controls the degree to which score differences affect the weights: when When the weights are linearly correlated with the standardized scores, when... At that time, the weight of high-scoring institutions is amplified (e.g.) When the scoring is doubled, the weight of the institution becomes quadrupled, which is suitable for scenarios where authoritative nodes need to be highlighted (such as the field of bacterial poisoning in a certain region). It can be dynamically adjusted through domain knowledge, for example, by taking from the domain of emerging infectious diseases. Balancing new data with traditional authority. Let be the feedback vector for institution i (e.g., a 0-1 judgment on the correctness of knowledge, or a confidence score). To ensure that feedback from highly professional rating agencies dominates the aggregated verification results.

[0057] In an embodiment of this scheme, suppose that in the sub-map of institution node A (a local hospital in a certain region): bacterial poisoning (Number of entities), bacterial poisoning Therefore, fungal poisoning The types of mushrooms include 12 wild mushroom species, 8 toxins, and 15 symptoms (mushroom poisoning). Assuming the calculation yields bacterial poisoning (Approaching full entropy); Under the above assumptions, the relation edges include 15 types, such as toxin and target, bacterial species and toxin, symptoms and toxin (e.g., bacterial poisoning). (Graph density 0.72, relational entropy, bacterial poisoning) Therefore, fungal poisoning ; Assumption Bacterial poisoning Bacterial poisoning at node B of the organization. Standardized bacterial poisoning Bacterial poisoning If we take Then bacterial poisoning Bacterial poisoning This means that the feedback weight of institutions in a certain region accounts for 87%, achieving an authoritative response to regional expertise.

[0058] After joining the communication network, an institutional node sends a knowledge graph retrieval request containing domain classification labels to the central node based on its locally stored institutional subgraph. The central node filters and returns the corresponding central knowledge graph subset based on the domain features in the request. The institutional node semantically aligns its local subgraph with the central knowledge graph subset, identifies structural differences or missing content through entity matching and entity relationship comparison, generates a verification application with domain annotations based on the identification results, and submits it to the central node. Subsequently, the central node evaluates the entity coverage, relationship complexity, and domain specialization of the institutional subgraph, expands the coverage of the central knowledge graph subset based on the evaluation results, and sends it to the institutional node. The institutional node then modifies its local institutional subgraph based on the received central knowledge graph subset.

[0059] Specifically, when an institutional node semantically aligns its local sub-graph with the subset returned by the center, it employs named entity recognition and synonym mapping technologies (such as UMLS semantic network) to match entity names, types, and relational predicates one by one. In this embodiment, it is assumed that the institutional node compares the synonym relationships between the *Boletus luridus* entity in its local graph and the *Amanita phalloides* in the center graph, while also comparing whether there are any missing relationships between toxins and hepatocyte damage. After identifying the differences, the institutional node generates a verification request with domain annotations, such as annotating the missing entity attributes in the domains of bacterial poisoning and toxin classification, and attaching a description of the clinical data source for the differences.

[0060] Specifically, the central node selects other institutional sub-graphs from within the communication network that match the corresponding field of the institutional node and have high scores. For example, it selects the bacterial poisoning-related graphs of three top-tier hospitals in a certain region, extracts region-specific content not included in the central graph (such as the local case association between *Agaricus subnigricans* and rhabdomyolysis), and forms a multi-backup recommendation set to send to the institutional node. The institutional node combines the multi-backup content with the central subset to supplement and modify the local graph, such as adding toxin metabolic pathways of locally unique bacterial species. At the same time, it preloads backup data of commonly used fields into the local cache, so that subsequent queries for similar knowledge do not require repeated requests to the central node, reducing the knowledge acquisition latency in the distributed network (communication network) and accurately improving the professional level of the institutional node.

[0061] When an institutional node joins the communication network, the central node collects its geographical location information and establishes a mapping relationship by combining the institutional node's professional field and the professional score of the professional field. The central node obtains the mapping relationship related to the geographical location information of each institutional node, infers the correlation between geographical location and professional field based on the professional score, and establishes a connection between geographical location and professional field based on the inference result. When selecting a verification node, the central node obtains the associated geographical location based on the professional field included in the verification application, and expands the institutional nodes as verification nodes based on the obtained geographical location.

[0062] In this embodiment of the scheme, it is assumed that the central node statistics show that the total professional scores of institutions in the field of "bacterial poisoning" account for 78% of the national total, with a calculated correlation strength of 0.85. The correlation strengths of neighboring regions such as regions B and C are 0.42 and 0.38, respectively, thus mapping "bacterial poisoning" to the geographical locations of regions A, B, and C. When an institution in region A submits a verification request for "classification of phytotoxicant," the central node selects institutions in region A with a correlation strength exceeding 0.7 for "bacterial poisoning" (such as an affiliated hospital of a medical university in region A or a disease control center in region A) as core nodes. Auxiliary nodes are then expanded to regions B and C based on geographical proximity (such as the infectious disease department of a hospital in region B). This ultimately forms a set of verification nodes containing 5 institutions from region A and 2 institutions from region B. The verification nodes are weighted according to the complexity of the related entities and entity relationships involved in "bacterial poisoning," with the comprehensive weight of local institutions in region A reaching 82%, ensuring that regional expertise dominates the verification results. In this way, the "bacterial poisoning" verification node is extended to geographically adjacent regions B and C, avoiding verification failures due to geographical limitations. For example, if a sudden bacterial poisoning event occurs in region A, and the local communication network is congested due to the large number of poisoned people, medical resources from regions B and C can be quickly called upon to participate in the knowledge verification.

[0063] After the central node summarizes the feedback information from each institutional node, it extracts the summary result record with the attached modification confirmation application, identifies the institutional node that sent the modification confirmation application, and calculates and updates the contribution value of the institutional node based on the summary result. When the central node adjusts the weight of the feedback information, it adjusts the weight according to the difference in contribution value of each institutional node. When the central node expands the subset of the central knowledge graph, it adjusts the expansion range according to the contribution value. The expansion range includes expansion in terms of the number of entities, types, and entity relationships.

[0064] Specifically, when the central node calculates the contribution value of the institutional nodes, it takes the professional field involved in the aggregated feedback information as the evaluation field, the geographical location of the institutional node as the evaluation location, obtains the professional scores of other institutional nodes in the same region and evaluation field as the evaluation location, and adjusts the contribution value of the current institutional node by floating up or down based on the average of the scores.

[0065] More specifically, based on the construction of a basic calculation framework for contribution values ​​related to geographical location, the impact of region on contribution values ​​is quantified through three dimensions: professional resource abundance, institutional collaboration density, and professional capability score, forming a basic adjustment model. On this basis, a dual matching mechanism combining field and region is introduced. The contribution value is then recalibrated using the average professional score of institutions in the same field and region, forming a progressive evaluation logic between basic calculation, regional benchmarking, and dynamic adjustment.

[0066] In this embodiment of the solution, it is assumed that knowledge types based on feedback information (such as cardiovascular disease, tumor diagnosis and treatment) are automatically classified, or that institutions manually label their professional fields. The administrative region (e.g., a region in region D, a region in region E) is determined by the institution's registered address, IP address, or terminal GPS information. Only federated nodes belonging to the same assessment field and geographical location as the current institution are included, excluding interference from institutions across fields or regions.

[0067] A sliding window mechanism is used to update the average professional ratings of institutions in the same field and region in real time, avoiding the lag of historical data (e.g., refreshing the average every 7 days). If the current institution's rating is greater than or equal to the regional average, the contribution value is amplified by a positive coefficient (reflecting the institution's authority in the field); if the current institution's rating is less than the regional average, the contribution value is finely adjusted by a negative coefficient (incentivizing improved data quality).

[0068] Assuming the base contribution value weight is calculated using three dimensions and accounts for 70% of the total weight; and the regional benchmarking adjustment weight is derived from comparing the average scores and accounts for 30% of the total weight; the total weight = base weight × (1 + regional benchmarking adjustment coefficient). When the number of institutions in a region is less than 5, the adjustment weight is automatically reduced (e.g., reduced to 15%) to avoid small sample bias; for niche areas such as rare diseases, a scarcity protection coefficient is set to increase the adjustment threshold for scores of institutions in the same region (e.g., allowing score differences within ±0.15 to remain unchanged).

[0069] The central node divides the knowledge graph into domain grids based on geographical location and associated professional fields. It then establishes consensus communication groups by acquiring institutional nodes within these grids. These institutional nodes select a group leader node based on their contribution values ​​and jointly define voting rules for data updates within the group. Members of the consensus group aggregate subgraphs related to their professional fields from their local knowledge graph to generate a domain consensus subgraph. Before submitting a modification confirmation request to the central node, each member of the consensus communication group sends the request to other members within the group for multiple rounds of voting. The group leader node collects feedback from all parties and coordinates conflicts until a consensus is reached. Finally, the group leader node sends the consensus result to the central node, which modifies the central knowledge graph based on the consensus result.

[0070] Specifically, the central node divides the consensus group into basic, advanced, and specialized layers. Each institutional node analyzes the coverage of entity types in its local data within its professional field, as well as the complexity of relationships between entities, to generate personalized knowledge requirements. When expanding the central knowledge graph, the central node prioritizes expanding the entity coverage of basic layer institutional nodes, expands advanced layer institutional nodes by balancing entity coverage with relationship complexity, and expands specialized layer institutional nodes by increasing the proportion of entity types in the professional field and increasing the complexity of relationships between entities.

[0071] More specifically, when a basic-level institutional node identifies a complex case during the diagnosis and treatment process (such as a difficult disease involving collaboration across multiple professional fields, a rare disease classification, or a case where the diagnosis and treatment process requires association with more than 5 entity relationships), it automatically retrieves the professional field data of the specialized-level institutional node, including the proportion of difficult disease entity types in that field (e.g., the lung cancer gene mutation classification coverage rate of a certain oncology hospital reaches 90%) and the complexity of the relationships between entities (e.g., the diagnosis and treatment path for lung cancer brain metastasis at this hospital involves 6 entity associations such as radiotherapy and targeted therapy), and generates a precise referral suggestion containing a list of recommended institutions and a corresponding diagnosis and treatment relationship map (e.g., it is recommended to refer to Oncology Hospital A, whose lung cancer brain metastasis diagnosis and treatment path covers 7 gene mutation types and 3 combination treatment options). After completing diagnosis or referral in the updated domain minimap, the institutional nodes will feed back data such as the accuracy of entity coverage in actual diagnosis (e.g., the accuracy of diagnosis of common diseases) and the efficiency of relationship matching (e.g., the fit between the diagnosis and treatment process of complex cases and the map) to the consensus group. The consensus group leader node will then dynamically adjust the knowledge update focus of each level based on the feedback from the entire group (e.g., if the basic layer institutions in a certain region have insufficient coverage of diabetes entities, increase the priority of expanding universal knowledge in that domain).

[0072] After receiving the verification information, the central node or institutional node parses the entity and entity relationship types and generates a structured request. The central node compares the verification content with the entity coverage and relationship complexity of the knowledge graph, and checks for logical conflicts and redundancy. The institutional node evaluates the adaptability based on local data. Multi-dimensional verification of logical, relational, and regional adaptability is performed. If logical conflicts or regional incompatibility exist, the central node or institutional node immediately returns the verification information with specific reasons. If the verification passes, the central node summarizes the verification results to form a consensus report and updates the node weights.

[0073] The above descriptions are merely embodiments of the present invention, and common knowledge regarding specific structures and characteristics is not elaborated upon here. It should be noted that those skilled in the art can make various modifications and improvements without departing from the structure of the present invention, and these should also be considered within the scope of protection of the present invention. These modifications and improvements will not affect the effectiveness of the present invention or the practicality of the patent. The scope of protection claimed in this application should be determined by the content of its claims, and the specific embodiments described in the specification can be used to interpret the content of the claims.

Claims

1. A biomedical knowledge extraction, fusion, and sharing method based on federated graph data, characterized in that: Includes the following steps: S10: Deploy localized institutional nodes within participating medical institutions and establish a communication network architecture between each institutional node and the central node; Institutional nodes jointly construct a central knowledge graph based on the stored biomedical knowledge, and the central node stores the central knowledge graph; S20: The institution node obtains the local patient's diagnosis and treatment data, extracts entities and entity relationships of the same modality from the diagnosis and treatment data to construct graph data, and then aligns each diagnosis and treatment data according to the entity relationships in the graph data to generate an institution subgraph. When the entities or entity relationships extracted from the newly added diagnosis and treatment data cannot be matched in the institution subgraph, the institution node takes the entities or entity relationships that cannot be matched as new graph data, adds the entities or entity relationships that have structural relationships with the new graph data to the new graph data, obtains the entities and entity relationships related to the entities in the new graph data from the central knowledge graph as verification content from the central node, compares the verification content with the new graph data, modifies the institution subgraph according to the comparison results, and generates a verification application based on the modification content and sends it to the central node. S30: After receiving the verification request, the central node sends the verification request to other institutional nodes; after receiving the verification request, the institutional nodes compare the content of the verification request with their local institutional subgraph, generate verification feedback information based on the comparison results, and send it to the central node; the central node summarizes the feedback information and modifies the central knowledge graph based on the summary results. S40: The organization node updates its local organization subgraph based on the monitoring results of the entity and entity relationship monitoring center knowledge graph in the local organization subgraph. The central node classifies entities by domain based on biomedical knowledge, and calculates the professional scores of institutional nodes in different domains based on the types, number, and relationships between entities in the institutional sub-graph. When summarizing feedback information, the central node obtains the professional scores of institutional nodes in the corresponding domains of the verification applications, and adjusts the weight of feedback information based on the comparison of the professional scores of each institutional node. After joining the communication network, an institutional node sends a knowledge graph retrieval request containing domain classification labels to the central node based on its locally stored institutional subgraph. The central node filters and returns the corresponding central knowledge graph subset based on the domain features in the request. The institutional node semantically aligns its local subgraph with the central knowledge graph subset, identifies structural differences or missing content through entity matching and entity relationship comparison, generates a verification application with domain annotations based on the identification results, and submits it to the central node. Subsequently, the central node evaluates the entity coverage, relationship complexity, and domain specialization of the institutional subgraph, expands the coverage of the central knowledge graph subset based on the evaluation results, and sends it to the institutional node. The institutional node then modifies its local institutional subgraph based on the received central knowledge graph subset.

2. The biomedical knowledge extraction, fusion, and sharing method based on federated graph data according to claim 1, characterized in that: When an institutional node joins the communication network, the central node collects its geographical location information and establishes a mapping relationship by combining the institutional node's professional field and the professional score of the professional field; the central node obtains the mapping relationship related to the geographical location information of each institutional node, infers the correlation between geographical location and professional field based on the professional score, and establishes a connection between geographical location and professional field based on the inference result; When selecting a verification node, obtain the associated geographical location based on the professional field included in the verification application, and expand the organization nodes as verification nodes based on the obtained geographical location.

3. The biomedical knowledge extraction, fusion, and sharing method based on federated graph data according to claim 2, characterized in that: After the central node summarizes the feedback information from each institutional node, it extracts the summary result record with the attached modification confirmation application, identifies the institutional node that sent the modification confirmation application, calculates and updates the contribution value of the institutional node based on the summary result; when the central node adjusts the weight of the feedback information, it adjusts the weight according to the difference in the contribution value of each institutional node. When a central node expands a subset of the central knowledge graph, the expansion range is adjusted based on the contribution value. The expansion range includes expansions in terms of the number of entities, types, and entity relationships.

4. The biomedical knowledge extraction, fusion, and sharing method based on federated graph data according to claim 3, characterized in that: When the central node calculates the contribution value of the institutional nodes, it takes the professional field involved in the aggregated feedback information as the evaluation field, the geographical location of the institutional node as the evaluation location, obtains the professional scores of other institutional nodes in the same region and evaluation field as the evaluation location, and adjusts the contribution value of the current institutional node by floating up or down based on the average of the scores.

5. The biomedical knowledge extraction, fusion, and sharing method based on federated graph data according to claim 3, characterized in that: The central node divides the knowledge graph into domain grids based on geographical location and associated professional fields. It then establishes consensus communication groups by acquiring institutional nodes within these grids. These institutional nodes select a group leader node based on their contribution values ​​and jointly define voting rules for data updates within the group. Members of the consensus group aggregate subgraphs related to their professional fields from their local knowledge graph to generate a domain consensus subgraph. Before submitting a modification confirmation request to the central node, each member of the consensus communication group sends the request to other members for multiple rounds of voting. The group leader node collects feedback from all parties and coordinates conflicts until a consensus is reached. Finally, the group leader node sends the consensus result to the central node, which modifies the central knowledge graph based on the consensus result.

6. The biomedical knowledge extraction, fusion, and sharing method based on federated graph data according to claim 5, characterized in that: The central node divides the consensus group into basic, advanced, and specialized layers. Each institutional node analyzes the coverage of entity types in its local data within its professional field, as well as the complexity of relationships between entities, to generate personalized knowledge requirements. When expanding the central knowledge graph, the central node prioritizes expanding the entity coverage of basic layer institutional nodes, expands advanced layer institutional nodes by balancing entity coverage with relationship complexity, and expands specialized layer institutional nodes by increasing the proportion of entity types in the professional field and increasing the complexity of relationships between entities.

7. The biomedical knowledge extraction, fusion, and sharing method based on federated graph data according to claim 6, characterized in that: After receiving the verification information, the central node or institutional node parses the entity and entity relationship types and generates a structured request. The central node compares the verification content with the entity coverage and relationship complexity of the knowledge graph, and checks for logical conflicts and redundancy. The institutional node evaluates the adaptability based on local data. Multi-dimensional verification of logical, relational, and regional adaptability is performed. If logical conflicts or regional incompatibility exist, the central node or institutional node immediately returns the verification information with specific reasons. If the verification passes, the central node summarizes the verification results to form a consensus report and updates the node weights.

8. A biomedical knowledge extraction, fusion, and sharing system based on federated graph data, characterized in that: The biomedical knowledge extraction, fusion and sharing method based on federated graph data as described in any one of claims 1-7 was used.

Citation Information

Patent Citations

  • Cross-mechanism medical knowledge graph representation learning method and system

    CN116821375A

  • Simulation system intelligent decision-making method and system based on knowledge graph and federated learning

    CN120087794A