Cross-institution patient follow-up visit data integration method and system based on knowledge graph

By constructing a five-layer architecture model using knowledge graphs, the problem of heterogeneity in patient follow-up data across institutions was solved, achieving unified standardization and data integration, thereby improving the continuity of medical services and the scientific nature of decision-making.

CN121456831AActive Publication Date: 2026-02-03川北医学院附属医院
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
CN202610004487.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-01-05
Publication Date
2026-02-03
Estimated Expiration
2046-01-05

AI Technical Summary

Technical Problem

Due to heterogeneity and lack of unified standards, cross-institutional patient follow-up data is difficult to share and collaboratively apply, affecting the continuity of medical services and the scientific nature of decision-making.

Method used

A knowledge graph-based multi-source heterogeneous data integration method is adopted. A five-layer architecture model is constructed to standardize and associate data, including heterogeneous data adaptation, entity parsing, semantic enhancement, dynamic association and attribute mapping, to establish direct and indirect relationships between entities and form a unified patient health record.

Benefits of technology

It enables centralized acquisition and efficient collaboration of cross-institutional data, improves the accuracy and completeness of data integration, and provides reliable support for medical decision-making.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121456831A_ABST
    Figure CN121456831A_ABST
Patent Text Reader

Abstract

The invention discloses a cross-institution patient follow-up visit data integration method and system based on a knowledge graph, and belongs to the field of medical data processing. The method comprises the following steps: firstly, acquiring various types of follow-up data such as cross-institution electronic medical records, inspection reports and follow-up records, then constructing a multi-source heterogeneous data integration model containing five layers of architectures such as a heterogeneous data adaptation layer and an entity analysis layer, and carrying out standardized processing on the data after training; then extracting entities, relationships and feature information in the data, establishing direct and indirect association relationships between the entities, and fusing to form a unified patient health record; and finally, receiving the patient identification information, matching and organizing and presenting corresponding data according to preset logic. According to the scheme, the differences of the cross-mechanism data in formats, fields and semantics are effectively eliminated, centralized acquisition of the data is realized, the accuracy and integrity of data integration are improved, and reliable support is provided for medical decision and cross-mechanism medical cooperation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of medical data processing, and in particular to a method and system for integrating cross-institutional patient follow-up data based on knowledge graphs. Background Technology

[0002] In the healthcare field, cross-institutional diagnosis and treatment and long-term follow-up have become important models for improving the quality of medical services. An increasing number of patients are transferring between different medical institutions to receive diagnosis, treatment, and follow-up services due to their medical needs. This process generates a large amount of data related to patient follow-up, encompassing various types such as electronic medical records, examination reports, follow-up records, rehabilitation training trajectory data, patient self-management logs, and cross-institutional referral handover vouchers. Currently, each medical institution collects and stores this data according to its own business processes and management needs, forming its own independent data management system. The industry has gradually recognized the importance of cross-institutional data integration. Some institutions are attempting to promote data sharing by establishing data transmission interfaces and formulating basic data collection standards. Some preliminary data integration technologies have also begun to be applied in local scenarios. However, overall, data remains fragmented, with multiple sources of data stored and maintained independently. The systematic nature and completeness of data integration need to be improved.

[0003] With the deepening of cross-institutional medical collaboration, existing data management and integration models have gradually revealed significant technical bottlenecks. Due to inconsistent information system construction standards across different medical institutions, and the lack of unified specifications for data storage formats, field definitions, and semantic representations, cross-institutional patient follow-up data exhibits significant heterogeneity, making it difficult to directly interoperate and collaboratively apply data from different sources. Simultaneously, existing technologies fail to effectively extract entity, relational, and feature information from the data, making it impossible to establish direct and indirect relationships between entities. This leaves patient data scattered across different institutions in an isolated state, hindering the formation of unified and complete patient health records. Furthermore, the lack of a dedicated integration model for cross-institutional, multi-source, heterogeneous follow-up data results in a lack of standardized data processing procedures. When users need to retrieve cross-institutional patient follow-up data, they cannot quickly and centrally obtain the required information, impacting not only the continuity and efficiency of medical services but also hindering the scientific nature of medical decision-making, failing to meet the needs of cross-institutional medical collaboration for complete and unified patient data. Summary of the Invention

[0004] The purpose of this invention is to overcome the shortcomings of the prior art and provide a method and system for cross-institutional patient follow-up data integration based on knowledge graphs.

[0005] The objective of this invention is achieved through the following technical solution: A knowledge graph-based method for integrating cross-institutional patient follow-up data is provided, comprising the following steps: S1. Obtain patient follow-up data across institutions, including electronic medical records, examination reports, follow-up records, rehabilitation training trajectory data, patient self-management log data, and cross-institutional referral handover voucher data; S2. Construct a multi-source heterogeneous data integration model. This model has a five-layer architecture, including a heterogeneous data adaptation layer, an entity parsing layer, a semantic enhancement layer, a dynamic association layer, and an attribute mapping layer. After training the model, the acquired patient follow-up data is standardized based on the model to unify the data format, field definition, and semantic expression. S3. By combining semantic mapping technology with the dynamic association layer rules of the model, entity information, relationship information and feature information are extracted from the standardized data, and direct and indirect association relationships between entities are established. The associations are then integrated to form a unified patient health record. S4. Receive patient identification information, match the corresponding patient follow-up data based on the unified patient health record and model attribute mapping layer rules, and organize and present the data according to the preset logic to achieve centralized acquisition of cross-institutional data.

[0006] Furthermore, step S1 includes: S1.1. Determine the source institutions and data transmission protocols for cross-institutional patient follow-up data; S1.2. Collect patient follow-up data from the source institution according to the established data transmission protocol. The collection methods include real-time interface connection, timed batch import, and offline file transfer. S1.3. Perform integrity and validity checks on the collected patient follow-up data. Integrity checks include checking for missing core fields and verifying the integrity of data records. Validity checks include checking for data format compliance and verifying logical consistency. S1.4. Retain the follow-up data of patients who pass the verification, and temporarily store the follow-up data of patients who fail the verification after marking the abnormality type, for subsequent supplementary collection or correction processing.

[0007] Furthermore, step S2 includes: S2.1. Construct a multi-source heterogeneous data integration model, configure each layer to be connected through adaptation components, encoding modules and feature enhancement components, specify the heterogeneous data adaptation layer to process different types of follow-up data, configure the entity parsing layer to include entity feature extraction and classification sub-modules, set the semantic enhancement layer to enhance semantic features through multi-layer neural networks, configure the dynamic association layer to include association rules, derivation and strength calculation units, and specify the attribute mapping layer to be responsible for entity attribute definition and matching; S2.2. Annotated cross-institutional multi-source heterogeneous medical data is used to construct training and validation sets. A joint loss function is constructed, which consists of heterogeneous data adaptation loss function, entity matching loss function and semantic association loss function. The initial learning rate, iteration batch and model convergence threshold are set. The parameters of each layer are adjusted alternately through the backpropagation algorithm to optimize the adaptation accuracy, parsing accuracy and association matching degree in turn until the overall adaptation accuracy of the validation set is stable. S2.3. Based on the model structure, standardization rules are formulated, and data standardization is completed through format adaptation, entity encoding, semantic enhancement and attribute matching.

[0008] Furthermore, step S3 includes: S3.1. Extract entity information, relation information and feature information from the standardized patient follow-up data. Entity information includes the unified code and attribute value of each entity. Relationship information includes the direct association type and indirect association clue between entities. Feature information includes time-series features, text features and structured features. S3.2. Based on the dynamic association layer rules of the multi-source heterogeneous data integration model, the improved cosine similarity algorithm is used to calculate the semantic similarity of entities in the patient follow-up data of different institutions. Combined with the dynamic weight of the association strength calculation unit, a similarity threshold is set as the association judgment condition to establish the direct association relationship between entities. S3.3. Through the indirect association derivation unit of the dynamic association layer, based on the attribute association clues, temporal association clues and event association clues between entities, the indirect association relationship between entities is deduced, the indirect association strength is calculated and compared with the preset strength threshold, and the valid indirect association relationship is retained. S3.4. Based on the established direct and effective indirect relationships, the follow-up data of patients from different institutions and of different types corresponding to the same patient are linked and integrated. The integration process maintains the consistency of entity unified coding, attribute information and relationship. According to the classification of the entity parsing layer and the structure of the attribute mapping layer, the data is organized into a unified patient health record.

[0009] Furthermore, step S4 includes: S4.1. Receive patient identification information from the data retrieval request. The patient identification information includes the patient's unique number, identity information, and entity unique code. S4.2. Based on the dynamic association layer rules of the patient identification information and the multi-source heterogeneous data integration model, the corresponding patient follow-up related data are traversed and matched in the unified patient health record. First, the directly related data corresponding to the core patient entity is matched, and then all related data are matched based on the indirect association relationship. S4.3. Classify and organize the matched patient follow-up data according to time order, data type and correlation strength. Data types include electronic medical records, examination reports, follow-up records, rehabilitation training tracks, self-management logs and referral handover vouchers; S4.4. Present patient follow-up data in the sorted order, simultaneously display the direct and indirect relationships between data, and support centralized acquisition and linked viewing of data across institutions.

[0010] Furthermore, in step S2.1, a modular encapsulation design is adopted to construct each layer of the multi-source heterogeneous data integration model. The heterogeneous data adaptation layer and the entity parsing layer are configured to achieve bidirectional data transmission through a data interface adaptation component. The entity parsing layer and the semantic enhancement layer are configured to be connected through an entity vector encoding module. The semantic enhancement layer and the dynamic association layer are configured to establish a mapping through a semantic feature enhancement component. The dynamic association layer and the attribute mapping layer are configured to be associated through a dynamic weight allocation module. It is stipulated that the components of each layer communicate through a standardized data interface.

[0011] Furthermore, in step S2.3, when performing format adaptation processing on the cross-institutional referral handover voucher data, the fixed and variable fields in the voucher are parsed through the structured data adaptation component, and the field mapping relationship is completed according to the requirements of the entity attribute definition unit, converting the unstructured voucher remarks information into a standardized text format; the temporal features of the rehabilitation training trajectory data are extracted through the sliding window submodule of the temporal data adaptation component, and the text information of the patient self-management log data is processed through the word segmentation and stop word filtering submodule of the text data adaptation component.

[0012] Furthermore, in step S3.2, the improved cosine similarity algorithm introduces an entity attribute weight factor and a time decay factor. The entity attribute weight factor is determined by the attribute type matching unit of the attribute mapping layer, and the time decay factor is dynamically adjusted according to the timestamp differences between rehabilitation training trajectory data and follow-up records, thereby improving the accuracy of association matching of data across different time spans.

[0013] Furthermore, in step S3.3, the indirect association derivation unit adopts a multi-step reasoning mechanism. First, it establishes initial indirect association clues based on the referral events in the cross-institutional referral handover voucher data. Then, it combines the behavioral characteristics in the patient self-management log data and the progress characteristics in the rehabilitation training trajectory data to gradually deduce the multi-level indirect association relationships between entities. The rationality of each step of the reasoning result is verified by the semantic feature fusion module of the semantic enhancement layer.

[0014] A cross-institutional patient follow-up data integration system based on knowledge graph is provided. The system includes a data acquisition module, a data processing module, a data association module, and a data retrieval module. The data acquisition module is used to acquire patient follow-up data from multiple institutions, including electronic medical records, examination reports, follow-up records, rehabilitation training trajectory data, patient self-management log data, and cross-institutional referral handover voucher data. The data processing module is used to construct a multi-source heterogeneous data integration model and standardize the patient follow-up data. The model is set as a five-layer architecture, with each layer connected through adapter components. The model training unit completes training based on the training set, validation set, and joint loss function. The data association module is used to extract data information and establish associations through semantic mapping technology combined with model rules, and integrate them to form a unified patient health record. The data retrieval module is used to receive patient identification information, match and organize the corresponding patient follow-up data based on the health record and model rules, and present the data.

[0015] The beneficial effects of this invention are: (1) Based on the five-layer architecture and standardized processing of multi-source heterogeneous data integration, the differences in format, fields and semantics of cross-institutional follow-up data are eliminated, a unified patient health record is formed, and cross-institutional data is centrally acquired; (2) By leveraging data integrity and validity verification mechanisms and entity association rules, the accuracy and integrity of data integration can be improved, providing reliable data support for medical decision-making and follow-up management; (3) The entity association design and structured presentation of knowledge graphs break down cross-institutional data barriers, simplify the data retrieval process, and support efficient cross-institutional medical collaboration. Attached Figure Description

[0016] Figure 1 A flowchart illustrating the steps of a cross-institutional patient follow-up data integration method based on knowledge graphs; Figure 2 A flowchart illustrating the steps of a knowledge graph-based method for integrating cross-institutional patient follow-up data is provided for this embodiment. Figure 3 This is a schematic diagram of the knowledge graph entity and attribute architecture provided for an embodiment. Figure 4 This is a schematic diagram of the multi-source heterogeneous data integration model architecture provided for an embodiment. Detailed Implementation

[0017] The technical solution of the present invention will be clearly and completely described below with reference to the embodiments. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0018] Example 1 See Figure 1 This paper presents a knowledge graph-based method for integrating cross-institutional patient follow-up data, which includes the following steps: S1. Obtain patient follow-up data across institutions, including electronic medical records, examination reports, follow-up records, rehabilitation training trajectory data, patient self-management log data, and cross-institutional referral handover voucher data; S2. Construct a multi-source heterogeneous data integration model. This model has a five-layer architecture, including a heterogeneous data adaptation layer, an entity parsing layer, a semantic enhancement layer, a dynamic association layer, and an attribute mapping layer. After training the model, the acquired patient follow-up data is standardized based on the model to unify the data format, field definition, and semantic expression. S3. By combining semantic mapping technology with the dynamic association layer rules of the model, entity information, relationship information and feature information are extracted from the standardized data, and direct and indirect association relationships between entities are established. The associations are then integrated to form a unified patient health record. S4. Receive patient identification information, match the corresponding patient follow-up data based on the unified patient health record and model attribute mapping layer rules, and organize and present the data according to the preset logic to achieve centralized acquisition of cross-institutional data.

[0019] Step S1 includes: S1.1. Determine the source institutions and data transmission protocols for cross-institutional patient follow-up data; S1.2. Collect patient follow-up data from the source institution according to the established data transmission protocol. The collection methods include real-time interface connection, timed batch import, and offline file transfer. S1.3. Perform integrity and validity checks on the collected patient follow-up data. Integrity checks include checking for missing core fields and verifying the integrity of data records. Validity checks include checking for data format compliance and verifying logical consistency. S1.4. Retain the follow-up data of patients who pass the verification, and temporarily store the follow-up data of patients who fail the verification after marking the abnormality type, for subsequent supplementary collection or correction processing.

[0020] Step S2 includes: S2.1. Construct a multi-source heterogeneous data integration model, configure each layer to be connected through adaptation components, encoding modules and feature enhancement components, specify the heterogeneous data adaptation layer to process different types of follow-up data, configure the entity parsing layer to include entity feature extraction and classification sub-modules, set the semantic enhancement layer to enhance semantic features through multi-layer neural networks, configure the dynamic association layer to include association rules, derivation and strength calculation units, and specify the attribute mapping layer to be responsible for entity attribute definition and matching; S2.2. Annotated cross-institutional multi-source heterogeneous medical data is used to construct training and validation sets. A joint loss function is constructed, which consists of heterogeneous data adaptation loss function, entity matching loss function and semantic association loss function. The initial learning rate, iteration batch and model convergence threshold are set. The parameters of each layer are adjusted alternately through the backpropagation algorithm to optimize the adaptation accuracy, parsing accuracy and association matching degree in turn until the overall adaptation accuracy of the validation set is stable. S2.3. Based on the model structure, standardization rules are formulated, and data standardization is completed through format adaptation, entity encoding, semantic enhancement and attribute matching.

[0021] Step S3 includes: S3.1. Extract entity information, relation information and feature information from the standardized patient follow-up data. Entity information includes the unified code and attribute value of each entity. Relationship information includes the direct association type and indirect association clue between entities. Feature information includes time-series features, text features and structured features. S3.2. Based on the dynamic association layer rules of the multi-source heterogeneous data integration model, the improved cosine similarity algorithm is used to calculate the semantic similarity of entities in the patient follow-up data of different institutions. Combined with the dynamic weight of the association strength calculation unit, a similarity threshold is set as the association judgment condition to establish the direct association relationship between entities. S3.3. Through the indirect association derivation unit of the dynamic association layer, based on the attribute association clues, temporal association clues and event association clues between entities, the indirect association relationship between entities is deduced, the indirect association strength is calculated and compared with the preset strength threshold, and the valid indirect association relationship is retained. S3.4. Based on the established direct and effective indirect relationships, the follow-up data of patients from different institutions and of different types corresponding to the same patient are linked and integrated. The integration process maintains the consistency of entity unified coding, attribute information and relationship. According to the classification of the entity parsing layer and the structure of the attribute mapping layer, the data is organized into a unified patient health record.

[0022] Step S4 includes: S4.1. Receive patient identification information from the data retrieval request. The patient identification information includes the patient's unique number, identity information, and entity unique code. S4.2. Based on the dynamic association layer rules of the patient identification information and the multi-source heterogeneous data integration model, the corresponding patient follow-up related data are traversed and matched in the unified patient health record. First, the directly related data corresponding to the core patient entity is matched, and then all related data are matched based on the indirect association relationship. S4.3. Classify and organize the matched patient follow-up data according to time order, data type and correlation strength. Data types include electronic medical records, examination reports, follow-up records, rehabilitation training tracks, self-management logs and referral handover vouchers; S4.4. Present patient follow-up data in the sorted order, simultaneously display the direct and indirect relationships between data, and support centralized acquisition and linked viewing of data across institutions.

[0023] In step S2.1, a modular encapsulation design is used to construct each layer of the multi-source heterogeneous data integration model. The heterogeneous data adaptation layer and the entity parsing layer are configured to achieve bidirectional data transmission through a data interface adaptation component. The entity parsing layer and the semantic enhancement layer are configured to be connected through an entity vector encoding module. The semantic enhancement layer and the dynamic association layer are configured to establish a mapping through a semantic feature enhancement component. The dynamic association layer and the attribute mapping layer are configured to be associated through a dynamic weight allocation module. It is stipulated that the components of each layer communicate through a standardized data interface.

[0024] In step S2.3, when performing format adaptation processing on the cross-institutional referral handover voucher data, the fixed and variable fields in the voucher are parsed through the structured data adaptation component, and the field mapping relationship is completed according to the requirements of the entity attribute definition unit. The unstructured voucher remarks information is converted into a standardized text format. The temporal features of the rehabilitation training trajectory data are extracted through the sliding window submodule of the temporal data adaptation component, and the text information of the patient self-management log data is processed through the word segmentation and stop word filtering submodule of the text data adaptation component.

[0025] In step S3.2, the improved cosine similarity algorithm introduces an entity attribute weight factor and a time decay factor. The entity attribute weight factor is determined by the attribute type matching unit of the attribute mapping layer, and the time decay factor is dynamically adjusted according to the timestamp differences between rehabilitation training trajectory data and follow-up records, thereby improving the accuracy of association matching of data across different time spans.

[0026] In step S3.3, the indirect association derivation unit adopts a multi-step reasoning mechanism. First, it establishes initial indirect association clues based on the referral events in the cross-institutional referral handover voucher data. Then, it combines the behavioral characteristics in the patient self-management log data and the progress characteristics of the rehabilitation training trajectory data to gradually deduce the multi-level indirect association relationships between entities. The rationality of each step of the reasoning result is verified by the semantic feature fusion module of the semantic enhancement layer.

[0027] In some embodiments, a knowledge graph-based cross-institutional patient follow-up data integration system includes: a data acquisition module, a data processing module, a data association module, and a data retrieval module. The data acquisition module is used to acquire patient follow-up data from multiple institutions, including electronic medical records, examination reports, follow-up records, rehabilitation training trajectory data, patient self-management log data, and cross-institutional referral handover voucher data. The data processing module is used to construct a multi-source heterogeneous data integration model and standardize the patient follow-up data. The model is set as a five-layer architecture, with each layer connected through adapter components. The model training unit completes training based on the training set, validation set, and joint loss function. The data association module is used to extract data information and establish associations through semantic mapping technology combined with model rules, and integrate them to form a unified patient health record. The data retrieval module is used to receive patient identification information, match and organize the corresponding patient follow-up data based on the health record and model rules, and present the data.

[0028] Example 2 This embodiment discloses a method and system for integrating cross-institutional patient follow-up data based on knowledge graphs. Through a series of steps such as standardization processing and correlation fusion, it achieves efficient integration and centralized acquisition of cross-institutional patient follow-up data, such as... Figure 2 As shown, the specific steps are as follows: S1. Data Acquisition: Data acquisition is a fundamental step in the integration of patient follow-up data across institutions. Collecting complete and effective relevant data from different sources provides data support for subsequent processing.

[0029] S1.1 Determine the data source and transmission protocol: First, it is necessary to clarify the scope of the source institutions for cross-institutional patient follow-up data. These source institutions cover different types of healthcare service providers, and their data storage and management models vary. Based on this, a corresponding data transmission protocol should be determined for each source institution. This protocol must meet the security, stability, and compatibility requirements for cross-institutional data transmission, ensuring smooth data transfer between different institutions' systems and avoiding data loss or transmission failures.

[0030] S1.2 Collect relevant data on patient follow-up across institutions: Patient follow-up data is collected from various source institutions according to the established data transmission protocol. The collection methods include real-time interface integration, scheduled batch import, and offline file transfer. The appropriate collection method can be selected based on the data volume, real-time requirements, and system capacity of different institutions.

[0031] Real-time interface integration is suitable for scenarios requiring timely data acquisition. By establishing a dedicated interface between the source organization's system and the data integration system, real-time synchronous data transmission is achieved, ensuring that data can be quickly collected into the integration system after it is generated. Scheduled batch import is suitable for situations with large data volumes and low real-time requirements. Data files are periodically exported from the source organization according to preset time intervals and then imported into the integration system, avoiding excessive pressure on the system due to real-time transmission of large amounts of data. Offline file transfer is suitable for some organizations that cannot achieve real-time interface integration or scheduled import. Data files are obtained through mobile storage devices or other means and then manually imported into the integration system.

[0032] S1.3 Data Validation Processing: The collected patient follow-up data were verified for completeness and validity to ensure that the data met the requirements for subsequent processing.

[0033] Completeness verification includes core field missing checks and data record integrity verification. Core field missing checks refer to checking each data entry for missing key core fields. These core fields are essential elements representing patient follow-up information and cannot be missing. Data record integrity verification checks whether the data records are complete and free from defects, breaks, or other issues, ensuring that each data entry fully reflects the information related to a follow-up visit.

[0034] Validation includes data format compliance checks and logical consistency verification. Data format compliance checks verify whether the data format conforms to preset standard format requirements, such as whether date formats, encoding formats, and numerical formats are standardized. Logical consistency verification checks whether there are logical contradictions between data, such as whether the follow-up time sequence for the same patient is reasonable, or whether the test results are consistent with the diagnostic conclusions.

[0035] S1.4 Verification Result Processing: Retain patient follow-up data that passes validation and incorporate it into subsequent data processing workflows. For patient follow-up data that fails validation, clearly label the anomaly type, including missing core fields, incorrect data format, logical contradictions, incomplete records, etc., and then temporarily store the anomaly-labeled data.

[0036] The temporarily stored abnormal data will be used for subsequent supplementary collection or correction processing. Corresponding processing measures will be taken for different types of abnormalities: if a core field is missing, a supplementary collection request can be initiated to the source institution to obtain the missing field information; if the data format is incorrect or there is a logical contradiction, the source institution can be notified to correct the data, or the data can be automatically corrected in the integration system according to the preset correction rules (provided that the data accuracy is ensured); if the record is incomplete and cannot be supplemented or corrected, this part of the data will be archived separately and will not be included in the subsequent integration process.

[0037] S2. Model Building and Data Standardization: Model building and data standardization are crucial steps in achieving cross-institutional data integration. By constructing a multi-source heterogeneous data integration model, the collected valid data is standardized to eliminate data differences and lay the foundation for data association and fusion.

[0038] S2.1 Constructing a multi-source heterogeneous data integration model: A multi-source heterogeneous data integration model is constructed. This model has a five-layer architecture, including a heterogeneous data adaptation layer, an entity parsing layer, a semantic enhancement layer, a dynamic association layer, and an attribute mapping layer. Each layer adopts a modular encapsulation design to ensure the independence and scalability of each layer's functions.

[0039] The layers are connected and interact with each other through specific components: the heterogeneous data adaptation layer and the entity parsing layer achieve bidirectional data transmission through the data interface adaptation component; the entity parsing layer and the semantic enhancement layer are connected through the entity vector encoding module; the semantic enhancement layer and the dynamic association layer establish a mapping through the semantic feature enhancement component; the dynamic association layer and the attribute mapping layer are associated through the dynamic weight allocation module; at the same time, it is stipulated that the components of each layer communicate through standardized data interfaces to ensure the consistency and compatibility of data transmission.

[0040] The heterogeneous data adaptation layer handles different types of follow-up data and includes three sub-components: a text data adaptation component, a time-series data adaptation component, and a structured data adaptation component. The text data adaptation component includes word segmentation and stop word filtering sub-modules for processing patient self-management log data; the time-series data adaptation component includes a sliding window sub-module for processing rehabilitation training trajectory data; and the structured data adaptation component includes fixed field and variable field parsing sub-modules for processing structured or semi-structured data such as cross-institutional referral handover vouchers, electronic medical records, examination reports, and follow-up records.

[0041] The entity parsing layer includes entity feature extraction and classification sub-modules, and sets up four sub-units: patient entity parsing unit, medical institution entity parsing unit, medical data entity parsing unit, and medical event entity parsing unit. These are used to extract and classify patient entities, medical institution entities, medical data entities, and medical event entities, respectively. Each entity parsing unit extracts the key features of the corresponding entity through the feature extraction sub-module, and then classifies the entity through the classification sub-module. Finally, a unique and unified code is assigned to each entity.

[0042] The semantic enhancement layer strengthens semantic features through a multi-layer neural network and consists of three sub-modules: a context semantic encoding module, a cross-domain semantic transfer module, and a semantic feature fusion module. The context semantic encoding module converts text data into vector form for easier subsequent processing; the cross-domain semantic transfer module enables semantic adaptation between different institutions and data types, eliminating semantic differences; and the semantic feature fusion module fuses multi-dimensional semantic features and verifies the rationality of subsequent inference results.

[0043] The dynamic association layer includes association rules, derivation and strength calculation units, and sets up three sub-units: association rule unit, indirect association derivation unit, and association strength calculation unit. The association rule unit is used to formulate the judgment rules for direct association between entities; the indirect association derivation unit is used to deduce the indirect association relationship between entities through multi-step reasoning; the association strength calculation unit includes a dynamic weight adjustment sub-module, which is used to calculate the strength of the association relationship between entities.

[0044] The attribute mapping layer is responsible for defining and matching entity attributes. It has three sub-units: entity attribute definition unit, attribute type matching unit, and attribute value standardization unit. The entity attribute definition unit is used to define the attribute information of various entities and store it in the form of key-value pairs. The attribute type matching unit is used to determine the matching relationship between different entity attributes and to clarify the weight factor of entity attributes. The attribute value standardization unit is used to unify the format of attribute values ​​of various entities and ensure the consistency of attribute information.

[0045] S2.2 Training a multi-source heterogeneous data integration model: Training and validation sets are constructed using labeled, multi-source, heterogeneous medical data from various institutions. The labeling process must ensure the accuracy and consistency of the data annotations, including information such as entity categories, entity attributes, and relationships between entities. The training set is used to learn the model parameters, while the validation set is used to verify the model's training effectiveness. The ratio of training to validation sets is determined based on the total amount and distribution of data, ensuring that both sets comprehensively cover medical data from different types and institutions.

[0046] A joint loss function is constructed, consisting of a heterogeneous data adaptation loss function, an entity matching loss function, and a semantic association loss function. The heterogeneous data adaptation loss function measures the model's error during data adaptation, the entity matching loss function measures the model's error during entity recognition and matching, and the semantic association loss function measures the model's error during semantic association establishment. By using the joint loss function, the training effects of each stage of the model can be comprehensively considered, ensuring the optimization of the overall model performance.

[0047] Training parameters such as initial learning rate, iteration batch size, and model convergence threshold are set. The initial learning rate controls the step size of model parameter updates, the iteration batch size determines the number of training iterations, and the model convergence threshold determines whether the model has completed training. The parameters of each layer are alternately adjusted using the backpropagation algorithm. In each iteration, the weight parameters of each layer are adjusted in reverse based on the error calculated by the joint loss function, thereby optimizing the fitting accuracy, parsing accuracy, and association matching degree in sequence.

[0048] During training, after each iteration, the model's performance is evaluated using the validation set, and the overall fit accuracy of the validation set is calculated. If the overall fit accuracy does not reach the model convergence threshold, the next iteration of training continues; if the overall fit accuracy reaches or exceeds the model convergence threshold and remains stable across multiple consecutive iterations without significant improvement, the model training is considered complete, training is stopped, and the trained model parameters are saved.

[0049] S2.3 Data Standardization Processing: Based on the structure of the trained multi-source heterogeneous data integration model, standardized rules are formulated. These rules cover multiple aspects such as data format, field definition, and semantic expression, ensuring that the standardized data can meet the requirements of subsequent association and fusion.

[0050] In accordance with standardized rules, the patient follow-up data that has passed the verification is standardized. The process includes four steps: format adaptation, entity encoding, semantic enhancement, and attribute matching.

[0051] The format adaptation process is handled by the heterogeneous data adaptation layer, which employs corresponding adaptation methods for different data types: For cross-institutional referral and handover voucher data, the structured data adaptation component parses the fixed and variable fields in the voucher, completes the field mapping relationships according to the requirements of entity attribute definition units, and transforms unstructured voucher remarks into a standardized text format; for rehabilitation training trajectory data, the sliding window submodule of the time-series data adaptation component extracts time-series features and organizes the data into a unified time-series data format; for patient self-management log data, the word segmentation and stop word filtering submodule of the text data adaptation component processes the text information, removes invalid information, and unifies the text format; for electronic medical records, examination reports, follow-up records, and other data, the structured data adaptation component performs format regularization to ensure that field arrangement, data type, etc., meet standardization requirements.

[0052] The entity coding process is completed by the entity parsing layer, which performs entity recognition and extraction on various types of data after format adaptation. Entity features are extracted through the feature extraction submodule of each entity parsing unit, and then the entities are classified through the classification submodule. Finally, a unique unified code is assigned to each entity. The code includes an organization identification segment and an internal number segment to ensure the uniqueness and traceability of the entity.

[0053] The semantic enhancement step is completed by the semantic enhancement layer, which converts the text data in the entity-encoded data into vector form through the context semantic encoding module, uses the cross-domain semantic transfer module to eliminate semantic differences between data from different institutions, and then uses the semantic feature fusion module to fuse multi-dimensional semantic features to enhance the semantic expression of the data and improve the data's recognizability.

[0054] The attribute matching process is completed by the attribute mapping layer. Based on the attribute information defined by the entity attribute definition unit, the attribute type matching unit establishes the matching relationship between entity attributes in the data and standard attributes, clarifies the attribute weight factors, and then the attribute value standardization unit unifies attribute values ​​of different formats into a standard format to ensure the consistency of attribute information.

[0055] S3. Data Association and Fusion: Data association and fusion is a key step in achieving cross-institutional data integration. By extracting key information from the data and establishing relationships between entities, patient follow-up data scattered across different institutions can be integrated to form a unified patient health record.

[0056] S3.1 Extract key information from the data: Information was extracted from standardized patient follow-up data, including three categories: entity information, relationship information, and feature information.

[0057] Entity information includes a unified code and attribute values ​​for each entity. The unified code is a unique code assigned by the entity parsing layer, and the attribute values ​​are the standardized attribute information after the attribute mapping layer. The attribute values ​​differ for different types of entities: the attribute values ​​for patient entities include basic information such as age, gender, and medical history; the attribute values ​​for medical institution entities include information such as institution type, region (excluding specific location names), and service area; the attribute values ​​for medical data entities include information such as data generation time, data content, and data type; and the attribute values ​​for medical event entities include information such as event occurrence time, event type, and participating entities.

[0058] Relationship information includes direct association types and indirect association clues between entities. Direct association types refer to the explicit relationships between entities, such as the "ownership" relationship between a patient and an electronic medical record, or the "generation" relationship between a medical institution and medical data. Indirect association clues refer to information that can indirectly reflect the relationships between entities, such as the association clue between behavioral characteristics in a patient's self-management log and the progress characteristics of rehabilitation training trajectory data, or the association clue between referral events in an inter-institutional referral handover certificate and different institutions.

[0059] Feature information includes time-series features, text features, and structured features. Time-series features refer to time-related features contained in the data, such as the time series features of rehabilitation training trajectory data and the time interval features of follow-up records. Text features refer to semantic features and keyword features contained in text-type data. Structured features refer to field features and data type features contained in structured data.

[0060] S3.2 Establish direct relationships between entities: Based on the dynamic association layer rules of the multi-source heterogeneous data integration model, an improved cosine similarity algorithm is used to calculate the semantic similarity of entities in patient follow-up data from different institutions. The improved cosine similarity algorithm introduces entity attribute weighting factors and time-decay factors. The entity attribute weighting factor is determined by the attribute type matching unit of the attribute mapping layer, assigning different weights according to the importance of different attributes to entity association. The time-decay factor is dynamically adjusted based on the timestamp differences between rehabilitation training trajectory data and follow-up records; the larger the timestamp difference, the smaller the time-decay factor, thereby reducing the impact of data with large time spans on association matching and improving the accuracy of association matching for data with different time spans.

[0061] By combining the dynamic weights of the association strength calculation unit, a similarity threshold is set as the association determination condition. The dynamic weights are adjusted according to factors such as the type of data and the credibility of the source institution, and the similarity threshold is determined based on the validation results during the model training process to ensure that effective and invalid associations can be accurately distinguished.

[0062] The calculated semantic similarity of entities is compared with a similarity threshold. If the semantic similarity is greater than or equal to the similarity threshold, it is determined that there is a direct relationship between the two entities, and the relationship type and strength are recorded. If the semantic similarity is less than the similarity threshold, it is determined that there is no direct relationship between the two entities, and no relationship is recorded.

[0063] S3.3 Establish indirect relationships between entities: The indirect association derivation unit of the dynamic association layer uses a multi-step reasoning mechanism to deduce the indirect association relationships between entities. First, initial indirect association clues are established based on referral events in the cross-institutional referral handover voucher data. Referral events involve multiple entities such as patients, transferring institutions, and receiving institutions. Through the association relationships between these entities, the basic framework of indirect associations is initially constructed.

[0064] By combining behavioral characteristics from patient self-management logs with progress characteristics from rehabilitation training trajectory data, multi-level indirect relationships between entities can be gradually deduced. For example, rehabilitation training behaviors recorded in patient self-management logs can be linked to corresponding training progress in rehabilitation training trajectory data, which in turn can be linked to corresponding medical advice, and further linked to entities such as the medical institutions or medical personnel who provided the advice; or, based on the time series of patient follow-up records from different institutions, the trend of changes in the patient's condition can be linked to the trend of changes in the condition, which in turn can be linked to corresponding examination reports, diagnostic conclusions, and other data.

[0065] During each step of the reasoning process, the semantic feature fusion module of the semantic enhancement layer verifies the rationality of the reasoning result. If the reasoning result conforms to semantic logic and business rules, the next step of reasoning continues. If the reasoning result has logical contradictions or does not conform to business rules, the derivation of that indirect correlation clue is terminated, and reasoning anomaly information is recorded.

[0066] The strength of indirect associations is calculated, taking into account factors such as the association hierarchy between entities and the degree of feature matching. The calculated indirect association strength is compared with a preset strength threshold. If the indirect association strength is greater than or equal to the preset strength threshold, the indirect association is retained, and the association type, association strength, and reasoning path are recorded. If the indirect association strength is less than the preset strength threshold, the indirect association is discarded and no association record is made.

[0067] S3.4 Integration and fusion to form a unified patient health record: Based on the established direct and effective indirect relationships, follow-up data from different institutions and patient types corresponding to the same patient are linked and merged. During the fusion process, the consistency of entity coding, attribute information, and relationships is strictly maintained to ensure that the merged data does not conflict or contradict.

[0068] Based on the classification of the entity parsing layer and the structure of the attribute mapping layer, the fused data is organized and structured to form a unified patient health record. The patient health record is centered on the patient entity and stored in a structured manner according to data type, time sequence, and other dimensions. Data types include electronic medical records, examination reports, follow-up records, rehabilitation training tracks, self-management logs, and referral handover vouchers. Each data type contains corresponding specific data content, associated entity information, and relationship information.

[0069] During the integration process, if discrepancies are found in the same type of data from different institutions, priority should be determined by considering factors such as the data generation time, the credibility of the source institution, and the completeness and accuracy of the data. Data with higher priority should be selected as the primary data, while other discrepancies should be retained as supplementary information to ensure the comprehensiveness and accuracy of patient health records.

[0070] S4. Data Retrieval and Presentation: Data retrieval and presentation is the final step in integrating patient follow-up data across institutions. It aims to provide users with a convenient and efficient way to obtain data and support centralized viewing and correlation analysis of data across institutions.

[0071] S4.1 receives patient identification information: The system receives data retrieval requests initiated by users. These requests contain patient identification information, including a unique patient number, identity verification information, and a unified entity code. This identification information is crucial for matching relevant patient data and ensuring accurate location of the target patient.

[0072] The system verifies the legality of the received patient identification information, including whether the format of the identification information meets the requirements and whether there are any invalid characters. If the identification information passes the verification, it proceeds to the subsequent data matching stage; if the identification information fails the verification, an error message is returned to the user, requiring the user to provide valid patient identification information again.

[0073] S4.2 Match the corresponding patient follow-up data: Based on the dynamic association layer rules of the patient identification information and the multi-source heterogeneous data integration model, the system traverses and matches corresponding patient follow-up data within the unified patient health record. The matching process employs a hierarchical matching strategy, first matching the directly associated data corresponding to the core patient entity. This directly associated data includes electronic medical records, examination reports, follow-up records, and other data directly associated with the patient entity, which are quickly located using unified entity coding.

[0074] Then, based on the indirect relationship, expand the matching of all related data. According to the established indirect relationship, traverse and deduce the data corresponding to other entities associated with the core patient entity, such as data related to medical institutions, medical events, rehabilitation training plans, etc., to ensure that the matched data fully covers the patient's follow-up information in various institutions.

[0075] During the matching process, the matching path and correlation strength of the data are recorded to facilitate subsequent data organization and presentation. If no corresponding patient follow-up data is found, a message indicating that the data does not exist is returned to the user.

[0076] S4.3 Categorize and organize the matched data: The matched patient follow-up data is categorized and organized according to time sequence, data type, and correlation strength. When organized by time sequence, the data is sorted from earliest to latest based on the timestamp of its generation, allowing users to easily view the time evolution of patient follow-up data. When organized by data type, the data is divided into electronic medical records, examination reports, follow-up records, rehabilitation training tracks, self-management logs, and referral handover vouchers, with data within each category then arranged chronologically. When organized by correlation strength, the data is sorted from highest to lowest correlation strength; data with higher correlation strength is more relevant to the patient's core needs and is presented to the user first.

[0077] During the classification and organization process, duplicate and invalid data are removed. Duplicate data refers to data with completely identical content or consistent information, while invalid data refers to data that is found to contain errors or has no practical significance after verification, ensuring that the organized data is concise and effective.

[0078] S4.4 presents patient follow-up data: Patient follow-up data is presented in a categorized and organized order, using a structured presentation method that clearly displays the basic information, core content, and related information of the data. Basic information includes the data generating institution, generation time, and data type; core content includes specific field information and values ​​(excluding specific privacy-sensitive values); related information includes the direct and indirect relationships between this data and other data, visually demonstrating the relationship paths and strength.

[0079] Users can click on related information to view other related data, supporting cross-organizational data viewing and enabling data traceability and collaborative analysis. The system also provides a data export function, allowing users to export the presented data into standard format files for later use.

[0080] In some embodiments, to implement the above-described method for integrating cross-institutional patient follow-up data based on knowledge graphs, a corresponding cross-institutional patient follow-up data integration system is deployed. This system includes a data acquisition module, a data processing module, a data association module, and a data retrieval module. The modules cooperate with each other to jointly complete the integration and retrieval of cross-institutional patient follow-up data.

[0081] The data acquisition module is used to acquire patient follow-up data from across institutions. This data includes electronic medical records, examination reports, follow-up records, rehabilitation training trajectory data, patient self-management log data, and cross-institutional referral handover voucher data. This module comprises a source institution docking unit, a data collection unit, a data verification unit, and an anomaly data processing unit. The source institution docking unit identifies the data source institution and data transmission protocol, establishing a communication connection with the source institution. The data collection unit collects data according to the established transmission protocol and collection method. The data verification unit verifies the integrity and validity of the collected data. The anomaly data processing unit handles anomaly data that fails verification.

[0082] The data processing module is used to construct a multi-source heterogeneous data integration model and standardize patient follow-up data. The model is set as a five-layer architecture, with each layer connected through adaptation components. This module includes a model building unit, a model training unit, and a data standardization unit. The model building unit is used to construct the five-layer architecture of the multi-source heterogeneous data integration model and the sub-components of each layer, defining the connection relationships and data interaction methods of each layer. The model training unit is used to complete model training based on the training set, validation set, and joint loss function, and optimize model parameters. The data standardization unit is used to formulate standardization rules based on the model structure, and perform format adaptation, entity encoding, semantic enhancement, and attribute matching processing on the validated data.

[0083] The data association module uses semantic mapping technology combined with model rules to extract data information and establish relationships, ultimately forming a unified patient health record. This module includes an information extraction unit, a direct association establishment unit, an indirect association establishment unit, and a data fusion unit. The information extraction unit extracts entity, relationship, and feature information from the standardized data. The direct association establishment unit uses an improved cosine similarity algorithm to establish direct relationships between entities. The indirect association establishment unit establishes indirect relationships between entities through a multi-step reasoning mechanism. The data fusion unit, based on the established relationships, fuses related data from the same patient to form a unified patient health record.

[0084] The data retrieval module receives patient identification information, matches and organizes corresponding patient follow-up data based on health records and model rules, and presents the data. This module includes an identification information receiving unit, a data matching unit, a data organization unit, and a data presentation unit. The identification information receiving unit receives user-input patient identification information and verifies its validity; the data matching unit matches corresponding patient follow-up data from the patient's health record; the data organization unit categorizes and organizes the matched data according to preset rules; and the data presentation unit presents the data in the organized order, supporting associated viewing and export of the data.

[0085] The knowledge graph-based cross-institutional patient follow-up data integration method and system provided in this embodiment achieves standardized processing of follow-up data from different institutions and of different types by constructing a multi-source heterogeneous data integration model. This eliminates differences in data format, field definitions, and semantic representations, ensuring data consistency and compatibility. By establishing direct and indirect relationships between entities, patient follow-up data scattered across different institutions are effectively linked and merged to form a unified patient health record, achieving centralized acquisition of cross-institutional data.

[0086] During data processing, a multi-level verification and optimization mechanism ensures the integrity, validity, and accuracy of the data, providing reliable data support for subsequent medical decision-making and follow-up management. Simultaneously, the system adopts a modular design, with clearly defined functions and collaborative operation among modules, improving data processing efficiency and flexibility. It supports rapid data retrieval and cross-referencing, providing users with a convenient data experience. Furthermore, the solution enhances its stability and scalability through a clear model architecture, standardized training processes, and standardized processing steps, meeting the cross-institutional data integration needs in different scenarios. The centralized acquisition of cross-institutional data is directly achieved through standardized processing and correlation fusion, breaking down data barriers between different institutions. This allows users to easily access complete patient follow-up data from various institutions, providing strong support for cross-institutional medical collaboration.

[0087] Example 3 In some specific implementations, refer to Figure 3 The knowledge graph architecture in this embodiment includes an entity layer and an attribute layer. The two layers support each other and, combined with the multi-source heterogeneous data integration process, realize the systematic integration and efficient utilization of patient follow-up data across institutions. The entire implementation process revolves around the functional implementation and technical connection of the two-layer architecture.

[0088] The entity layer encompasses four types of entities: patient entities, medical institution entities, medical data entities, and medical event entities. The construction and extraction of each entity rely on the entity parsing layer of the multi-source heterogeneous data integration model. During the data acquisition phase, electronic medical records, examination reports, follow-up records, rehabilitation training trajectory data, patient self-management log data, and cross-institutional referral handover voucher data are collected from different institutions through real-time interface integration, scheduled batch imports, or offline file transfers. After completeness and validity verification, qualified data is retained for entity construction. The entity parsing layer extracts key features of each entity through a feature extraction submodule, and then completes entity classification through a classification submodule, assigning a unique and unified code to each entity to ensure its uniqueness and traceability. Specifically, medical data entities correspond to the six types of follow-up-related data collected, with each data entry serving as an instance of a medical data entity, associated with the corresponding generating institution (medical institution entity) and the associated patient (patient entity). Medical event entities cover specific medical behaviors such as referrals, examinations, and follow-ups, associating with participating patient entities, medical institution entities, and corresponding medical data entities.

[0089] The attribute layer includes direct relationships, indirect associations, relational attributes, and entity attributes, while also covering key attribute information such as data type and generation time. Its construction and improvement are integrated throughout the entire process of data standardization and association fusion. Entity attributes are defined and matched by the attribute mapping layer. The attributes of patient entities include basic information such as age, gender, and medical history; the attributes of medical institution entities include information such as institution type and service scope; the attributes of medical data entities include information such as data type, generation time, and data content; and the attributes of medical event entities include information such as event type, occurrence time, and participating entities. All attributes are processed through attribute value standardization units, with a unified format and expression standard. The establishment of direct and indirect relationships is dominated by a dynamic association layer. Direct relationships are determined based on the semantic similarity of entities calculated using an improved cosine similarity algorithm. The algorithm introduces entity attribute weighting factors and time-series decay factors. The entity attribute weighting factors are determined by the attribute type matching unit of the attribute mapping layer, while the time-series decay factor is dynamically adjusted according to data timestamp differences to ensure the accuracy of entity associations across different time spans and with varying attribute importance. Indirect relationships are derived through a multi-step reasoning mechanism. Starting with referral events in cross-institutional referral handover data as initial clues, and combining behavioral characteristics from patient self-management logs with progress characteristics from rehabilitation training trajectories, multi-level indirect relationships are gradually constructed. The rationality of each inference result is verified by the semantic feature fusion module of the semantic enhancement layer, ultimately retaining valid indirect relationships that meet the required strength. Relationship attributes describe the characteristics of direct and indirect relationships, including association strength, association type, and reasoning path, providing a basis for assessing the credibility of data associations.

[0090] The implementation of the entity and attribute layers relies on the collaborative work of a five-layer architecture of a multi-source heterogeneous data integration model. During model construction, a modular encapsulation design is adopted. The heterogeneous data adaptation layer processes different types of data through dedicated components, providing a data source with a unified format for the entity layer. The entity parsing layer directly supports entity extraction and classification. The semantic enhancement layer strengthens the semantic expression of entities and features through multi-layer neural networks, providing semantic support for establishing relationships in the attribute layer. The dynamic association layer focuses on constructing and calculating the strength of direct and indirect relationships in the attribute layer. The attribute mapping layer is responsible for defining, matching, and standardizing entity attributes and relationship attributes. During model training, labeled cross-institutional multi-source heterogeneous medical data is used to construct training and validation sets. A joint loss function is used to optimize adaptation accuracy, parsing accuracy, and association matching degree, ensuring that the model can accurately support the construction of the entity and attribute layers.

[0091] In the data association and fusion stage, based on the unified coding of the entity layer and the association relationship of the attribute layer, data from different institutions and of different types corresponding to the same patient are integrated. Following the classification of the entity parsing layer and the structure of the attribute mapping layer, a unified patient health record is formed. Within the health record, various entities at the entity layer form a network of connections through direct and indirect relationships at the attribute layer. Attributes such as data type and generation time serve as criteria for retrieval and sorting, supporting efficient data retrieval. Upon receiving patient identification information (including the patient's unique number, identity information, and unified entity coding), the health record is traversed based on the association rules of the attribute layer. First, directly related data corresponding to the patient entity is matched. Then, the matching range is expanded through indirect associations. The results are categorized and presented according to time order, data type, and association strength, simultaneously displaying the direct and indirect relationships between data, supporting centralized acquisition and associated viewing of cross-institutional data.

[0092] This architecture clarifies the data subjects through the entity layer and standardizes entity characteristics and relationships through the attribute layer. Combined with the model's standardized processing and association fusion capabilities, it effectively solves the problems of heterogeneous and loosely related data across institutions. The structured design of entities and attributes makes data integration more logical, and the clear division of various entities and attributes improves the efficiency of data retrieval and utilization, providing reliable data support for cross-institutional medical collaboration while ensuring the standardization of the data integration process and the accuracy of the results.

[0093] Example 4 In some specific implementations, refer to Figure 4 This embodiment constructs a multi-source heterogeneous data integration model with a five-layer architecture, and combines standardized processing and correlation fusion processes to achieve systematic integration and centralized retrieval of patient follow-up data across institutions. During the data acquisition phase, the source institutions and data transmission protocols for cross-institutional patient follow-up data were first identified. Following these protocols, methods such as real-time interface integration, scheduled batch import, or offline file transfer were used to collect electronic medical records, examination reports, follow-up records, rehabilitation training trajectory data, patient self-management logs, and cross-institutional referral handover documents. After collection, the data underwent integrity and validity verification. Integrity verification included checking for missing core fields and verifying the completeness of data records. Validity verification included checking for data format compliance and logical consistency. Data that passed verification was retained for subsequent processing, while data that failed was marked as abnormal and temporarily stored for later supplementary collection or correction.

[0094] In the model building phase, a modular encapsulation design is adopted to build a five-layer architecture. Each layer is connected through adapter components, coding modules, and feature enhancement components, and all components are required to communicate through standardized data interfaces.

[0095] The heterogeneous data adaptation layer is equipped with three dedicated components: a text data adaptation component containing word segmentation and stop word filtering submodules, specifically for processing patient self-management log data; a time-series data adaptation component containing a sliding window submodule, responsible for processing rehabilitation training trajectory data; and a structured data adaptation component containing fixed field and variable field parsing submodules, used to process cross-institutional referral handover voucher data, electronic medical records, examination reports, and follow-up records, achieving overall data format adaptation and unifying text, sequence, and structured formats. The entity parsing layer contains four parsing units: the patient entity parsing unit, the medical institution entity parsing unit, the medical data entity parsing unit, and the medical event entity parsing unit. Each unit is equipped with an entity feature extraction and classification sub-module, which outputs patient entities, medical institution entities, medical data entities, and referral / examination / follow-up event entities, respectively, and assigns a unique and unified code to each entity to complete the entity recognition and coding work. The semantic enhancement layer consists of three core modules: the context semantic encoding module, which converts text into vectors; the cross-domain semantic transfer module, which adapts semantic data from different institutions; and the semantic feature fusion module, which is responsible for multi-dimensional feature fusion and subsequent inference result verification. The semantic features are enhanced through multi-layer neural networks to improve feature recognition. The dynamic association layer comprises three functional units: the association rule unit uses an improved cosine similarity algorithm to determine direct associations; the indirect association derivation unit performs multi-step reasoning based on attributes, time series, and event clues; and the association strength calculation unit is responsible for calculating the association strength and performing threshold filtering to establish effective direct and indirect association relationships. The attribute mapping layer is set up with three units: the entity attribute definition unit defines entity attributes in the form of key-value pairs, the attribute type matching unit determines the entity attribute weight factor, and the attribute value standardization unit unifies the attribute value format, thus completing the binding between entities and relation attributes.

[0096] During model training, labeled, multi-source, heterogeneous medical data from various institutions were used to construct training and validation sets. The labels included entity categories, attributes, and relationships between entities. A joint loss function was constructed, consisting of a heterogeneous data adaptation loss function, an entity matching loss function, and a semantic association loss function. An initial learning rate, iteration batch size, and model convergence threshold were set. The parameters of each layer were alternately adjusted using the backpropagation algorithm to sequentially optimize adaptation accuracy, parsing accuracy, and association matching degree until the overall adaptation accuracy of the validation set stabilized. Model training was then completed, and the parameters were saved. Standardization rules were established based on the model structure, and the validated data underwent standardization processing: the heterogeneous data adaptation layer performed format adaptation and assigned data to corresponding components for processing according to data type; the entity parsing layer extracted and encoded entities from the adapted data, assigning unique and unified codes; the semantic enhancement layer converted text data into vectors, eliminating cross-institutional semantic differences and fusing multi-dimensional features; and the attribute mapping layer completed entity attribute definition, matching, and standardization, ultimately unifying the data format, field definitions, and semantic representations.

[0097] In the data association and fusion stage, entity information, relationship information, and feature information are first extracted from the standardized data. Entity information includes the unified code and attribute values ​​of each entity; relationship information includes direct association types and indirect association clues between entities; and feature information covers temporal features, textual features, and structured features. Based on dynamic association layer rules, the association rule unit uses an improved cosine similarity algorithm that introduces entity attribute weight factors and temporal decay factors to calculate the semantic similarity of entities in data from different institutions. Combined with the dynamic weights of the association strength calculation unit and a preset similarity threshold, direct association relationships between entities are established. The indirect association derivation unit uses referral events in cross-institutional referral handover voucher data as initial clues, and combines the behavioral characteristics of patient self-management logs and the progress characteristics of rehabilitation training trajectories to conduct multi-step reasoning to deduce indirect association relationships. The rationality of each step of reasoning results is verified by the semantic feature fusion module of the semantic enhancement layer, the indirect association strength is calculated and compared with a preset threshold, and valid indirect association relationships are retained. Based on the established direct and effective indirect relationships, data from different institutions and of different types corresponding to the same patient are linked and integrated. The integration process maintains the consistency of entity coding, attribute information and relationships. According to the classification of the entity parsing layer and the structure of the attribute mapping layer, a unified patient health record is formed.

[0098] In the data retrieval and presentation phase, the system receives patient identification information from user data retrieval requests, including the patient's unique ID, identity verification information, and entity unified code. After validating the identification information, it iterates through and matches data in the unified patient health record based on model attribute mapping layer rules and dynamic association layer rules. First, it matches directly related data corresponding to core patient entities, then expands the matching based on indirect relationships to all relevant data. The matching results are categorized and organized by time order, data type, and association strength. Data types include electronic medical records, examination reports, follow-up records, rehabilitation training tracks, self-management logs, and referral handover vouchers. The data is presented in a structured manner according to the organized order, synchronously displaying direct and indirect relationships between data, supporting user-linked viewing, and enabling centralized access to cross-institutional data.

[0099] This implementation process effectively solves the problems of adapting and semantically unifying multi-source heterogeneous data through the collaborative work of the model's five-layer architecture. Standardized processing eliminates differences in format and field definitions among data from different institutions, and the unified patient health record formed by association and fusion breaks down cross-institutional data barriers. The modular design of each layer's components improves the model's scalability and maintainability, the application of entity coding and association rules ensures the accuracy of data integration, and the structured retrieval and presentation method provides convenience for data use, making it suitable for various cross-institutional medical data integration scenarios.

[0100] The above description is merely a preferred embodiment of the present invention. It should be understood that the present invention is not limited to the forms disclosed herein and should not be construed as excluding other embodiments. It can be used in various other combinations, modifications, and environments, and can be altered within the scope of the concept described herein through the above teachings or related technologies or knowledge. Modifications and variations made by those skilled in the art that do not depart from the spirit and scope of the present invention should be within the protection scope of the appended claims.

Claims

1. A method for integrating cross-institutional patient follow-up data based on knowledge graphs, characterized in that, Includes the following steps: S1. Obtain patient follow-up data across institutions, including electronic medical records, examination reports, follow-up records, rehabilitation training trajectory data, patient self-management log data, and cross-institutional referral handover voucher data; S2. Construct a multi-source heterogeneous data integration model. This model has a five-layer architecture, including a heterogeneous data adaptation layer, an entity parsing layer, a semantic enhancement layer, a dynamic association layer, and an attribute mapping layer. After training the model, the acquired patient follow-up data is standardized based on the model to unify the data format, field definition, and semantic expression. S3. By combining semantic mapping technology with the dynamic association layer rules of the model, entity information, relationship information and feature information are extracted from the standardized data, and direct and indirect association relationships between entities are established. The associations are then integrated to form a unified patient health record. S4. Receive patient identification information, match the corresponding patient follow-up data based on the unified patient health record and model attribute mapping layer rules, and organize and present the data according to the preset logic to achieve centralized acquisition of cross-institutional data.

2. The method according to claim 1, characterized in that, Step S1 includes: S1.

1. Determine the source institutions and data transmission protocols for cross-institutional patient follow-up data; S1.

2. Collect patient follow-up data from the source institution according to the established data transmission protocol. The collection methods include real-time interface connection, timed batch import, and offline file transfer. S1.

3. Perform integrity and validity checks on the collected patient follow-up data. Integrity checks include checking for missing core fields and verifying the integrity of data records. Validity checks include checking for data format compliance and verifying logical consistency. S1.

4. Retain the follow-up data of patients who pass the verification, and temporarily store the follow-up data of patients who fail the verification after marking the abnormality type, for subsequent supplementary collection or correction processing.

3. The method according to claim 1, characterized in that, Step S2 includes: S2.

1. Construct a multi-source heterogeneous data integration model, configure each layer to be connected through adaptation components, encoding modules and feature enhancement components, specify the heterogeneous data adaptation layer to process different types of follow-up data, configure the entity parsing layer to include entity feature extraction and classification sub-modules, set the semantic enhancement layer to enhance semantic features through multi-layer neural networks, configure the dynamic association layer to include association rules, derivation and strength calculation units, and specify the attribute mapping layer to be responsible for entity attribute definition and matching; S2.

2. Annotated cross-institutional multi-source heterogeneous medical data is used to construct training and validation sets. A joint loss function is constructed, which consists of heterogeneous data adaptation loss function, entity matching loss function and semantic association loss function. The initial learning rate, iteration batch and model convergence threshold are set. The parameters of each layer are adjusted alternately through the backpropagation algorithm to optimize the adaptation accuracy, parsing accuracy and association matching degree in turn until the overall adaptation accuracy of the validation set is stable. S2.

3. Based on the model structure, standardization rules are formulated, and data standardization is completed through format adaptation, entity encoding, semantic enhancement and attribute matching.

4. The method according to claim 1, characterized in that, Step S3 includes: S3.

1. Extract entity information, relation information and feature information from the standardized patient follow-up data. Entity information includes the unified code and attribute value of each entity. Relationship information includes the direct association type and indirect association clue between entities. Feature information includes time-series features, text features and structured features. S3.

2. Based on the dynamic association layer rules of the multi-source heterogeneous data integration model, the improved cosine similarity algorithm is used to calculate the semantic similarity of entities in the patient follow-up data of different institutions. Combined with the dynamic weight of the association strength calculation unit, a similarity threshold is set as the association judgment condition to establish the direct association relationship between entities. S3.

3. Through the indirect association derivation unit of the dynamic association layer, based on the attribute association clues, temporal association clues and event association clues between entities, the indirect association relationship between entities is deduced, the indirect association strength is calculated and compared with the preset strength threshold, and the valid indirect association relationship is retained. S3.

4. Based on the established direct and effective indirect relationships, the follow-up data of patients from different institutions and of different types corresponding to the same patient are linked and integrated. The integration process maintains the consistency of entity unified coding, attribute information and relationship. According to the classification of the entity parsing layer and the structure of the attribute mapping layer, the data is organized into a unified patient health record.

5. The method according to claim 1, characterized in that, Step S4 includes: S4.

1. Receive patient identification information from the data retrieval request. The patient identification information includes the patient's unique number, identity information, and entity unique code. S4.

2. Based on the dynamic association layer rules of the patient identification information and the multi-source heterogeneous data integration model, the corresponding patient follow-up related data are traversed and matched in the unified patient health record. First, the directly related data corresponding to the core patient entity is matched, and then all related data are matched based on the indirect association relationship. S4.

3. Classify and organize the matched patient follow-up data according to time order, data type and correlation strength. Data types include electronic medical records, examination reports, follow-up records, rehabilitation training tracks, self-management logs and referral handover vouchers; S4.

4. Present patient follow-up data in the sorted order, simultaneously display the direct and indirect relationships between data, and support centralized acquisition and linked viewing of data across institutions.

6. The method according to claim 3, characterized in that, In step S2.1, a modular encapsulation design is used to construct each layer of the multi-source heterogeneous data integration model. The heterogeneous data adaptation layer and the entity parsing layer are configured to achieve bidirectional data transmission through a data interface adaptation component. The entity parsing layer and the semantic enhancement layer are configured to be connected through an entity vector encoding module. The semantic enhancement layer and the dynamic association layer are configured to establish a mapping through a semantic feature enhancement component. The dynamic association layer and the attribute mapping layer are configured to be associated through a dynamic weight allocation module. It is stipulated that the components of each layer communicate through a standardized data interface.

7. The method according to claim 3, characterized in that, In step S2.3, when performing format adaptation processing on the cross-institutional referral handover voucher data, the fixed and variable fields in the voucher are parsed through the structured data adaptation component, and the field mapping relationship is completed according to the requirements of the entity attribute definition unit. The unstructured voucher remarks information is converted into a standardized text format. The temporal features of the rehabilitation training trajectory data are extracted through the sliding window submodule of the temporal data adaptation component, and the text information of the patient self-management log data is processed through the word segmentation and stop word filtering submodule of the text data adaptation component.

8. The method according to claim 4, characterized in that, In step S3.2, the improved cosine similarity algorithm introduces an entity attribute weight factor and a time decay factor. The entity attribute weight factor is determined by the attribute type matching unit of the attribute mapping layer, and the time decay factor is dynamically adjusted according to the timestamp differences between rehabilitation training trajectory data and follow-up records, thereby improving the accuracy of association matching of data across different time spans.

9. The method according to claim 4, characterized in that, In step S3.3, the indirect association derivation unit adopts a multi-step reasoning mechanism. First, it establishes initial indirect association clues based on the referral events in the cross-institutional referral handover voucher data. Then, it combines the behavioral characteristics in the patient self-management log data and the progress characteristics of the rehabilitation training trajectory data to gradually deduce the multi-level indirect association relationships between entities. The rationality of each step of the reasoning result is verified by the semantic feature fusion module of the semantic enhancement layer.

10. A cross-institutional patient follow-up data integration system based on knowledge graphs, characterized in that, The system includes a data acquisition module, a data processing module, a data association module, and a data retrieval module; The data acquisition module is used to acquire patient follow-up data from multiple institutions, including electronic medical records, examination reports, follow-up records, rehabilitation training trajectory data, patient self-management log data, and cross-institutional referral handover voucher data. The data processing module is used to construct a multi-source heterogeneous data integration model and standardize the patient follow-up data. The model is set as a five-layer architecture, with each layer connected through adapter components. The model training unit completes training based on the training set, validation set, and joint loss function. The data association module is used to extract data information and establish associations through semantic mapping technology combined with model rules, and integrate them to form a unified patient health record. The data retrieval module is used to receive patient identification information, match and organize the corresponding patient follow-up data based on the health record and model rules, and present the data.

Citation Information

Patent Citations

  • A multi-source information fusion knowledge map representation learning method based on adaptive weight

    CN109033129A

  • Multi-modal content understanding method and system based on knowledge graph

    CN120372538A

  • Financial business compliance risk identification method based on knowledge graph reasoning

    CN120494960A

  • Large language model training method and system based on knowledge graph enhancement

    CN120578770A

  • Knowledge graph construction method and system based on large language model

    CN120633803A