Hemodialysis data management method and system based on whole course of disease multi-dimensional data integration

By constructing a full-process dialysis data chain and using blockchain technology, the problem of inaccurate data matching in the hemodialysis system has been solved, enabling collaborative analysis of multi-dimensional data, improving the efficiency and accuracy of dialysis research, and meeting compliance requirements.

CN122024985APending Publication Date: 2026-05-12AFFILIATED HUSN HOSPITAL OF FUDAN UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
AFFILIATED HUSN HOSPITAL OF FUDAN UNIV
Filing Date
2026-04-02
Publication Date
2026-05-12

AI Technical Summary

Technical Problem

Existing hemodialysis systems cannot achieve unified timestamp alignment across systems, resulting in clinical, equipment, and laboratory data from different sources not being accurately matched to the corresponding single dialysis session. This leads to a missing sample transfer chain, making it difficult to support precision medicine research and the implementation of personalized dialysis protocols.

Method used

By constructing a full-course dialysis data chain with the patient's unique identifier as the main index, using blockchain technology to record the flow nodes of biological samples, and combining a precise mapping mechanism between omics data and dialysis sessions, automatic matching and association of multi-dimensional data can be achieved, and a longitudinal analysis mechanism for sample, omics, and clinical data can be established.

Benefits of technology

It enables traceability and interpretability of multi-dimensional data in hemodialysis scenarios, improves data integration efficiency, reduces research costs, ensures data reproducibility and compliance, and supports dialysis adequacy assessment, complication risk prediction, and personalized treatment plan optimization.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122024985A_ABST
    Figure CN122024985A_ABST
Patent Text Reader

Abstract

The invention relates to a hemodialysis data management method and system based on whole course multi-dimensional data integration. The method comprises the following steps: acquiring multi-source hemodialysis data of a patient after the patient sees a doctor for the first time, and processing the multi-source hemodialysis data into patient dialysis data under a uniform timestamp; a single dialysis session is used as a basic disease course unit, and a unique dialysis session identifier is allocated to each basic disease course unit, so that a whole disease course dialysis data chain with a unique identifier of a patient as a root node and the dialysis session identifier as a core anchor point is constructed; based on the incidence relation between the sample and the omics, the multi-modal data of the sample dimension and the omics dimension are hooked to the whole course dialysis data chain, so that data tracing based on the whole course dialysis data chain is achieved, and longitudinal modeling analysis is conducted based on data, obtained through tracing, of a plurality of basic course units. According to the method, traceability and interpretable cooperation of clinical, equipment, samples and omics parameters on the same whole course timeline in a hemodialysis scene is effectively realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of medical big data technology, and in particular to a method and system for managing hemodialysis data based on the integration of multidimensional data throughout the entire disease course. Background Technology

[0002] The evaluation of hemodialysis treatment effectiveness, early warning of complications, and adjustment of personalized treatment plans rely heavily on long-term data comparisons of the same patient. Currently, patient-related data in the industry is scattered and stored in heterogeneous systems such as HIS (Hospital Information System), LIS (Laboratory Information System), dialysis equipment, biobanks, and omics testing platforms, which has the following drawbacks: The existing system has not achieved unified timestamp alignment of cross-system data. Clinical, equipment and laboratory data from different sources cannot be accurately matched to the corresponding single dialysis session. Insufficient correlation of underlying data leads to insufficient accuracy of data source for longitudinal disease progression analysis. Currently, biological samples are only labeled with the basic sampling time on transfer devices and managed independently from clinical course data. The lack of a sample transfer chain makes it impossible to exclude the interference of confounding factors such as sampling timing, number of freeze-thaw cycles, and batch testing in omics detection results, which makes it difficult to support the reproducibility of precision medicine research. Existing dialysis management systems only provide statistical displays of equipment parameters and individual clinical indicators for a single dialysis session. They do not establish a mapping relationship between omics data and corresponding dialysis sessions and biological samples, making it impossible to achieve joint longitudinal analysis of clinical phenotypes, equipment parameters, and omics characteristics. This makes it difficult to discover the underlying molecular mechanisms of dialysis-related complications and efficacy differences, and also fails to support the training of personalized early warning models. Summary of the Invention

[0003] The technical problem to be solved by the present invention is to provide a hemodialysis data management method and system based on the integration of multidimensional data throughout the entire course of the disease, which can effectively realize the traceability and interpretability of clinical, equipment, sample and omics parameters on the same time axis throughout the entire course of the disease in the hemodialysis scenario.

[0004] The technical solution adopted by this invention to solve its technical problem is: to provide a hemodialysis data management method based on the integration of multidimensional data throughout the entire disease course, including: Acquire multi-source hemodialysis data of patients since their first visit, including multimodal data from clinical, sample, equipment, and omics dimensions; Using the patient's unique identifier as the primary index, the multi-source hemodialysis data is deduplicated and correlated, and then converted to a unified timestamp to form patient dialysis data. Using a single dialysis session as the basic disease process unit, a unique dialysis session identifier is assigned to each basic disease process unit, thereby constructing a first full-process dialysis data chain with the patient's unique identifier as the root node and the dialysis session identifier as the core anchor point. Based on multimodal data of the sample dimension, a sample flow chain is generated for each biological sample. A digital record node is created for each flow node in the sample flow chain. The digital record node is attached to the first full-course dialysis data chain to obtain the second full-course dialysis data chain. Based on its correspondence with the biological samples, the multimodal data of the omics dimension is mapped to the second full-course dialysis data chain to obtain the third full-course dialysis data chain; Using the third full-course dialysis data chain, data from several basic disease units are traced and obtained, and longitudinal modeling analysis is performed based on the traced data.

[0005] Furthermore, the construction of the first full-course dialysis data chain, with the patient's unique identifier as the root node and the dialysis session identifier as the core anchor point, includes: For any given basic disease process unit, the patient's dialysis data before, during, and after dialysis is expanded in the time domain, and each is formed into a dialysis data sub-chain and bound to the corresponding dialysis session identifier; The dialysis data sub-chains are linked to the corresponding patient's unique identifier in chronological order to form the first full-course dialysis data chain.

[0006] Furthermore, the predialysis multimodal data includes laboratory indicators, vital signs, and medication information.

[0007] Furthermore, when the digital recording node is attached to the first full-course dialysis data chain, it is attached to the corresponding dialysis data sub-chain according to the laboratory indicators generated by the biological sample detection corresponding to the digital recording node.

[0008] Furthermore, the multimodal data in the dialysis includes the dialysis device number and the corresponding real-time operating parameters, which include blood flow rate, dialysate flow rate, transmembrane pressure, ultrafiltration volume, and alarm events.

[0009] Furthermore, the post-dialysis multimodal data includes complication records, weight changes, blood pressure recovery status, and medical order adjustment information.

[0010] Furthermore, it also includes: Obtain all non-dialysis events of the patient since the first visit and convert them to a unified timestamp; For any of the aforementioned non-dialysis events, construct a corresponding type identifier and full event data; In chronological order, the type identifier and the full event data are inserted into the gap between two core anchor points adjacent to their time.

[0011] Furthermore, the circulation nodes include collection, storage, processing, detection, and destruction, and the sample circulation chain is built on blockchain.

[0012] This invention also provides a hemodialysis data management system based on multidimensional data integration throughout the entire disease course, comprising: The processing module is used to acquire multi-source hemodialysis data of patients since their first visit. Using the patient's unique identifier as the main index, the multi-source hemodialysis data is deduplicated and correlated, and then converted to a unified timestamp to form patient dialysis data. The multi-source hemodialysis data includes multimodal data in clinical, sample, equipment, and omics dimensions. A construction module is used to assign a unique dialysis session identifier to each basic dialysis session as a basic disease unit, thereby constructing a first full-course dialysis data chain with the patient's unique identifier as the root node and the dialysis session identifier as the core anchor point. The association module is used to generate a sample flow chain for each biological sample based on multimodal data of the sample dimension, create a digital record node for each flow node in the sample flow chain, and attach the digital record node to the first full-course dialysis data chain to obtain the second full-course dialysis data chain; according to its correspondence with the biological sample, the multimodal data of the omics dimension is mapped to the second full-course dialysis data chain to obtain the third full-course dialysis data chain; The analysis module is used to utilize the third full-course dialysis data chain to trace and obtain multi-dimensional full-course hemodialysis data around a single dialysis session, and to perform longitudinal modeling analysis based on the traced data.

[0013] Furthermore, it also includes a privacy protection and hierarchical authorization module, used to perform field-level desensitization, differential privacy noise addition, and role-based access control on the patient dialysis data.

[0014] Beneficial effects By adopting the above-mentioned technical solution, the present invention has the following advantages and positive effects compared with the prior art: This invention constructs a blockchain-based digital record of all circulation nodes for each biological sample, and simultaneously links the circulation record to the entire dialysis data chain. Combined with a precise mapping mechanism between omics data and corresponding biological samples and dialysis sessions, it achieves real-time linkage and matching of sample entities, circulation process information, omics test results, and corresponding dialysis clinical data. This avoids mismatches at the architectural level and effectively ensures the reproducibility of precision medicine research results. This invention relies on a three-layer full-course dialysis data chain with a single dialysis session as the core anchor point. It can effectively realize one-click association and retrieval of multi-dimensional data from clinical, equipment, sample, and omics perspectives. It eliminates the need for manual cross-system organization and matching of scattered data, and can directly support longitudinal analysis needs such as dialysis adequacy assessment, complication risk prediction, and personalized dialysis protocol optimization. This significantly reduces the data preparation cost for clinical and scientific research and significantly shortens the research cycle. This invention achieves automatic matching and integration of multi-source dialysis-related data by combining standardized data adaptation rules with deduplication, association, and unified timestamp alignment mechanisms of the patient's unique identifier master index. This significantly improves data integration efficiency, reduces data redundancy, and fundamentally solves the problem of existing cross-system data not being accurately associated. This invention meets the compliance protection requirements of health and medical data through field-level anonymization, differential privacy noise addition, and role-based fine-grained access control mechanisms. At the same time, relying on the immutable evidence storage characteristics of blockchain, it realizes trusted evidence storage of the entire process of data access, sample operation, and data modification, supports ethical auditing and full-link traceability of risk events, and solves the problems of insufficient compliance and difficulty in audit traceability in existing systems. Attached Figure Description

[0015] Figure 1 This is a schematic diagram of the time axis of the multimodal data link according to the first embodiment of the present invention; Figure 2 This is a flowchart of the first embodiment of the present invention; Figure 3 This is a system architecture diagram of the second embodiment of the present invention; Figure 4 This is a schematic diagram of the federated learning collaboration module according to the second embodiment of the present invention; Figure 5 This is a schematic diagram of the nested sandbox containers according to the second embodiment of the present invention. Detailed Implementation

[0016] The present invention will be further illustrated below with reference to specific embodiments. It should be understood that these embodiments are for illustrative purposes only and are not intended to limit the scope of the invention. Furthermore, it should be understood that after reading the teachings of this invention, those skilled in the art can make various alterations or modifications to the invention, and these equivalent forms also fall within the scope defined by the appended claims.

[0017] The first embodiment of the present invention relates to a hemodialysis data management method based on the integration of multidimensional data throughout the entire course of the disease, which aims to solve the problems of difficulty in integrating multi-source heterogeneous data in the field of hemodialysis, lack of traceability and correlation between samples and clinical data, and insufficient compliance of scientific research sharing.

[0018] The main steps include the following: Data acquisition and processing; Construction of a multimodal data chain covering the entire disease course; Analysis was conducted based on a multimodal data chain covering the entire disease course.

[0019] Step 1: Data acquisition and processing can be performed using the following methods: Raw data from hospital information systems, laboratory information systems, medical imaging systems, and dialysis machine IoT terminals are collected via the HL7 / FHIR protocol interface. The original data is structurally transformed and standardized in encoding. The patient's unique identifier is used as the main index to complete the deduplication and association of cross-system data, generating a unified structured dataset.

[0020] Step 2: Construction of a multimodal data chain covering the entire disease course. This involves vertically aligning diagnostic and treatment events, dialysis records, biosample information, and omics data using a unified timestamp as the axis to construct a multimodal data chain covering the entire disease course. Furthermore, a unique code is assigned to each biosample, and the sample flow chain is recorded, enabling traceable association between sample data and corresponding omics data and clinical disease course chains. More specifically, based on the strong rhythmicity and equipment-driven characteristics of the hemodialysis treatment process, a multimodal data chain construction method centered on the dialysis session can be adopted, rather than using a general disease event node alignment method. This method includes the following important mechanisms: 1. Definition of dialysis session anchor points: Each hemodialysis treatment process is used as a basic disease unit. A unique session identifier (Session ID) is generated for each dialysis session. This identifier is bound to the patient's unique master index and is used to carry all relevant data before, during and after the dialysis session. 2. Multimodal data stratified acquisition and time alignment: Around the anchor point of the dialysis session, the following data modalities were acquired and aligned in chronological order: Before dialysis: Laboratory indicators (such as Scr, BUN, K) + Ca 2+ Hb), vital signs, and medication information; During dialysis: Real-time operating parameters of the dialysis machine, including high-frequency time-series data such as blood flow, dialysate flow, transmembrane pressure, ultrafiltration volume, and alarm events; Post-dialysis: Records of complications, weight changes, blood pressure recovery, and adjustments to medical orders; Each modal data is vertically bound to the dialysis session identifier through a unified timestamp mechanism, forming a multimodal data sub-chain around a single dialysis session; 3. Cross-session vertical splicing mechanism: By linking multiple dialysis sessions of the same patient in chronological order, a multimodal data chain covering the entire disease course is constructed with the patient's unique identifier as the root node and the dialysis session as the core anchor point, thereby realizing the continuous expression of the patient's dialysis life cycle. 4. Non-dialysis event supplementation mechanism: For non-dialysis events such as hospitalization, infection, surgery, and transplant evaluation, these events are inserted between adjacent dialysis sessions using event type identifiers to ensure the integrity of the disease progression chain and medical continuity.

[0021] like Figure 1 As shown, by assigning a globally unique identifier to a single complete dialysis treatment cycle (i.e., a dialysis session), this identifier can be used as a "link" to bind all the scattered data involved in a single dialysis session and match it with the upper-level patient master index. The globally unique identifier can adopt the naming rule of [TXH + date + patient master index suffix].

[0022] A single dialysis session covers a period not only from the time a dialysis patient starts to the time they stop using the dialysis machine, but also the entire closed loop from the start of the dialysis process to the complete archiving of data. Specifically, this includes the doctor issuing the dialysis prescription, the patient signing in at the hospital, pre-dialysis blood collection, starting the dialysis machine, real-time data collection from the equipment throughout the dialysis process, recording of adverse events / interventions during dialysis, post-dialysis blood collection, stopping the dialysis machine, uploading of nursing records, feedback of pre-dialysis / post-dialysis test / omics results, and verification and archiving of all data.

[0023] Existing routine dialysis records only store scattered fields filled in by nurses. This implementation method, however, uses dialysis session IDs to link all heterogeneous data scattered across different systems. If a patient is temporarily disconnected from the dialysis machine due to complications during the same dialysis session and then resumes treatment, it is classified as the same dialysis session and uses the same ID. If the doctor prescribes a new dialysis session to make up for the interruption at another time, it is considered a new session, thus avoiding data mismatch.

[0024] The data chain constructed in this way is naturally suited to the high-frequency, highly constrained, and device-driven characteristics of hemodialysis, unlike general full-course data alignment methods based on event similarity. When performing subsequent omics differential analysis, simply locking onto a specific dialysis session ID allows direct mapping to the device parameters, intervention plan, and omics change results for that session, completely eliminating the problem of "not knowing which dialysis session the omics results correspond to or being unable to pinpoint the cause of the difference." Furthermore, when anonymizing for research, even after removing the privacy information bound to the patient's master index, the full-dimensional data of a single dialysis session remains completely interconnected, complying with privacy compliance requirements without compromising the analytical value of the data.

[0025] In some preferred embodiments, the three stages of pre-dialysis, during dialysis, and post-dialysis can be distinguished by rigid time boundary anchoring and mandatory stage labels, and data sub-chains can be constructed for each stage to align with clinical workflows. The stage division can be based on the dialysis machine start / stop confirmation time as an absolute benchmark; these two timestamps cannot be manually modified, avoiding human error in stage division. Specifically, pre-dialysis is defined as the time from the dialysis prescription issuance to the start / stop confirmation time; during dialysis is defined as the time from the start / stop confirmation time to the end / stop confirmation time; and post-dialysis is defined as the time from the end / stop confirmation time to the completion of data archiving for the current dialysis session.

[0026] All nodes in each stage will be forcibly tagged with the corresponding stage label and bound to the four types of core data generated in that stage, preventing data cross-stage concatenation. Specifically, the following methods can be used: 1. The front-end carries the full-dimensional data generated in this stage, including: Clinical dimensions: Dialysis prescription for this procedure, predialysis signs (weight, blood pressure, coagulation function, etc.), and preoperative assessment records; Sample dimensions: Sampling records, sample numbers, and storage information for predialysis blood / urine samples; Equipment-related aspects: Dialysis equipment self-inspection results, batch numbers of consumables used this time (dialyzer, tubing, dialysate), and pre-flushing completion record; Verification rule: If the node time is later than the machine access time, the system will automatically intercept and report an error to avoid mislabeling data in the middle / after stage as the front stage.

[0027] The intermediate section carries the dynamic time-series data for this stage, including: Equipment dimension: Time-series operating parameters (venous pressure, transmembrane pressure, actual ultrafiltration rate, conductivity, etc.) every 5 minutes, and time / type / processing records of all alarms; Clinical dimensions: adverse events during dialysis (hypotension, convulsions, etc.), records of intervention measures, and records of temporary prescription adjustments; Sample dimensions: Information on blood and waste fluid samples temporarily collected during the procedure. All nodes are strictly sorted by timestamp to fully restore the dynamic changes throughout the dialysis process. Even if the patient is briefly taken off the machine to continue dialysis, as long as it belongs to the same dialysis session, it will still be classified as the dialysis session and will not be split into other stages.

[0028] 3. Transparent section: All associated data generated after mounting and unmounting: Clinical dimensions: post-dialysis signs, patient self-assessment, efficacy evaluation, and next dialysis appointment record; Sample dimensions: Collection / storage information of postdialysis blood samples and end-stage dialysis waste fluid samples; Omics dimension: The test results and omics detection / analysis results of all pre-dialysis / dialysis / post-dialysis samples will be automatically associated with the nodes of the corresponding sampling stages; Verification rule: Omics results that are not bound to samples from the corresponding stage cannot be added to the subchain to avoid data mismatch.

[0029] The aforementioned sample-omics-clinical association mechanism does not exist independently of the disease course data chain, but rather forms a one-to-one synergistic relationship with the entire disease course multimodal data chain. This can be achieved through the following specific implementation methods: 1. Sample generation and dialysis session binding mechanism: For blood, plasma or other biological samples collected during hemodialysis, a corresponding dialysis session identifier is bound at the time of sample generation, so that each biological sample has a clear clinical time context and treatment context at the generation stage.

[0030] 2. Sample transfer chain and data chain synchronization mechanism: Each sample is assigned a physical QR code identifier, and the flow nodes of sample collection, storage, processing, detection and destruction are recorded in the blockchain; at the same time, the digital record node corresponding to the sample is attached to the corresponding dialysis session node in the full-process multimodal data chain to realize the synchronous update of sample physical flow and disease process data version.

[0031] 3. Back-link mapping mechanism for omics data: Once a sample generates corresponding proteomics, metabolomics, or genomics data, the omics data is back-mapped to the corresponding dialysis session node through the sample's unique identifier, and further embedded into the patient's multimodal data chain throughout the entire disease course, so that the omics features can be interpreted as biological response results under specific dialysis stages and treatment parameter conditions.

[0032] 4. Vertical analysis of collaborative mechanisms: Based on the above correlation methods, the multidimensional analysis and intelligent modeling engine can perform longitudinal modeling analysis on the changing trends of omics characteristics with dialysis protocol adjustments, complication occurrences, or long-term prognosis across multiple dialysis sessions for patients.

[0033] Through the above mechanism, traceable and interpretable collaboration of "sample-omics-clinical-equipment" across the same entire disease course timeline is achieved in hemodialysis scenarios, unlike general medical data management solutions that only anonymize or share medical data. This collaborative mechanism firmly binds the four data categories of "sample-omics-clinical-equipment" through a two-layer index architecture, preventing mismatches and omissions. The upper layer uses the patient master index to anchor specific patients, and the middle layer uses the dialysis session ID to anchor specific dialysis sessions. All four data categories are bound to these two IDs. Because the sample QR code is already bound to the corresponding dialysis session ID when it is generated, each blockchain transfer record of the sample is automatically synchronized to the corresponding stage node of the dialysis session in the entire disease course data chain: for example, the collection record of pre-dialysis samples is automatically attached to the pre-dialysis sub-chain node of the dialysis session, and the test results of post-dialysis samples are automatically attached to the post-dialysis sub-chain node.

[0034] Step 3: Analyze the data based on the full-course multimodal data chain. Machine learning algorithms can be used to perform dialysis adequacy assessments, complication risk predictions, and personalized dialysis prescription recommendations. All data sources for analysis are based on the full-course multimodal data chain, anchored by the patient's unique master index. The data automatically aggregates all regular dialysis sessions for the patient, aligning them along both dialysis frequency and natural time sequence, and automatically removing interfering data from temporary or emergency dialysis.

[0035] like Figure 2 As shown, in addition to the three main steps mentioned above, the entire process may also include one or more of the following steps: The dataset is anonymized by performing field-level desensitization, differential privacy noise addition, and role-based access control. The differential privacy noise addition can adopt an adaptive budget allocation strategy, dynamically adjusting the ε value based on the sensitivity of the data fields. Provide containerized isolated computing services to researchers who have passed ethical approval in a sandbox analysis environment, and support tiered authorization transactions for data packets; it may also include: performing vulnerability scanning and signature verification on algorithm images uploaded by researchers before task execution; and collecting system call audit logs in real time during task execution and uploading them to a regulatory ledger; By using blockchain technology, data access logs, analysis task parameters, and result download records are written into an immutable regulatory ledger, enabling full-process compliance auditing.

[0036] The second embodiment of the present invention relates to a hemodialysis data management system based on the integration of multidimensional data throughout the entire disease course, used to implement the above-mentioned hemodialysis data management method, specifically including: The multi-source data acquisition module is used to collect raw clinical, imaging, laboratory, and equipment operation data from the Hospital Information System (HIS), Laboratory Information System (LIS), Medical Image Storage and Transmission System (PACS), and dialysis machine IoT terminals via the HL7 / FHIR protocol interface. The data standardization and master index matching module is used to perform structured transformation and coding standardization of the original data according to the HL7 / FHIR and OMOP-CDM standards, and to complete cross-system data deduplication and association with the patient's unique identifier as the master index, generating a unified structured dataset; The full-course multimodal data chain construction module is used to vertically align diagnosis and treatment events, dialysis records, biological sample information and omics data in a unified structured dataset with timestamps as the axis, forming a scalable multimodal disease course chain; The Sample-Omics-Clinical Association Module is used to assign a unique code to each biological sample and record the sample flow chain, enabling traceable association between sample data and corresponding omics data and clinical disease progression chains. The multidimensional analysis and intelligent modeling engine is used to perform dialysis adequacy assessment, complication risk prediction, and personalized dialysis prescription recommendation based on a multimodal disease progression chain using machine learning algorithms, and outputs the analysis results.

[0037] In some preferred embodiments, such as Figure 3 As shown, it also includes: The privacy protection and graded authorization module is used to perform field-level desensitization, differential privacy noise addition, and role-based access control on the unified structured dataset and multimodal pathology chain, generating a desensitized dataset that meets ethical review requirements. The data sharing and transaction module is used to provide containerized and isolated computing services for de-identified datasets to researchers who have passed ethical approval in a sandbox analysis environment, and supports hierarchical authorization transactions for data packets; The regulatory audit chain module uses blockchain technology to write data access logs, analysis task parameters, and result download records into an immutable regulatory ledger, enabling full-process compliance auditing.

[0038] The multi-source data acquisition module also includes an IoT gateway unit for edge caching and breakpoint resumption of real-time operating parameters of the dialysis machine. The data standardization and master index matching module further includes a data quality repair unit for probabilistic graphical model completion and consistency verification of missing or conflicting fields.

[0039] The full-course multimodal data chain construction module adopts an scalable key-time sequence storage structure, supporting the dynamic addition of new data modalities under a unique patient identifier. The multidimensional analysis and intelligent modeling engine includes: an XGBoost-based dialysis adequacy index calculation sub-engine; an LSTM-based acute complication prediction sub-engine; and a reinforcement learning-based personalized dialysis prescription recommendation sub-engine.

[0040] like Figure 4 As shown, the sample-omics-clinical association module binds the sample entity to the digital circulation chain through a dual coding mechanism of QR code and blockchain, enabling real-time synchronization of sample location and data version.

[0041] The privacy protection and tiered authorization module employs a differential privacy algorithm, in which... , To provide maximum data utility within a quantifiable privacy budget. For example... Figure 5 As shown, the sandbox analysis environment is based on Docker-in-Docker container nesting technology, allocating an independent computing namespace and temporary data volume to each researcher, which are automatically destroyed after the task is completed. The data sharing and transaction module has built-in smart contracts for automatically executing data usage licenses, fee settlements, and result returns on the blockchain, supporting a pay-per-use model for multi-center research. The regulatory audit chain module uses the Hyperledger Fabric consortium blockchain and separates and stores data access logs from different hospital nodes through channel isolation technology to meet cross-hospital data compliance requirements.

[0042] In some implementations, the system further includes a federated learning collaboration module, used to collaboratively train a global complication prediction model based on local anonymized datasets from multiple hospitals without requiring out-of-domain communication, exchanging only encrypted gradient parameters. The federated learning collaboration module employs a hybrid mechanism of homomorphic encryption and differential privacy to ensure that individual patient information cannot be inferred during gradient uploading.

[0043] Example 1: Deployment and Application of a Single-Center Hemodialysis Full-Cycle Data Management Platform in a Tertiary Hospital The platform in this embodiment is deployed in the hemodialysis center of a top-tier hospital in East China. The specific deployment and operation are as follows: The platform is deployed on a hardware foundation consisting of a dual-socket Intel Xeon Gold 6248R server cluster, configured with 512GB of memory and 40TB of NVMe SSD storage, and equipped with a 10GbE fiber optic network to ensure data transmission efficiency. The software environment uses the CentOS 8.2 operating system, based on Kubernetes 1.24 and Docker 20.10 for containerized scheduling, and includes a built-in HyperledgerFabric 2.4 blockchain notarization module, a PostgreSQL 14 relational database, and an Apache Nifi 1.16 data flow component.

[0044] During the platform's operation phase, multi-source heterogeneous data is incrementally collected daily through the HL7 / FHIR standard interface, including 12,000 HIS medical records, 35,000 LIS laboratory records, 8,000 PACS image indexes, and 15GB of dialysis machine operation logs. In the data association process, medical insurance number + ID card hash is used as the unified master index, and a probabilistic graphical model is used to complete missing fields, achieving a cross-system data matching accuracy of 99.3%.

[0045] After data integration, the platform automatically constructs a full-cycle multimodal data chain for each patient. Taking patient P0001 as an example, its data timeline covers from March 1, 2021 to April 30, 2023, integrating a total of 420 diagnosis and treatment events, 1,152 dialysis records, 86 sample records, and 1.2TB of genomic VCF files. All data are arranged in order according to the anchor point of the dialysis session.

[0046] At the analytical application level, the platform's built-in multidimensional analysis module can support multi-scenario modeling: the AUC for predicting hypotension during dialysis using the XGBoost model reaches 0.91, the accuracy of predicting hypotension within 12 hours using the LSTM model reaches 89%, and the error for recommending ultrafiltration volume using the reinforcement learning model is ≤3.2%.

[0047] The platform's privacy protection module performs automated de-identification processing on 18 sensitive fields, including age and phone number, and employs a differential privacy mechanism. , The data utility loss is only 5.7%, balancing data security and availability.

[0048] In the context of scientific research data sharing, researchers submitted a proposal on "the association between gut microbiota and inflammatory factors in hemodialysis patients." With ethics approval number 2023-03-017, they obtained Level 2 access authorization and could run the R4.2 image for analysis within the sandbox environment provided by the platform. The sandbox environment was automatically destroyed after 72 hours, and the audit chain record of the entire operation process, with hash value 0x9a7f…, was written to block height 318472, making the entire process traceable.

[0049] In the sample-omics-clinical association application scenario, scanning the QR code of sample number S000185 can locate the sample stored in the 3rd column of the 5th layer of the -80℃ refrigerator, corresponding to the version number v2.1 of the blockchain record. The detection data of this sample is compared with the IL-6 genotype rs1800795 of patient P0001, and a significant correlation result of p<0.001 is obtained.

[0050] Example 2: Multi-center Hemodialysis Research Collaboration Application Based on Federated Learning This example illustrates a multi-center research collaboration scenario across institutions, involving three medical institutions: East China Hospital, West China Hospital, and Zhongshan Hospital.

[0051] During the collaborative modeling process, all three participating parties trained LSTM complication prediction models based on locally deployed hemodialysis data management platforms and anonymized full-course data. They only uploaded homomorphically encrypted gradient parameters to the trusted coordinator, and no original individual data flowed out of the local institutions.

[0052] After receiving the encrypted gradients uploaded by the three parties, the coordinator completes the aggregation operation, updates the global model parameters, and distributes them to each participant for iterative training. After three rounds of iteration, the prediction AUC of the global model is improved by 0.05 compared with the single-center model. There is no cross-domain transmission of raw data throughout the process, which avoids the data security risks of multi-center cooperation.

[0053] During the modeling process, differential privacy noise is added to the aggregated gradient, with parameters set to... , This can effectively prevent the risk of inferring sensitive individual information through gradients and meet compliance requirements.

Claims

1. A method for managing hemodialysis data based on the integration of multidimensional data throughout the entire disease course, characterized in that, include: Acquire multi-source hemodialysis data of patients since their first visit, including multimodal data from clinical, sample, equipment, and omics dimensions; Using the patient's unique identifier as the primary index, the multi-source hemodialysis data is deduplicated and correlated, and then converted to a unified timestamp to form patient dialysis data. Using a single dialysis session as the basic disease process unit, a unique dialysis session identifier is assigned to each basic disease process unit, thereby constructing a first full-process dialysis data chain with the patient's unique identifier as the root node and the dialysis session identifier as the core anchor point. Based on multimodal data of the sample dimension, a sample flow chain is generated for each biological sample. A digital record node is created for each flow node in the sample flow chain. The digital record node is attached to the first full-course dialysis data chain to obtain the second full-course dialysis data chain. Based on its correspondence with the biological samples, the multimodal data of the omics dimension is mapped to the second full-course dialysis data chain to obtain the third full-course dialysis data chain; Using the third full-course dialysis data chain, data from several basic disease units are traced and obtained, and longitudinal modeling analysis is performed based on the traced data.

2. The hemodialysis data management method according to claim 1, characterized in that, The construction of the first full-course dialysis data chain, with the patient's unique identifier as the root node and the dialysis session identifier as the core anchor, includes: For any given basic disease process unit, the patient's dialysis data before, during, and after dialysis is expanded in the time domain, and each is formed into a dialysis data sub-chain and bound to the corresponding dialysis session identifier; The dialysis data sub-chains are linked to the corresponding patient's unique identifier in chronological order to form the first full-course dialysis data chain.

3. The hemodialysis data management method according to claim 2, characterized in that, The predialysis multimodal data includes laboratory indicators, vital signs, and medication information.

4. The hemodialysis data management method according to claim 3, characterized in that, When the digital recording node is attached to the first full-course dialysis data chain, it is attached to the corresponding dialysis data sub-chain according to the laboratory indicators generated by the biological sample detection corresponding to the digital recording node.

5. The hemodialysis data management method according to claim 2, characterized in that, The multimodal data in the dialysis includes the dialysis device number and the corresponding real-time operating parameters, which include blood flow rate, dialysate flow rate, transmembrane pressure, ultrafiltration volume, and alarm events.

6. The hemodialysis data management method according to claim 2, characterized in that, The post-dialysis multimodal data includes complication records, weight changes, blood pressure recovery, and medical order adjustments.

7. The hemodialysis data management method according to claim 1, characterized in that, Also includes: Obtain all non-dialysis events of the patient since the first visit and convert them to a unified timestamp; For any of the aforementioned non-dialysis events, construct a corresponding type identifier and full event data; In chronological order, the type identifier and the full event data are inserted into the gap between two core anchor points adjacent to their time.

8. The hemodialysis data management method according to claim 1, characterized in that, The circulation nodes include collection, storage, processing, detection, and destruction, and the sample circulation chain is built on blockchain.

9. A hemodialysis data management system based on the integration of multidimensional data throughout the entire disease course, characterized in that, include: The processing module is used to acquire multi-source hemodialysis data of patients since their first visit. Using the patient's unique identifier as the main index, the multi-source hemodialysis data is deduplicated and correlated, and then converted to a unified timestamp to form patient dialysis data. The multi-source hemodialysis data includes multimodal data in clinical, sample, equipment, and omics dimensions. A construction module is used to assign a unique dialysis session identifier to each basic dialysis session as a basic disease unit, thereby constructing a first full-course dialysis data chain with the patient's unique identifier as the root node and the dialysis session identifier as the core anchor point. The association module is used to generate a sample flow chain for each biological sample based on multimodal data of the sample dimension, create a digital record node for each flow node in the sample flow chain, and attach the digital record node to the first full-course dialysis data chain to obtain the second full-course dialysis data chain. Based on its correspondence with the biological samples, the multimodal data of the omics dimension is mapped to the second full-course dialysis data chain to obtain the third full-course dialysis data chain; The analysis module is used to utilize the third full-course dialysis data chain to trace and obtain multi-dimensional full-course hemodialysis data around a single dialysis session, and to perform longitudinal modeling analysis based on the traced data.

10. The hemodialysis data management system according to claim 9, characterized in that, It also includes a privacy protection and hierarchical authorization module, which is used to perform field-level desensitization, differential privacy noise addition, and role-based access control on the patient dialysis data.