Multi-center data fusion and quality control system and method for chronic diseases
By standardizing access to multi-source data, implementing multi-dimensional quality control, ensuring privacy and security in storage, and facilitating cross-institutional data sharing, the technical bottlenecks in multi-center data processing for chronic diseases have been resolved. This has enabled efficient, secure, and deep data integration, thereby improving the data quality and analytical value of multi-center studies.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-05-01
- Publication Date
- 2026-07-24
AI Technical Summary
In existing technologies, multi-center data processing for chronic diseases suffers from problems such as inconsistent data formats, incomplete quality control, privacy and security risks, insufficient data fusion, untimely feedback of quality control results, and low storage and management efficiency, which seriously restrict the development of multi-center studies and the mining of data value.
By employing a multi-source data standardization access module, a multi-dimensional data quality control module, a privacy and security fusion storage module, a cross-hospital data collaborative sharing module, and a quality control result visualization and feedback module, we can achieve standardized collection, multi-dimensional quality control, secure storage, compliant sharing, and in-depth fusion analysis of chronic disease data, and establish a closed-loop quality control result feedback mechanism.
It has achieved standardized access and unified format for chronic disease data, improved data quality and security, ensured privacy protection, enabled efficient sharing and deep integration of data across hospitals, and promoted continuous optimization of data quality.
Smart Images

Figure CN122455210A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the fields of medical data processing, multi-center clinical research, and medical data quality control standardization, specifically to a multi-center data fusion and quality control system and method for chronic diseases. Background Technology
[0002] Liver disease, diabetes, and hypertension are three major chronic diseases with high incidence rates, long durations, and numerous complications, making them key concerns in my country's public health field. Conducting multi-center clinical studies on these three chronic diseases and mining the value of massive amounts of cross-institutional clinical data is crucial for revealing disease pathogenesis, optimizing treatment plans, and formulating public health prevention and control strategies. However, the current medical industry faces numerous technical bottlenecks and industry pain points in the processing and application of multi-center data for these three chronic diseases, severely restricting the conduct of multi-center studies and the mining of data value.
[0003] First, the data formats for the three major chronic diseases are inconsistent, and there is a lack of standardized access specifications. Electronic medical record systems, laboratory testing systems, and follow-up management systems in different medical institutions are independently developed, resulting in significant differences in data storage formats. For example, liver function indicators (ALT, AST, bilirubin, etc.) for liver disease have issues with inconsistent naming conventions and reference ranges within hospitals; fasting blood glucose and glycated hemoglobin (HbA1c) monitoring data for diabetes are stored in various formats, including structured, semi-structured, and unstructured; and systolic / diastolic blood pressure records and blood pressure classification standards for hypertension vary across hospitals. Furthermore, there is no unified standard for accessing multi-source data (electronic medical records, laboratory reports, imaging diagnostic results, follow-up data, and medication records), making direct data exchange between hospitals impossible and requiring extensive manual processing, resulting in extremely low efficiency.
[0004] Secondly, the data quality is inconsistent, lacking a multi-dimensional quality control mechanism specific to the three major chronic diseases. Clinical data for these diseases contains numerous missing, outlier, duplicate, and logically contradictory values. For example, glycated hemoglobin test data for diabetic patients is missing, blood pressure monitoring values for hypertensive patients exceed the physiologically reasonable range, and there are duplicate entries in the inpatient medical records of liver disease patients. Existing data quality control methods are mostly general manual checks, lacking specific quality control rules tailored to the characteristics of the three major chronic diseases. Furthermore, quality control only targets a single data dimension, failing to achieve multi-dimensional intelligent quality control of data completeness, accuracy, consistency, timeliness, and standardization. The quality control results are highly subjective, with a high rate of missed detections, and cannot meet the high-quality data requirements of multi-center studies.
[0005] Third, cross-hospital data sharing poses privacy and security risks and lacks a compliant collaborative sharing mechanism. The three major chronic disease data contain patients' personal privacy information and core clinical data, which are sensitive medical data. Current cross-hospital data sharing mostly involves offline copying or simple transmission, failing to adopt privacy protection technologies that comply with the "Personal Information Protection Law" and the "Regulations on Medical Record Management of Medical Institutions," easily leading to the leakage of patient privacy. Furthermore, the lack of a tiered authorization mechanism for cross-hospital data collaborative sharing means that there is no clear control over data access, use, and modification between medical institutions, resulting in poor traceability of data use.
[0006] Fourth, multi-center data fusion lacks dedicated models, resulting in insufficient data value mining. Existing data fusion methods are general data splicing methods that do not construct dedicated fusion analysis models based on the disease characteristics of liver disease, diabetes, and hypertension. This makes it impossible to achieve deep fusion of data from different dimensions of these three chronic diseases (such as imaging and laboratory data for liver disease, blood glucose and medication data for diabetes, and blood pressure and complication data for hypertension). At the same time, the fusion of multi-source heterogeneous data only stays at the data level, failing to achieve fusion at the feature and knowledge levels, and thus cannot provide valuable data analysis results for multi-center studies.
[0007] Fifth, data quality control results are not fed back in a timely manner and lack visual management tools. Currently, after data quality control is completed, only a text-based quality control report is output, without visualizing the results. Medical institutions cannot quickly locate data quality problems. At the same time, there is no closed-loop feedback and correction mechanism for quality control results. Problematic data discovered during quality control cannot be promptly pushed to the data reporting institution for correction, resulting in the inability to continuously improve data quality.
[0008] Sixth, multi-center data storage lacks integration, resulting in low data management efficiency. Data on the three major chronic diseases from different medical institutions are stored in a scattered manner, without establishing a unified, secure, and integrated storage system. The data storage media and formats vary, making it impossible to achieve centralized management and efficient retrieval of cross-institutional data. At the same time, the lack of data version management and backup mechanisms increases the risk of data loss and corruption, affecting the continuity of multi-center research. Summary of the Invention
[0009] The purpose of this invention is to provide a multi-center data fusion and quality control system and method for chronic diseases, so as to solve the problems existing in the prior art.
[0010] To achieve the above objectives, the present invention provides the following technical solution: a multi-center data fusion and quality control system for chronic diseases, including a multi-source data standardization access module, a multi-dimensional data quality control module, a privacy and security fusion storage module, a cross-hospital data collaborative sharing module, a chronic disease-specific data fusion analysis module, and a quality control result visualization and feedback module;
[0011] The multi-source data standardization access module enables the standardized collection and access of multi-source heterogeneous data for the three major chronic diseases, and outputs unified structured clinical data for the three major chronic diseases.
[0012] The multi-dimensional data quality control module performs multi-dimensional intelligent quality control on the standardized data of the three major chronic diseases, and outputs high-quality quality-controlled data.
[0013] The privacy and security fusion storage module encrypts and fusions the multi-center data of the three major chronic diseases after quality control, so as to achieve secure data storage and efficient management.
[0014] The cross-hospital data collaboration and sharing module enables compliant cross-hospital sharing of sensitive data for the three major chronic diseases, while also enabling full-process traceability of data usage;
[0015] The chronic disease-specific data fusion and analysis module enables in-depth fusion and analysis of multi-center data for the three major chronic diseases, and uncovers the correlation patterns between the data.
[0016] The quality control result visualization and feedback module enables the visualization and closed-loop feedback correction of quality control results, thereby promoting continuous improvement in data quality.
[0017] Furthermore, the multi-source data standardization access module includes a chronic disease data element standard setting unit, a multi-source data acquisition unit, a heterogeneous data conversion unit, and a data standardization verification unit;
[0018] The chronic disease data element standard setting unit formulates exclusive data element standards based on national medical data standards and combined with the characteristics of the three major chronic diseases, clarifying the name, definition, data type, value range, unit, and coding rules of the data elements; the multi-source data acquisition unit supports multiple acquisition methods to realize batch and real-time acquisition of multi-source heterogeneous raw data of the three major chronic diseases; the heterogeneous data conversion unit performs format conversion and structuring processing on the heterogeneous raw data, outputting structured data that conforms to the chronic disease data element standards; the data standardization verification unit performs standardization verification on the converted structured data, and after passing the verification, it is transmitted to the multi-dimensional data quality control module.
[0019] Furthermore, the multi-dimensional data quality control module includes a chronic disease-specific quality control rule base construction unit, a data integrity quality control unit, a data accuracy quality control unit, a data consistency quality control unit, a data timeliness and standardization quality control unit, and a problem data processing unit;
[0020] The chronic disease-specific quality control rule base construction unit combines the three major chronic disease clinical diagnosis and treatment guidelines to build a multi-dimensional quality control rule base, which supports dynamic updates; the data integrity quality control unit uses a missing value imputation weight calculation formula to identify and process missing values; the data accuracy quality control unit uses an outlier determination formula to identify and verify outliers; the problem data processing unit classifies and processes various types of problem data and performs secondary quality control, and after passing the quality control, the data is transmitted to the privacy and security fusion storage module.
[0021] Furthermore, the outlier determination formula is as follows: ,in The data value to be detected. The mean of the dataset containing the data element. The standard deviation of the dataset containing the data element. The outlier determination coefficient is 2 to 3; the formula for calculating the missing value imputation weight is as follows: ,in For the first Filling weights for each reference data source, For the first Quality ratings of reference data sources For the first The similarity coefficient between each reference data source and the data source containing the missing values. For reference, the number of data sources, This serves as an index for the reference data source.
[0022] Furthermore, the privacy and security integrated storage module includes a data encryption unit, a distributed integrated storage unit, a data indexing and retrieval unit, a data version management unit, and an off-site backup unit;
[0023] The data encryption unit performs end-to-end encryption and desensitization processing on the data of the three major chronic diseases; the distributed fusion storage unit adopts a distributed cloud storage architecture to realize the classified and fusion storage of data; the data indexing and retrieval unit constructs a multi-dimensional retrieval index to realize fast data retrieval; the data version management unit realizes version recording and backtracking of data operations.
[0024] The off-site backup unit establishes an off-site multi-copy backup mechanism to ensure reliable data storage.
[0025] Furthermore, the cross-institute data collaboration and sharing module includes a hierarchical authorization management unit, a privacy computing sharing unit, a data access and use unit, and a data use traceability unit;
[0026] The hierarchical authorization management unit constructs a role-based hierarchical authorization management mechanism to achieve refined control of data permissions; the privacy computing sharing unit combines privacy computing technology to achieve "usable but invisible" sharing of data on the three major chronic diseases; the data access and use unit provides authorized users with standardized data access and use channels; and the data use traceability unit constructs a full-process data use traceability mechanism to form a permanently stored traceability log.
[0027] Furthermore, the chronic disease-specific data fusion and analysis module includes a chronic disease data feature extraction unit, a chronic disease-specific fusion model construction unit, a multi-center data fusion and analysis unit, and a fusion result output unit;
[0028] The chronic disease data feature extraction unit extracts multiple types of features from the three major chronic disease data to form disease features and common features; the chronic disease-specific fusion model construction unit constructs a deep fusion analysis model specific to the three major chronic diseases respectively; the multi-center data fusion analysis unit uses the feature fusion similarity formula of the three major chronic disease data to realize the fusion of data features; the fusion result output unit outputs and stores the fusion analysis results after structured processing.
[0029] Furthermore, the similarity formula for the fusion of the three major chronic disease data features is as follows: ,in To achieve similarity in data feature fusion among medical institutions, , For the feature vectors of different medical institutions, , For the corresponding eigenvalues, The number of feature dimensions, , Index for medical institutions, For feature indexing.
[0030] Furthermore, the quality control result visualization and feedback module includes a quality control result statistics unit, a quality control result visualization display unit, a problem data push unit, a correction data review unit, and a data quality continuous optimization unit;
[0031] The quality control result statistics unit performs multi-dimensional statistical analysis on the quality control results to form a statistical data set; the quality control result visualization display unit realizes the visualization display and query of quality control results in various forms; the problem data push unit pushes problem data to the corresponding medical institutions in categories; the corrected data review unit conducts a second quality control review on the corrected data; and the data quality continuous optimization unit mines the patterns of quality problems, optimizes the quality control rule base, and establishes a data quality evaluation system.
[0032] The integration and quality control methods for a multicenter data fusion and quality control system for chronic diseases include the following steps:
[0033] Step 1: Standardized access to multi-source data, formulate data element standards for the three major chronic diseases, collect heterogeneous raw data from multiple sources and perform standardized conversion and verification, and output unified structured clinical data for the three major chronic diseases;
[0034] Step 2: Multi-dimensional data quality control. Construct three dedicated quality control rule bases for chronic diseases, perform multi-dimensional intelligent quality control on standardized data, identify and process problematic data through algorithms, and output high-quality quality-controlled data after secondary quality control.
[0035] Step 3: Privacy and security integrated storage. After quality control, the data is encrypted, de-identified, and distributed for integrated storage. A data indexing and backup mechanism is built to achieve secure data storage and efficient management.
[0036] Step 4: Cross-institute data collaboration and sharing, hierarchical authorization for entities participating in multi-center research, and compliant cross-institute data sharing by combining privacy computing technology, while tracing the entire data usage process;
[0037] Step 5: Chronic disease-specific data fusion analysis, extract the three major chronic disease data features, construct a specific fusion model, achieve deep fusion of multi-center data through feature fusion similarity formula, explore data correlation patterns and output fusion analysis results;
[0038] Step 6: Visualize and provide feedback on quality control results. Compile and visualize the quality control results, push problematic data to medical institutions and review and correct the data, and optimize the quality control rule base to achieve continuous improvement in data quality.
[0039] Compared with the prior art, the beneficial effects of the present invention are:
[0040] This invention establishes specific data element standards and access specifications for liver disease, diabetes, and hypertension. Through heterogeneous data conversion technology, it transforms multi-source heterogeneous raw data into unified structured data, solving the problem of inconsistent data formats, naming, and coding among different medical institutions. This enables standardized collection and access of data for the three major chronic diseases, laying the foundation for multi-center data fusion.
[0041] This invention constructs a multi-dimensional quality control mechanism for three major chronic diseases, significantly improving data quality: Combining clinical diagnosis and treatment guidelines and disease characteristics of the three major chronic diseases, it constructs a dedicated multi-dimensional intelligent quality control rule base, designs core algorithm formulas such as outlier judgment and missing value imputation weights, and realizes full-dimensional intelligent quality control of data integrity, accuracy, consistency, timeliness, and standardization. It replaces traditional manual quality control methods, improves the objectivity and efficiency of quality control, effectively solves quality problems such as missing, abnormal, and duplicate data in the three major chronic diseases, and outputs high-quality data that meets the requirements of multi-center clinical research.
[0042] By employing privacy protection and distributed storage technologies, this invention ensures data storage security and efficient management: It performs end-to-end encryption and de-identification processing on data from three major chronic diseases, combines a distributed cloud storage architecture to achieve integrated data storage, and constructs data indexing, version management, and off-site backup mechanisms. This not only protects patient privacy and the data security of medical institutions, but also enables efficient retrieval, version backtracking, and reliable storage of multi-center data, avoiding the risk of data loss or damage.
[0043] By combining privacy-preserving computation and hierarchical authorization technologies, this invention enables compliant and collaborative data sharing across hospitals: It constructs a role-based hierarchical authorization management mechanism and combines privacy-preserving computation technologies such as federated learning and secure multi-party computation to achieve "usable but invisible" cross-hospital sharing of sensitive data for three major chronic diseases. This breaks down data silos between medical institutions and establishes a full-process traceability mechanism for data use, ensuring compliant data use and preventing data abuse and leakage. This provides a secure and compliant cross-hospital data sharing channel for multi-center research.
[0044] This invention constructs dedicated fusion models for three chronic diseases to achieve deep fusion analysis of multi-center data: Specifically, it develops dedicated data fusion analysis models for the disease characteristics of liver disease, diabetes, and hypertension, and designs data feature fusion similarity formulas. This achieves deep fusion from the data layer to the feature layer and then to the knowledge layer, uncovering clinical correlation patterns between the data of the three chronic diseases, rather than simply splicing general data. This enhances the targeting and depth of multi-center data fusion, providing valuable data analysis results for multi-center clinical research and treatment plan optimization for the three chronic diseases.
[0045] Establishing a closed-loop quality control result feedback and correction mechanism to promote continuous optimization of data quality: This invention enables multi-dimensional visualization of the quality control results of data for the three major chronic diseases, allowing users to quickly grasp the data quality status and specific problems; at the same time, it constructs a closed-loop problem data feedback and correction mechanism, pushing problem data to medical institutions in real time for review and correction, and dynamically optimizing the data quality by combining the data quality evaluation system and quality control rule base, thereby promoting medical institutions to improve the quality of data reporting and achieving continuous optimization of multi-center data quality for the three major chronic diseases. Attached Figure Description
[0046] Figure 1 This is a system module diagram of the present invention;
[0047] Figure 2 This is a flowchart of the method of the present invention;
[0048] Figure 3 This is a schematic diagram of the multi-source data standardization access module of the present invention;
[0049] Figure 4 This is a schematic diagram of the multi-dimensional data quality control module of the present invention;
[0050] Figure 5 This is a schematic diagram of the privacy and security fusion storage module of the present invention. Detailed Implementation
[0051] The technical solutions of the present invention will be clearly and completely described below with reference to the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of the present invention.
[0052] Please see Figure 1-5 This invention provides a multi-center data fusion and quality control system for chronic diseases, including a multi-source data standardization access module, a multi-dimensional data quality control module, a privacy and security fusion storage module, a cross-hospital data collaborative sharing module, a chronic disease-specific data fusion analysis module, and a quality control result visualization and feedback module.
[0053] The multi-source data standardization access module is used to realize the standardized collection and access of multi-source heterogeneous data of the three major chronic diseases in medical institutions at all levels, formulate data access specifications and data element standards for liver disease, diabetes and hypertension, standardize and convert raw data of different formats and structures, and output unified structured clinical data of the three major chronic diseases.
[0054] The multi-dimensional data quality control module is based on the disease characteristics and clinical diagnosis and treatment guidelines of the three major chronic diseases. It constructs a dedicated multi-dimensional intelligent quality control rule library to perform multi-dimensional quality control on the standardized data of the three major chronic diseases in terms of completeness, accuracy, consistency, timeliness, and standardization. Through algorithms, it realizes the automatic identification and processing of missing values, outliers, duplicate values, and logically contradictory values, and outputs high-quality quality-controlled data.
[0055] The privacy and security fusion storage module adopts privacy protection technology and distributed storage architecture to encrypt and fuse multi-center data of three major chronic diseases after quality control, establishes data version management and off-site backup mechanism, and realizes secure data storage, efficient retrieval and reliable management.
[0056] The cross-institutional data collaboration and sharing module constructs a hierarchical authorization mechanism for cross-institutional data collaboration and sharing. Combined with privacy computing technology, it realizes the "usable but not visible" sharing of sensitive data of the three major chronic diseases, provides compliant cross-institutional data access and use channels for multi-center research, and realizes full-process traceability of data use.
[0057] The chronic disease-specific data fusion and analysis module constructs a data fusion and analysis model specifically for the disease characteristics of liver disease, diabetes, and hypertension. It achieves deep fusion of feature layers and knowledge layers of multi-source, multi-dimensional, and cross-hospital data for the three chronic diseases, explores the correlation patterns between data, and provides data analysis results for multi-center clinical research.
[0058] The quality control results visualization and feedback module enables multi-dimensional visualization of the quality control results for the three major chronic diseases, and constructs a closed-loop quality control results feedback and correction mechanism. It pushes problematic data discovered during quality control to the data reporting medical institutions in real time, enabling timely correction of problematic data and continuous improvement of data quality.
[0059] The multi-source data standardization access module specifically includes a chronic disease data element standard setting unit, a multi-source data acquisition unit, a heterogeneous data conversion unit, and a data standardization verification unit;
[0060] The Chronic Disease Data Element Standardization Unit, based on national medical data standards such as the "Basic Dataset for Electronic Medical Records" and the "Basic Dataset for Clinical Laboratory Tests," and combining the disease characteristics and clinical treatment guidelines for liver disease, diabetes, and hypertension, has formulated specific data element standards for three major chronic diseases. These standards clearly define the name, definition, data type, value range, unit, and coding rules for each data element. Specifically, the liver disease data element covers patient basic information, liver function indicators, hepatitis virus biomarkers, liver imaging characteristics, medication records, and follow-up data; the diabetes data element covers patient basic information, blood glucose levels, glycated hemoglobin, insulin levels, diabetic complications, medication records, and dietary and exercise follow-up data; and the hypertension data element covers patient basic information, blood pressure levels, blood pressure classification, hypertension complications, medication records, and lifestyle follow-up data.
[0061] Multi-source data acquisition unit: Supports connection to electronic medical record systems, laboratory examination systems, image archiving and communication systems (PACS), follow-up management systems, and hospital information systems (HIS) of medical institutions at all levels. It provides acquisition methods for structured, semi-structured, and unstructured data, including direct database connection acquisition, API interface acquisition, and file upload acquisition, to achieve batch and real-time acquisition of multi-source heterogeneous raw data for the three major chronic diseases.
[0062] Heterogeneous Data Conversion Unit: Performs format conversion and structuring processing on the collected heterogeneous raw data. For unstructured data, it uses Natural Language Processing (NLP) technology for entity recognition, relation extraction, and structuring conversion; for semi-structured data, it unifies the data format and maps fields; and for data of the same type with different naming conventions, it unifies the naming and encoding according to the chronic disease data element standard, converting all heterogeneous data into unified structured data that conforms to the three major chronic disease data element standards.
[0063] Data standardization verification unit: Performs standardization verification on the converted structured data to verify whether the data conforms to the value range, data type, and coding rules of the chronic disease data element standard. Data that does not conform to the standard is marked and fed back to the data acquisition unit, where the data reporting agency makes corrections. After the verification is passed, standardized clinical data of the three major chronic diseases are output and transmitted to the multi-dimensional data quality control module.
[0064] The multi-dimensional data quality control module specifically includes a chronic disease-specific quality control rule base construction unit, a data integrity quality control unit, a data accuracy quality control unit, a data consistency quality control unit, a data timeliness and standardization quality control unit, and a problem data processing unit. This module is designed with two algorithm formulas, namely the outlier judgment formula and the missing value imputation weight calculation formula, to achieve intelligent quality control of data accuracy and integrity.
[0065] The Chronic Disease-Specific Quality Control Rule Base Construction Unit: Combining clinical treatment guidelines and physiological characteristics of liver disease, diabetes, and hypertension, a multi-dimensional quality control rule base is constructed specifically for these three chronic diseases, with rules for five quality control dimensions: completeness, accuracy, consistency, timeliness, and standardization. The completeness rule specifies the mandatory fields for each data element in the three chronic diseases (e.g., glycated hemoglobin is mandatory for diabetic patients); the accuracy rule sets the physiologically reasonable range for data based on medical common sense (e.g., the physiologically reasonable range for systolic blood pressure in hypertensive patients is 60-250 mmHg); the consistency rule requires the same data element from the same patient to remain consistent across different data sources (e.g., the patient's name and ID number are consistent in electronic medical records and follow-up data); the timeliness rule specifies the time requirements for data reporting and updating (e.g., follow-up data for the three chronic diseases must be reported within 72 hours after follow-up); and the standardization rule requires data to conform to the naming, coding, and unit requirements of chronic disease data element standards. The rule base also supports dynamic updates, allowing for the addition, modification, and deletion of quality control rules based on updates to clinical treatment guidelines and the needs of multi-center studies.
[0066] Data integrity quality control unit: Based on the integrity rules in the quality control rule base, it performs integrity checks on the standardized data of the three major chronic diseases, identifies missing values in the data, and calculates the missing rate of each data element; for missing values, it performs weighted imputation using the missing value imputation weight calculation formula, as follows:
[0067]
[0068] in: For the first Each reference data source assigns a weight to imputation of missing values, with values ranging from [0,1]. ; For the first The quality score of each reference data source is comprehensively evaluated based on the accuracy, consistency, and standardization of the data, and the value range is [0,1]. For the first The similarity coefficient between each reference data source and the data source containing the missing value is calculated by combining the patient's basic information and disease diagnosis and treatment information, and the value range is [0,1]. The number of reference data sources used to participate in missing value imputation; For reference data source index, =1,2,3,...,n.
[0069] After calculating the imputation weights of each reference data source according to the above formula, the weighted average method is used to impute numerical missing values, and the corresponding values of the reference data source with the highest weight are selected for imputation of categorical missing values; datasets with missing rates exceeding the preset threshold are marked and fed back to the problem data processing unit.
[0070] Data accuracy quality control unit: Based on the accuracy rules in the quality control rule base, it performs accuracy checks on the data of the three major chronic diseases, and identifies outliers in the data using an outlier determination formula, as follows:
[0071]
[0072] in The data value to be detected. The mean of the dataset containing the data element. The standard deviation of the dataset containing the data element. The outlier determination coefficient is set according to the clinical characteristics of different data elements of the three major chronic diseases, with a value of 2 to 3. The value is 2 for indicators with small physiological fluctuations, such as glycated hemoglobin, and 3 for indicators with large physiological fluctuations, such as random blood glucose.
[0073] When the data value to be tested meets the above formula, it is judged as an outlier; at the same time, a second verification is performed in combination with the physiologically reasonable range rules in the quality control rule base. Outliers that exceed the physiologically reasonable range are directly marked, and outliers that are within the physiologically reasonable range but meet Formula 1 are manually reviewed and confirmed; the identified outliers are transmitted to the problem data processing unit for processing.
[0074] Data Consistency Quality Control Unit: Based on the consistency rules in the quality control rule base, it performs multi-dimensional consistency checks on the data of the three major chronic diseases, including the consistency of basic information of the same patient, the numerical consistency of the same data element in different data sources, and the logical consistency of disease diagnosis and test indicators (such as the logical consistency between diabetes diagnosis results and glycated hemoglobin indicators). It uses data comparison algorithms to compare multi-source data field by field, identifies inconsistent data and marks it, and transmits it to the problem data processing unit.
[0075] Data Timeliness and Standardization Quality Control Unit: Based on the timeliness rules in the quality control rule base, it checks whether the reporting time and update time of the three major chronic disease data meet the requirements, identifies data with timeliness issues such as overdue reporting and failure to update in a timely manner; based on the standardization rules, it checks whether the naming, coding, and units of the data conform to the chronic disease data element standard, identifies non-standard data and marks it, and transmits it to the problem data processing unit.
[0076] Problem Data Processing Unit: Classifies and processes identified missing values, outliers, inconsistent values, time-sensitive values, and standardization-related values. Processing methods include automatic filling, automatic correction, manual review, and data removal. Secondary quality control is performed on the automatically processed data to ensure that the processed data meets the quality control requirements. After passing the quality control, high-quality post-quality control data for the three major chronic diseases is output and transmitted to the privacy and security fusion storage module.
[0077] The privacy and security converged storage module specifically includes a data encryption unit, a distributed converged storage unit, a data indexing and retrieval unit, a data version management unit, and an off-site backup unit;
[0078] Data encryption unit: Performs end-to-end encryption processing on the multi-center data of the three major chronic diseases after quality control. It uses symmetric encryption algorithms to encrypt and store patients' privacy-sensitive data (name, ID number, mobile phone number, address, etc.) and uses asymmetric encryption algorithms to encrypt cross-module data transmission. At the same time, it performs de-identification processing on patients' privacy-sensitive data and uses de-identification technology to replace the patient's unique identifier with an anonymous identifier, achieving dual protection of privacy data through encryption and de-identification processing.
[0079] Distributed fused storage unit: Adopting a distributed cloud storage architecture, a fused storage system for multi-center data of three major chronic diseases is constructed. The data is classified and stored according to "liver disease dataset", "diabetes dataset" and "hypertension dataset", and further classified by medical institution, data type and data collection time. The cross-hospital data of the three major chronic diseases are fused and stored in distributed storage nodes, realizing the physical dispersion and logical centralization of data. This not only improves the scalability of data storage, but also avoids data loss caused by single node failure.
[0080] Data Indexing and Retrieval Unit: A multi-dimensional retrieval index is constructed for the three major chronic disease data that are integrated and stored. The index dimensions include disease type, patient basic information, test indicators, treatment time, medical institution, etc. A combination of full-text search and precise search is adopted to achieve rapid retrieval and accurate location of multi-center data, meeting the data analysis needs of multi-center studies.
[0081] Data Version Management Unit: Establish a version management mechanism for the three major chronic disease data, record the version of data addition, modification and deletion operations, record the time, operator and operation content of each data operation, realize the data version traceability and recovery; at the same time, assign numbers to the data versions, clarify the applicable scenarios of each version of data, and ensure the consistency of data versions used in multi-center studies.
[0082] Off-site backup unit: Construct an off-site multi-copy backup mechanism to synchronously back up the three major chronic disease data in the integrated storage to storage nodes in different regions. The backup frequency is divided into real-time backup and scheduled backup. Real-time backup is used for core clinical data, and daily scheduled backup is used for non-core data. At the same time, the integrity of the backup data is checked regularly to ensure the availability of the backup data and avoid the risk of data loss or damage.
[0083] The cross-institute data collaboration and sharing module specifically includes a hierarchical authorization management unit, a privacy computing sharing unit, a data access and usage unit, and a data usage traceability unit;
[0084] Tiered Authorization Management Unit: Construct a role-based tiered authorization management mechanism, dividing data usage roles into different levels such as research administrators, clinical researchers, and medical institution data administrators, and assigning different access, usage, and modification permissions for the three major chronic diseases to different roles; at the same time, conduct qualification reviews of medical institutions participating in multi-center studies, and only those that pass the review can obtain data sharing permissions, so as to achieve refined control of data permissions.
[0085] Privacy Computation Sharing Unit: Combining privacy computing technologies such as federated learning and secure multi-party computation, this unit enables "usable but invisible" cross-hospital sharing of sensitive data related to the three major chronic diseases. It keeps the data of the three major chronic diseases of each medical institution locally, and only transmits the characteristic information of the data to the multi-center research platform through privacy computing technology. This enables joint analysis of cross-hospital data without leaking the original data, breaking down data silos and protecting patient privacy and the data security of medical institutions.
[0086] Data Access and Usage Unit: Provides authorized users and medical institutions with standardized channels for cross-hospital data access and usage, supports online data query, statistical analysis, and batch extraction, and performs secondary anonymization processing on extracted data to ensure privacy and security during data use; provides customized data extraction services for the specific needs of multi-center studies, improving the targeting and efficiency of data sharing.
[0087] Data Usage Traceability Unit: Construct a full-process traceability mechanism for cross-hospital data usage for the three major chronic diseases, and record the data access time, accesser, access content, purpose of use, and results of use throughout the process to form a data usage traceability log; the traceability log is permanently stored and can be queried and audited at any time to ensure the compliant use of data and avoid data abuse and leakage.
[0088] The chronic disease-specific data fusion and analysis module specifically includes a chronic disease data feature extraction unit, a chronic disease-specific fusion model construction unit, a multi-center data fusion and analysis unit, and a fusion result output unit. This module designs an algorithm formula for the fusion similarity formula of three major chronic disease data features, which is used to realize the feature layer fusion of multi-source data of the three major chronic diseases.
[0089] Chronic disease data feature extraction unit: Feature extraction is performed on the multi-center data of the three major chronic diseases after quality control. For numerical features, normalization is used to extract quantitative features, for categorical features, one-hot encoding is used to extract classification features, and for textual features, word vector model is used to extract semantic features. The unit also extracts the disease features of liver disease, diabetes and hypertension and the common features of cross-hospital data to provide a feature foundation for data fusion analysis.
[0090] Chronic Disease-Specific Fusion Model Construction Unit: Three dedicated data fusion and analysis models are constructed for liver disease, diabetes, and hypertension, respectively, addressing their disease characteristics and multi-center research needs. The liver disease fusion model focuses on fusing liver function test data, liver imaging features, and hepatitis virus biomarker data; the diabetes fusion model focuses on fusing blood glucose monitoring data, glycated hemoglobin data, complication data, and medication data; and the hypertension fusion model focuses on fusing blood pressure monitoring data, blood pressure classification data, complication data, and lifestyle data. These fusion models are constructed using deep learning algorithms, achieving deep fusion from the data layer to the feature layer and then to the knowledge layer.
[0091] Multi-center data fusion analysis unit: Based on the extracted features of the three major chronic diseases and a dedicated fusion model, it uses a similarity formula for the fusion of features of the three major chronic diseases to achieve feature-level fusion of cross-hospital and multi-source data. The formula is as follows:
[0092]
[0093] in: For the first The medical institution and the first The similarity of the three major chronic disease data features of the medical institutions is calculated, with a value range of [0,1]. The closer the value is to 1, the higher the feature similarity and the better the fusion. For the first Three major chronic disease data feature vectors of a medical institution For the first Feature vectors of three major chronic disease data from a medical institution; For the first The first medical institution 1 eigenvalue, For the first The first medical institution One eigenvalue; The number of dimensions for the three major chronic disease data features; , Index of participating healthcare institutions. , For feature index, =1,2,3,..., .
[0094] The similarity of data features among medical institutions is calculated based on the above formula. Features with high similarity are fused and aggregated, while features with low similarity are analyzed differently. At the same time, a knowledge layer fusion of multi-center data is achieved through a chronic disease-specific fusion model to uncover the correlation patterns between data, such as the correlation between liver function indicators and imaging features in liver disease patients, the correlation between blood glucose levels and medication in diabetic patients, and the correlation between blood pressure data and complications in hypertensive patients, providing data analysis results for multi-center clinical studies.
[0095] The fusion result output unit: It structures the results of multi-center data fusion analysis and outputs data in the form of statistical reports, feature association maps, and disease diagnosis and treatment pattern analysis reports. The fusion results are transmitted to the quality control result visualization and feedback module and stored in the privacy and security fusion storage module to provide data support for multi-center studies.
[0096] The quality control result visualization and feedback module specifically includes a quality control result statistics unit, a quality control result visualization display unit, a problem data push unit, a correction data review unit, and a data quality continuous optimization unit.
[0097] Quality control result statistics unit: Performs multi-dimensional statistical analysis on the quality control results of the multi-dimensional data quality control module. Statistical indicators include the missing rate, outlier rate, duplicate rate, consistency pass rate, standardization pass rate, and overall data quality pass rate of each data element for the three major chronic diseases. At the same time, it is grouped and statistically analyzed by medical institution, disease type, and data collection time to form a quality control result statistical data set.
[0098] The quality control results visualization unit uses visualization technology to present the statistical data of quality control results in an intuitive form. The display formats include bar charts, line charts, pie charts, heat maps, and data quality dashboards. It enables multi-dimensional visualization queries of quality control results and supports precise queries by medical institution, disease type, quality control dimension, and time range, allowing users to quickly grasp the overall quality status and specific quality issues of multi-center data for the three major chronic diseases.
[0099] Problem Data Push Unit: Constructs a closed-loop problem data feedback and correction mechanism. It classifies the unprocessed problem data identified by the multi-dimensional data quality control module according to medical institutions and pushes it to the data administrators of the corresponding medical institutions in real time through system messages, SMS, emails, etc., while also pushing specific information about the problem data and correction requirements.
[0100] The data review unit conducts a second quality control review on the revised data reported by medical institutions. The review includes whether the revision meets the data element standards for the three major chronic diseases, whether the original data quality problems have been resolved, and whether new quality problems have been generated. If the review is passed, the revised data will be updated to the privacy and security fusion storage module. If the review fails, the revision opinions will be fed back to the medical institutions, requiring them to make revisions again.
[0101] The Data Quality Continuous Optimization Unit conducts statistical analysis on the results of previous quality control and correction processes to uncover common patterns and high-frequency issues in data quality problems. It optimizes the quality control rule base of the multi-dimensional data quality control module for high-frequency quality issues and provides data quality improvement suggestions to medical institutions. It also establishes a multi-center data quality evaluation system for three major chronic diseases, regularly evaluating and ranking the data quality of various medical institutions to promote improvements in data reporting quality and achieve continuous data quality optimization.
[0102] The integration and quality control methods for a multicenter data fusion and quality control system for chronic diseases include the following steps:
[0103] Step 1: Standardized access to multi-source data, formulate data element standards for the three major chronic diseases, collect heterogeneous raw data from multiple sources and perform standardized conversion and verification, and output unified structured clinical data for the three major chronic diseases;
[0104] Step 2: Multi-dimensional data quality control. Construct three dedicated quality control rule bases for chronic diseases, perform multi-dimensional intelligent quality control on standardized data, identify and process problematic data through algorithms, and output high-quality quality-controlled data after secondary quality control.
[0105] Step 3: Privacy and security integrated storage. After quality control, the data is encrypted, de-identified, and distributed for integrated storage. A data indexing and backup mechanism is built to achieve secure data storage and efficient management.
[0106] Step 4: Cross-institute data collaboration and sharing, hierarchical authorization for entities participating in multi-center research, and compliant cross-institute data sharing by combining privacy computing technology, while tracing the entire data usage process;
[0107] Step 5: Chronic disease-specific data fusion analysis, extract the three major chronic disease data features, construct a specific fusion model, achieve deep fusion of multi-center data through feature fusion similarity formula, explore data correlation patterns and output fusion analysis results;
[0108] Step 6: Visualize and provide feedback on quality control results. Compile and visualize the quality control results, push problematic data to medical institutions and review and correct the data, and optimize the quality control rule base to achieve continuous improvement in data quality.
[0109] Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions described in the foregoing embodiments or make equivalent substitutions for some of the technical features. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.
Claims
1. A multi-center data fusion and quality control system for chronic diseases, characterized by: It includes a multi-source data standardization access module, a multi-dimensional data quality control module, a privacy and security fusion storage module, a cross-hospital data collaborative sharing module, a chronic disease-specific data fusion analysis module, and a quality control result visualization and feedback module; The multi-source data standardization access module enables the standardized collection and access of multi-source heterogeneous data for the three major chronic diseases, and outputs unified structured clinical data for the three major chronic diseases. The multi-dimensional data quality control module performs multi-dimensional intelligent quality control on the standardized data of the three major chronic diseases, and outputs high-quality quality-controlled data. The privacy and security fusion storage module encrypts and fusions the multi-center data of the three major chronic diseases after quality control, so as to achieve secure data storage and efficient management. The cross-hospital data collaboration and sharing module enables compliant cross-hospital sharing of sensitive data for the three major chronic diseases, while also enabling full-process traceability of data usage; The chronic disease-specific data fusion and analysis module enables in-depth fusion and analysis of multi-center data for the three major chronic diseases, and uncovers the correlation patterns between the data. The quality control result visualization and feedback module enables the visualization and closed-loop feedback correction of quality control results, thereby promoting continuous improvement in data quality.
2. The chronic disease multicenter data fusion and quality control system according to claim 1, characterized in that: The multi-source data standardization access module includes a chronic disease data element standard setting unit, a multi-source data acquisition unit, a heterogeneous data conversion unit, and a data standardization verification unit; The chronic disease data element standard setting unit formulates exclusive data element standards based on national medical data standards and combined with the characteristics of the three major chronic diseases, clarifying the name, definition, data type, value range, unit, and coding rules of the data elements; the multi-source data acquisition unit supports multiple acquisition methods to realize batch and real-time acquisition of multi-source heterogeneous raw data of the three major chronic diseases; the heterogeneous data conversion unit performs format conversion and structuring processing on the heterogeneous raw data, outputting structured data that conforms to the chronic disease data element standards; the data standardization verification unit performs standardization verification on the converted structured data, and after passing the verification, it is transmitted to the multi-dimensional data quality control module.
3. The chronic disease multicenter data fusion and quality control system according to claim 1, characterized in that: The multi-dimensional data quality control module includes a chronic disease-specific quality control rule base construction unit, a data integrity quality control unit, a data accuracy quality control unit, a data consistency quality control unit, a data timeliness and standardization quality control unit, and a problem data processing unit; The chronic disease-specific quality control rule base construction unit combines the three major chronic disease clinical diagnosis and treatment guidelines to build a multi-dimensional quality control rule base, which supports dynamic updates; the data integrity quality control unit uses a missing value imputation weight calculation formula to identify and process missing values; the data accuracy quality control unit uses an outlier determination formula to identify and verify outliers. The problem data processing unit classifies and processes various types of problem data and performs secondary quality control. After passing the quality control, the data is transmitted to the privacy and security fusion storage module.
4. The chronic disease multicenter data fusion and quality control system according to claim 3, characterized in that: The outlier determination formula is as follows: ,in The data value to be detected. The mean of the dataset containing the data element. The standard deviation of the dataset containing the data element. The outlier determination coefficient is 2 to 3; the formula for calculating the missing value imputation weight is as follows: ,in For the first Filling weights for each reference data source, For the first Quality ratings of reference data sources For the first The similarity coefficient between each reference data source and the data source containing the missing values. For reference, the number of data sources, This serves as an index for the reference data source.
5. The chronic disease multicenter data fusion and quality control system according to claim 1, characterized in that: The privacy and security integrated storage module includes a data encryption unit, a distributed integrated storage unit, a data indexing and retrieval unit, a data version management unit, and an off-site backup unit; The data encryption unit performs end-to-end encryption and desensitization processing on the data of the three major chronic diseases; the distributed fusion storage unit adopts a distributed cloud storage architecture to realize the classified and fusion storage of data; the data indexing and retrieval unit constructs a multi-dimensional retrieval index to realize fast data retrieval; the data version management unit realizes version recording and backtracking of data operations. The off-site backup unit establishes an off-site multi-copy backup mechanism to ensure reliable data storage.
6. The chronic disease multicenter data fusion and quality control system according to claim 1, characterized in that: The cross-institute data collaboration and sharing module includes a hierarchical authorization management unit, a privacy computing sharing unit, a data access and usage unit, and a data usage traceability unit; The hierarchical authorization management unit constructs a role-based hierarchical authorization management mechanism to achieve refined control of data permissions; the privacy computing sharing unit combines privacy computing technology to achieve "usable but invisible" sharing of data on the three major chronic diseases; the data access and use unit provides authorized users with standardized data access and use channels; and the data use traceability unit constructs a full-process data use traceability mechanism to form a permanently stored traceability log.
7. The chronic disease multicenter data fusion and quality control system according to claim 1, characterized in that: The chronic disease-specific data fusion and analysis module includes a chronic disease data feature extraction unit, a chronic disease-specific fusion model construction unit, a multi-center data fusion and analysis unit, and a fusion result output unit. The chronic disease data feature extraction unit extracts multiple types of features from the three major chronic disease data to form disease features and common features; the chronic disease-specific fusion model construction unit constructs a deep fusion analysis model specific to the three major chronic diseases respectively; the multi-center data fusion analysis unit uses the feature fusion similarity formula of the three major chronic disease data to realize the fusion of data features; the fusion result output unit outputs and stores the fusion analysis results after structured processing.
8. The chronic disease multicenter data fusion and quality control system according to claim 7, characterized in that: The similarity formula for fusing the data features of the three major chronic diseases is as follows: ,in To achieve similarity in data feature fusion among medical institutions, , For the feature vectors of different medical institutions, , For the corresponding eigenvalues, The number of feature dimensions, , Index for medical institutions, For feature indexing.
9. The chronic disease multicenter data fusion and quality control system according to claim 1, characterized in that: The quality control result visualization and feedback module includes a quality control result statistics unit, a quality control result visualization display unit, a problem data push unit, a correction data review unit, and a data quality continuous optimization unit. The quality control result statistics unit performs multi-dimensional statistical analysis on the quality control results to form a statistical data set; the quality control result visualization display unit realizes the visualization display and query of the quality control results in various forms; The problem data push unit categorizes and pushes problem data to the corresponding medical institutions; The data correction auditing unit performs a second quality control audit on the corrected data; the data quality continuous optimization unit mines the patterns of quality problems, optimizes the quality control rule base, and establishes a data quality evaluation system.
10. The fusion and quality control method of the chronic disease multicenter data fusion and quality control system according to any one of claims 1-9, characterized in that: Includes the following steps: Step 1: Standardized access to multi-source data, formulate data element standards for the three major chronic diseases, collect heterogeneous raw data from multiple sources and perform standardized conversion and verification, and output unified structured clinical data for the three major chronic diseases; Step 2: Multi-dimensional data quality control. Construct three dedicated quality control rule bases for chronic diseases, perform multi-dimensional intelligent quality control on standardized data, identify and process problematic data through algorithms, and output high-quality quality-controlled data after secondary quality control. Step 3: Privacy and security integrated storage. After quality control, the data is encrypted, de-identified, and distributed for integrated storage. A data indexing and backup mechanism is built to achieve secure data storage and efficient management. Step 4: Cross-institute data collaboration and sharing, hierarchical authorization for entities participating in multi-center research, and compliant cross-institute data sharing by combining privacy computing technology, while tracing the entire data usage process; Step 5: Chronic disease-specific data fusion analysis, extract the three major chronic disease data features, construct a specific fusion model, achieve deep fusion of multi-center data through feature fusion similarity formula, explore data correlation patterns and output fusion analysis results; Step 6: Visualize and provide feedback on quality control results. Compile and visualize the quality control results, push problematic data to medical institutions and review and correct the data, and optimize the quality control rule base to achieve continuous improvement in data quality.