Personal health record data processing method and system under metadata driving
Through metadata-driven personal health record data standards and ETL processes, combined with data warehouses and system parallel computing framework, the shortcomings of existing systems in data specifications and processing methods are solved, efficient integration and unified management of multi-source heterogeneous health data are achieved, and intelligent and personalized decision-making support for health management is improved.
Patent Information
- Application Number
- CN202510190639.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-20
- Publication Date
- 2025-06-17
AI Technical Summary
Existing systems have shortcomings in data specifications, processing methods and system architecture, and it is difficult to effectively integrate, exchange and utilize multi-source heterogeneous health data, affecting the intelligent and personalized decision-making support of health management.
By building metadata-driven personal health profile data standards and specifications, designing standardized forms, integrating heterogeneous data sources using ETL processes, using data warehouses for data mining and analysis, and ensuring efficient data processing and security through the system parallel computing framework and security and privacy guarantee module.
It realizes effective integration and unified management of multi-source heterogeneous health data, improves data quality and consistency, supports efficient storage and distributed processing of structured and unstructured data, and provides personalized health management and disease prediction models to ensure efficient and automation of processing processes.
Smart Images

Figure CN120164562A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the fields of computer big data processing and blockchain technology. The present invention proposes a method and system for processing personal health record data based on a data warehouse driven by metadata. Background Art
[0002] With the development of social economy and the improvement of people's living standards, the concept of proactive health has gradually emerged, and personal health management has received increasing attention. Proactive health emphasizes that individuals achieve disease prevention, early intervention, and personalized treatment by long-term monitoring and maintaining their own health data. In this context, the Personal Health Record (PHR) has emerged as the times require. The personal health record covers various information such as an individual's physical examination records, medical service records, and daily health monitoring data. These data usually come from different medical institutions, health devices, and management platforms, showing the characteristics of diverse data sources and heterogeneous formats. As an important tool for recording and integrating individual health data, the personal health record plays a crucial role in improving the efficiency and quality of health management.
[0003] Given the characteristics of diverse data sources and heterogeneous formats of PHR, how to effectively integrate, exchange, and utilize these data has become the core issue in the field of health big data processing. Efficient data integration and exchange can not only break data silos and achieve data interconnection and interoperability, but also provide comprehensive and accurate decision-making support for medical services and health management.
[0004] Among them, one of the key technologies for data integration and exchange is ETL (Extract, Transform, Load). The ETL process includes data extraction (Extract), transformation (Transform), and loading (Load). Its core purpose is to extract, standardize, clean, and load data from different data sources into a target data warehouse. The ETL process enables multi-source heterogeneous data to be stored in a unified format, providing high-quality data support for subsequent analysis and applications.
[0005] However, there are many deficiencies in existing data processing systems. First, the data specifications are not unified. There are significant differences in the data formats, naming conventions, coding methods, etc. adopted by various medical institutions and health platforms, resulting in difficulties in seamless data integration. Second, existing data processing methods are often relatively single, failing to select appropriate ETL technologies according to specific scenarios to effectively cover multiple links such as data cleaning, data transformation, data reduction, and data integration. This lack of flexibility and adaptability in the processing method has led to a decline in data quality, which in turn affects subsequent data analysis, model training, and decision support. In addition, in terms of system architecture design, many data processing systems lack flexibility and scalability, unable to adapt to the processing requirements of massive heterogeneous health data, and it is also difficult to achieve efficient processing and analysis of real-time data.
[0006] Against this background, how to address the deficiencies of existing systems in terms of data specifications, processing methods, and system architecture, achieve efficient integration, exchange, and utilization of multi-source health data, and provide intelligent and personalized decision support for health management has become an urgent problem to be solved in the field of personal health record data processing. Summary of the Invention
[0007] This application provides a method and system for processing personal health record data based on a data warehouse driven by metadata to overcome the problems raised in the background art.
[0008] Specifically, the technical solution of the present invention is realized through the following basic concepts:
[0009] A method for processing personal health record data based on a data warehouse driven by metadata, characterized by including the following steps:
[0010] Step S110: Construct a shared document file for personal health record data standards and specifications, design corresponding chapter templates, item templates, and component templates, determine specific data elements and coding rules, and design standardized forms for each heterogeneous data source according to the specifications to ensure the uniformity of data entry for each data source;
[0011] Step S120: Integrate data from multiple heterogeneous data sources through the ETL process, and use data extraction, transformation, and loading technologies to ensure that data from different systems can be stored in a unified format and support subsequent analysis;
[0012] Step S130: Utilize the integrated data in the data warehouse and apply various analysis tools and algorithms to perform data mining, trend analysis, and health assessment on personal health records;
[0013] Step S140: Present the analysis results to the user in the form of charts or reports.
[0014] Preferably, the specification designs chapter templates with themes such as public health, medical services, health management, biological and omics data, environmental data, etc. The first-level directories include, but are not limited to, resident health records, electronic medical records, medication records, health examinations, lifestyle behaviors, psychological assessments, genomes, proteomes, microbiota, descriptions of living environments, environmental factor impacts, etc. The heterogeneous data sources include, but are not limited to, resident health record management systems, various medical business systems, various open platforms of embedded institutions, various wearable devices, etc.
[0015] Preferably, the data ETL process includes:
[0016] Step S210: Select a suitable extraction method according to the real-time nature and update frequency of the data;
[0017] Step S220: Perform processing operations such as cleaning, transformation, and integration on the extracted data to eliminate the differences between the original data set and the standard data form, and ensure data quality and consistency;
[0018] Step S230: Adopt an incremental loading method and introduce timestamp and event marking technologies to ensure that only the data that has changed since the last update is loaded in a single query, improve the efficiency of data loading, and maintain the timeliness of the data.
[0019] Preferably, the data processing operations include, but are not limited to:
[0020] Step S310: Eliminate outliers and invalid data in the data through rule verification and data cleaning;
[0021] Step S320: Unify the format and standardize the original data to ensure that the fields and types are consistent;
[0022] Step S330: Use feature screening and data simplification methods to retain the key information in the data;
[0023] Step S340: Integrate multi-source data through data association and fusion technologies to generate a standardized integration form.
[0024] Preferably, the data analysis process is supported by data mining and data analysis technologies, and the supported services include, but are not limited to, related services in aspects such as health management and services, public health and disease management, industry innovation and R & D, supervision and policy support, etc.
[0025] A personal health record data processing system based on a data warehouse driven by metadata, characterized in that the personal health record data integration system is sequentially connected to a data source layer, a data transmission layer, a data storage layer, a programming model layer, a data analysis layer, and a data application layer, and includes a system parallel computing framework (1), an archive data processing module (2), a security and privacy protection module (3), and a system monitoring and operation and maintenance management module (4).
[0026] Preferably, the system parallel computing framework (1) includes:
[0027] ETL toolkit (11): responsible for extracting, transforming, and loading multi-source health data;
[0028] Data storage medium (12): supports structured and unstructured data storage and provides an efficient and reliable access interface;
[0029] Distributed computing framework (13): realizes distributed storage and processing of large-scale data and supports batch and real-time processing;
[0030] Intelligent analysis and decision-making platform (14): conducts health data analysis and provides personalized health management and disease prediction models;
[0031] Task scheduling system (15): coordinates and schedules each task to ensure the efficiency and automation of the processing flow;
[0032] Containerization and microservice module (16): modular deployment and management to ensure system scalability and convenient maintenance.
[0033] Preferably, the archive data processing module (2) includes a data collection and transmission module (21), a data processing and integration module (22), a data storage and management module (23), a data analysis and intelligent decision-making module (24), a visualization and report generation module (25), etc.
[0034] Preferably, the security and privacy protection module (3) includes:
[0035] Privacy protection module (31): combines a certain privacy protection mechanism, such as data encryption, permission control, etc., to ensure the privacy of personal health data during the collection, storage, and use processes;
[0036] Log and monitoring module (32): responsible for recording and monitoring data access, operation records, and abnormal behaviors in the system to ensure the transparency and traceability of data use and timely discover security risks.
[0037] Preferably, the system monitoring and operation and maintenance management module (4) includes:
[0038] Performance Monitoring Module (41): Real-time monitors the operating status and performance metrics of the system to ensure the efficiency of the data processing process and the stable operation of the system;
[0039] Fault Handling Module (42): Automatically detects and handles system faults, reduces downtime through warning and recovery mechanisms, and ensures the continuity and reliability of the system.
[0040] Advantages of the present invention: The method and system for processing personal health record data driven by metadata and based on a data warehouse proposed by the present invention establish a unified health data standard specification, design first-level directories such as resident health records, electronic medical records, medication records, and genomes with public health, medical services, health management, biological and omics data, environmental data, etc. as themes, and ensure the effective integration of heterogeneous data sources such as resident health record management systems, medical business systems, and wearable devices. The system architecture includes a data source layer, a data transmission layer, a data storage layer, a programming model layer, a data analysis layer, and a data application layer, covering a parallel computing framework, an archive data processing module, a security and privacy protection module, and a system monitoring module, realizing the extraction, cleaning, transformation, and integration of multi-source data, supporting the efficient storage and distributed processing of structured and unstructured data, and providing personalized services and decision-making support for different entities through an intelligent analysis and decision-making platform to ensure the efficiency and automation of the processing process. Description of the Drawings
[0041] In order to more clearly illustrate the technical solutions in the embodiments of this specification or the prior art, the following will briefly describe the drawings required for use in the description of the embodiments of this specification or the prior art. Obviously, the following only shows the drawings required for some embodiments of this application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0042] Figure 1 provided by the first embodiment of this specification
[0043] Figure 1 is a schematic flow chart of the method for processing personal health record data;
[0044] Figure 2 is a schematic flow chart of the data ETL process;
[0045] Figure 3 is a schematic flow chart of the data processing operation;
[0046] Figure 4 is a system block diagram of the personal health record data processing system. Detailed Embodiments
[0047] To enable those skilled in the art to better understand the technical solutions in this specification, the technical solutions in the embodiments of this specification will be clearly and completely described below in conjunction with the accompanying drawings. Obviously, the embodiments involved in the specific implementation manners are only a part of the embodiments of this application, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the specific implementation manners without creative efforts shall fall within the protection scope of this application.
[0048] The terms used in the embodiments of this application are only for the purpose of describing specific embodiments, and are not intended to limit this application. The singular forms "a" and "the" used in this application and the appended claims are also intended to include the plural forms, unless the context clearly indicates otherwise. It should also be understood that the term "and / or" used herein refers to and includes any or all possible combinations of one or more of the associated listed items.
[0049] The first embodiment of this specification (hereinafter referred to as "Embodiment 1") provides a method for accessing health data based on blockchain. The execution subjects of Embodiment 1 include, but are not limited to, terminals or servers (including blockchain node servers) or operating systems or application programs. That is, the execution subjects can be diverse and can be set, used, or transformed according to needs. In addition, a third-party application program can assist the execution subject in executing Embodiment 1. For example, the method in Embodiment 1 can be executed by a server, and a corresponding application program can be installed on a terminal (which can be held by a user). Data transmission can be carried out between the terminal or the application program and the server to assist the server in executing the method in Embodiment 1.
[0050] As Figure 1 and Figure 2 shown, the method for processing personal health record data driven by metadata provided in Embodiment 1 includes:
[0051] S101:
[0052] Please refer to Figures 1 to 3 , in an embodiment of the present invention, a method for processing personal health record data driven by metadata includes the following steps:
[0053] Step S110: Construct a shared document file for the personal health record data standard specification, design corresponding chapter templates, item templates, and component templates, and design a standardized form for each heterogeneous data source according to the specification to ensure the uniformity of data entry for each data source;
[0054] Step S120: Integrate data from multiple heterogeneous data sources through an ETL process to ensure that the data is stored in a unified format and construct a complete personal health record;
[0055] Step S130: Apply a variety of analysis tools and algorithms to perform data mining, trend analysis, and health assessment on the personal health record;
[0056] Step S140: Present the analysis results to the user in the form of charts or reports.
[0057] Through the above technical solution, this embodiment realizes the effective integration and unified management of multi-source heterogeneous health data. The constructed standardized data specification provides a clear framework for the integration of heterogeneous data sources, ensuring the consistency and comparability of the input and storage of various types of health data, and improving the data quality. The application of the ETL process solves the heterogeneity problem in the data cleaning, transformation, and loading processes, ensuring the seamless integration and unified storage of data between different systems. The analysis and mining of the integrated data in the data warehouse can efficiently perform health trend assessment and personalized health decision support, and finally visually display the analysis results to the user through visualization tools, thereby improving the intelligence and convenience of health management.
[0058] In a possible implementation manner, the specification designs chapter templates with themes such as public health, medical services, health management, biological and omics data, environmental data, etc. The first-level directories include, but are not limited to, resident health records, electronic medical records, medication records, health examinations, lifestyle behaviors, psychological assessments, genomes, proteomes, microbiota, descriptions of living environments, environmental factor impacts, etc. The heterogeneous data sources include, but are not limited to, resident health record management systems, various medical business systems, various open platforms of embedded institutions, various wearable devices, etc.
[0059] Shared documents are the digitization of paper clinical reports generated in daily medical activities, enabling machine-readable, reducing the storage, transmission, and reading of cumbersome paper documents, and enabling semantic-level exchange and sharing between heterogeneous systems. The shared document specification is based on a mature foreign general architecture and, on the premise of meeting the actual needs of Chinese health information shared documents, uses data elements and data sets to standardize and constrain the data elements in the health information shared documents, uses the template library constraint as a means to normatively describe the specific business content of the health information shared documents, and uses value domain codes as standards to normatively record the coded data elements in the health information shared documents, thereby clearly showing the business context of specific application documents and the mutual relationship between data elements, and supporting higher-level semantic interoperability.
[0060] The shared document specification standardizes the document architecture and document templates. The document architecture defines the format of exchanged data, encoding rules, error handling mechanisms, etc., which can be designed with reference to international health data exchange standards such as HL7 V3 and HL7 FHIR. The present invention does not limit this aspect. The document template, aiming at domestic medical activities and business scenarios, further restricts the business structure and semantics on the basis of the document architecture, so as to form a reusable and nested template library. According to the shared document structure, templates can be divided into document templates, chapter templates, item templates, and component templates. The document template restricts the chapters and items that a certain type of clinical document should contain according to specific business requirements; the chapter template is a further subdivision of the document template; the item template is a further refined expression of the concepts appearing in the chapter template and can be defined nestedly; the component template is the smallest constituent unit of the document template, which is used to describe and restrict in detail the identifier, name, definition, data type, representation format, allowed values, etc. of data elements, and can be analogized to a single field in a form, responsible for specifying the specific data format and valid value range of this field.
[0061] It should be noted that through systematic literature research and collation, this embodiment has deeply explored the personal health record data sources and data attribution themes, and developed detailed chapter templates, item templates, and component templates. As a comprehensive record of individual health data, the personal health record covers multi-dimensional information such as physical examinations, medical treatments, and health management, providing comprehensive support for health management. On this basis, this embodiment selects key data themes related to health management to form a systematic chapter design, including five major categories: public health, medical services, health management, biological and omics data, and environmental data. These themes fully reflect the extensive connotations and diverse data sources of the personal health record. In the chapter design, the item templates and component templates are divided according to further health information requirements. The first-level items include:
[0062] (1) Resident health record, which records the individual's basic personal information and health history information, and can be specifically divided into secondary items such as basic information, disease history, allergy history, family health history, and current medical history;
[0063] (2) Electronic case record, which records detailed medical data during the medical treatment process, and can be specifically divided into secondary items such as diagnosis information, medical images, patient medical records, and hospitalization records;
[0064] (3) Medication record, which records the usage of drugs associated with a single case;
[0065] (4) Health detection, which records daily health monitoring data, and can be specifically divided into 17 secondary items of common health indicators such as blood pressure and heart rate;
[0066] (5) Lifestyle behaviors, recording daily habits related to health, can be specifically divided into secondary entries such as eating habits, exercise records, sleep quality, lifestyle, etc.;
[0067] (6) Psychological assessment, recording the mental health status of an individual, can be specifically divided into secondary entries such as depressive symptoms, anxiety levels, life satisfaction, etc.;
[0068] (7) Genome, proteome, microbiota, recording personal biological and omics data for disease risk prediction and personalized health management;
[0069] (8) Description of living environment, recording the basic information of an individual's living environment;
[0070] (9) Influence of environmental factors, reflecting the potential impact of the environment on health.
[0071] To ensure the integrity and consistency of data integration, the present invention makes a detailed division of the data attribution and clarifies the heterogeneous data source types, specifically including:
[0072] (1) Resident health record management system; its data adopts standardized medical data formats such as HL7, ICD coding, etc., covering information such as health examinations, diagnoses, treatment histories, etc. The data is usually updated regularly and has a high degree of content structuring;
[0073] (2) Various medical business systems, such as the hospital's electronic medical record system (EMR), prescription management system, inspection and examination system, etc.; their data formats may be inconsistent, with different coding systems (such as ICD-10, CPT), and may contain a large amount of unstructured data (such as doctor's notes, imaging data, etc.);
[0074] (3) Various embedded institutional open platforms, such as health management platforms, community health service platforms, etc.; these platforms usually integrate a variety of device data, which may include medical activity monitoring data and environmental monitoring data. The formats are not unified and involve a large amount of sensor data. The data volume is huge and updated frequently;
[0075] (4) Various wearable devices, such as smart bracelets, smart watches, etc., collecting health behavior data such as heart rate, steps, sleep data, etc. These data are usually real-time streaming data, and the formats may be custom JSON or CSV formats.
[0076] These data sources not only ensure the comprehensiveness of health data, but also guarantee the standardization and usability of the data through the refined design of entries and component templates, ensuring that they can effectively meet the needs of personal health management and medical services.
[0077] Meanwhile, taking the secondary entry of "Basic Information" as an example, the specific design of the component template in the PHR sharing document specification is further illustrated. The data elements covered under the "Basic Information" entry of the present invention include: unique identifier, file number, name, ID number, ID type code, date of birth, work unit, personal phone number, contact person's name, contact person's phone number, permanent residence type code, household registration type code, ethnic group code, blood type code, education level code, political status code, occupation code, etc. Among them, the permanent residence type code is a data element with an identifier of residence_type_code, which is used to identify the permanent residence type of an individual and reflect the attributes or categories of their actual place of residence. This data element adopts the S3 data type (S3 indicates that this data element is character type and belongs to the form of a code table, and the specific coding values are preferably from a predefined coding table). Its representation format is N2 (N represents numeric characters, and 2 indicates that the length of the data element is 2-digit numbers). The value range codes include: 01 represents urban residents, 02 represents rural residents, 03 represents urban floating population, 04 represents rural floating population, and 99 represents unknown. Correspondingly, the component template of the permanent residence type code should complete the definition of the above content and embed it into the PHR sharing document specification to achieve standardized data storage and sharing.
[0078] In a possible implementation manner, the data ETL process includes:
[0079] Step S210: Select an appropriate extraction method according to the real-time nature and update frequency of the data;
[0080] Step S220: Perform processing operations such as cleaning, transforming, and integrating the extracted data to eliminate the differences between the original data set and the standard data form, and ensure data quality and consistency;
[0081] Step S230: Adopt the method of incremental loading, introduce timestamp and event marking technologies, ensure that only the data that has changed since the last update is loaded in a single query, and improve the efficiency of data loading and maintain the timeliness of the data.
[0082] In a possible implementation manner, the data processing operations include but are not limited to:
[0083] Step S310: Eliminate outliers and invalid data in the data through rule verification and data cleaning;
[0084] Step S320: Uniformly format and standardize the original data to ensure that the fields and types are consistent;
[0085] Step S330: Use feature screening and data simplification methods to retain the key information in the data;
[0086] Step S340: Integrate multi-source data through data association and fusion technologies to generate a standardized integrated form.
[0087] Among them, data extraction refers to the process of obtaining the required data from multiple data sources and importing it into a data processing system. Data extraction can be carried out based on two methods: documents or intermediate libraries. The data extraction method based on documents is to regularly generate and read standardized documents from different data sources, and import the data in these documents into the target system as needed. It is suitable for the import of data sources that are not frequently updated or historical archive data, such as static information like residents' health records and electronic medical records. The advantage of this method is that it is relatively simple to implement, easy to transfer standardized data between heterogeneous systems, especially with high adaptability in cross-institutional data exchange.
[0088] Another method is data extraction based on an intermediate library, which establishes an intermediate data warehouse to monitor and synchronize the data updates between the data source and the target system in real time. This method is suitable for data that needs to be updated frequently and processed in real time, such as health monitoring data, medication records, and daily health behaviors. Data extraction based on an intermediate library can effectively process dynamically updated data. Through incremental loading and event-driven mechanisms, it ensures the timeliness and consistency of data, while avoiding the repeated extraction of full-volume data and improving the data processing efficiency.
[0089] According to the characteristics of the data in the personal health record, for dynamic data that needs to be continuously updated, such as public health management and health assessment, the extraction method based on an intermediate library is adopted. For relatively static or infrequently changing omics data such as genomics, proteomics, and microbiota, the method based on documents is used for regular extraction and import. This design of selecting appropriate extraction methods according to different data characteristics ensures the flexibility and efficiency of the system.
[0090] At the same time, data transformation refers to the process of converting the original data from its original format or structure into a standardized form that meets the system requirements for subsequent analysis and processing. The process of data transformation usually involves various data processing operations, aiming to improve the quality and consistency of data. Commonly used processing operations include data cleaning, data reduction, data transformation, and data integration, etc.
[0091] Data cleaning aims to remove noise and outliers from data, such as operations like null value filling, duplicate data deletion, and correction of data with inconsistent formats. Through rule verification and consistency checks, it can be ensured that the data meets the accuracy requirements specified by the system before storage. For handling missing values, after classifying the data, the most suitable method for filling missing values can be selected according to the characteristics of different types of data. For quantitative data (such as weight, blood pressure, etc.), prediction models based on deep learning, such as regression or LSTM algorithms, can be used to predict and fill the missing values by leveraging the correlation of data with potential association patterns or historical data, which can effectively handle numerical missing values. For categorical data (such as disease types, allergy history, etc.), an intelligent filling mechanism based on knowledge graphs and reasoning is adopted, and reasonable inferences are made by analyzing the patient's health status and other known information. For example, by combining the patient's weight, blood sugar, and family medical history, it can be intelligently inferred that the patient may have diabetes or hypertension. For some data that is missing under specific circumstances, the system can also use external data sources, such as combining public health databases, regional health statistics data, etc., to provide a basis for filling the missing health indicators, especially in cases where certain group or environmental factors are missing.
[0092] The purpose of data reduction is to simplify the data structure, reduce the data dimension and redundancy, so as to improve the data processing efficiency and reduce the storage and computing burden. By methods such as feature selection, dimensionality reduction techniques (such as principal component analysis PCA), and data aggregation, the most representative feature information in the data is retained. For example, in the process of processing health behavior data, the system automatically identifies key indicators closely related to health status, such as weight, eating habits, exercise frequency, etc., through a rule-based feature selection algorithm, and then effectively compresses the above data in combination with data reduction techniques, thus optimizing the subsequent data storage and analysis process. Innovatively, the data reduction tool of the present invention can be optimized and adjusted according to different data analysis requirements to meet the requirements of automation and intelligence. Specifically, this system can also adopt an adaptive dimensionality reduction method based on deep learning to dynamically identify the characteristics of different data and automatically adjust the dimensionality reduction strategy according to the changes in the data and the analysis target. For example, for different combinations of input health data, the system automatically selects appropriate dimensionality reduction algorithms (such as autoencoders, variational autoencoders, etc.), and adjusts the number of layers and compression ratio of dimensionality reduction in real time through training the model to adapt to the changing data characteristics and analysis requirements. This method can effectively improve the accuracy and storage efficiency of data processing, and is especially suitable for processing large-scale health data from multiple heterogeneous data sources.
[0093] Data transformation refers to the conversion of data in terms of format or encoding to conform to the standards of the target system. Common operations include unit conversion, data type conversion, encoding mapping, etc. In the personal health record data processing system, data transformation usually involves uniformly converting health data from different data sources into a standard format. For example, the system maps disease codes from different medical systems to a unified disease classification standard (such as the ICD-10 coding standard), or unifies the body temperature data output by different devices into Celsius or Fahrenheit to ensure the availability of data for subsequent analysis.
[0094] Data integration is the final step of data transformation. By fusing heterogeneous data from different data sources, a standardized integrated form is generated. Data integration usually involves operations such as primary key matching, foreign key association, and data table merging to ensure the effective integration of multi-source data within the same system. For example, after the electronic medical record data from the medical business system and the health monitoring data from wearable devices are fused by associating personal identifiers, a complete personal health record can be generated to support subsequent health assessment and personalized medical services.
[0095] In one embodiment, considering the uneven reliability and integrity of heterogeneous data from different sources, a reference data, i.e., the complete validity, needs to be defined for each integrated form to illustrate the completeness or effectiveness of the integrated data. Among them, assuming that the data sources of the integrated form data are N, through primary key matching, the primary key matching completion degree is N1, and N1 = n1 / N, where n1 is the number of successfully matched ones. If the data from all N sources can be matched, then n1 = N and the final value of N1 is 1, indicating that all heterogeneous data sources are effectively available; further referring to the foreign key associated data, the foreign key association degree is N2. If the number of existing foreign key associated data is n2, then the foreign key association completeness N2 = n2 / N. If all the foreign key associated data exist, then N2 = 1, indicating a relatively high completeness of heterogeneous data. If only a small amount of foreign key associated data exists, then n2 < N; the final form data completeness = N1 + N2. The higher this data value is and the closer it is to 2, the more complete, effective, and reliable the data is.
[0096] In addition, data loading refers to the process of writing processed data into a target system, which is mainly divided into two methods: full-load loading and incremental loading. Full-load loading is applicable to scenarios of initial or small-scale data. However, when the data volume is large and the updates are frequent, the efficiency is relatively low. Incremental loading only loads the data that has changed since the last update, significantly improving the efficiency. In view of the characteristics of heterogeneous data sources, the incremental loading technology can be adaptively improved. Specifically, on the basis of using technologies such as timestamps or event markers to accurately identify and load the changed data by marking the change status of the data, the system also introduces an intelligent data source identification mechanism and a dynamic synchronization strategy, enabling the system to dynamically adjust the granularity and frequency of incremental loading according to the characteristics of different data sources to achieve more efficient incremental data processing, which is particularly suitable for the dynamic data management and multi-source data integration in the personal health record system.
[0097] It should be noted that the present invention does not limit the specific implementation manner of the ETL (Extract, Transform, Load) process, which can be implemented through programming languages or completed by existing ETL tools, such as Kettle, Apache NiFi, Sqoop, etc.
[0098] Through the above technical solutions, this embodiment provides an efficient ETL method, which optimizes data processing by selecting appropriate extraction methods and incremental loading to ensure the efficiency and timeliness of data loading. At the same time, the relevant data processing operations also improve the quality and consistency of the data, realizing the standardized integration of multi-source data and providing reliable data support for subsequent analysis and applications.
[0099] In a possible implementation manner, the data analysis process is supported by data mining and data analysis technologies, and the supported services include but are not limited to related services in aspects such as health management and services, public health and disease management, industry innovation and research and development, supervision and policy support, etc.
[0100] It should be noted that through systematic literature research and in-depth collation, this embodiment comprehensively explores the business areas supported by personal health records in data analysis, covering multiple aspects such as health management and services, public health and disease management, industry innovation and R & D, supervision and policy support, etc. Business in the category of health management and services includes but is not limited to health monitoring and assessment, chronic disease prediction, personalized medicine recommendation, optimal allocation of medical resources, etc.; business in the category of public health and disease management involves prevention and control of infectious diseases, prevention and control of occupational diseases, and health risk assessment, etc.; business in the category of industry innovation and R & D, especially in the pharmaceutical field, can be applied to links such as drug demand prediction and new drug R & D, and in the biomedical field, it supports the application of technologies such as genome sequencing and precision medicine; business in the category of supervision and policy support covers aspects such as supervision of medical service quality, drug quality control, and dynamic monitoring of population health, providing data support for national policy formulation and public health management.
[0101] Please refer to Figure 4 , in an embodiment of the present invention, a personal health record data processing system based on a data warehouse driven by metadata, the personal health record data integration system is sequentially connected to a data source layer, a data transmission layer, a data storage layer, a programming model layer, a data analysis layer, and a data application layer, and includes a system parallel computing framework (1), an archive data processing module (2), a security and privacy protection module (3), and a system monitoring and operation and maintenance management module (4).
[0102] Through the above technical solution, this embodiment provides a personal health record data processing system based on a data warehouse, which adopts a metadata-driven method to achieve efficient data processing under a multi-level architecture. The system parallel computing framework (1) improves the speed and scalability of data processing, ensuring the efficient processing of large-scale health data; the archive data processing module (2) is responsible for data cleaning, transformation, integration, and storage of personal health records, ensuring the integrity and consistency of data; the security and privacy protection module (3) ensures data security and privacy protection through technical means such as encryption and access control; the system monitoring and operation and maintenance management module (4) is responsible for real-time monitoring of the system operation status and providing operation and maintenance management functions to ensure the stability and efficient operation of the system. Through the collaborative work of each module, this embodiment effectively improves the processing efficiency and data quality of personal health records.
[0103] In a possible implementation manner, the system parallel computing framework (1) includes:
[0104] ETL tool library (11): responsible for extraction, transformation, and loading of multi-source health data;
[0105] Data storage medium (12): supports structured and unstructured data storage, and provides an efficient and reliable access interface;
[0106] Distributed computing framework (13): Achieves distributed storage and processing of large-scale data, supporting batch and real-time processing;
[0107] Intelligent analysis and decision-making platform (14): Conducts health data analysis and provides personalized health management and disease prediction models;
[0108] Task scheduling system (15): Coordinates and schedules various tasks to ensure the efficiency and automation of the processing flow;
[0109] Containerization and microservices module (16): Enables modular deployment and management, ensuring system scalability and ease of maintenance.
[0110] In a possible implementation, the system parallel computing framework (1) ensures the efficient processing and analysis of health data by integrating a variety of advanced technology tools. Specifically, the ETL tool library (11) uses Sqoop to be responsible for extracting, transforming, and loading health data from multi-source data. The efficient data migration ability of Sqoop can achieve seamless interaction of large-scale data between relational databases and the Hadoop ecosystem, which is particularly suitable for batch data processing scenarios. In terms of data storage (12), the system integrates Hadoop HDFS and PostgreSQL. The former is dedicated to distributed storage of large-scale unstructured data, and the latter is used to manage structured data to ensure data consistency and support for complex queries. In the distributed computing framework (13), Apache Spark is selected as the core computing engine, and its in-memory computing ability significantly improves the efficiency of batch and real-time data processing. At the same time, it supports complex analysis tasks such as machine learning and graph computing to meet the processing requirements of multi-dimensional health data. To achieve intelligent analysis and decision-making, the system uses Python and R languages to build an intelligent analysis platform (14), leveraging its rich health data analysis libraries to provide personalized health management and disease prediction models to cope with diverse data application scenarios. The task scheduling system (15) selects Apache Airflow to automate the coordination of various processing tasks. The powerful dependency management function of Airflow ensures the sequential execution and smooth scheduling of each task. To enhance the scalability and ease of maintenance of the system, the system realizes containerized deployment through Docker and combines Kubernetes for the orchestration and management of microservices. The containerized architecture (16) ensures the independence of modules, and each module can be flexibly upgraded and extended, thus ensuring the stability and efficiency of the system under high concurrency and large data volumes. Through the combination of the above tools and technologies, the system framework of this embodiment not only improves the performance and efficiency of health data processing but also has excellent scalability, maintainability, and flexibility through modular and automated design, making it suitable for the processing and analysis application scenarios of large-scale health data.
[0111] In a possible implementation, the file data processing module (2) includes modules such as a data acquisition and transmission module (21), a data processing and integration module (22), a data storage and management module (23), a data analysis and intelligent decision-making module (24), and a visualization and report generation module (25).
[0112] In a possible implementation, the security and privacy protection module (3) includes:
[0113] Privacy protection module (31): Combining a certain privacy protection mechanism, such as data encryption, permission control, etc., to ensure the privacy of personal health data during the processes of collection, storage, and use;
[0114] Log and monitoring module (32): Responsible for recording and monitoring data access, operation records, and abnormal behaviors in the system, ensuring the transparency and traceability of data usage, and promptly discovering security risks.
[0115] It should be noted that in this implementation, the security and privacy protection module (3) is not limited to basic privacy protection and monitoring mechanisms, but also includes more comprehensive security policies to address complex network security threats. The privacy protection module (31) not only supports data encryption and permission control, but also can combine more advanced privacy protection technologies such as homomorphic encryption and differential privacy to ensure that sensitive information is not leaked during the process of multi-party sharing or analysis of data. In addition, the module integrates dynamic access control policies, adjusts access permissions in real time according to user roles and behaviors, thereby enhancing the flexibility and security of the system. In the log and monitoring module (32), distributed log recording and real-time monitoring technologies are adopted, which can not only record every data access operation in detail, but also conduct real-time detection and early warning of potential security threats. This module can identify and respond to potential network attacks and abnormal operation behaviors through the use of abnormal behavior detection algorithms, ensure the traceability and integrity of data, and can also perform automated responses and block high-risk operations when necessary. Through the combination of multiple measures, this module provides strong security and privacy protection for the entire system, ensuring the security and legal use of personal health data in a multi-party environment.
[0116] In a possible implementation, the system monitoring and operation and maintenance management module (4) includes:
[0117] Performance monitoring module (41): Real-time monitoring of the running status and performance indicators of the system to ensure the efficiency of the data processing process and the stable operation of the system;
[0118] Fault handling module (42): Automatically detecting and handling system faults, reducing downtime through early warning and recovery mechanisms, and ensuring the continuity and reliability of the system.
[0119] It should be noted that the system monitoring and operation and maintenance management module (4) in this embodiment not only provides basic performance monitoring and fault handling functions, but also combines intelligent operation and maintenance technologies to improve the stability and operation efficiency of the system. The performance monitoring module (41) can perform real-time monitoring and data analysis on key performance indicators such as CPU, memory, network bandwidth, and I / O of each component of the system, and generate dynamic performance reports by introducing distributed monitoring tools (such as Prometheus) and a log collection system. This module also supports custom threshold alarms and automatically notifies the operation and maintenance personnel to take corresponding measures in a timely manner. In addition, the fault handling module (42) integrates intelligent fault detection algorithms and can identify potential performance bottlenecks and abnormal situations. By combining machine learning technologies, this module can predict the probability of a fault occurring and automatically trigger preventive maintenance measures to reduce the risk of system crashes. The fault handling module is also equipped with an automated fault recovery mechanism that can quickly restore critical services when a system fault occurs, reduce downtime, and ensure business continuity. This module not only improves the reliability and efficiency of the system, but also significantly reduces the complexity of manual intervention and improves the overall automation level of operation and maintenance management.
[0120] The above are only the preferred embodiments of the present invention. It should be pointed out that for those of ordinary skill in the art, without departing from the inventive concept of the present invention, several modifications and improvements can be made, and these all belong to the protection scope of the present invention.
[0121] The user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, research data, etc.) involved in the above embodiments are all information and data that have been authorized by the user or fully authorized by all parties, and the collection, use, and processing of relevant data need to comply with the relevant laws, regulations, and standards of relevant countries and regions.
[0122] The above description is only an illustration of the preferred embodiments of the present application and the technical principles applied. Those skilled in the art should understand that the scope of disclosure involved in the present application is not limited to the technical solutions formed by the specific combination of the above technical features, but also covers other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the above disclosure concept. For example, a technical solution formed by mutually replacing the above features with equivalent technical features disclosed (but not limited to) in the present application.
[0123] Moreover, although the operations are depicted in a particular order, this should not be construed as requiring that the operations be performed in the particular order shown or in sequential order. In certain circumstances, multitasking and parallel processing may be advantageous. Similarly, although several specific implementation details are included in the foregoing description, these should not be construed as limiting the scope of the present application. Certain features that are described in the context of separate embodiments may also be implemented in combination in a single embodiment. Conversely, the various features that are described in the context of a single embodiment may also be implemented separately or in any suitable sub-combination in multiple embodiments.
[0124] Although the subject matter has been described in language specific to structural features and / or methodological logical acts, it should be understood that the subject matter defined is not necessarily limited to the specific features or acts described above. Rather, the specific features and acts described above are merely example forms of implementation. With regard to the apparatus in the above embodiments, the specific manner in which each module performs operations has been described in detail in the embodiments related to the method, and will not be elaborated here.
Claims
1. A metadata-driven personal health record data processing method based on a data warehouse, characterized in that: The following steps are involved: Step S110: construct a shared document file of the personal health record data standard specification, design the corresponding chapter template, entry template, component template, determine the specific data element and coding rules, design a standardized form for each heterogeneous data source according to the specification, and ensure the unified data entry of each data source; Step S120: Integrate data from multiple heterogeneous data sources through the ETL process, using data extraction, transformation, and loading technology to ensure that data from different systems can be stored in a unified format and support subsequent analysis; Step S130: using the integrated data in the data warehouse, applying a variety of analysis tools and algorithms to perform data mining, trend analysis and health assessment on personal health records; Step S140: present the analysis results to the user in the form of charts or reports.
2. The method for processing personal health records based on a data warehouse and driven by metadata according to claim 1 is characterized in that: The specification designs chapter templates based on the themes of public health, medical services, health management, biological and omics data, and environmental data. The first-level directory includes residents' health records, electronic medical records, medication records, health tests, living behaviors, psychological assessments, genomes, proteomes, microbiomes, living environment descriptions, and environmental factors. The heterogeneous data sources include residents' health record management systems, various medical business systems, various embedded institution open platforms, and various wearable devices.
3. The metadata-driven personal health record data processing method based on a data warehouse according to claim 1 is characterized in that: The data ETL process includes: Step S210: Select a suitable extraction method according to the real-time nature and update frequency of the data; Step S220: cleaning, converting, integrating and other processing operations are performed on the extracted data to eliminate the differences between the original data set and the standard data form to ensure data quality and consistency; Step S230: Adopting the incremental loading method, introducing timestamp and event marking technology, ensure that a single query only loads data that has changed since the last update, thereby improving the efficiency of data loading and maintaining the timeliness of data.
4. The metadata-driven personal health record data processing method based on a data warehouse according to claim 3 is characterized in that: The data processing operations include but are not limited to: Step S310: Eliminate abnormal values and invalid data in the data through rule verification and data cleaning; Step S320: unify and standardize the format of the original data to ensure that the fields and types are consistent; Step S330: using feature screening and data simplification methods to retain key information in the data; Step S340: Integrate multi-source data through data association and fusion technology to generate a standardized integrated form.
5. The metadata-driven personal health record data processing method based on data warehouse according to claim 1 is characterized in that: The data analysis process is technically supported by data mining and data analysis, and the supported businesses include health management and services, public health and disease management, industry innovation and R&D, supervision and policy support and other related businesses.
6. A metadata-driven personal health record data processing system based on a data warehouse, characterized in that: The personal health record data integration system sequentially connects the data source layer, data transmission layer, data storage layer, programming model layer, data analysis layer and data application layer, and includes a system parallel computing framework, an archive data processing module, a security and privacy protection module, and a system monitoring and operation and maintenance management module.
7. The metadata-driven personal health record data processing system based on data warehouse according to claim 6 is characterized in that: The system parallel computing framework includes: ETL tool library: responsible for the extraction, conversion, and loading of multi-source health data; Data storage media: supports structured and unstructured data storage, and provides efficient and reliable access interfaces; Distributed computing framework: realizes distributed storage and processing of large-scale data, supporting batch and real-time processing; Intelligent analysis and decision-making platform: conduct health data analysis and provide personalized health management and disease prediction models; Task scheduling system: coordinate and schedule various tasks to ensure efficient and automated processing; Containerization and microservice modules: modular deployment and management to ensure system scalability and ease of maintenance.
8. The metadata-driven personal health record data processing system based on data warehouse according to claim 6 is characterized in that: The archive data processing module includes modules such as data acquisition and transmission module, data processing and integration module, data storage and management module, data analysis and intelligent decision-making module, visualization and report generation module, etc.
9. The metadata-driven personal health record data processing system based on data warehouse according to claim 6 is characterized in that: The security and privacy protection module includes: Privacy protection module: Combined with some privacy protection mechanisms, such as data encryption and permission control, to ensure the privacy of personal health data during collection, storage and use; Log and monitoring module: responsible for recording and monitoring data access, operation records and abnormal behaviors in the system, ensuring the transparency and traceability of data use, and timely discovering security risks.
10. The metadata-driven personal health record data processing system based on data warehouse according to claim 6, characterized in that: The system monitoring and operation and maintenance management module includes: Performance monitoring module: real-time monitoring of the system's operating status and performance indicators to ensure the efficiency of the data processing process and the stable operation of the system; Fault handling module: automatically detects and handles system faults, reduces downtime through early warning and recovery mechanisms, and ensures system continuity and reliability.
Citation Information
Patent Citations
Method of standardizing and promoting health medical big data
CN111125061A
Data processing method for full-life-cycle resident health archive
CN119441330A
Scalable analytics engineering platform for intelligent insights in healthcare
DE202025100262U1
Cited By
DTU point table configuration and analysis method based on template compression and topology partition
CN121172974A