Healthcare data management method, system, electronic device and storage medium
By implementing hierarchical and categorized management of health and medical data, the problem of large data volume but insufficient data quality has been solved. This has enabled refined management of data during collection, preprocessing, mining, and storage, thereby improving the quality and accuracy of the data.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- CHINA ACADEMY OF INFORMATION & COMM
- Filing Date
- 2022-07-04
- Publication Date
- 2026-07-24
AI Technical Summary
The lack of standards, norms, and quality control in health and medical big data has led to a problem of "large quantity but poor quality" in data collection, processing, mining, and storage.
A hierarchical classification management approach is adopted to classify and categorize health and medical data at each stage of collection, preprocessing, mining, storage, and quality verification. By establishing a master index and data model for patient and disease information, data cleaning, aggregation, code value matching, and data mining are carried out. Statistical and machine learning algorithms are used to analyze data quality, and multiple classifications and storage processes are performed to ultimately form usable data that meets the classification standards.
It has achieved hierarchical and classified management of the entire lifecycle of medical and health data, improved the quality and accuracy of data, ensured that data meets hierarchical and classified standards in each processing stage, and improved the quality of the final data used.
Smart Images

Figure CN115274122B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of health and medical big data processing technology, such as a method, system, electronic device and storage medium for managing health and medical data. Background Technology
[0002] Currently, in 2016, the State Council issued the "Guiding Opinions on Promoting and Regulating the Standardized Application and Development of Health and Medical Big Data," requiring the acceleration of the application of health and medical big data in industry governance, clinical research, public health, new business formats and models; in 2018, the National Health Commission issued the "National Health and Medical Big Data Standards, Security and Service Management Measures (Trial)," requiring the safe and standardized application of health and medical big data to fully unleash its value.
[0003] With the government's increased investment in health and medical big data, and the continuous strengthening of the information technology infrastructure of the health system, a vast amount of data has been accumulated in areas such as electronic medical records, health records, population information, and medical insurance records. Due to the large quantity, wide scope, and good extrapolation capabilities of health and medical big data, it effectively supports smart healthcare services such as medical artificial intelligence, chronic disease management, and precision treatment, becoming a crucial cornerstone for the development of digital healthcare.
[0004] In the process of implementing the embodiments of this application, at least the following problems were found in the related technology:
[0005] Due to the lack of standards, specifications, and quality control appropriate to the field, health and medical big data suffers from a problem of "large quantity but poor quality" in terms of data collection, processing, mining, storage, and application. Summary of the Invention
[0006] To provide a basic understanding of some aspects of the disclosed embodiments, a brief summary is given below. This summary is not intended as a general commentary, nor is it intended to identify key / important components or describe the scope of protection of these embodiments, but rather as a prelude to the detailed description that follows.
[0007] This application provides a method, system, electronic device, and storage medium for managing health and medical data, in order to improve the quality of health and medical big data.
[0008] In some embodiments, methods for managing health and medical data include:
[0009] Data is collected and then first-level classified to obtain the first-level classification results; the classification process is based on the data source and the data security level.
[0010] The first hierarchical classification result is preprocessed to obtain the preprocessed result; the required data in the preprocessed result is then subjected to a second hierarchical classification to obtain the second hierarchical classification result; all the data in the preprocessed result is then subjected to a third hierarchical classification to obtain the third hierarchical classification result; wherein, during the data preprocessing process, a master index based on patient information and / or a data model based on disease information are established, and the preprocessed result includes a subject dataset based on patient information and / or a data model based on disease information.
[0011] The second classification result is subjected to data mining processing, and then the data mining result is subjected to a fourth classification to obtain the fourth classification result; wherein, the data mining result includes the disease type and probability of an individual or multiple individuals, and / or the disease site and probability.
[0012] The third and fourth classification results are subjected to a first data storage process, and the data after the first storage process is subjected to a fifth classification to obtain the fifth classification result.
[0013] Based on the data quality requirements, the fifth classification result is subjected to data quality verification processing, and then the data quality verification result is subjected to a sixth classification to obtain the sixth classification result.
[0014] The sixth classification result is then subjected to a second data storage process to form usable data.
[0015] Optionally, data collection includes: importing medical and health data into a preset data source according to a preset import method; wherein the preset data source includes relational databases, big data systems, and real-time data source interfaces; the preset method includes offline data import, or single-table or batch data import, or automatic scheduled import, or full and incremental data import; establishing a unique data identifier for the imported medical and health data; mapping semantically similar but differently expressed words to slogan words; and providing standard definitions for data elements, data indicators, and data indicator dimensions.
[0016] Optionally, the first hierarchical classification result is preprocessed, including:
[0017] The first hierarchical classification results are cleaned according to the set rules; the set rules include one or more of the following: checking field type, maximum value, minimum value, maximum string length, minimum string length, missing values, and numerical precision; the data cleaning includes one or more of the following: performing one or more operations such as null value imputation, deduplication, and field filtering, and performing discretization processing on continuous data and sparse processing on classified data;
[0018] The data cleaning results are then aggregated. This aggregation includes one or more of the following: associating identical entities from multiple data sources, removing redundant attributes, detecting data value conflicts and providing corresponding handling; performing multi-table joins, using methods such as left join, right join, full join, and inner join; aggregating data according to custom rules; filtering data according to custom rules; replacing all or some fields in the data stream; and splitting composite fields in the data stream according to corresponding criteria and placing the splitting results into corresponding new columns.
[0019] The results of data aggregation are subjected to code value matching; the code value matching process includes one or more of the following: standardizing the code values of drugs, diseases, surgeries, tests, examinations, fees, institutions and departments; performing standard-to-standard code value mapping matching; and performing intelligent recommendations based on an artificial intelligence engine.
[0020] A master index based on patient information is established based on the code value matching results to obtain a topic dataset based on patient information. Establishing a master index based on patient information includes one or more of the following: performing rule-based patient master index identification and hierarchical management of the accuracy of the patient master index; performing patient master index identification based on an artificial intelligence model.
[0021] And / or,
[0022] A data model based on disease information is established based on the code value matching results; the process of establishing the data model includes one or more of the following: configuring and generating metadata templates, and synchronizing metadata information based on the metadata templates; mapping the medical and health organization data model to a standard data model; and copying the template of the same medical and health organization data model.
[0023] Optionally, the second hierarchical classification result is subjected to data mining processing, including one or more of the following:
[0024] By statistically analyzing the results of the second classification, the disease type and its probability for an individual or multiple individuals, and / or the disease location and its probability can be obtained.
[0025] The second classification result is classified using machine learning algorithms to obtain the disease type and probability of an individual or multiple individuals, and / or the disease location and probability.
[0026] Optionally, data storage of the hierarchical classification results includes: storing the hierarchical classification results according to the data type and storage requirements of the hierarchical classification results, and according to preset storage performance; wherein the hierarchical classification results include the third hierarchical classification results and the fourth hierarchical classification results, or the hierarchical classification results include the sixth hierarchical classification results; the data types include: relational data, text data, image data, structured data, and semi-structured data; the storage requirements include: file storage, object storage, off-site backup, and relational database; the preset storage performance includes high availability and horizontal scalability.
[0027] Optionally, based on data quality requirements, the fifth-level classification results undergo data quality verification processing, including: obtaining data quality requirements, which include data management objectives, quality assessment rules, and quality assessment schemes; obtaining data verification indicators, which include consistency, accuracy, completeness, standardization, correlation, and custom indicators; and performing data quality verification processing on the fifth-level classification results based on the data quality requirements and the data verification indicators. The data quality verification process includes: determining data connotation rules, and performing connotation analysis on the fifth-level classification results based on the connotation rules; the connotation rules represent the medical logic between data.
[0028] Optionally, the management methods for healthcare data may also include one or more of the following:
[0029] Perform data anonymization processing on the data;
[0030] Based on the hierarchical classification of the data, corresponding hierarchical classification information security protection should be implemented for the data.
[0031] Conduct risk monitoring and risk assessment on the data;
[0032] Issue a safety warning.
[0033] In some embodiments, the health and medical data management system includes a data acquisition management module, a data processing management module, a data mining management module, a data storage management module, and a data quality management module;
[0034] The data acquisition and management module is used to collect data and perform a first-level classification of the collected data to obtain a first-level classification result; wherein, the classification process is a process of classifying according to the data source and classifying according to the data security level;
[0035] The data processing and management module is used to preprocess the first hierarchical classification result to obtain the preprocessed result; to perform a second hierarchical classification on the required data in the preprocessed result to obtain the second hierarchical classification result; and to perform a third hierarchical classification on all the data in the preprocessed result to obtain the third hierarchical classification result. During the data preprocessing process, a master index based on patient information and / or a data model based on disease information are established, and the preprocessed result includes a subject dataset based on patient information and / or a data model based on disease information.
[0036] The data mining management module is used to perform data mining processing on the second hierarchical classification result, and then perform a fourth hierarchical classification on the data mining processing result to obtain the fourth hierarchical classification result; wherein, the data mining processing result includes the disease type and probability of an individual or multiple individuals, and / or the disease site and probability.
[0037] The data storage management module is used to perform a first data storage process on the third and fourth classification results, and to perform a fifth classification on the data after the first storage process to obtain the fifth classification result.
[0038] The data quality management module is used to perform data quality verification processing on the fifth classification result according to data quality requirements, and then perform a sixth classification on the data quality verification result to obtain the sixth classification result.
[0039] The data storage management module is also used to perform a second data storage process on the sixth-level classification results to form usable data.
[0040] In some embodiments, the electronic device includes a processor and a memory storing program instructions, the processor being configured to execute the health and medical data management method provided in the foregoing embodiments when executing the program instructions.
[0041] In some embodiments, the storage medium stores program instructions that, when executed, perform the health and medical management method provided in the foregoing embodiments.
[0042] The health and medical data management method, system, electronic device, and storage medium provided in this application can achieve the following technical effects:
[0043] After each data processing step—data acquisition, data preprocessing, data mining, data storage, and data quality verification—data is managed through hierarchical classification. This allows for the adjustment of the processing results from one data processing step to the next. Except for data acquisition, all input data in each processing step is hierarchically classified, achieving hierarchical classification management throughout the entire lifecycle of healthcare data. This ensures that data in each processing step conforms to data classification standards, ultimately improving the quality of usable data. Furthermore, in the health and medical data management method provided in this application, the results of data preprocessing are divided into demand data, which is then sent to the data mining process. The results of the data mining process, along with all the data after preprocessing, are classified and categorized. The classification and categorization results are then stored. This further achieves refined classification and categorization management of health and medical data, which is more conducive to the accuracy of the classification and categorization of the stored data and improves the quality of the final usable data. Simultaneously, since the third and fourth classification and categorization results are stored during the data storage and processing, performing a classification and categorization of the stored data before data quality verification, and simultaneously classifying the third and fourth classification and categorizing the results, facilitates the data quality verification process to verify data that meets the classification and categorization standards, obtaining data quality verification results that better conform to the classification and categorization standards, and improving the quality of the final usable data.
[0044] The above general description and the description below are exemplary and illustrative only and are not intended to limit this application. Attached Figure Description
[0045] One or more embodiments are illustrated by way of example with reference to the accompanying drawings. These illustrative descriptions and drawings do not constitute a limitation on the embodiments. Elements having the same reference numerals in the drawings are considered similar elements, and wherein:
[0046] Figure 1 This is a flowchart illustrating a method for managing health and medical data provided in an embodiment of this application;
[0047] Figure 2 This is a schematic diagram of a data acquisition process provided in an embodiment of this application;
[0048] Figure 3 This is a schematic diagram of a data preprocessing procedure provided in an embodiment of this application;
[0049] Figure 4 This is a schematic diagram of a data quality verification process provided in an embodiment of this application;
[0050] Figure 5 This is a schematic diagram of a health and medical data management system provided in an embodiment of this application;
[0051] Figure 6 This is a schematic diagram of an electronic device provided in an embodiment of this application. Detailed Implementation
[0052] To provide a more detailed understanding of the features and technical content of the embodiments of this application, the implementation of the embodiments of this application will be described in detail below with reference to the accompanying drawings. The accompanying drawings are for illustrative purposes only and are not intended to limit the embodiments of this application. In the following technical description, for ease of explanation, several details are used to provide a full understanding of the disclosed embodiments. However, one or more embodiments may still be implemented without these details. In other cases, well-known structures and devices may be simplified in their depiction to simplify the drawings.
[0053] The terms "first," "second," etc., in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate for the embodiments of this application described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion.
[0054] Unless otherwise stated, the term "multiple" means two or more.
[0055] In this embodiment, the character " / " indicates that the objects before and after it are in an "or" relationship. For example, A / B means: A or B.
[0056] The term "and / or" describes an association between objects, indicating that three relationships can exist. For example, A and / or B means: A or B, or A and B.
[0057] The hierarchical classification in this application includes classifying data according to its source and classifying data according to its security level. Different categories of data have different sources; different levels of data have different security levels.
[0058] The following provides an example of how data sources and data security levels are defined:
[0059] Data sources can be categorized into six types: personal attributes, health status, medical applications, medical payments, health resources, and public health. As shown in Table 1, data security levels can be classified into five levels: those that can be used publicly, those that can be accessed on a large scale, those that can be accessed on a medium scale, those that can be accessed on a small scale, and those that can be accessed only on a very small scale under strict control.
[0060] Table 1 Data Security Levels and Corresponding Usage Scope
[0061]
[0062]
[0063] The data sources and data security levels mentioned above are merely illustrative examples and do not constitute a specific limitation on the health and medical data management methods provided in the embodiments of this application. Those skilled in the art can determine the category of data sources and the data security level that conforms to the actual situation based on the actual classification requirements of the data sources and the actual classification requirements of the data security level (scope of use).
[0064] Figure 1 This is a flowchart illustrating a method for managing health and medical data provided in an embodiment of this application.
[0065] Combination Figure 1 As shown, the methods for managing health and medical data include:
[0066] S101. Collect data and perform the first-level classification on the collected data to obtain the first-level classification result.
[0067] The classification and grading process involves classifying data according to its source and grading it according to its security level.
[0068] For example, data can be categorized into six data sources: personal attributes, health status, medical applications, medical payments, health resources, and public health; data can also be classified according to the grading standards shown in Table 1.
[0069] S102. Perform data preprocessing on the first classification result to obtain the preprocessed result; perform a second classification on the required data in the preprocessed result to obtain the second classification result; and perform a third classification on all the data in the preprocessed result to obtain the third classification result.
[0070] In the data preprocessing process, a master index based on patient information and / or a data model based on disease information are established. The results of the preprocessing include a subject dataset based on patient information and / or a data model based on disease information.
[0071] For example, the main index based on patient information can be the patient's name, ID number, medical insurance card number, or hospital management number. The subject dataset based on patient information includes, but is not limited to, patient gender, age, place of origin, home address, past medical history, and physiological parameters.
[0072] Disease-based data models include various parameters associated with a particular disease. For example, a diabetes-based data model includes parameters related to diabetes such as blood glucose levels and parameters related to unexplained weight loss.
[0073] The aforementioned required data refers to data that needs to be processed through data mining. The required data will vary depending on the actual application scenario, and those skilled in the art can determine the corresponding required data based on the actual application scenario.
[0074] Since the preprocessing results include subject datasets based on patient information and / or data models based on disease information, the second and third classifications are processes of classifying according to data source, classifying according to the category of subject dataset and / or data model, and classifying according to data security level.
[0075] S103. Perform data mining processing on the second classification result, and then perform a fourth classification on the data mining result to obtain the fourth classification result.
[0076] The results of data mining include the disease type and probability of an individual or multiple individuals, and / or the disease location and probability.
[0077] After data mining processing of the second-level classification results, the sensitivity of the data changes because the results contain the disease type and probability of individuals or multiple individuals, and / or the disease site and probability. For example, obtaining the probability of someone having diabetes increases the sensitivity of this information. In this case, performing a fourth-level classification on the results of data mining is more conducive to making health and medical data more in line with the classification standards and improving the quality of health and medical data.
[0078] Optionally, data mining processing is performed on the second-level classification results, including one or more of the following:
[0079] By statistically analyzing the results of the second-level classification, we can obtain the disease type and its probability for an individual or multiple individuals, and / or the disease location and its probability.
[0080] The second-level classification results are classified using machine learning algorithms to obtain the disease type and probability of an individual or multiple individuals, and / or the disease location and probability.
[0081] The statistical analysis of the second-level classification results includes, but is not limited to: calculating the variance of the second-level classification results, calculating the covariance matrix of the second-level classification results, and calculating the standard deviation of the second-level classification results.
[0082] The second-level classification results are classified using machine learning algorithms, including but not limited to: classifying the second-level classification results using text analysis algorithms, classifying / clustering algorithms, using regression algorithms, using machine recommendation algorithms, and using association analysis algorithms.
[0083] Since the results of data mining include the disease type and probability of individuals or multiple individuals, and / or the disease location and probability, the fourth classification is a process of classifying according to the data source, according to the category of the subject dataset and / or data model, according to the disease type and / or disease location, and according to the data security level.
[0084] Furthermore, feature engineering can be performed on the second-level classification results, and then machine learning algorithms can be used to classify the feature-engineered results to obtain the disease type and probability of an individual or multiple individuals, and / or the disease location and probability.
[0085] The feature engineering processing of the second-level classification results includes, but is not limited to: feature discretization processing of the second-level classification results, random pre-sampling processing of the second-level classification results, and feature vector segmentation processing of the second-level classification results; and the feature engineering processing algorithm can be reused in algorithm engineering.
[0086] In practical applications, the data mining process may also include one or more of the following:
[0087] The algorithm model results are evaluated using multiple evaluation metrics, and the existing model results are displayed using multiple visualization methods.
[0088] The model creation process is guided through templates, case studies, and tutorial-style explanations of algorithms.
[0089] The algorithm can be customized by writing it directly or by calling an interface in multiple programming languages such as Python, Java, and R, and is compatible with secondary development languages; it can also view the log records of the model training task execution process.
[0090] This can improve the user experience for maintenance personnel.
[0091] S104. Perform the first data storage processing on the third and fourth classification results, and then perform the fifth classification on the data after the first storage processing to obtain the fifth classification result.
[0092] The fifth level of classification involves classifying data according to its source, the subject dataset and / or data model, the type of disease and / or the location of the disease, the encryption and access information of health and medical data, and the data security level.
[0093] Optionally, the third and fourth classification results undergo a first data storage process, including: storing the third and fourth classification results according to their data types and storage requirements, based on preset storage performance; wherein the data types include: relational data, text data, image data, structured data, and semi-structured data; the storage requirements include: file storage, object storage, off-site backup, and relational database; and the preset storage performance includes high availability and horizontal scalability.
[0094] The high availability mentioned above does not refer to a specific numerical value of "high," but rather to the fact that high availability (HA) means improving the availability of systems and applications by minimizing downtime caused by routine maintenance operations (planned) and sudden system crashes (unplanned).
[0095] In actual storage processes, storage performance may also include one or more of the following:
[0096] It can perform fast read, write, and query operations on massive amounts of data;
[0097] It is capable of performing batch read and write operations and real-time write operations with a cluster throughput of over 100MB / s.
[0098] It can store columnar data and perform millisecond-level queries and writes;
[0099] It can store row data and back up metadata off-site;
[0100] It can automatically and elastically scale the cluster.
[0101] S105. Based on the data quality requirements, perform data quality verification processing on the fifth-level classification results, and then perform a sixth-level classification on the data quality verification results to obtain the sixth-level classification results.
[0102] The sixth level of classification involves categorizing data by source, subject dataset and / or data model, disease type and / or affected area, encryption and retrieval information of health and medical data, data quality verification results, and data security level. Encryption information includes the type and volume of encrypted data; retrieval information includes the type, frequency, and user of the accessed data.
[0103] Optionally, data quality verification processing is performed on the fifth-level classification results, including: storing the fifth-level classification results according to the data type and storage requirements of the fifth-level classification results, and storing them according to preset storage performance; wherein, the data types include: relational data, text data, image data, structured data, and semi-structured data; the storage requirements include: file storage, object storage, off-site backup, and relational database; and the preset storage performance includes high availability and horizontal scalability.
[0104] S106. Perform a second data storage process on the sixth-level classification results to form usable data.
[0105] Optionally, the sixth-level classification results undergo a first data storage process, including: storing the sixth-level classification results according to the data type and storage requirements of the sixth-level classification results, based on preset storage performance; wherein the data types include: relational data, text data, image data, structured data, and semi-structured data; the storage requirements include: file storage, object storage, off-site backup, and relational database; and the preset storage performance includes high availability and horizontal scalability.
[0106] After each data processing step—data acquisition, data preprocessing, data mining, data storage, and data quality verification—data is managed through hierarchical classification. This allows for the adjustment of the processing results from one data processing step to the next. Except for data acquisition, all input data in each processing step is hierarchically classified, achieving hierarchical classification management throughout the entire lifecycle of healthcare data. This ensures that data in each processing step conforms to data classification standards, ultimately improving the quality of usable data. Furthermore, in the health and medical data management method provided in this application, the results of data preprocessing are divided into demand data, which is then sent to the data mining process. The results of the data mining process, along with all the data after preprocessing, are classified and categorized. The classification and categorization results are then stored. This further achieves refined classification and categorization management of health and medical data, which is more conducive to the accuracy of the classification and categorization of the stored data and improves the quality of the final usable data. Simultaneously, since the third and fourth classification and categorization results are stored during the data storage and processing, performing a classification and categorization of the stored data before data quality verification, and simultaneously classifying the third and fourth classification and categorizing the results, facilitates the data quality verification process to verify data that meets the classification and categorization standards, obtaining data quality verification results that better conform to the classification and categorization standards, and improving the quality of the final usable data.
[0107] Figure 2 This is a schematic diagram of a data acquisition process provided in an embodiment of this application.
[0108] Combination Figure 2 As shown, the collected data includes:
[0109] S201. Import medical and health data into the preset data source according to the preset import method.
[0110] The preset data sources include relational databases, big data systems, and real-time data source interfaces; the preset methods include offline data import, single-table or batch data import, automatic scheduled import, or full and incremental data import.
[0111] The aforementioned relational databases include, but are not limited to: MySQL, Oracle, SQL Server, and the domestic database DM.
[0112] Big data systems include, but are not limited to: Hive, HDFS, MongoDB, and Postgres.
[0113] Real-time data source interfaces include, but are not limited to: Kafka, Oracle CDC, MySQL binlog, SQLserverCDC, and RabbitMQ.
[0114] In some specific applications, the types of imported healthcare data include structured data, unstructured data, and semi-structured data.
[0115] During the data import process, files can also be aggregated. The types of aggregated files include, but are not limited to: ftp, excel and csv.
[0116] S202. Establish a unique identifier for the imported medical and health data.
[0117] In the later data preprocessing process, this unique data identifier facilitates the establishment of a master index based on patient information, or a data model based on disease type.
[0118] S203. Map words with the same meaning but different expressions to slogan words.
[0119] For example, in place of origin information, "boy" and "male" can be mapped to "male". Existing semantic analysis algorithms can be used to achieve the above mapping process, which will not be elaborated here.
[0120] This mapping facilitates subsequent hierarchical classification processes, as well as data preprocessing, data mining, data storage, and data quality verification processes.
[0121] S204 provides standard definitions for data elements, data metrics, and data metric dimensions.
[0122] Among them, data element refers to the unique identifier of the data; data indicator refers to the data indicators of medical and health data, such as the normal blood glucose concentration range and the abnormal blood glucose concentration range; data indicator dimension refers to the standard unit of the data indicator, such as blood glucose concentration uniformly using 1g / ml.
[0123] Data collection can be achieved using the methods described above.
[0124] In some specific applications, it can also provide functions such as visual configuration for data acquisition source and target ends, management of synchronization tasks, and task monitoring; provide visual mapping of data fields on the source and target ends; and provide visual implementation of data element management, data indicator management, data standard dimension management, and data dictionary. This facilitates monitoring and maintenance of the health and medical data management system by its administrators, improving the user experience.
[0125] Figure 3This is a schematic diagram of a data preprocessing process provided in an embodiment of this application.
[0126] Combination Figure 3 As shown, data preprocessing is performed on the first-level classification results, including:
[0127] S301. Perform data cleaning on the first-level classification results according to the set rules.
[0128] The rules set include one or more of the following: checking field type, maximum value, minimum value, maximum string length, minimum string length, missing values, and numerical precision.
[0129] The data cleaning process includes one or more of the following: performing one or more operations such as null value imputation, deduplication, and field filtering; discretizing continuous data; and sparsening categorical data.
[0130] S302. Perform data aggregation on the results of data cleaning.
[0131] Data aggregation includes one or more of the following:
[0132] For the same entity associated with multiple data sources, remove redundant attributes, detect data value conflicts, and provide corresponding handling.
[0133] Perform multi-table joins, with join methods including left join, right join, full join, and inner join; aggregate data based on custom rules;
[0134] Filter the data according to custom rules;
[0135] Replace all or part of the fields in the data stream; split the composite fields in the data stream according to the corresponding criteria and place the split results into the corresponding new columns.
[0136] S303. Perform code value matching on the results of data aggregation.
[0137] The code value matching process includes one or more of the following: standardizing the code values of drugs, diseases, surgeries, tests, examinations, fees, institutions, and departments; performing standard-to-standard code value mapping and matching; and making intelligent recommendations based on an artificial intelligence engine.
[0138] S304. Establish a master index based on patient information for the code value matching results to obtain a topic dataset based on patient information.
[0139] This includes establishing a master index based on patient information, including one or more of the following:
[0140] Perform rule-based identification of patient master indexes and classify and manage the accuracy of patient master indexes.
[0141] Perform patient master index identification based on an artificial intelligence model.
[0142] S305. Establish a data model based on disease information based on the code value matching results.
[0143] The process of building a data model includes one or more of the following:
[0144] Configure and generate metadata templates, and synchronize metadata information based on the metadata templates;
[0145] Mapping healthcare organization data models to standard data models;
[0146] Copy the template of the same healthcare organization data model.
[0147] In specific applications, the data preprocessing process may include only the step of establishing a data model based on disease information, excluding the step of obtaining a main dataset based on patient information; or, the data preprocessing process may include only the step of obtaining a main dataset based on patient information, excluding the step of establishing a data model based on disease information; or, the data preprocessing process may include both the step of establishing a data model based on disease information and the step of establishing a data model based on disease information.
[0148] In the above embodiments, the data preprocessing process is only illustrated by the example of a step that can simultaneously include the step of establishing a data model based on disease information and the step of establishing a data model based on disease information. Those skilled in the art can determine the data preprocessing process that meets the actual needs according to actual requirements.
[0149] In specific applications, the data preprocessing process may also include one or more of the following:
[0150] Perform full and incremental task scheduling, customize task execution cycles, and enable data flow between different data sources;
[0151] Configure custom rules using SQL, Java, or other programming languages;
[0152] By defining custom data quality verification functions during data processing, we can quickly configure and verify data rules during the data processing process.
[0153] Provide unified data services, enabling queries on single or multiple tables in a SQL-like format and returning data that meets the specified conditions;
[0154] It enables lifecycle management of service application programming interfaces (APIs) and allows for the visual generation and management of APIs.
[0155] Perform report analysis on the service API.
[0156] Intelligent tagging of data, including tag model creation, tag processing, and derivative tag management; and editing, viewing, and deletion functions developed using SQL.
[0157] Data tag operation: conduct full lifecycle management of tags, such as tag online / offline, and public management of tag assets;
[0158] Provide data tagging services, encapsulate tagging services in an API service manner, and make them available for internal and external applications to use, thereby improving the user experience of the health and medical data management system administrators;
[0159] Conduct data tag analysis to analyze tag production and usage, and clarify the total number of tags, total number of APIs, API performance, etc., in order to improve the user experience of the health and medical data management system administrators.
[0160] Figure 4 This is a schematic diagram of a data quality verification process provided in an embodiment of this application.
[0161] Combination Figure 4 As shown, based on data quality requirements, data quality verification processing is performed on the fifth-level classification results, including:
[0162] S401. Obtain data quality requirements.
[0163] Data quality requirements include data management objectives, quality assessment rules, and quality assessment schemes.
[0164] In practical applications, based on the application scenarios and data management goals of medical and health data, authorized users can propose quality management solutions and customize quality assessment rules and schemes.
[0165] Data management objectives, quality assessment rules, and quality assessment schemes can be pre-stored in the database. When data quality verification processing begins, the data quality requirements can be obtained by reading the database; alternatively, the data quality requirements can be obtained in real time in response to user input.
[0166] This step provides users with a visual interface, enabling them to customize data quality rule categories and manage data quality rule versions, thus achieving quality rule management.
[0167] S402. Obtain data verification indicators.
[0168] Data quality verification metrics include consistency, accuracy, completeness, standardization, relevance, and custom metrics.
[0169] Data quality verification metrics can be stored in a database. When data quality verification is required, the metrics can be obtained by reading data from the database; alternatively, they can be obtained in real time in response to user input.
[0170] S403. Based on data quality requirements and data verification indicators, perform data quality verification processing on the fifth-level classification results.
[0171] The data quality verification process includes: determining the data connotation rules, and performing connotation analysis on the fifth-level classification results based on the connotation rules; the connotation rules represent the medical logic between the data.
[0172] This enables quality verification processing for health and medical data.
[0173] In practical applications, data quality verification of the fifth-level classification results may also include one or more of the following:
[0174] Set verification rules that conform to logical data, and perform rule verification;
[0175] Conduct data quality verification task management, such as executing data quality audit tasks, configuring scheduling information, setting the execution cycle of audit tasks and executing scheduling, (real-time) monitoring of data quality audit tasks, and viewing historical task execution status;
[0176] Perform quality control tasks regularly or irregularly, and generate relevant issue reports;
[0177] When data quality issues arise, it can issue quality alerts, trace the underlying mechanisms, and push messages in various forms.
[0178] By leveraging technologies such as medical knowledge graphs and artificial intelligence, we can conduct in-depth analysis of the quality of medical data and improve the quality of data content.
[0179] Record quality issues in the verification results and generate a scoring report; process and aggregate the data to form a results analysis and generate an impact report;
[0180] Based on the inspection results, a quality result analysis is conducted, including a summary data list, an error summary list, quality score analysis, inspection rule analysis, and problem fluctuation analysis. Suggestions for improving the quality of the problematic data are then generated to guide the data quality improvement work.
[0181] The rules for managing data content are configured and managed in a visual manner.
[0182] The above provides an illustrative description of each data processing step in the management methods for health and medical data. In practical applications, the management methods for health and medical data also include data security management processes to prevent, monitor, and provide early warnings against external attacks.
[0183] For example, data security management may include one or more of the following:
[0184] Perform data anonymization processing on the data;
[0185] Based on the hierarchical classification of the data, corresponding hierarchical classification information security protection should be implemented for the data.
[0186] Conduct risk monitoring and risk assessment on the data;
[0187] Issue a safety warning.
[0188] Data anonymization can include: setting anonymization encryption rules, algorithms, and tasks before performing data anonymization; or static anonymization, such as in a non-production environment, where the anonymized data is converted and then extracted into an anonymized database.
[0189] Based on the data's hierarchical classification, corresponding hierarchical classification and protection of the data may include: determining the information security protection level corresponding to the classification results (e.g., the aforementioned first, second, third, fourth, fifth, and sixth hierarchical classification results); and protecting the data corresponding to the hierarchical classification results using the protection mechanisms corresponding to the information security protection levels. In the process of determining the information security protection level corresponding to the classification results, artificial intelligence algorithms such as entity recognition and text parsing can be used to determine the information security protection level; alternatively, a visual interface can be provided to respond to user input and obtain the information security protection level corresponding to the classification results.
[0190] Data risk monitoring and risk assessment can include: scanning for security items based on a set security thesaurus and security rules, scanning specified data sources, and identifying sensitive information; or, building a privacy model through proactive privacy protection technology to monitor, assess, proactively alert, and trace responsibility for data at risk of privacy leakage.
[0191] Security alerts can include: monitoring the security status of data throughout its lifecycle and issuing warnings when security risks are detected.
[0192] Figure 5 This is a schematic diagram of a health and medical data management system provided in an embodiment of this application.
[0193] Combination Figure 5As shown, the health and medical data management system includes a data acquisition management module 51, a data processing management module 52, a data mining management module 53, a data storage management module 54, and a data quality management module 55.
[0194] The data acquisition and management module 51 is used to collect data and perform the first-level classification of the collected data to obtain the first-level classification result; the classification process is a process of classifying according to the data source and classifying according to the data security level.
[0195] The data processing and management module 52 is used to preprocess the first hierarchical classification results to obtain the preprocessed results; to perform a second hierarchical classification on the required data in the preprocessed results to obtain the second hierarchical classification results; and to perform a third hierarchical classification on all the data in the preprocessed results to obtain the third hierarchical classification results. During the data preprocessing process, a master index based on patient information and a data model based on disease information are established, and the preprocessed results include a topic dataset based on patient information.
[0196] The data mining management module 53 is used to perform data mining processing on the first part of the data in the second classification result, and then perform a fourth classification on the data mining result to obtain the fourth classification result; wherein, the data mining result includes the disease type and probability of an individual or multiple individuals, and / or the disease location and probability.
[0197] The data storage management module 54 is used to perform the first data storage processing on the third and fourth classification results, and to perform the fifth classification on the data after the first storage processing to obtain the fifth classification result.
[0198] The data quality management module 55 is used to perform data quality verification processing on the fifth-level classification results according to data quality requirements, and then perform a sixth-level classification on the data quality verification results to obtain the sixth-level classification results.
[0199] The data storage management module 54 is also used to perform a second data storage process on the sixth-level classification results to form usable data.
[0200] Optionally, the data acquisition and management module 51 includes an import unit, an identifier establishment unit, a mapping unit, and a definition unit. The import unit is used to import medical and health data from preset data sources according to preset import methods. The preset data sources include relational databases, big data systems, and real-time data source interfaces. The preset methods include offline data import, single-table or batch data import, automatic scheduled import, or full and incremental data import. The identifier establishment unit is used to establish unique identifiers for the imported medical and health data. The mapping unit is used to map semantically similar but differently expressed words to slogan words. The definition unit is used to provide standard definitions for data elements, data indicators, and data indicator dimensions.
[0201] Optionally, the data processing management module 52 includes a data cleaning unit, a data aggregation unit, a code value matching unit, and a main index unit and / or a model building unit;
[0202] The data cleaning unit is used to clean the first-level classification results according to the set rules. The set rules include one or more of the following: checking field type, maximum value, minimum value, maximum string length, minimum string length, missing values, and numerical precision; data cleaning includes one or more of the following: performing one or more operations such as null value imputation, deduplication, and field filtering, and performing discretization processing on continuous data and sparse processing on classified data.
[0203] The data aggregation unit is used to aggregate the results of data cleaning. Data aggregation includes one or more of the following: associating the same entities from multiple data sources, removing redundant attributes, detecting data value conflicts and providing corresponding processing; performing multi-table joins, with join methods including left join, right join, full join, and inner join; aggregating data according to custom rules; filtering data according to custom rules; replacing all or part of the fields in the data stream; and splitting composite fields in the data stream according to corresponding criteria and placing the split results into corresponding new columns.
[0204] The code value matching unit is used to perform code value matching on the results of data aggregation; the code value matching process includes one or more of the following: standardizing and annotating the code values of drugs, diseases, surgeries, tests, examinations, fees, institutions and departments; performing standard-to-standard code value mapping matching; and performing intelligent recommendations based on an artificial intelligence engine;
[0205] The master index unit is used to build a master index based on patient information from the results of code value matching, and to obtain a subject dataset based on patient information. Building a master index based on patient information includes one or more of the following: performing rule-based patient master index identification and hierarchical management of the accuracy of the patient master index; performing patient master index identification based on artificial intelligence models.
[0206] The model building unit is used to build a data model based on disease information. The data model building process includes one or more of the following: configuring and generating metadata templates, and synchronizing metadata information based on the metadata templates; mapping the healthcare organization data model to a standard data model; and copying the template of the same healthcare organization data model.
[0207] Optionally, the data mining management module 53 includes one or more of the following: a statistical analysis unit and a machine algorithm processing unit; the statistical analysis unit is used to obtain the disease type and probability of an individual or multiple individuals, and / or the disease location and probability, by statistically analyzing the second-level classification results; the machine algorithm processing unit is used to classify the second-level classification results by machine learning algorithms to obtain the disease type and probability of an individual or multiple individuals, and / or the disease location and probability.
[0208] Optionally, the data storage management module 54 is specifically used to store the hierarchical classification results according to the data type of the hierarchical classification results and the storage requirements of the hierarchical classification results, in accordance with preset storage performance; wherein, the hierarchical classification results include the third hierarchical classification results and the fourth hierarchical classification results, or the hierarchical classification results include the sixth hierarchical classification results; the data types include: relational data, text data, image data, structured data, and semi-structured data; the storage requirements include: file storage, object storage, off-site backup, and relational database; the preset storage performance includes high availability and horizontal scaling.
[0209] Optionally, the data quality management module 55 includes a first acquisition unit, a second acquisition unit, and a quality verification unit; the first acquisition unit is used to acquire data quality requirements, which include data management objectives, quality assessment rules, and quality assessment schemes; the second acquisition unit is used to acquire data verification indicators, which include consistency, accuracy, completeness, standardization, relevance, and custom indicators; the quality verification unit is used to perform data quality verification processing on the fifth-level classification results based on the data quality requirements and data verification indicators; wherein, the data quality verification processing includes: determining data connotation rules, and performing connotation analysis on the fifth-level classification results based on the connotation rules; the connotation rules represent the medical logic between data.
[0210] Optionally, the medical and health data management system may also include one or more of the following: a data anonymization module, a security protection module, a risk assessment module, and an early warning module; the data anonymization module is used to perform data anonymization processing; the security protection module is used to perform corresponding hierarchical classification information security protection for the data according to its classification; the risk assessment module is used to monitor and assess the risks of the data; and the early warning module is used to issue security warnings.
[0211] like Figure 6As shown in the embodiment of this application, an electronic device includes:
[0212] The processor 61 and memory 62 may also include a communication interface 63 and a bus 64. The processor 61, communication interface 63, and memory 62 can communicate with each other via the bus 64. The communication interface 63 can be used for information transmission. The processor 61 can invoke logical instructions stored in the memory 62 to execute the health and medical data management method provided in the foregoing embodiments.
[0213] Furthermore, the logical instructions in the aforementioned memory 62 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium.
[0214] The memory 62, as a computer-readable storage medium, can be used to store software programs and computer-executable programs, such as program instructions / modules corresponding to the methods in the embodiments of this application. The processor 61 executes functional applications and data processing by running the software programs, instructions, and modules stored in the memory 62, thereby implementing the methods in the above-described method embodiments.
[0215] The memory 62 may include a program storage area and a data storage area. The program storage area may store the operating system and applications required for at least one function; the data storage area may store data created based on the use of the terminal device. Furthermore, the memory 62 may include high-speed random access memory and may also include non-volatile memory.
[0216] This application provides a computer-readable storage medium storing computer-executable instructions configured to execute the health and medical data management method provided in the foregoing embodiments.
[0217] This application provides a computer program product, which includes a computer program stored on a computer-readable storage medium. The computer program includes program instructions, which, when executed by a computer, cause the computer to perform the health and medical data management method provided in the foregoing embodiments.
[0218] The aforementioned computer-readable storage medium may be a transient computer-readable storage medium or a non-transitory computer-readable storage medium.
[0219] The technical solutions of this application embodiment can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes one or more instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods in this application embodiment. The aforementioned storage medium can be a non-transitory storage medium, including: USB flash drive, portable hard drive, read-only memory (ROM), random access memory (RAM), magnetic disk, or optical disk, and other media capable of storing program code; it can also be a transient storage medium.
[0220] The foregoing description and accompanying drawings fully illustrate embodiments of this application to enable those skilled in the art to practice them. Other embodiments may include structural, logical, electrical, procedural, and other changes. The embodiments represent only possible variations. Individual components and functions are optional unless explicitly required, and the order of operations may vary. Parts and features of some embodiments may be included in or replace parts and features of other embodiments. Moreover, the terminology used in this application is for describing embodiments only and is not intended to limit the claims. As used in the description of embodiments and claims, the singular forms “a,” “an,” and “the” are intended to equally include the plural forms unless the context clearly indicates otherwise. Additionally, when used in this application, the terms “comprise” and its variations “comprises” and / or “comprising” refer to the presence of stated features, integrals, steps, operations, elements, and / or components, but do not exclude the presence or addition of one or more other features, integrals, steps, operations, elements, components, and / or groups thereof. Unless otherwise specified, an element defined by the phrase "comprising a..." does not exclude the presence of other identical elements in the process, method, or apparatus that includes that element. In this document, each embodiment may focus on describing the differences from other embodiments, and similar or identical parts between embodiments can be referred to mutually. For methods, products, etc., disclosed in the embodiments, if they correspond to the method section disclosed in the embodiments, then the relevant parts can be referred to the description of the method section.
[0221] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of the embodiments of this application. Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.
[0222] The methods and products (including but not limited to devices and equipment) disclosed in the embodiments herein can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For instance, the division of units may be merely a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. In addition, the mutual coupling or direct coupling or communication connection shown or discussed may be through some interfaces, and the indirect coupling or communication connection of devices or units may be electrical, mechanical, or other forms. Units described as separate components may or may not be physically separate, and components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to implement this embodiment according to actual needs. Furthermore, the functional units in the embodiments of this application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit.
[0223] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to embodiments of this application. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code, which contains one or more executable instructions for implementing a specified logical function. In some alternative implementations, the functions marked in the blocks may occur in a different order than those shown in the drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. Each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, may be implemented using a dedicated hardware-based system that performs the specified function or action, or using a combination of dedicated hardware and computer instructions.
Claims
1. A method for managing health and medical data, characterized in that, include: Data is collected and then first-level classified to obtain the first-level classification results; the classification process is based on the data source and the data security level. The first hierarchical classification result is preprocessed to obtain a preprocessed result; the required data in the preprocessed result is then subjected to a second hierarchical classification to obtain a second hierarchical classification result; and all the data in the preprocessed result is then subjected to a third hierarchical classification to obtain a third hierarchical classification result. During the data preprocessing process, a master index based on patient information and / or a data model based on disease information are established. The preprocessed result includes a subject dataset based on patient information and / or a data model based on disease information. The second and third hierarchical classifications are based on data source, subject dataset and / or data model categories, and data security level. The second classification result is subjected to data mining processing, and then the data mining result is subjected to a fourth classification to obtain the fourth classification result; wherein, the data mining result includes the disease type and probability of an individual or multiple individuals, and / or the disease site and probability. The third and fourth classification results are subjected to a first data storage process, and the data after the first storage process is subjected to a fifth classification to obtain the fifth classification result. Based on the data quality requirements, the fifth classification result is subjected to data quality verification processing, and then the data quality verification result is subjected to a sixth classification to obtain the sixth classification result. The sixth classification result is then subjected to a second data storage process to form usable data.
2. The management method according to claim 1, characterized in that, Data collected includes: Import medical and health data into a preset data source according to a preset import method; wherein, the preset data source includes relational databases, big data systems and real-time data source interfaces; the preset import method includes offline data import, or single table or batch data import, or automatic timed import, or full and incremental data import; Establish a unique identifier for the imported medical and health data; Map words with the same meaning but different expressions to slogan words; It provides standard definitions for data elements, data metrics, and data metric dimensions.
3. The management method according to claim 1, characterized in that, Data preprocessing is performed on the first hierarchical classification results, including: The first hierarchical classification results are cleaned according to the set rules; the set rules include one or more of the following: checking field type, maximum value, minimum value, maximum string length, minimum string length, missing values, and numerical precision; the data cleaning includes one or more of the following: performing one or more operations such as null value imputation, deduplication, and field filtering, and performing discretization processing on continuous data and sparse processing on classified data; The data cleaning results are then aggregated. This aggregation includes one or more of the following: associating identical entities from multiple data sources, removing redundant attributes, detecting data value conflicts and providing corresponding handling; performing multi-table joins, using methods such as left join, right join, full join, and inner join; aggregating data according to custom rules; filtering data according to custom rules; replacing all or some fields in the data stream; and splitting composite fields in the data stream according to corresponding criteria and placing the splitting results into corresponding new columns. The results of data aggregation are subjected to code value matching; the code value matching process includes one or more of the following: standardizing the code values of drugs, diseases, surgeries, tests, examinations, fees, institutions and departments; performing standard-to-standard code value mapping matching; and performing intelligent recommendations based on an artificial intelligence engine. A master index based on patient information is established based on the code value matching results to obtain a topic dataset based on patient information. Establishing a master index based on patient information includes one or more of the following: performing rule-based patient master index identification and hierarchical management of the accuracy of the patient master index; performing patient master index identification based on an artificial intelligence model. And / or, A data model based on disease information is established based on the code value matching results; the process of establishing the data model includes one or more of the following: configuring and generating a metadata template, and synchronizing metadata information based on the metadata template; mapping the medical and health organization data model to a standard data model; and copying the template of the same medical and health organization data model.
4. The management method according to claim 1, characterized in that, Data mining processing is performed on the second hierarchical classification results, including one or more of the following: By statistically analyzing the results of the second classification, the disease type and its probability for an individual or multiple individuals, and / or the disease location and its probability can be obtained. The second classification result is classified using machine learning algorithms to obtain the disease type and probability of an individual or multiple individuals, and / or the disease location and probability.
5. The management method according to claim 1, characterized in that, The hierarchical classification results are stored, including: Based on the data type of the classification results and the storage requirements of the classification results, the classification results are stored according to the preset storage performance. The classification results include the third classification results and the fourth classification results, or the classification results include the sixth classification results; The data types include: relational data, text data, image data, structured data, and semi-structured data; the storage requirements include: file storage, object storage, off-site backup, and relational databases; the preset storage performance includes high availability and horizontal scaling; high availability means minimizing downtime caused by routine maintenance operations and sudden system crashes.
6. The management method according to claim 1, characterized in that, Based on data quality requirements, the fifth-level classification results undergo data quality verification processing, including: Obtain data quality requirements, which include data management objectives, quality assessment rules, and quality assessment schemes; Obtain data validation metrics, which include consistency, accuracy, completeness, standardization, relevance, and custom metrics; Based on the data quality requirements and the data verification indicators, the fifth-level classification results are subjected to data quality verification processing; wherein, the data quality verification processing process includes: determining data connotation rules, and performing connotation analysis on the fifth-level classification results based on the connotation rules; the connotation rules represent the medical logic between data.
7. The management method according to any one of claims 1 to 6, characterized in that, It also includes one or more of the following: Perform data anonymization processing on the data; Based on the hierarchical classification of the data, corresponding hierarchical classification information security protection should be implemented for the data. Conduct risk monitoring and risk assessment on the data; Issue a safety warning.
8. A management system for health and medical data, characterized in that, include: The data acquisition and management module is used to collect data and perform initial classification and grading of the collected data to obtain the first-level classification results. The classification and grading process is based on the data source and the data security level. The data processing and management module is used to preprocess the first hierarchical classification result to obtain the preprocessed result; to perform a second hierarchical classification on the required data in the preprocessed result to obtain the second hierarchical classification result; and to perform a third hierarchical classification on all the data in the preprocessed result to obtain the third hierarchical classification result. In the data preprocessing process, a master index based on patient information and / or a data model based on disease information are established, and the preprocessed result includes a subject dataset based on patient information and / or a data model based on disease information. The data mining management module is used to perform data mining processing on the second hierarchical classification results, and then perform a fourth hierarchical classification on the data mining processing results to obtain the fourth hierarchical classification results; wherein, the data mining processing results include the disease type and probability of an individual or multiple individuals, and / or the disease location and probability. The data storage management module is used to perform a first data storage process on the third and fourth classification results, and to perform a fifth classification on the data after the first storage process to obtain the fifth classification result. The data quality management module is used to perform data quality verification processing on the fifth-level classification result according to data quality requirements, and then perform a sixth-level classification on the data quality verification result to obtain the sixth-level classification result. The data storage management module is also used to perform a second data storage process on the sixth-level classification results to form usable data.
9. An electronic device comprising a processor and a memory storing program instructions, characterized in that, The processor is configured to perform the health and medical data management method as described in any one of claims 1 to 7 when executing the program instructions.
10. A storage medium storing program instructions, characterized in that, The program instructions execute the health and medical data management method as described in any one of claims 1 to 7 during runtime.