Data management and reinforcement method based on medical security data warehouse
By implementing data governance and reinforcement methods in the medical insurance data warehouse, inconsistent problems in data aggregation methods, data verification rules, data exports, etc. have been solved, and data quality improvement, application capabilities enhancement and data sharing have been achieved, ensuring the security and efficient use of medical insurance data.
Patent Information
- Application Number
- CN202411949136.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-27
- Publication Date
- 2025-05-16
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The existing medical insurance data warehouses have problems such as inconsistent data aggregation channels, inconsistent data verification rules, inconsistent data exports, inconsistent data statistics sources, inconsistent data query sources, and inconsistent data sharing bases, resulting in low data quality, insufficient application capabilities, and unsatisfactory data sharing progress.
A data governance and reinforcement method based on medical insurance data warehouse is proposed, including data extraction, processing, storage, classification calculation, access control, permission control and data display steps. Through unified data aggregation entrance, unified data quality verification rules, unified data export, unified data statistics source, unified data query source and unified data sharing base, the standardization, standardization and efficient utilization of data are achieved.
Through this method, the data source is standardized, the data aggregation path is unified, the data quality is improved, the data export is unified, the statistical accuracy is improved, the common data query is unified, the data sharing ability is improved, and the security of medical insurance data is ensured.
Smart Images

Figure CN120011457A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of medical insurance data warehouses, and in particular to a data governance reinforcement method based on medical insurance data warehouses. Background Art
[0002] A data warehouse is a subject-oriented, integrated, non-volatile data collection used to support the formulation of management decisions. In the field of medical insurance, the medical insurance data warehouse also has such characteristics. It can integrate medical insurance-related data from multiple sources (such as medical institutions, medical insurance payment systems, etc.). The purpose of integration is to form a unified data platform for easy management and analysis, thereby providing a basis for management decisions related to medical insurance. However, the existing medical data warehouse has the following shortcomings and needs further optimization:
[0003] (1) Inconsistent data aggregation methods
[0004] At present, all the construction units of the provincial medical insurance information platform subsystem have the need for data aggregation, but they all aggregate data on their own according to needs. As a result, the channels and calibers of data aggregation are inconsistent. At the same time, some data are aggregated repeatedly, which puts pressure on the storage resources of the provincial data center and the application of data aggregation products, and is not conducive to the rational use of data center resources.
[0005] (2) Inconsistent data verification rules
[0006] At present, there is no comprehensive and unified data quality verification rule for the provincial data collected, and the data quality is not strictly tested according to the requirements. There is no data quality report issued to various cities, which is not conducive to the development of data governance in various cities. As a result, the quality of the data collected by the province to the country is not high, and the data governance work in the country is not ranked high.
[0007] (3) Inconsistent data export
[0008] At present, the work of data query, data statistical analysis, operation monitoring, etc. of each subsystem of the provincial medical security information platform is done by the construction units of each subsystem to collect data and statistical data by themselves. The inconsistent application export of data and the inconsistent statistical caliber of data have resulted in low application capabilities of related subsystems. At the same time, there is no centralized data middle-end service for providing data applications to the outside world in the province. The batch data query and export of some public services to the outside world are based on the production database to directly provide data services to the outside world, which has caused application pressure on the production environment of the provincial data center and local data centers.
[0009] (4) Inconsistent statistical sources
[0010] Currently, the monthly and quarterly reports submitted to the National Bureau are generated in the report library temporarily expanded by the data center of each city. The data statistical sources of the reports are scattered in various cities. Due to the special needs of report generation, the RDS report library expanded by various cities does not meet the stable report statistical needs in terms of data storage resources and data generation efficiency. The data statistical source for generating relevant indicators of national reports urgently needs a data source with high data service capabilities that is concentrated in the entire province.
[0011] (5) Data query sources are not unified
[0012] Currently, the data support sources for provincial common data queries are not unified, whether in terms of service capabilities for common query needs or data statistical sources and statistical calibers, and it is impossible to provide high-level data application capabilities for common data downloads and data desensitization.
[0013] (6) Data sharing base is not unified
[0014] The data sharing work is unable to provide comprehensive and accurate data sharing capabilities due to inconsistent data support bases, inconsistent data grading and classification standards, and inconsistent data quality verification, resulting in unsatisfactory progress in provincial medical insurance data sharing work. Summary of the invention
[0015] In order to solve the existing problems, the present invention provides a data governance reinforcement system method for a medical security data warehouse, and the specific scheme is as follows:
[0016] A data governance reinforcement method based on a medical insurance data warehouse includes the following steps:
[0017] S1, obtains data collected from different data sources and performs data extraction operations based on established rules;
[0018] S2, after the data to be processed is accessed from the data obtained in step S1, data processing is performed;
[0019] S3, after the data is processed, it is stored in the data warehouse and classified and calculated;
[0020] S4, transfer the data calculated in step S3 to the data service layer to implement access control, authority control and transmission encryption to control the data call permissions of other systems, and can download data in batches, or encapsulate the data into a standard API as needed;
[0021] S5 displays decision-making layer data and realizes various customized data report analysis needs and personalized data application needs.
[0022] Preferably, the data source types in step S1 include structured data and unstructured data, and the acquired data include files and message middleware stored in relational databases based on Oracle, MySQL and PostgreSQL, as well as data transmission via an API interface.
[0023] Preferably, the data processing in step S2 includes offline data processing and real-time data processing, specifically, data extraction and analysis. Specifically, offline data processing relies on the big data development and governance platform DataWorks to extract data; real-time data processing uses the data transmission service DTS to perform streaming analysis on database logs.
[0024] Preferably, step S3 includes storage and calculation of offline data, as well as storage and calculation of real-time data; wherein, offline data storage and offline calculation adopt the offline big data computing platform MaxCompute offline computing engine to realize the storage and calculation of batch structured data, and the offline big data warehouse is divided into STG, ODS, and ADS; real-time calculation and storage adopt Flink to flexibly and efficiently realize data parallel processing, and the real-time big data warehouse provides the ADS layer.
[0025] Preferably, storing and calculating data in step S3 specifically includes the following steps:
[0026] S31, in the data buffer layer STG of the data model of the data warehouse management center, the original data is integrated, and the full amount of data is stored in partitions based on time;
[0027] S32, in the data operation layer ODS, the data in the data buffer layer STG is transported, quality controlled, cleaned and managed to form the full amount of data that can be used by the subject; at the same time, a physical deletion data synchronization processing mechanism is established in the production library of the data center in each city to ensure the uniformity of the data;
[0028] S33, the data after the subject design of the data operation layer ODS is stored in the common model layer CDM, and the subject design uses LDM modeling to process the data; at the same time, according to the characteristics of medical insurance business, data monitoring direction and trend, a multi-dimensional indicator analysis model is designed to establish a multi-type common indicator library;
[0029] S34, through the data application layer (ADS), develops various thematic application data to provide carrier support for enabling innovative applications of medical insurance data.
[0030] Preferably, step S5 displays the decision-making layer data through the visualization large-screen tool dataV; uses the Alibaba Cloud intelligent reporting tool QuickBI to realize various customized data report analysis needs; and realizes personalized data application needs based on the Alibaba Cloud product technology system.
[0031] The present invention also discloses a computer-readable storage medium, on which a computer program is stored. After the computer program is run, any of the above methods is executed.
[0032] The present invention also discloses a computer system, including a processor and a storage medium, wherein a computer program is stored on the storage medium, and the processor reads and runs the computer program from the storage medium to execute any of the methods described above.
[0033] The beneficial effects of the present invention are:
[0034] (1) Standardize data sources and unify data aggregation channels
[0035] The data resource platform unifies the data aggregation entrance, and integrates the medical insurance data source, external data source and other department data sources through the unified entrance. According to the real-time necessity of the data, the data is extracted to the cache layer through real-time data synchronization and offline data synchronization to facilitate unified management.
[0036] (2) Improve data quality and unify data quality verification rules
[0037] The data in the cache layer is placed in the data operation layer and verified according to the province's unified data quality verification rules. Data without problems is placed in the common data model layer. Problematic data is uniformly governed and a problem data report is provided. After unified data governance, data quality is improved to better meet national data quality requirements and control when collecting data.
[0038] (3) Standardize data export and unify data export of medical insurance platforms
[0039] According to actual business needs and data usage requirements, we uniformly provide three types of interface services, including API interface, database interface and file interface, to meet actual needs and unify data exports, so as to better control data and provide data security.
[0040] (4) Improve statistical accuracy and unify statistical sources of report data
[0041] The statistical source of report data uniformly uses data from the data application layer. After data governance, the data quality of these data is improved, and the statistical report information data is more accurate. In addition, all cities in the province use the same data source to reduce the differences in report data caused by the use of different data in various cities, so that the report data can be better analyzed.
[0042] (5) Improve the quality and efficiency of data query and unify common data query
[0043] The common data query functions for all cities in the province are developed in a unified manner to avoid differences in queried data due to different statistical calibers in various cities, so that the queried data can be better analyzed.
[0044] (6) Improve data sharing and exchange, and unify shared data sources
[0045] Unify the data sharing source and the data sharing export, based on service management, through encryption, signing, desensitizing, hierarchical authorization and other methods, on the basis of security and controllability, share the medical insurance data to the Provincial Data Resources Bureau through the interface to realize the opening of data service capabilities. As a platform open to the society, it has the ability to support Internet-level concurrent responses.
[0046] (7) Ensure the security of medical insurance data and unify data security management
[0047] According to the data classification and classification criteria, privacy desensitization protection and security confidentiality protection are strengthened, a security protection system is established, and security protection capabilities are improved. Mainly including: data permission control, data desensitization protection, data encryption circulation, data application traces, etc. BRIEF DESCRIPTION OF THE DRAWINGS
[0048] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0049] Figure 1 is a flow chart of the method of the present invention;
[0050] Figure 2 The figure is a block diagram of the system structure principle based on the method of the present invention. DETAILED DESCRIPTION
[0051] In order to make the purpose, technical solution and advantages of the embodiments of the present invention clearer, the technical solution in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0052] like Figure 1 , which is a data governance reinforcement method for a medical insurance data warehouse, comprising the following steps:
[0053] S1, obtains data collected and accessed from different data sources, and performs data extraction operations based on established rules.
[0054] Among them, data source types include structured data and unstructured data, and the acquired data includes data stored in relational databases based on Oracle, MySQL and PostgreSQL, as well as files and message middleware that realize data transmission through API interfaces.
[0055] S2, after the data to be processed is accessed from the data obtained in step S1, data processing is performed.
[0056] Among them, data processing includes offline data processing and real-time data processing. Specifically, data is extracted and analyzed. Specifically, offline data processing relies on the big data development and governance platform DataWorks to extract data; real-time data processing uses the data transmission service DTS to perform streaming analysis on database logs.
[0057] S3, after the data is processed, it is stored in the data warehouse and classified and calculated, including the storage and calculation of offline data and the storage and calculation of real-time data.
[0058] Among them, offline data storage and offline computing use the offline big data computing platform MaxCompute offline computing engine to realize the storage and computing of batch structured data. The offline big data warehouse is divided into STG, ODS, and ADS. Real-time computing and storage use Flink to flexibly and efficiently realize data parallel processing. The real-time big data warehouse provides the ADS layer.
[0059] The storage and calculation of data in step S3 specifically includes the following steps:
[0060] S31, in the data buffer layer STG of the data model of the data warehouse management center, the original data is integrated, and the full amount of data is stored in partitions based on time;
[0061] S32, in the data operation layer ODS, the data in the data buffer layer STG is transported, quality controlled, cleaned and managed to form the full amount of data that can be used by the subject; at the same time, a physical deletion data synchronization processing mechanism is established in the production library of the data center in each city to ensure the uniformity of the data;
[0062] S33, the data after the subject design of the data operation layer ODS is stored in the common model layer CDM, and the subject design uses LDM modeling to process the data; at the same time, according to the characteristics of medical insurance business, data monitoring direction and trend, a multi-dimensional indicator analysis model is designed to establish a multi-type common indicator library;
[0063] S34, through the data application layer (ADS), develops various thematic application data to provide carrier support for enabling innovative applications of medical insurance data.
[0064] S4, transfers the data calculated in step S3 to the data service layer to implement access control, authority control and transmission encryption to control the data call permissions of other systems, and can download data in batches, or encapsulate the data into a standard API as needed.
[0065] S5, display the decision-making data, and realize various customized data report analysis needs and personalized data application needs. Display the decision-making data through the visualization large screen tool dataV; use Alibaba Cloud intelligent reporting tool QuickBI to realize various customized data report analysis needs; realize personalized data application needs based on Alibaba Cloud product technology system.
[0066] The present invention also discloses a computer-readable storage medium and a computer system, wherein a computer program is stored on the computer-readable storage medium, and after the computer program is run, the method described above is executed. A computer system includes a processor and a storage medium, wherein the computer program is stored on the storage medium, and the processor reads and runs the computer program from the storage medium to execute the method described above.
[0067] like Figure 2 A system based on a data governance reinforcement method for a medical insurance data warehouse includes a data resource integration center, a data warehouse management center, a data asset management center, a data security management center, a data quality governance center, a data application service center, and a visual data monitoring center.
[0068] 1. Data resource integration center: provides access to data resources. The data access function is capable of extracting data from different data sources (structured data, unstructured data) according to specified rules. The extracted data can be stored in two categories: on-site or off-site. The extracted data can provide input for data conversion, or it can be directly processed or loaded.
[0069] It provides the capabilities of data collection, data conversion and cleaning, data loading, and data export for mainstream data sources. It can exchange data with column-based databases, traditional databases, data warehouses, local and remote files, etc. It supports streaming collection of log files and data analysis, and supports ETL extraction in the form of message queues.
[0070] The data resource integration center includes a data source management module, a real-time data synchronization module, an offline data synchronization module, and an integrated scheduling module.
[0071] 1. Data source management module: used for visual operation, configuration and task management of various data sources in the process of data resource integration. It supports the storage and management of various types of data in a wide range. The main sources of data include medical insurance data sources such as basic information center, user center, policy center, insurance participation center, collection center, settlement center, procurement center, electronic voucher center, and other data sources shared by other departments.
[0072] 2. Real-time data synchronization module: For some data applications with high real-time requirements, it provides real-time data processing services, collects real-time and quasi-real-time data directly from the source system in the form of message queues, and processes and processes the real-time data through the real-time data processing framework, providing real-time and quasi-real-time data services for the front-end application system.
[0073] 3. Offline data synchronization module: By establishing an offline data batch collection channel, the data synchronization requirements from the data source layer to the big data warehouse are met. The data synchronization frequency is recommended to be T+1. The data synchronization method is to realize data batch collection through ETL tools, and to complete the warehousing after realizing big data offline calculation processing through big data offline calculation tools.
[0074] Since there is data dependency between various business subsystems, in order to reduce possible problems in the data governance process, it is necessary to formulate a time sequence for data integration based on the specific circumstances of each business subsystem.
[0075] Historical stock data is synchronized once and collected into the data warehouse in batches; daily incremental T+1 data is synchronized to the data warehouse regularly according to the data update time.
[0076] 4. Integrated scheduling module: including real-time and offline scheduling task monitoring, trigger-type task monitoring, and unified management of data-integrated scheduling tasks. It is used to build task modules according to user needs, configure scheduling strategies, provide equal interval scheduling and timed scheduling, and has two startup modes: automatic operation and manual execution, thus forming a unified scheduling management module. In the scheduling list, you can view the log information and running status of the scheduling, and you can also edit, view, delete, pause, and enable the scheduling.
[0077] 2. Data Warehouse Management Center:
[0078] The big data warehouse is the core carrier of the construction of the data center content end. It is used to unify data standards and indicator calibers, unify the medical insurance business data needed by the application layer, and organize and store them in a certain way. Through the construction of the big data warehouse, the data standards and indicator calibers are unified, the data quality is guaranteed, and the decoupling between the business layer data model and the big data application layer data model is achieved. At the same time, it can avoid duplication of construction and save storage space.
[0079] The data warehouse management center includes data model layering module, data model subject domain division module, data modeling module and indicator library design module;
[0080] 1. Data model layering module: used to organize and store the medical insurance data warehouse in a certain way. The data model layering module divides the medical insurance data warehouse into four layers, namely the data buffer layer STG, the operation data layer ODS, the common model layer CDM and the application data layer ADS.
[0081] (1) Data buffer layer STG: It is the first storage area in the process of business data flow. The data warehouse extracts data from the data sources of various business systems and loads it into the data buffer layer, thus laying a solid foundation for business data cleaning, quality control and other processes. The data cache layer can improve the convenience of ETL processing and the efficiency of ETL execution.
[0082] The data buffer layer STG should meet the following requirements:
[0083] 1) The table structure is consistent or similar to that of the source system;
[0084] 2) Capable of quickly deleting data;
[0085] 3) Ability to quickly query detailed data;
[0086] 4) Each record includes the insertion and update time fields of the data;
[0087] 5) Each record includes job instances for inserting and updating data.
[0088] (2) The operational data layer (ODS) integrates the data generated by each isolated business system, extracts, cleans, and transmits it, and then loads it into this layer to form a unified, global data set. The data in the ODS layer is isomorphic to the data in the source system, with the purpose of simplifying the subsequent data processing. In terms of data granularity, the data granularity of the ODS layer is the finest. The operational data layer is an area where data in the cache layer is cleaned and governed based on data standards and quality verification rules, using data development and governance tools. It is also an area where current and historical data are stored.
[0089] It is necessary to perform quality control work such as denoising, deduplication, and dirty data removal at this layer, standardize and normalize detailed data, and record and mark problematic data (dirty data). The operational data layer should meet the following requirements:
[0090] 1) The table structure should be consistent or similar to that of the source system;
[0091] 2) Capable of supporting rapid data deletion;
[0092] 3) Support detailed data query function;
[0093] 4) Capable of storing all historical data;
[0094] 5) Capable of date-based, system-based rapid deletion and data extraction;
[0095] 6) Each record includes the insertion and update time fields of the data;
[0096] 7) Each record includes job instances for inserting and updating data.
[0097] (3) Common model layer CDM: It is used to store the data after the theme design of the operation data layer. The theme design uses LDM modeling to process the data.
[0098] The general model layer should meet the following requirements:
[0099] 1) Use LDM to perform topic modeling on the data;
[0100] 2) Ability to identify subordinate topics from table names;
[0101] 3) Each record includes the data insertion and update time fields;
[0102] 4) Each record includes the job instance number of the inserted and updated data.
[0103] In the actual construction process, the general model layer will be further subdivided into the detailed data layer (DWD) and the data summary layer (DWS).
[0104] Among them, the detailed data layer DWD:
[0105] Maintain the same data granularity as the operational data layer ODS, provide certain data quality assurance, process data based on ODS, and provide cleaner data; at the same time, in order to improve the usability of the data detail layer, this layer will use some dimension degradation techniques. When a dimension does not have any data required by the data warehouse, it can be degraded to the fact table to reduce the association between the fact table and the dimension table. For example: There is no need to use a dimension table to store a large dimension such as the order id, but the order id is very important when we generally perform data analysis, so we redundantly store the order id in the fact table. This dimension is a degenerate dimension.
[0106] This layer designs zipper tables and flow tables as needed. Zipper tables are defined for the way tables store data in data warehouse design. As the name suggests, zipper records history, recording all the changes of a thing from its beginning to its current state.
[0107] The flow table stores a user's change records. For example, in a flow table, each modification record of a user will be recorded in the daily data, but there is only one record in the zipper table.
[0108] This is a granularity issue that needs to be taken into account when designing a zipper table. We can of course set the granularity to be smaller, usually days are enough.
[0109] The data aggregation layer DWS is divided by subject and generates wide tables to provide data support for subsequent business queries, OLAP analysis and data distribution. This layer may use various types of databases to meet the needs of specific application scenarios. This layer has the clearest requirements. Data dimensions and analysis results are designed according to business needs. Therefore, data in this layer can be directly connected to OLAP analysis, or directly used or displayed by data applications.
[0110] (4) Application Data Layer ADS:
[0111] It provides data services for applications and stores data converted and processed from the data aggregation layer. The data structure in this layer is independently designed by each application vendor according to the needs of the data application theme. This layer may use various types of databases to meet the needs of specific application scenarios. This layer has the clearest requirements. The data dimensions and analysis results are designed according to business needs. Therefore, the data in this layer can be directly connected to OLAP analysis, and can also be directly used or displayed by data applications.
[0112] 1. Data model subject domain division module: Subject domain is a collection of closely related data subjects. It is divided and abstracted according to the perspective of business demand analysis. According to the business focus, these data subjects are divided into different subject domains. A subject is a scope for integrating, classifying and analyzing the data of a certain analysis object in various systems in production at a higher level. It is an abstract concept, and each subject corresponds to a macro analysis field. According to the business process, a subject domain is abstracted from one business process.
[0113] 2. The data modeling module is used to establish and maintain a set of effective workflows and specifications to ensure that different logical data model designers can operate in accordance with a unified standard.
[0114] 3. Indicator library design module: Indicators are usually organized and managed with business as the center to support various complex business scenarios and analysis needs. The indicator library processes and converts data for business purposes to generate various indicator items for specific business needs. Modeling and development based on specific business scenarios and business needs are the basis for realizing dynamic configuration and generation of reports.
[0115] 3. Data Asset Management Center:
[0116] According to the hierarchical structure and subject domain division of data integration, various objects at each layer, such as tables, stored procedures, indexes, data links, functions and packages, are managed; regular attention is paid to the operation of collection tasks to locate and resolve some anomalies that may occur during the collection process; data authority control and access data encryption and desensitization are performed to improve data security.
[0117] The data asset management center includes a metadata management module and an asset catalog management module.
[0118] 1. Metadata management module:
[0119] According to the hierarchical structure and subject domain division of data integration, it is necessary to implement the management of various objects at each layer, such as tables, stored procedures, indexes, data links, functions and packages. Clearly represent the data flow between the hierarchical structures, the relationship between the objects, and the information of various data services provided to the outside. The metadata content involves all data links of the entire big data resource platform, including data collection, layer-by-layer processing and auditing, and the processing from data services to the final application presentation. Metadata management runs through the entire process and achieves effective interaction with each link. Metadata management includes metadata definition, query, maintenance, inspection, analysis, lineage management and data maps.
[0120] 2. Asset Catalog Management Module
[0121] The main function of asset catalog management is to use metadata to describe the characteristics of information resources, form unified and standardized catalog content, and form a catalog information database through effective organization and management of catalog content, providing information resource discovery and positioning services for the aggregation and sharing of information resources and support for applications.
[0122] All government information resources are organized and managed in accordance with unified standards and specifications, and directory content query and retrieval services are provided to users through the directory system based on the directory information base. Through the construction of the directory system, the information resources of each business department are catalogued and dynamically managed, making it easier to fully grasp the overall information resource status of each department.
[0123] 4. Data Security Management Center:
[0124] It is used to standardize a unified data security use system, establish a data security management center, standardize the safe use of data, build a unified data security sharing center across the province, and ensure that medical insurance data sharing, transmission and sharing meet data security management requirements. The overall architecture design is based on the two key mechanisms of cloud platform data security and resource isolation to ensure the maturity, stability, security and reliability of the business.
[0125] The data security management center must establish complete information security management measures, rely on the data security management system, strengthen the security control of data applications, follow the ethical standards and information security level protection standards of the medical insurance industry in both clinical research and patient services, only provide the minimum data set required for the business, and conduct access audits.
[0126] The data security management center includes a data classification module, a data authority control module and a data desensitization module.
[0127] 1. Data classification module:
[0128] Data classification aims to comprehensively sort out data assets and determine the corresponding data level, which is a necessary prerequisite and foundation for implementing effective data classification management. Data classification management is the basic work for establishing a unified and complete data lifecycle security protection framework, which can provide support for the medical insurance system to formulate targeted data security control measures.
[0129] Medical insurance data levels are divided into three levels from high to low: core level, important level, and general level.
[0130] Core level: medical insurance data related to key areas of national security, the lifeline of the national economy, important livelihoods and major public interests, as well as other medical insurance data determined after evaluation. Once such data is illegally used or shared, it may directly affect political security.
[0131] Important level: Medical insurance data that may affect national security, economic operation, social stability, public health and safety after being leaked, tampered with or damaged. Medical insurance data that only affects organizations and individuals is generally not considered important data.
[0132] General level: medical insurance data other than core level and important level.
[0133] ①Important data:
[0134] User account password information, digital certificate (private key) information.
[0135] The service population covers the entire city. Medical insurance-related diagnosis and treatment services, personal information, medical insurance settlement, two-fixed institution settlement information, insurance participation information, and collection and payment information data.
[0136] ②Core data:
[0137] The service population covers the entire province. Medical insurance-related diagnosis and treatment services, personal information, medical insurance settlement, two-fixed institution settlement information, insurance information, and collection information data.
[0138] 2. Data desensitization module:
[0139] Based on the classification and grading results of sensitive fields in corporate customers, individual users, and business data in the communications industry, an automated data classification and grading management capability is formed; based on the automated detection and discovery technology of sensitive data, privacy protection processing of the return results of user data requests is achieved; based on the desensitizing function library with multi-modality and multi-business requirements, a hardware-level high-speed dynamic / static desensitization capability is formed for data sets involving sensitive data threats.
[0140] By establishing basic rules such as sensitive data keyword library, sensitive data combination rule library, and sensitive information processing rule library, sensitive fields in the original data can be processed without affecting the accuracy of data analysis results, thereby reducing data sensitivity and reducing the risk of personal privacy leakage. Common data desensitization methods mainly include:
[0141] 1) Data replacement
[0142] Replace the real value with a fixed imaginary value set;
[0143] 2) Reverse inference
[0144] Find mappings that may infer sensitive fields from certain fields and desensitize these fields;
[0145] 3) Offset and rounding
[0146] By randomly shifting digital data, offset rounding ensures the approximate authenticity of the range while maintaining data security.
[0147] 4) Mask shielding
[0148] Masking is a powerful tool for desensitizing some information in account data.
[0149] 5) Flexible encoding
[0150] When special desensitization rules are needed, flexible encoding can be performed to meet various possible desensitization rules.
[0151] 6) Invalidation
[0152] Sensitive data can be desensitized by truncating, encrypting, hiding, etc., so that it is no longer of use value.
[0153] 7) Randomization
[0154] Replace the true value with random data, keeping the randomness of the replacement value to simulate the authenticity of the sample.
[0155] 8) Occlusion
[0156] Refers to replacing part of the sensitive data with masking symbols (such as "X, *"), so that part of the sensitive data remains public. This method can largely desensitize while maintaining the original data appearance.
[0157] 3. Data permission control module:
[0158] Through centralized authorization, fine-grained authorization and other forms, the authorized scope of IP addresses and accounts of data exports is strictly controlled to provide strict data security isolation. The security management center has a convenient permission management function, provides a visual application approval process, and can audit and manage permissions, which improves data security and facilitates data permission management.
[0159] Apply for permissions online: Select the data table you need permissions for and quickly apply online, changing the original offline mode of contacting administrators to improve work efficiency.
[0160] Permission audit / return: Administrators can quickly and easily view the corresponding personnel of database table permissions and conduct audit management. Users can also actively return permissions that are no longer needed.
[0161] Permission approval management: The previous mode of direct authorization by administrators is changed to an approval authorization mode, providing a visual and process-based management authorization mechanism, and the approval process can be traced back afterwards.
[0162] In the Security Management Center module, you can view global data table permissions within the organization, manage table permissions, and apply for / approve data table permissions.
[0163] The data center supports multi-dimensional management and control of members of the organization through methods such as organization Owner account, AccessKey and AccessSecret.
[0164] 5. Data Quality Governance Center:
[0165] Used to configure data quality check rules, establish data quality check tasks, regularly check data quality, provide daily data quality operation result reports, and manage dirty data.
[0166] The data quality governance center includes a data quality rule configuration module, a data quality task configuration module, a data quality report module, and a dirty data management module.
[0167] 1. Data quality rule configuration module:
[0168] The information platform can configure quality rules and establish a quality rule library. The quality rule library is used to store data quality inspection rules defined by users according to data quality standards. The rules cover the following aspects:
[0169] 1) Integrity check: including table record missing check and field value missing check;
[0170] 2) Consistency check: including field value type, length, format, code and rationality check;
[0171] 3) Accuracy check: record the rationality of the business, check the field value range, and check whether the characters are garbled;
[0172] 4) Timeliness check: Check whether the data is collected, processed and output in a timely manner;
[0173] 5) Others: National data inspection rules.
[0174] 2. Data quality task configuration module:
[0175] To ensure data quality, a data quality detection scheduling task is established, including:
[0176] 1) Establish a daily data inspection mechanism to conduct quality inspections on collected and managed data;
[0177] 2) Establish a data problem alarm and timely response mechanism.
[0178] 3. Data quality reporting module:
[0179] The data quality report provides a daily data quality report for each data table. According to the definition of quality rules, the data quality report is also divided into table-level rule quality reports and field-level quality reports.
[0180] (1) Table-level quality rules:
[0181] The system will calculate the output time of the table and the number of data items in each table by default. Users can set the rule range for the output time and the number of records. The set rules will be displayed on the chart as a red straight line or curve.
[0182] The association relationship between the master and slave tables will display the associated fields of the master and slave tables, as well as the number of data entries in the associated fields of the current slave table that do not belong to the master table.
[0183] (2) Field-level quality rules:
[0184] The field-level quality rules will display all the field quality verification rules that have been configured in the table, including the field name, verification rule, the number and percentage of items that meet the rules, the number and percentage of items that do not meet the rules, and a link to view error data.
[0185] 4. Dirty data management module:
[0186] Click View Dirty Data in the rule list to enter the dirty data viewing page. The first column displays the column name and data of the current quality rule. The following columns display the data in the order of the fields in the table. Users can download the current dirty data.
[0187] 6. Data Application Service Center:
[0188] A service interface used to provide data empowerment to the outside world, and includes a data service management module and a data service application module.
[0189] The data middle platform provides data-enabled service interfaces to the outside world through the data service center. Depending on the service type and data type, it mainly provides three types of interface services, including API interface, database interface and file interface.
[0190] API interface:
[0191] The API interface is the data access and calling method used by the data warehouse (data center) to access external applications or data. The API interface should meet the following requirements:
[0192] 1) API generation has basic capabilities such as paging and filtering;
[0193] 2) Support HTTP / HTTPS standard interface protocol;
[0194] 3) Support POST and GET requests;
[0195] Database interface:
[0196] The database interface is the database access service provided by the data warehouse (data center) to the outside world and should meet the following requirements:
[0197] 1) Support offline data services, and authorize offline tables in the big data warehouse to users, so that they can further build data application layers based on business needs;
[0198] 2) Support real-time data services, provide streaming data and OLAP data to users, and support the streaming data needs and ad hoc query and analysis needs of business systems;
[0199] 3) Support real-time data warehouses to collect data in the form of real-time data (mainly streaming data) and offline data, support real-time data warehouses to provide data in the form of streaming data, and support standard database interface queries;
[0200] 4) Support real-time, offline data authorization management, and support table-level authorization (preferably support field-level authorization).
[0201] Data file interface:
[0202] The data file interface is the data access and calling method used by the data warehouse (data center) for external application or data access needs. The data file interface should meet the following requirements:
[0203] 1) Support the ability to provide external data services in the form of files;
[0204] 2) Supports multiple mainstream data file storage including FTP and OSS as data sources;
[0205] 3) Support mainstream file types such as CSV and TEXT;
[0206] 4) Supports setting basic data file settings such as character set, compression status, row and column separators, etc.
[0207] 5) Support viewing file transfer status, including transferring, completed, failed and reasons, etc.
[0208] 1. Data service management module
[0209] Control data call permissions of other systems through access control, permission control and transmission encryption.
[0210] 2. Data service application module
[0211] (1) Data collection application module
[0212] The data of core business systems, basic information centers, insurance participation centers, payment centers, settlement centers, procurement centers, and group payment centers are collected by the National Bureau, and the T+1 data source is obtained from the data warehouse ODS layer to ensure data consistency.
[0213] (2) Internal data application module
[0214] Based on the relevant data requirements of the internal subsystems of the provincial medical insurance information platform, corresponding data application topics are generated to meet the needs of data applications of business subsystems such as public services, operation monitoring, intelligent supervision, and group payment.
[0215] (3) External shared application modules
[0216] Based on the service management model, the data sharing and exchange center shares medical insurance data with the Provincial Data Resources Bureau and other departments and units in a safe and controllable manner through encryption, signing, desensitization, hierarchical authorization, etc. At the same time, according to the needs of medical insurance business, relevant data information is obtained from external departments such as public security, education, health, human resources and social security, and the Data Resources Bureau.
[0217] (4) Data statistics application module
[0218] The data statistics application includes the national report module and the data query module. The national report module generates various situations such as the basic medical insurance participants, the employees’ basic medical insurance participants and special personnel, the collection of employees’ basic medical insurance premiums, and the medical expenses of employees’ basic medical insurance by selecting the annual quarter, administrative division, and insurance type; the data query module displays the details of employees’ insurance participation, residents’ insurance participation, medical insurance settlement, electronic voucher usage rate, mobile payment, etc. within the time range through information query.
[0219] 7. Visual data monitoring center:
[0220] It is used to achieve large-screen panoramic display of key data information, realize data cockpit-style display and management, and visualize data monitoring information such as the province's data aggregation trajectory, integration status, shared analysis, etc.
[0221] Those skilled in the art will further appreciate that the various illustrative logic blocks, modules, circuits, and algorithm steps described in conjunction with the embodiments disclosed herein may be implemented as electronic hardware, computer software, or a combination of the two. To clearly illustrate this interchangeability of hardware and software, various illustrative components, blocks, modules, circuits, and steps are generally described above in terms of their functionality. Whether such functionality is implemented as hardware or software depends on the specific application and the design constraints imposed on the overall system. The technician may implement the described functionality in different ways for each specific application, but such implementation decisions should not be interpreted as resulting in a departure from the scope of the present invention.
[0222] The various illustrative logic blocks, modules, and circuits described in conjunction with the embodiments disclosed herein may be implemented or performed with a general purpose processor, a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA) or other programmable logic device, discrete gate or transistor logic, discrete hardware components, or any combination thereof designed to perform the functions described herein. A general purpose processor may be a microprocessor, but in the alternative, the processor may be any conventional processor, controller, microcontroller, or state machine. The processor may also be implemented as a combination of computing devices, such as a combination of a DSP and a microprocessor, a plurality of microprocessors, one or more microprocessors in cooperation with a DSP core, or any other such configuration.
[0223] The steps of the method or algorithm described in conjunction with the embodiments disclosed herein may be embodied directly in hardware, in a software module executed by a processor, or in a combination of the two. The software module may reside in a RAM memory, a flash memory, a ROM memory, an EPROM memory, an EEPROM memory, a register, a hard disk, a removable disk, a CD-ROM, or any other form of storage medium known in the art. An exemplary storage medium is coupled to a processor so that the processor can read and write information from / to the storage medium. In an alternative, a storage medium may be integrated into a processor. The processor and the storage medium may reside in an ASIC. The ASIC may reside in a user terminal. In an alternative, the processor and the storage medium may reside in a user terminal as discrete components.
[0224] In one or more exemplary embodiments, the functions described may be implemented in hardware, software, firmware, or any combination thereof. If implemented as a computer program product in software, each function may be stored on or transmitted by a computer-readable medium as one or more instructions or codes. Computer-readable media include both computer storage media and communication media, including any medium that facilitates the transfer of a computer program from one place to another. Storage media may be any available medium that can be accessed by a computer. As an example and not limitation, such a computer-readable medium may include RAM, ROM, EEPROM, CD-ROM or other optical disk storage, disk storage or other magnetic storage device, or any other medium that can be used to carry or store the desired program code in the form of an instruction or data structure and can be accessed by a computer. Any connection is also properly referred to as a computer-readable medium. For example, if the software is transmitted from a website, a server, or other remote source using a coaxial cable, a fiber optic cable, a twisted pair, a digital subscriber line (DSL), or wireless technologies such as infrared, radio, and microwaves, the coaxial cable, fiber optic cable, twisted pair, DSL, or wireless technologies such as infrared, radio, and microwaves are included in the definition of the medium. Disk and disc as used herein include compact disc (CD), laser disc, optical disc, digital versatile disc (DVD), floppy disk and Blu-ray disc, wherein disk often reproduces data magnetically, while disc reproduces data optically with lasers. Combinations of the above should also be included within the scope of computer-readable media.
[0225] The previous description of the disclosure is provided to enable any person skilled in the art to make or use the disclosure. Various modifications to the disclosure will be apparent to those skilled in the art, and the general principles defined herein may be applied to other variations without departing from the spirit or scope of the disclosure. Thus, the disclosure is not intended to be limited to the examples and designs described herein, but should be granted the widest scope consistent with the principles and novel features disclosed herein.
[0226] Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent substitutions for some of the technical features therein; and these modifications or substitutions do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A data governance and reinforcement method based on medical insurance data warehouse, characterized in that: The following steps are involved: S1, obtains data collected from different data sources and performs data extraction operations based on established rules; S2, after the data to be processed is accessed from the data obtained in step S1, data processing is performed; S3, after the data is processed, it is stored in the data warehouse and classified and calculated; S4, transfer the data calculated in step S3 to the data service layer to implement access control, authority control and transmission encryption to control the data call permissions of other systems, and can download data in batches, or encapsulate the data into a standard API as needed; S5 displays decision-making layer data and realizes various customized data report analysis needs and personalized data application needs.
2. The method according to claim 1, characterized in that: The data source types in step S1 include structured data and unstructured data, and the acquired data include files and message middleware stored in relational databases based on Oracle, MySQL and PostgreSQL, as well as data transmission through API interfaces.
3. The method according to claim 1, characterized in that: The data processing in step S2 includes offline data processing and real-time data processing, specifically, data extraction and analysis. Specifically, offline data processing relies on the big data development and governance platform DataWorks to extract data; real-time data processing uses the data transmission service DTS to perform streaming analysis on database logs.
4. The method according to claim 1, characterized in that: Step S3 includes storage and calculation of offline data, and storage and calculation of real-time data; Among them, offline data storage and offline computing use the offline big data computing platform MaxCompute offline computing engine to achieve batch structured data storage and computing. The offline big data warehouse is divided into STG, ODS, and ADS; Flink is used for real-time computing and storage, which enables flexible and efficient data parallel processing. The real-time big data warehouse provides the ADS layer.
5. The method according to claim 4, characterized in that: The storage and calculation of data in step S3 specifically includes the following steps: S31, in the data buffer layer STG of the data model of the data warehouse management center, the original data is integrated, and the full amount of data is stored in partitions based on time; S32, in the data operation layer ODS, the data in the data buffer layer STG is transported, quality controlled, cleaned and managed to form the full amount of data that can be used by the subject; at the same time, a physical deletion data synchronization processing mechanism is established in the production library of the data center in each city to ensure the uniformity of the data; S33, the data after the subject design of the data operation layer ODS is stored in the common model layer CDM, and the subject design uses LDM modeling to process the data; at the same time, according to the characteristics of medical insurance business, data monitoring direction and trend, a multi-dimensional indicator analysis model is designed to establish a multi-type common indicator library; S34, through the data application layer (ADS), develops various thematic application data to provide carrier support for enabling innovative applications of medical insurance data.
6. The method according to claim 1, characterized in that: Step S5 displays the decision-making layer data through the visualization large-screen tool dataV; uses Alibaba Cloud's intelligent reporting tool QuickBI to meet various customized data report analysis needs; and implements personalized data application needs based on Alibaba Cloud's product technology system.
7. A computer-readable storage medium, characterized in that: A computer program is stored on the medium, and after the computer program is run, the method according to any one of claims 1 to 6 is executed.
8. A computer system, characterized in that: The method comprises a processor and a storage medium, wherein a computer program is stored in the storage medium, and the processor reads and runs the computer program from the storage medium to execute the method as claimed in any one of claims 1 to 6.
Citation Information
Patent Citations
Multi-source heterogeneous mass data acquisition and treatment system based on middleware of Internet of Things
CN117851389A
System and method for planning and implementing a data warehouse solution
US7092968B1
Smart agriculture AIOT distributed big data storage platform
WO2023004881A1
Cited By
Locomotive big data base framework and working method thereof
CN120670500A