Data management and reinforcement system for medical security data warehouse

By establishing a data governance reinforcement system, the problem of inconsistent data in the medical insurance data warehouse is solved, the standardized aggregation, verification and security management of data is realized, the data quality and application capabilities are improved, and the needs of unified management and analysis are met.

CN120011456AInactive Publication Date: 2025-05-16ANHUI YILIANZHONG TECH DEV CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202411948504.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-27
Publication Date
2025-05-16
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The existing medical insurance data warehouse has problems such as inconsistent data aggregation channels, verification rules, exports, statistical sources and sharing bases, resulting in low data quality and insufficient application capabilities, which cannot meet the needs of unified management and analysis.

Method used

Establish a data governance and reinforcement system for medical insurance data warehouses, including a data resource integration center, a data warehouse management center, a data asset management center, a data security management center, a data quality governance center and a visual data monitoring center. Through unified data aggregation, verification, export and security management, improve data quality and application capabilities.

Benefits of technology

It realizes the standardized aggregation of data sources, unifies data quality verification rules and exports, improves data statistics accuracy and sharing capabilities, ensures data security, and improves the resource utilization efficiency and management level of data centers.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120011456A_ABST
    Figure CN120011456A_ABST
Patent Text Reader

Abstract

The invention discloses a data governance reinforcement system for a medical security data warehouse. The system comprises a data resource integration center, a data warehouse management center, a data asset management center, a data safety management center, a data quality governance center, a data application service center and a visual data monitoring center. The data resource integration center provides convergence access of data resources; the data warehouse management center is used for uniformly converging data and carrying out hierarchical organization and storage; the data asset management center realizes management of each object of each layer and provides an information resource positioning service; the data security management center is used for establishing a data security management center and standardizing the security use of the data; the data quality management center is used for configuring a data quality inspection rule and managing dirty data; and the visual data monitoring center realizes large-screen panoramic display of data key information. The method has the advantages of unifying data aggregation ways, improving data quality, unifying data quality verification rules and medical insurance platform data exits and the like.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of medical insurance data warehouses, and in particular to a data governance and reinforcement system for medical insurance data warehouses. Background Art

[0002] A data warehouse is a subject-oriented, integrated, non-volatile data collection used to support the formulation of management decisions. In the field of medical insurance, the medical insurance data warehouse also has such characteristics. It can integrate medical insurance-related data from multiple sources (such as medical institutions, medical insurance payment systems, etc.). The purpose of integration is to form a unified data platform for easy management and analysis, thereby providing a basis for management decisions related to medical insurance. However, the existing medical data warehouse has the following shortcomings and needs further optimization:

[0003] (1) Inconsistent data aggregation methods

[0004] At present, all the construction units of the provincial medical insurance information platform subsystem have the need for data aggregation, but they all aggregate data on their own according to needs. As a result, the channels and calibers of data aggregation are inconsistent. At the same time, some data are aggregated repeatedly, which puts pressure on the storage resources of the provincial data center and the application of data aggregation products, and is not conducive to the rational use of data center resources.

[0005] (2) Inconsistent data verification rules

[0006] At present, there is no comprehensive and unified data quality verification rule for the provincial data collected, and the data quality is not strictly tested according to the requirements. There is no data quality report issued to various cities, which is not conducive to the development of data governance in various cities. As a result, the quality of the data collected by the province to the country is not high, and the data governance work in the country is not ranked high.

[0007] (3) Inconsistent data export

[0008] At present, the work of data query, data statistical analysis, operation monitoring, etc. of each subsystem of the provincial medical security information platform is done by the construction units of each subsystem to collect data and statistical data by themselves. The inconsistent application export of data and the inconsistent statistical caliber of data have resulted in low application capabilities of related subsystems. At the same time, there is no centralized data middle-end service for providing data applications to the outside world in the province. The batch data query and export of some public services to the outside world are based on the production database to directly provide data services to the outside world, which has caused application pressure on the production environment of the provincial data center and local data centers.

[0009] (4) Inconsistent statistical sources

[0010] Currently, the monthly and quarterly reports submitted to the National Bureau are generated in the report library temporarily expanded by the data center of each city. The data statistical sources of the reports are scattered in various cities. Due to the special needs of report generation, the RDS report library expanded by various cities does not meet the stable report statistical needs in terms of data storage resources and data generation efficiency. The data statistical source for generating relevant indicators of national reports urgently needs a data source with high data service capabilities that is concentrated in the entire province.

[0011] (5) Data query sources are not unified

[0012] Currently, the data support sources for provincial common data queries are not unified, whether in terms of service capabilities for common query needs or data statistical sources and statistical calibers, and it is impossible to provide high-level data application capabilities for common data downloads and data desensitization.

[0013] (6) Data sharing base is not unified

[0014] The data sharing work is unable to provide comprehensive and accurate data sharing capabilities due to inconsistent data support bases, inconsistent data grading and classification standards, and inconsistent data quality verification, resulting in unsatisfactory progress in provincial medical insurance data sharing work. Summary of the invention

[0015] In order to solve the existing problems, the present invention provides a data governance and reinforcement system for a medical insurance data warehouse. The specific scheme is as follows:

[0016] A data governance and reinforcement system for a medical security data warehouse, including a data resource integration center, a data warehouse management center, a data asset management center, a data security management center, a data quality governance center, a data application service center, and a visual data monitoring center;

[0017] The data resource integration center provides aggregated access to data resources;

[0018] The data warehouse management center is used to unify data standards and indicator calibers, gather the medical insurance business data required by the application layer, and organize and store them in layers in a certain way;

[0019] The data asset management center implements the management of various objects at each layer according to the hierarchical structure and subject domain division of data integration, and provides information resource discovery and positioning services;

[0020] The data security management center is used to standardize a unified data security use system, establish a data security management center, standardize the safe use of data, build a unified data security sharing center across the province, and ensure that the sharing, transmission and sharing of medical insurance data meet the data security management requirements;

[0021] The data quality governance center is used to configure data quality inspection rules, establish data quality inspection tasks, regularly inspect data quality, provide daily data quality operation result reports, and manage dirty data;

[0022] The data application service center is used to provide a data empowerment service interface to the outside, and includes a data service management module and a data service application module;

[0023] The visual data monitoring center is used to realize large-screen panoramic display of key data information, realize data cockpit-style display and management, and visualize data monitoring information such as data aggregation trajectory, integration status, shared analysis, etc. across the province.

[0024] The beneficial effects of the present invention are:

[0025] (1) Standardize data sources and unify data aggregation channels

[0026] The data resource platform unifies the data aggregation entrance, and integrates the medical insurance data source, external data source and other department data sources through the unified entrance. According to the real-time necessity of the data, the data is extracted to the cache layer through real-time data synchronization and offline data synchronization to facilitate unified management.

[0027] (2) Improve data quality and unify data quality verification rules

[0028] The data in the cache layer is placed in the data operation layer and verified according to the province's unified data quality verification rules. Data without problems is placed in the common data model layer. Problematic data is uniformly governed and a problem data report is provided. After unified data governance, data quality is improved to better meet national data quality requirements and control when collecting data.

[0029] (3) Standardize data export and unify data export of medical insurance platforms

[0030] According to actual business needs and data usage requirements, we uniformly provide three types of interface services, including API interface, database interface and file interface, to meet actual needs and unify data exports, so as to better control data and provide data security.

[0031] (4) Improve statistical accuracy and unify statistical sources of report data

[0032] The statistical source of report data uniformly uses data from the data application layer. After data governance, the data quality of these data is improved, and the statistical report information data is more accurate. In addition, all cities in the province use the same data source to reduce the differences in report data caused by the use of different data in various cities, so that the report data can be better analyzed.

[0033] (5) Improve the quality and efficiency of data query and unify common data query

[0034] The common data query functions for all cities in the province are developed in a unified manner to avoid differences in queried data due to different statistical calibers in various cities, so that the queried data can be better analyzed.

[0035] (6) Improve data sharing and exchange, and unify shared data sources

[0036] Unify the data sharing source and the data sharing export, based on service management, through encryption, signing, desensitizing, hierarchical authorization and other methods, on the basis of security and controllability, share the medical insurance data to the Provincial Data Resources Bureau through the interface to realize the opening of data service capabilities. As a platform open to the society, it has the ability to support Internet-level concurrent responses.

[0037] (7) Ensure the security of medical insurance data and unify data security management

[0038] According to the data classification and classification criteria, privacy desensitization protection and security confidentiality protection are strengthened, a security protection system is established, and security protection capabilities are improved. This mainly includes: data permission control, data desensitization protection, data encryption circulation, data application traces, etc. BRIEF DESCRIPTION OF THE DRAWINGS

[0039] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0040] Figure 1 It is a system structure principle block diagram of the present invention;

[0041] Figure 2 The present invention is a flow chart of the method. DETAILED DESCRIPTION

[0042] In order to make the purpose, technical solution and advantages of the embodiments of the present invention clearer, the technical solution in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0043] like Figure 1, a data governance and reinforcement system for a medical insurance data warehouse, including a data resource integration center, a data warehouse management center, a data asset management center, a data security management center, a data quality governance center, a data application service center and a visual data monitoring center.

[0044] 1. Data resource integration center: provides access to data resources. The data access function is capable of extracting data from different data sources (structured data, unstructured data) according to specified rules. The extracted data can be stored in two categories: on-site or off-site. The extracted data can provide input for data conversion, or it can be directly processed or loaded.

[0045] It provides the capabilities of data collection, data conversion and cleaning, data loading, and data export for mainstream data sources. It can exchange data with column-based databases, traditional databases, data warehouses, local and remote files, etc. It supports streaming collection of log files and data analysis, and supports ETL extraction in the form of message queues.

[0046] The data resource integration center includes a data source management module, a real-time data synchronization module, an offline data synchronization module, and an integrated scheduling module.

[0047] 1. Data source management module: used for visual operation, configuration and task management of various data sources in the process of data resource integration. It supports the storage and management of various types of data in a wide range. The main sources of data include medical insurance data sources such as basic information center, user center, policy center, insurance participation center, collection center, settlement center, procurement center, electronic voucher center, and other data sources shared by other departments.

[0048] 2. Real-time data synchronization module: For some data applications with high real-time requirements, it provides real-time data processing services, collects real-time and quasi-real-time data directly from the source system in the form of message queues, and processes and processes the real-time data through the real-time data processing framework, providing real-time and quasi-real-time data services for the front-end application system.

[0049] 3. Offline data synchronization module: By establishing an offline data batch collection channel, the data synchronization requirements from the data source layer to the big data warehouse are met. The data synchronization frequency is recommended to be T+1. The data synchronization method is to realize data batch collection through ETL tools, and to complete the warehousing after realizing big data offline calculation processing through big data offline calculation tools.

[0050] Since there is data dependency between various business subsystems, in order to reduce possible problems in the data governance process, it is necessary to formulate a time sequence for data integration based on the specific circumstances of each business subsystem.

[0051] Historical stock data is synchronized once and collected into the data warehouse in batches; daily incremental T+1 data is synchronized to the data warehouse regularly according to the data update time.

[0052] 4. Integrated scheduling module: including real-time and offline scheduling task monitoring, trigger-type task monitoring, and unified management of data-integrated scheduling tasks. It is used to build task modules according to user needs, configure scheduling strategies, provide equal interval scheduling and timed scheduling, and has two startup modes: automatic operation and manual execution, thus forming a unified scheduling management module. In the scheduling list, you can view the log information and running status of the scheduling, and you can also edit, view, delete, pause, and enable the scheduling.

[0053] 2. Data Warehouse Management Center:

[0054] The big data warehouse is the core carrier of the construction of the data center content end. It is used to unify data standards and indicator calibers, unify the medical insurance business data needed by the application layer, and organize and store them in a certain way. Through the construction of the big data warehouse, the data standards and indicator calibers are unified, the data quality is guaranteed, and the decoupling between the business layer data model and the big data application layer data model is achieved. At the same time, it can avoid duplication of construction and save storage space.

[0055] The data warehouse management center includes data model layering module, data model subject domain division module, data modeling module and indicator library design module;

[0056] 1. Data model layering module: used to organize and store the medical insurance data warehouse in a certain way. The data model layering module divides the medical insurance data warehouse into four layers, namely the data buffer layer STG, the operation data layer ODS, the common model layer CDM and the application data layer ADS.

[0057] (1) Data buffer layer STG: It is the first storage area in the process of business data flow. The data warehouse extracts data from the data sources of various business systems and loads it into the data buffer layer, thus laying a solid foundation for business data cleaning, quality control and other processes. The data cache layer can improve the convenience of ETL processing and the efficiency of ETL execution.

[0058] The data buffer layer STG should meet the following requirements:

[0059] 1) The table structure is consistent or similar to that of the source system;

[0060] 2) Capable of quickly deleting data;

[0061] 3) Ability to quickly query detailed data;

[0062] 4) Each record includes the insertion and update time fields of the data;

[0063] 5) Each record includes job instances for inserting and updating data.

[0064] (2) The operational data layer (ODS) integrates the data generated by each isolated business system, extracts, cleans, and transmits it, and then loads it into this layer to form a unified, global data set. The data in the ODS layer is isomorphic to the data in the source system, with the purpose of simplifying the subsequent data processing. In terms of data granularity, the data granularity of the ODS layer is the finest. The operational data layer is an area where data in the cache layer is cleaned and governed based on data standards and quality verification rules, using data development and governance tools. It is also an area where current and historical data are stored.

[0065] It is necessary to perform quality control work such as denoising, deduplication, and dirty data removal at this layer, standardize and normalize detailed data, and record and mark problematic data (dirty data). The operational data layer should meet the following requirements:

[0066] 1) The table structure should be consistent or similar to that of the source system;

[0067] 2) Capable of supporting rapid data deletion;

[0068] 3) Support detailed data query function;

[0069] 4) Capable of storing all historical data;

[0070] 5) Capable of date-based, system-based rapid deletion and data extraction;

[0071] 6) Each record includes the insertion and update time fields of the data;

[0072] 7) Each record includes job instances for inserting and updating data.

[0073] (3) Common model layer CDM: It is used to store the data after the theme design of the operation data layer. The theme design uses LDM modeling to process the data.

[0074] The general model layer should meet the following requirements:

[0075] 1) Use LDM to perform topic modeling on the data;

[0076] 2) Ability to identify subordinate topics from table names;

[0077] 3) Each record includes the data insertion and update time fields;

[0078] 4) Each record includes the job instance number of the inserted and updated data.

[0079] In the actual construction process, the general model layer will be further subdivided into the detailed data layer (DWD) and the data summary layer (DWS).

[0080] Among them, the detailed data layer DWD:

[0081] Maintain the same data granularity as the operational data layer ODS, provide certain data quality assurance, process data based on ODS, and provide cleaner data; at the same time, in order to improve the usability of the data detail layer, this layer will use some dimension degradation techniques. When a dimension does not have any data required by the data warehouse, it can be degraded to the fact table to reduce the association between the fact table and the dimension table. For example: There is no need to use a dimension table to store a large dimension such as the order id, but the order id is very important when we generally perform data analysis, so we redundantly store the order id in the fact table. This dimension is a degenerate dimension.

[0082] This layer designs zipper tables and flow tables as needed. Zipper tables are defined for the way tables store data in data warehouse design. As the name suggests, zipper records history, recording all the changes of a thing from its beginning to its current state.

[0083] The flow table stores a user's change records. For example, in a flow table, each modification record of a user will be recorded in the daily data, but there is only one record in the zipper table.

[0084] This is a granularity issue that needs to be taken into account when designing a zipper table. We can of course set the granularity to be smaller, usually days are enough.

[0085] The data aggregation layer DWS is divided by subject and generates wide tables to provide data support for subsequent business queries, OLAP analysis and data distribution. This layer may use various types of databases to meet the needs of specific application scenarios. This layer has the clearest requirements. Data dimensions and analysis results are designed according to business needs. Therefore, data in this layer can be directly connected to OLAP analysis, or directly used or displayed by data applications.

[0086] (4) Application Data Layer ADS:

[0087] It provides data services for applications and stores data converted and processed from the data aggregation layer. The data structure in this layer is independently designed by each application vendor according to the needs of the data application theme. This layer may use various types of databases to meet the needs of specific application scenarios. This layer has the clearest requirements. The data dimensions and analysis results are designed according to business needs. Therefore, the data in this layer can be directly connected to OLAP analysis, and can also be directly used or displayed by data applications.

[0088] 2. Data model subject domain division module: Subject domain is a collection of closely related data subjects. It is divided and abstracted according to the perspective of business demand analysis. According to the business focus, these data subjects are divided into different subject domains. A subject is a scope for integrating, classifying and analyzing the data of a certain analysis object in various systems in production at a higher level. It is an abstract concept, and each subject corresponds to a macro analysis field. According to the business process, a subject domain is abstracted from one business process.

[0089] 3. The data modeling module is used to establish and maintain a set of effective workflows and specifications to ensure that different logical data model designers can operate in accordance with a unified standard.

[0090] 4. Indicator library design module: Indicators are usually organized and managed with business as the center to support various complex business scenarios and analysis needs. The indicator library processes and converts data in a business-oriented manner to generate various indicator items for specific business needs. Modeling and development based on specific business scenarios and business needs are the basis for realizing dynamic configuration and generation of reports.

[0091] 3. Data Asset Management Center:

[0092] According to the hierarchical structure and subject domain division of data integration, various objects at each layer, such as tables, stored procedures, indexes, data links, functions and packages, are managed; regular attention is paid to the operation of collection tasks to locate and resolve some anomalies that may occur during the collection process; data authority control and access data encryption and desensitization are performed to improve data security.

[0093] The data asset management center includes a metadata management module and an asset catalog management module.

[0094] 1. Metadata management module:

[0095] According to the hierarchical structure and subject domain division of data integration, it is necessary to implement the management of various objects at each layer, such as tables, stored procedures, indexes, data links, functions and packages. Clearly represent the data flow between the hierarchical structures, the relationship between the objects, and the information of various data services provided to the outside. The metadata content involves all data links of the entire big data resource platform, including data collection, layer-by-layer processing and auditing, and the processing from data services to the final application presentation. Metadata management runs through the entire process and achieves effective interaction with each link. Metadata management includes metadata definition, query, maintenance, inspection, analysis, lineage management and data maps.

[0096] 2. Asset Catalog Management Module

[0097] The main function of asset catalog management is to use metadata to describe the characteristics of information resources, form unified and standardized catalog content, and form a catalog information database through effective organization and management of catalog content, providing information resource discovery and positioning services for the aggregation and sharing of information resources and support for applications.

[0098] All government information resources are organized and managed in accordance with unified standards and specifications, and directory content query and retrieval services are provided to users through the directory system based on the directory information base. Through the construction of the directory system, the information resources of each business department are catalogued and dynamically managed, making it easier to fully grasp the overall information resource status of each department.

[0099] 4. Data Security Management Center:

[0100] It is used to standardize a unified data security use system, establish a data security management center, standardize the safe use of data, build a unified data security sharing center across the province, and ensure that medical insurance data sharing, transmission and sharing meet data security management requirements. The overall architecture design is based on the two key mechanisms of cloud platform data security and resource isolation to ensure the maturity, stability, security and reliability of the business.

[0101] The data security management center must establish complete information security management measures, rely on the data security management system, strengthen the security control of data applications, follow the ethical standards and information security level protection standards of the medical insurance industry in both clinical research and patient services, only provide the minimum data set required for the business, and conduct access audits.

[0102] The data security management center includes a data classification module, a data authority control module and a data desensitization module.

[0103] 1. Data classification module:

[0104] Data classification aims to comprehensively sort out data assets and determine the corresponding data level, which is a necessary prerequisite and foundation for implementing effective data classification management. Data classification management is the basic work for establishing a unified and complete data lifecycle security protection framework, which can provide support for the medical insurance system to formulate targeted data security control measures.

[0105] Medical insurance data levels are divided into three levels from high to low: core level, important level, and general level.

[0106] Core level: medical insurance data related to key areas of national security, the lifeline of the national economy, important livelihoods and major public interests, as well as other medical insurance data determined after evaluation. Once such data is illegally used or shared, it may directly affect political security.

[0107] Important level: Medical insurance data that may affect national security, economic operation, social stability, public health and safety after being leaked, tampered with or damaged. Medical insurance data that only affects organizations and individuals is generally not considered important data.

[0108] General level: medical insurance data other than core level and important level.

[0109] ①Important data:

[0110] User account password information, digital certificate (private key) information.

[0111] The service population covers the entire city. Medical insurance-related diagnosis and treatment services, personal information, medical insurance settlement, two-fixed institution settlement information, insurance participation information, and collection and payment information data.

[0112] ②Core data:

[0113] The service population covers the entire province. Medical insurance-related diagnosis and treatment services, personal information, medical insurance settlement, two-fixed institution settlement information, insurance information, and collection information data.

[0114] 2. Data desensitization module:

[0115] Based on the classification and grading results of sensitive fields in corporate customers, individual users, and business data in the communications industry, an automated data classification and grading management capability is formed; based on the automated detection and discovery technology of sensitive data, privacy protection processing of the return results of user data requests is achieved; based on the desensitizing function library with multi-modality and multi-business requirements, a hardware-level high-speed dynamic / static desensitization capability is formed for data sets involving sensitive data threats.

[0116] By establishing basic rules such as sensitive data keyword library, sensitive data combination rule library, and sensitive information processing rule library, sensitive fields in the original data can be processed without affecting the accuracy of data analysis results, thereby reducing data sensitivity and reducing the risk of personal privacy leakage. Common data desensitization methods mainly include:

[0117] 1) Data replacement

[0118] Replace the real value with a fixed imaginary value set;

[0119] 2) Reverse inference

[0120] Find mappings that may infer sensitive fields from certain fields and desensitize these fields;

[0121] 3) Offset and rounding

[0122] By randomly shifting digital data, offset rounding ensures the approximate authenticity of the range while maintaining data security.

[0123] 4) Mask shielding

[0124] Masking is a powerful tool for desensitizing some information in account data.

[0125] 5) Flexible encoding

[0126] When special desensitization rules are required, flexible encoding can be performed to meet various possible desensitization rules.

[0127] 6) Invalidation

[0128] Sensitive data can be desensitized by truncating, encrypting, hiding, etc., so that it is no longer of use value.

[0129] 7) Randomization

[0130] Replace the true value with random data, keeping the randomness of the replacement value to simulate the authenticity of the sample.

[0131] 8) Occlusion

[0132] Refers to replacing part of the sensitive data with masking symbols (such as "X, *"), so that part of the sensitive data remains public. This method can largely desensitize while maintaining the original data appearance.

[0133] 3. Data permission control module:

[0134] Through centralized authorization, fine-grained authorization and other forms, the authorized scope of IP addresses and accounts of data exports is strictly controlled to provide strict data security isolation. The security management center has a convenient permission management function, provides a visual application approval process, and can audit and manage permissions, which improves data security and facilitates data permission management.

[0135] Apply for permissions online: Select the data table you need permissions for and quickly apply online, changing the original offline mode of contacting administrators to improve work efficiency.

[0136] Permission audit / return: Administrators can quickly and easily view the corresponding personnel of database table permissions and conduct audit management. Users can also actively return permissions that are no longer needed.

[0137] Permission approval management: The previous mode of direct authorization by administrators is changed to an approval authorization mode, providing a visual and process-based management authorization mechanism, and the approval process can be traced back afterwards.

[0138] In the Security Management Center module, you can view global data table permissions within the organization, manage table permissions, and apply for / approve data table permissions.

[0139] The data center supports multi-dimensional management and control of members of the organization through methods such as organization Owner account, AccessKey and AccessSecret.

[0140] 5. Data Quality Governance Center:

[0141] Used to configure data quality check rules, establish data quality check tasks, regularly check data quality, provide daily data quality operation result reports, and manage dirty data.

[0142] The data quality governance center includes a data quality rule configuration module, a data quality task configuration module, a data quality report module, and a dirty data management module.

[0143] 1. Data quality rule configuration module:

[0144] The information platform can configure quality rules and establish a quality rule library. The quality rule library is used to store data quality inspection rules defined by users according to data quality standards. The rules cover the following aspects:

[0145] 1) Integrity check: including table record missing check and field value missing check;

[0146] 2) Consistency check: including field value type, length, format, code and rationality check;

[0147] 3) Accuracy check: record the rationality of the business, check the field value range, and check whether the characters are garbled;

[0148] 4) Timeliness check: Check whether the data is collected, processed and output in a timely manner;

[0149] 5) Others: National data inspection rules.

[0150] 2. Data quality task configuration module:

[0151] To ensure data quality, a data quality detection scheduling task is established, including:

[0152] 1) Establish a daily data inspection mechanism to conduct quality inspections on collected and governed data;

[0153] 2) Establish a data problem alarm and timely response mechanism.

[0154] 3. Data quality reporting module:

[0155] The data quality report provides a daily data quality report for each data table. According to the definition of quality rules, the data quality report is also divided into table-level rule quality reports and field-level quality reports.

[0156] (1) Table-level quality rules:

[0157] The system will calculate the output time of the table and the number of data items in each table by default. Users can set the rule range for the output time and the number of records. The set rules will be displayed on the chart as a red straight line or curve.

[0158] The association relationship between the master and slave tables will display the associated fields of the master and slave tables, as well as the number of data entries in the associated fields of the current slave table that do not belong to the master table.

[0159] (2) Field-level quality rules:

[0160] The field-level quality rules will display all the field quality verification rules that have been configured in the table, including the field name, verification rule, the number and percentage of items that meet the rules, the number and percentage of items that do not meet the rules, and a link to view error data.

[0161] 4. Dirty data management module:

[0162] Click View Dirty Data in the rule list to enter the dirty data viewing page. The first column displays the column name and data of the current quality rule. The following columns display the data in the order of the fields in the table. Users can download the current dirty data.

[0163] 6. Data Application Service Center:

[0164] A service interface used to provide data empowerment to the outside world, and includes a data service management module and a data service application module.

[0165] The data middle platform provides data-enabled service interfaces to the outside world through the data service center. Depending on the service type and data type, it mainly provides three types of interface services, including API interface, database interface and file interface.

[0166] API interface:

[0167] The API interface is the data access and calling method used by the data warehouse (data center) to access external applications or data. The API interface should meet the following requirements:

[0168] 1) API generation has basic capabilities such as paging and filtering;

[0169] 2) Support HTTP / HTTPS standard interface protocol;

[0170] 3) Support POST and GET requests;

[0171] Database interface:

[0172] The database interface is the database access service provided by the data warehouse (data center) to the outside world and should meet the following requirements:

[0173] 1) Support offline data services, and authorize offline tables in the big data warehouse to users, so that they can further build data application layers based on business needs;

[0174] 2) Support real-time data services, provide streaming data and OLAP data to users, and support the streaming data needs and ad hoc query and analysis needs of business systems;

[0175] 3) Support real-time data warehouses to collect data in real-time (mainly streaming data) and offline data modes, support real-time data warehouses to provide data in streaming data modes, and support standard database interface queries;

[0176] 4) Support real-time, offline data authorization management, and support table-level authorization (preferably support field-level authorization).

[0177] Data file interface:

[0178] The data file interface is the data access and calling method used by the data warehouse (data center) for external application or data access needs. The data file interface should meet the following requirements:

[0179] 1) Support the ability to provide external data services in the form of files;

[0180] 2) Supports multiple mainstream data file storage including FTP and OSS as data sources;

[0181] 3) Support mainstream file types such as CSV and TEXT;

[0182] 4) Supports setting basic data file settings such as character set, compression status, row and column separators, etc.

[0183] 5) Support viewing file transfer status, including transferring, completed, failed and reasons, etc.

[0184] 1. Data service management module

[0185] Control data call permissions of other systems through access control, permission control and transmission encryption.

[0186] 2. Data service application module

[0187] (1) Data collection application module

[0188] The data of core business systems, basic information centers, insurance participation centers, payment centers, settlement centers, procurement centers, and group payment centers are collected by the National Bureau, and the T+1 data source is obtained from the data warehouse ODS layer to ensure data consistency.

[0189] (2) Internal data application module

[0190] Based on the relevant data requirements of the internal subsystems of the provincial medical insurance information platform, corresponding data application topics are generated to meet the needs of data applications of business subsystems such as public services, operation monitoring, intelligent supervision, and group payment.

[0191] (3) External shared application modules

[0192] Based on the service management model, the data sharing and exchange center shares medical insurance data with the Provincial Data Resources Bureau and other departments and units in a safe and controllable manner through encryption, signing, desensitization, hierarchical authorization, etc. At the same time, according to the needs of medical insurance business, relevant data information is obtained from external departments such as public security, education, health, human resources and social security, and the Data Resources Bureau.

[0193] (4) Data statistics application module

[0194] The data statistics application includes the national report module and the data query module. The national report module generates various situations such as the basic medical insurance participants, the employees’ basic medical insurance participants and special personnel, the collection of employees’ basic medical insurance premiums, and the medical expenses of employees’ basic medical insurance by selecting the annual quarter, administrative division, and insurance type; the data query module displays the details of employees’ insurance participation, residents’ insurance participation, medical insurance settlement, electronic voucher usage rate, mobile payment, etc. within the time range through information query.

[0195] 7. Visual data monitoring center:

[0196] It is used to achieve large-screen panoramic display of key data information, realize data cockpit-style display and management, and visualize data monitoring information such as the province's data aggregation trajectory, integration status, shared analysis, etc.

[0197] like Figure 2 A method for reinforcing a data governance system based on a medical insurance data warehouse includes the following steps:

[0198] S1, obtains data collected and accessed from different data sources, and performs data extraction operations based on established rules.

[0199] Among them, data source types include structured data and unstructured data, and the acquired data includes data stored in relational databases based on Oracle, MySQL and PostgreSQL, as well as files and message middleware that realize data transmission through API interfaces.

[0200] S2, after the data to be processed is accessed from the data obtained in step S1, data processing is performed.

[0201] Among them, data processing includes offline data processing and real-time data processing. Specifically, data is extracted and analyzed. Specifically, offline data processing relies on the big data development and governance platform DataWorks to extract data; real-time data processing uses the data transmission service DTS to perform streaming analysis on database logs.

[0202] S3, after the data is processed, it is stored in the data warehouse and classified and calculated, including the storage and calculation of offline data and the storage and calculation of real-time data.

[0203] Among them, offline data storage and offline computing use the offline big data computing platform MaxCompute offline computing engine to realize the storage and computing of batch structured data. The offline big data warehouse is divided into STG, ODS, and ADS. Real-time computing and storage use Flink to flexibly and efficiently realize data parallel processing. The real-time big data warehouse provides the ADS layer.

[0204] The storage and calculation of data in step S3 specifically includes the following steps:

[0205] S31, in the data buffer layer STG of the data model of the data warehouse management center, the original data is integrated, and the full amount of data is stored in partitions based on time;

[0206] S32, in the data operation layer ODS, the data in the data buffer layer STG is transported, quality controlled, cleaned and managed to form the full amount of data that can be used by the subject; at the same time, a physical deletion data synchronization processing mechanism is established in the production library of the data center in each city to ensure the uniformity of the data;

[0207] S33, the data after the subject design of the data operation layer ODS is stored in the common model layer CDM, and the subject design uses LDM modeling to process the data; at the same time, according to the characteristics of medical insurance business, data monitoring direction and trend, a multi-dimensional indicator analysis model is designed to establish a multi-type common indicator library;

[0208] S34, through the data application layer (ADS), develops various thematic application data to provide carrier support for enabling innovative applications of medical insurance data.

[0209] S4, transfers the data calculated in step S3 to the data service layer to implement access control, authority control and transmission encryption to control the data call permissions of other systems, and can download data in batches, or encapsulate the data into a standard API as needed.

[0210] S5, display the decision-making data, and realize various customized data report analysis needs and personalized data application needs. Display the decision-making data through the visualization large screen tool dataV; use Alibaba Cloud intelligent reporting tool QuickBI to realize various customized data report analysis needs; realize personalized data application needs based on Alibaba Cloud product technology system.

[0211] The present invention also discloses a computer-readable storage medium and a computer system, wherein a computer program is stored on the computer-readable storage medium, and after the computer program is run, the method described above is executed. A computer system includes a processor and a storage medium, wherein the computer program is stored on the storage medium, and the processor reads and runs the computer program from the storage medium to execute the method described above.

[0212] Those skilled in the art will further appreciate that the various illustrative logic blocks, modules, circuits, and algorithm steps described in conjunction with the embodiments disclosed herein may be implemented as electronic hardware, computer software, or a combination of the two. To clearly illustrate this interchangeability of hardware and software, various illustrative components, blocks, modules, circuits, and steps are generally described above in terms of their functionality. Whether such functionality is implemented as hardware or software depends on the specific application and the design constraints imposed on the overall system. The technician may implement the described functionality in different ways for each specific application, but such implementation decisions should not be interpreted as resulting in a departure from the scope of the present invention.

[0213] The various illustrative logic blocks, modules, and circuits described in conjunction with the embodiments disclosed herein may be implemented or performed with a general purpose processor, a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA) or other programmable logic device, discrete gate or transistor logic, discrete hardware components, or any combination thereof designed to perform the functions described herein. A general purpose processor may be a microprocessor, but in the alternative, the processor may be any conventional processor, controller, microcontroller, or state machine. The processor may also be implemented as a combination of computing devices, such as a combination of a DSP and a microprocessor, a plurality of microprocessors, one or more microprocessors in cooperation with a DSP core, or any other such configuration.

[0214] The steps of the method or algorithm described in conjunction with the embodiments disclosed herein may be embodied directly in hardware, in a software module executed by a processor, or in a combination of the two. The software module may reside in a RAM memory, a flash memory, a ROM memory, an EPROM memory, an EEPROM memory, a register, a hard disk, a removable disk, a CD-ROM, or any other form of storage medium known in the art. An exemplary storage medium is coupled to a processor so that the processor can read and write information from / to the storage medium. In an alternative, a storage medium may be integrated into a processor. The processor and the storage medium may reside in an ASIC. The ASIC may reside in a user terminal. In an alternative, the processor and the storage medium may reside in a user terminal as discrete components.

[0215] In one or more exemplary embodiments, the functions described may be implemented in hardware, software, firmware, or any combination thereof. If implemented as a computer program product in software, each function may be stored on or transmitted by a computer-readable medium as one or more instructions or codes. Computer-readable media include both computer storage media and communication media, including any medium that facilitates the transfer of a computer program from one place to another. Storage media may be any available medium that can be accessed by a computer. As an example and not limitation, such a computer-readable medium may include RAM, ROM, EEPROM, CD-ROM or other optical disk storage, disk storage or other magnetic storage device, or any other medium that can be used to carry or store the desired program code in the form of an instruction or data structure and can be accessed by a computer. Any connection is also properly referred to as a computer-readable medium. For example, if the software is transmitted from a website, a server, or other remote source using a coaxial cable, a fiber optic cable, a twisted pair, a digital subscriber line (DSL), or wireless technologies such as infrared, radio, and microwaves, the coaxial cable, fiber optic cable, twisted pair, DSL, or wireless technologies such as infrared, radio, and microwaves are included in the definition of the medium. Disk and disc as used herein include compact disc (CD), laser disc, optical disc, digital versatile disc (DVD), floppy disk and Blu-ray disc, wherein disk often reproduces data magnetically, while disc reproduces data optically with lasers. Combinations of the above should also be included within the scope of computer-readable media.

[0216] The previous description of the disclosure is provided to enable any person skilled in the art to make or use the disclosure. Various modifications to the disclosure will be apparent to those skilled in the art, and the general principles defined herein may be applied to other variations without departing from the spirit or scope of the disclosure. Thus, the disclosure is not intended to be limited to the examples and designs described herein, but should be granted the widest scope consistent with the principles and novel features disclosed herein.

[0217] Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent substitutions for some of the technical features therein; and these modifications or substitutions do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A data governance and reinforcement system for a medical insurance data warehouse, characterized by: Including data resource integration center, data warehouse management center, data asset management center, data security management center, data quality governance center, data application service center and visual data monitoring center; The data resource integration center provides aggregated access to data resources; The data warehouse management center is used to unify data standards and indicator calibers, gather the medical insurance business data required by the application layer, and organize and store them in layers in a certain way; The data asset management center implements the management of various objects at each layer according to the hierarchical structure and subject domain division of data integration, and provides information resource discovery and positioning services; The data security management center is used to standardize a unified data security use system, establish a data security management center, standardize the safe use of data, build a unified data security sharing center across the province, and ensure that the sharing, transmission and sharing of medical insurance data meet the data security management requirements; The data quality governance center is used to configure data quality inspection rules, establish data quality inspection tasks, regularly inspect data quality, provide daily data quality operation result reports, and manage dirty data; The data application service center is used to provide a data empowerment service interface to the outside, and includes a data service management module and a data service application module; The visual data monitoring center is used to realize large-screen panoramic display of key data information, realize data cockpit-style display and management, and visualize data monitoring information such as data aggregation trajectory, integration status, shared analysis, etc. across the province.

2. The system according to claim 1, characterized in that: The data resource integration center includes a data source management module, a real-time data synchronization module, an offline data synchronization module and an integrated scheduling module; The data source management module is used to perform visual operation, configuration and task management on each data source in the data resource integration process; The real-time data synchronization module is used to provide real-time data processing services, collect real-time and quasi-real-time data directly from the source system in the form of message queues, and process and process the real-time data through the real-time data processing framework to provide real-time and quasi-real-time data services for the front-end application system; The offline data synchronization module meets the data synchronization requirements from the data source layer to the big data warehouse by establishing an offline data batch collection channel, and completes the warehousing after the big data offline calculation processing through the big data offline calculation tool; The integrated scheduling module includes unified management of data integration scheduling tasks of real-time and offline scheduling task monitoring and triggered task monitoring, which is used to build task modules according to user needs, configure scheduling strategies, provide equal interval scheduling and scheduled scheduling, and form a unified scheduling management module.

3. The system according to claim 1, characterized in that: The data warehouse management center includes a data model layering module, a data model subject domain division module, a data modeling module and an indicator library design module; The data model hierarchical module is used to organize and store the medical insurance data warehouse in a hierarchical manner in a certain manner; The data model subject domain division module is used to abstractly classify business data subjects and divide them into different subject domains according to business demand analysis; The data modeling module is used to establish and maintain a set of effective workflows and specifications to ensure that different logical data model designers can operate in accordance with a unified caliber; The indicator library design module is used to process and convert data into business-oriented ones to generate various indicator items for specific business needs.

4. The system according to claim 3, characterized in that: The data model layering module divides the medical insurance data warehouse into four layers, namely, the data buffer layer STG, the operation data layer ODS, the common model layer CDM and the application data layer ADS; The data buffer layer STG is the first storage area in the business data flow process. The data warehouse extracts data from the data sources of each business system and loads it into the data buffer layer; The operational data layer (ODS) integrates the data generated by each isolated business system, extracts, cleans and transmits it, and then loads it into this layer to form a unified, global data set, thus simplifying the subsequent data processing work; The common model layer CDM is used to store data after the subject design of the operation data layer. The subject design uses LDM modeling to process data; and the common model layer CDM includes a detailed data layer DWD and a data summary layer DWS; wherein the detailed data layer DWD maintains the same data granularity as the operation data layer ODS, and provides a certain data quality guarantee, processes the data on the basis of ODS, and provides cleaner data; the data summary layer DWS generates wide tables according to subject division, and provides data support for subsequent business queries, OLAP analysis and data distribution; The application data layer ADS provides data services for applications and stores data converted and processed from the data aggregation layer. The data structure in this layer is independently designed by each application manufacturer according to the requirements of the data application theme.

5. The system according to claim 1, characterized in that: The data asset management center regularly monitors the operation of collection tasks, locates and resolves some anomalies that may occur during the collection process, and performs data authority control and encryption and desensitization of access data to improve data security; The data asset management center includes a metadata management module and an asset catalog management module; The metadata management module needs to implement various objects at each layer, including tables, stored procedures, indexes, data links, functions and packages, according to the hierarchical structure and subject domain division of data integration, and clearly represent the data flow between the hierarchical structures, the relationship between the objects, and the information of various data services provided to the outside; The asset catalog management module uses metadata to describe the characteristics of information resources to form unified and standardized catalog content. Through effective organization and management of catalog content, a catalog information library is formed to provide information resource discovery and positioning services for the aggregation and sharing of information resources and support for applications.

6. The system according to claim 1, characterized in that: The data security management center includes a data classification module, a data authority control module and a data desensitization module; The data classification module comprehensively sorts out the data assets and determines the corresponding data level; The data desensitization module is used to process sensitive fields in the original data without affecting the accuracy of the data analysis results by establishing basic rules of a sensitive data keyword library, a sensitive data combination rule library, and a sensitive information processing rule library, thereby reducing data sensitivity and reducing the risk of personal privacy leakage; The data permission control module strictly controls the IP address of the data export and the authorization scope of the account through centralized authorization and fine-grained authorization, and provides strict data security isolation.

7. The system according to claim 1, characterized in that: The data quality governance center includes a data quality rule configuration module, a data quality task configuration module, a data quality report module and a dirty data management module; The data quality rule configuration module is used to configure quality rules on the information platform and establish a quality rule library, which is used to store data quality inspection rules defined by users according to data quality standards; The data quality task configuration module is used to establish data quality detection scheduling tasks to ensure data quality; The data quality reporting module is used to provide a daily data quality operation result report for each data table.

8. The system according to claim 1, characterized in that: The data application service center provides three types of interface services, including API interface, database interface and file interface, according to different service types and data types; The API interface is the data access and calling method used by the data warehouse, i.e. the data center, for external applications or data access needs; The database interface is a database access service provided by the data warehouse to the outside world; The file interface data warehouse adopts data access and calling methods for external application or data access requirements.

9. The system according to claim 1, characterized in that: The data application service center includes a data service management module and a data service application module; The data application service center is used to control the data calling authority of other systems through access control, authority control and transmission encryption; The data service application module includes a data collection application module, an internal data application module, an external data sharing application module, and a data statistics application module; the data collection application module is used to track business data to the national bureau; the internal data application module generates corresponding data application topics based on the relevant data requirements of the internal subsystem of the provincial medical insurance information platform to meet the needs of data application of the business subsystem; the external data sharing application module is used to share medical insurance data to the provincial data resources bureau and external departments and units through encryption, signing, desensitizing, and hierarchical authorization. At the same time, according to the needs of medical insurance business, relevant data information is obtained from external departments; the data statistics application module is a national reporting module and a data query module.

Citation Information

Patent Citations

  • Big data platform for smart city construction

    CN112685385A

  • Method and system for managing data resources

    CN117632954A

  • Method for designing data warehouse in earth observation field

    CN118626469A

  • Data service method and system based on domestic autonomous controllable environment

    CN119127135A

  • Emergency resource sharing and exchange system

    US20200334605A1