A metadata-centered data management platform and implementation method
Patent Information
- Application Number
- CN202310946458.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-07-28
- Publication Date
- 2026-09-22
- Estimated Expiration
- 2043-07-28
AI Technical Summary
[0031]1、本发明创新性的提供了一种以元数据为核心的数据管理平台及实现方法,对元数据的采集进行了融合,采用“实时采集+周期性离线对比”的方式,对采集过程进行了优化,用到哪些表采集哪些表的元数据,降低了元数据采集的量,避免大量无用的表进入数据管理平台形成的数据沼泽。
Smart Images

Figure CN116910078B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of data processing technology, and in particular to a data management platform and implementation method with metadata as its core. Background Technology
[0002] The statements in this section are merely background information related to the present invention and do not necessarily constitute prior art.
[0003] Data governance is the management of the entire data lifecycle, including traditional data integration and storage processes such as data collection, cleaning, and transformation, as well as data asset catalogs, data standards, quality, security, data development, data services and applications. All business, technical and management activities carried out throughout the data lifecycle fall under the scope of data governance.
[0004] A data management platform is a tool designed to help clients implement data governance more effectively. A data management platform typically covers four main modules: data integration, data governance, data development, and data services.
[0005] Metadata serves as the cornerstone of data governance and a crucial tool for enterprise managers to manage data. Metadata management plays a vital role in the success and quality of data governance projects. Effective metadata management will undoubtedly make subsequent data governance and applications much smoother. If the metadata of a data management platform is poorly managed, the more data accumulated, the easier it is for a data swamp to form, and data will no longer be an asset but rather become the root of problems.
[0006] The inventors discovered that existing data management platforms have the following problems:
[0007] (1) There are generally two ways to collect metadata in the industry: one is periodic offline collection, and the other is real-time reading of the database metadata without local archiving. Periodic offline collection is limited by periodic scheduling and may result in inconsistent metadata. Real-time reading cannot track changes in metadata and cannot provide users with necessary reminders.
[0008] (2) Conventional lineage relationships mainly show the relationship between tables and the relationship between fields. However, when there is a problem with the interface or the task fails, the corresponding table needs to be found manually before lineage analysis can be performed, which is inefficient and prone to errors. Summary of the Invention
[0009] To address the shortcomings of existing technologies, this invention provides a data management platform and implementation method centered on metadata. It integrates metadata collection and adopts a "real-time collection + periodic offline comparison" approach to optimize the collection process, reduce the amount of metadata collected, and prevent a large number of useless tables from entering the data management platform and forming a data swamp.
[0010] To achieve the above objectives, the present invention adopts the following technical solution:
[0011] Firstly, the present invention provides a method for implementing a data management platform with metadata as its core.
[0012] A method for implementing a data management platform with metadata at its core includes the following processes:
[0013] The business database is registered with the data management platform through the business data source;
[0014] When adding an offline synchronization task, if an input table is selected, the task will directly connect to the business database to read the metadata information of the input table.
[0015] When saving offline tasks, the input table is included in the metadata collection scope of the business data source. If it is already in the collection scope, the metadata is compared to see if it has been changed. If there is no change, the version number of the metadata is saved to the task. If there is a change, the metadata of the data platform is updated to generate the latest version of the metadata of the input table.
[0016] According to the scheduling cycle, the collected metadata is periodically compared with the metadata of the business data source to see if they are consistent. If there are changes, a new version is generated and a metadata change reminder message is generated.
[0017] As a further limitation of the first aspect of the present invention, tasks and API interfaces are treated as entities, and a link relationship of sequentially connected tables, tasks, tables and APIs is constructed.
[0018] As a further limitation of the first aspect of the present invention, when registering the business data source, the required tables are selected for metadata collection, and after successful collection, table entities are constructed in the graph database.
[0019] When data warehouse data sources are modeled and saved, the corresponding tables in the data warehouse are automatically constructed into table entities in the graph database.
[0020] When a data synchronization task is launched and running, the data synchronization task entity is automatically created in the graph database, and the relationship between the input table and the task and the relationship between the task and the output table are established.
[0021] As a further limitation of the first aspect of the present invention, in the case of data development, when the development task is launched, a data development task entity is constructed based on the input table and output table information, and the relationship between the input table and the development task and the output table is constructed.
[0022] As a further limitation of the first aspect of the present invention, during API development, the API interface entity is automatically constructed, and the relationship between the data table and the API interface is automatically constructed based on the data table in the API interface.
[0023] Secondly, this invention provides a data management platform with metadata as its core.
[0024] A data management platform with metadata at its core includes at least:
[0025] The data governance module is configured to manage metadata using the metadata-centric data management platform implementation method described in the first aspect of this invention.
[0026] As a further limitation of the second aspect of the invention, it also includes a data planning module, configured to: perform orderly storage of data assets through unified management of data architecture, storage architecture, and data warehouse data sources.
[0027] As a further limitation of the second aspect of the invention, it also includes a data integration module, configured to: perform the aggregation of multi-source heterogeneous data and construct a data synchronization channel to enable rapid inbound of business data.
[0028] As a further limitation of the second aspect of the invention, it also includes a data development module configured to perform model design, self-service development, and API development tasks.
[0029] As a further limitation of the second aspect of the invention, it also includes a data service module, configured to: perform integrated data management across the entire data domain, and perform data sharing and circulation across departments, regions, and levels.
[0030] Compared with the prior art, the beneficial effects of the present invention are:
[0031] 1. This invention innovatively provides a data management platform and implementation method with metadata as the core. It integrates metadata collection and adopts a "real-time collection + periodic offline comparison" approach to optimize the collection process. It determines which tables' metadata to collect, reduces the amount of metadata collected, and avoids a large number of useless tables entering the data management platform, thus preventing a data swamp.
[0032] 2. This invention innovatively provides a data management platform and implementation method with metadata as the core, ensuring that users use real-time metadata to avoid discrepancies when performing maintenance tasks, while using periodic offline comparison function to dynamically monitor metadata, and will remind users of the changes as soon as they occur.
[0033] 3. This invention innovatively provides a data management platform and implementation method with metadata as the core. By expanding the lineage between tasks and API interfaces, it can more intuitively find relevant tasks and tables for common data management issues such as interface data anomalies and task execution anomalies, greatly improving the efficiency of problem analysis and processing. At the same time, it can better analyze the impact of anomalies on interfaces.
[0034] Advantages of additional aspects of the invention will be set forth in part in the description which follows, and in part will be obvious from the description, or may be learned by practice of the invention. Attached Figure Description
[0035] The accompanying drawings, which form part of this invention, are used to provide a further understanding of the invention. The illustrative embodiments of the invention and their descriptions are used to explain the invention and do not constitute an improper limitation of the invention.
[0036] Figure 1 The data management platform architecture diagram provided for this invention;
[0037] Figure 2 This is a schematic diagram of the improved metadata collection process provided by the present invention;
[0038] Figure 3 The lineage diagram provided by this invention includes added task and API entities. Detailed Implementation
[0039] The present invention will be further described below with reference to the accompanying drawings and embodiments.
[0040] It should be noted that the following detailed descriptions are exemplary and intended to provide further illustration of the invention. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains.
[0041] Where there is no conflict, the embodiments and features in the embodiments of the present invention can be combined with each other.
[0042] like Figure 1 As shown, this embodiment provides a data management platform with metadata as its core, including a data planning module, a data integration module, a data governance module, a data development module, and a data service module;
[0043] The data planning module is configured to achieve orderly storage of data assets through unified management of data architecture, storage architecture, and data warehouse data sources, thereby fully leveraging the value of data.
[0044] The data integration module is configured to primarily aggregate multi-source heterogeneous data, establish an efficient and stable data synchronization channel, and enable rapid data ingestion into the warehouse.
[0045] The data governance module includes:
[0046] Data Standards Unit: The data standards adhere to the principle of "design first, development later; standardize first, model later" and standardize the definitions of data elements, code sets, business terms, etc. to eliminate inconsistent understanding of data across systems and realize the concept of integrated development and governance.
[0047] Data Quality Unit: Automatically detects dirty data in the data warehouse that does not conform to quality rules, automatically generates quality reports and proactively issues alerts, and, combined with data lineage, can effectively prevent dirty data from spreading downstream.
[0048] Metadata Unit: Clarify the current status of data assets, help data governance personnel understand data sources, trace data flow, and improve the efficiency of troubleshooting during data synchronization, data governance, and data development.
[0049] Model design unit: Standardizes the data model development process, reduces the difficulty of data model development, and improves the efficiency of data model development.
[0050] The data development module includes:
[0051] Self-service development unit: Through a visual offline data development kit, data resources are transformed into new forms required by business needs, thereby extracting data value.
[0052] API Development Unit: Through visual configuration, it provides the ability to quickly build APIs, helping users create and publish data interfaces to realize the external release of data value.
[0053] The data service module is configured to provide integrated and intelligent data management across the entire data domain, help users intuitively grasp data assets, provide cross-departmental, cross-regional, and cross-level data sharing and circulation services, and create a comprehensive data resource system.
[0054] The Operations and Maintenance Center module is configured to assist task operations and maintenance personnel in offline task management and instance operations and maintenance, help operations and maintenance personnel improve operational efficiency, and facilitate intelligent operations and maintenance.
[0055] The tenant management module is configured to use tenant isolation technology to ensure data security and privacy, while improving resource utilization, reducing operating costs, and helping to reduce costs and increase efficiency.
[0056] To address the issues of inconsistent metadata and lack of awareness of changes, this embodiment integrates metadata collection, employing a "real-time collection + periodic offline comparison" approach. This overcomes the shortcomings of existing methods and optimizes the collection process, specifying which tables' metadata to collect, reducing the amount of metadata collected and preventing a large number of useless tables from entering the data management platform and creating a data swamp. Specifically, as shown below... Figure 2 As shown, the process includes the following:
[0057] S1.1: Register the business database through the business data source, but this step eliminates the usual metadata collection and maintenance, reducing manual maintenance.
[0058] S1.2: When adding an offline synchronization task, when selecting the input table (this refers to which table in the source database you want to synchronize to another database; for a data synchronization task, this is the input table), the task directly connects to the business database to read the table's metadata information and directly uses the actual metadata information of the database, effectively avoiding data inconsistency caused by untimely local metadata collection.
[0059] S1.3: When saving offline tasks, include the input table in the metadata of the business data source (a database may contain many tables, but we only focus on a small portion. We periodically collect and compare these tables. The collection scope refers to the tables included in the comparison scope). If the table is within the collection scope, compare the metadata (one copy is the metadata we previously collected and stored, and the other is the current metadata of the source database; this comparison is necessary to detect any changes) to see if any changes have occurred. If there are no changes, retain the version number of the metadata in the task; if there are changes, update the metadata of the data platform to generate the latest version of the metadata for that table.
[0060] S1.4: Periodically compare the collected metadata with the metadata of the business data source according to the scheduling cycle. If there are changes, update the new version and send a metadata change reminder to the user.
[0061] S1.5: Data development tasks and API development processes are the same as those described above and will not be listed again.
[0062] S1.6: The metadata management function allows you to view the metadata information and version change information of the corresponding table.
[0063] The above methods ensure that users use real-time metadata to avoid discrepancies when performing maintenance tasks, while the periodic offline comparison function dynamically monitors the metadata and alerts users to changes as soon as they occur.
[0064] In building a data management platform, to address the problem that traditional data lineage analysis only displays relationships between tables, making it difficult to locate and analyze problems, this embodiment creatively constructs a table→task→table→API link relationship as entities, better serving lineage analysis and impact analysis. Specifically, as follows... Figure 3 As shown, the process includes the following:
[0065] S2.1: When registering a business data source, select the tables to be used for metadata collection. After successful collection, construct the table entities in the graph database.
[0066] S2.2: When modeling and saving the data warehouse data source (a data management platform usually has a centralized storage to aggregate data from various business systems for easy analysis later, this library is the data warehouse data source), the corresponding tables in the data warehouse are automatically constructed into table entities in the graph database.
[0067] S2.3: When the data synchronization task is launched, the data synchronization task entity is automatically created in the graph database, and the relationship between the input table and the task and the relationship between the task and the output table are constructed.
[0068] S2.4: For situations involving data development, when a development task goes live, a data development task entity is constructed based on the input and output table information, and the relationships between the input table and the development task, as well as between the development task and the output table, are established.
[0069] S2.5: During API development, automatically construct the API interface entity and automatically construct the relationship between the data table and the API interface based on the data table in the API interface.
[0070] In this way, the desired clear data lineage relationship between tables, tasks, and interfaces is automatically formed during the data governance process. Through the new lineage relationship, it is easy to analyze which data anomaly was caused by which task and which table's data, and which downstream tables and API interfaces will be affected when a task is abnormal.
[0071] This invention effectively solves the problems of inconsistency between metadata and reality and the inability to perceive metadata changes by using a "real-time reading + periodic comparison" approach, making the entire data management process smoother and providing a clearer picture of metadata changes. By expanding the lineage between tasks and API interfaces, it can more intuitively locate relevant tasks and tables for common data management issues such as interface data anomalies and task execution anomalies, greatly improving the efficiency of problem analysis and processing, and enabling better analysis of the impact of anomalies on interfaces.
[0072] This invention is applicable not only to data management platforms, but also to the construction methods of data governance tools. Furthermore, the core innovation lies in the materialization of tasks and APIs within the lineage diagram, and is not limited to the lineage diagram style shown in this embodiment.
[0073] The above description is merely a preferred embodiment of the present invention and is not intended to limit the invention. Various modifications and variations can be made to the present invention by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.
Claims
1. A method for implementing a data management platform with metadata as its core, characterized in that, Includes the following processes: The business database is registered with the data management platform through the business data source; When adding an offline synchronization task, if an input table is selected, the task will directly connect to the business database to read the metadata information of the input table. When saving offline tasks, the input table is included in the metadata collection scope of the business data source. If it is already in the collection scope, the metadata is compared to see if it has been changed. If there is no change, the version number of the metadata is saved to the task. If there is a change, the metadata of the data platform is updated to generate the latest version of the metadata of the input table. According to the scheduling cycle, the collected metadata is periodically compared with the metadata of the business data source to see if they are consistent. If there are changes, a new version is generated and a metadata change reminder message is generated. Treat tasks and API interfaces as entities, and construct a chain relationship between tables, tasks, tables, and APIs that are connected in sequence; When registering a business data source, select the tables you need to collect metadata. After successful collection, construct the table entities in the graph database. When data warehouse data sources are modeled and saved, the corresponding tables in the data warehouse are automatically constructed into table entities in the graph database. When the data synchronization task is launched and running, the data synchronization task entity is automatically created in the graph database, and the relationship between the input table and the task and the relationship between the task and the output table are constructed. During API development, the API interface entity is automatically constructed, and the relationship between the data table and the API interface is automatically constructed based on the data table in the API interface.
2. The data management platform implementation method with metadata as its core as described in claim 1, characterized in that, For scenarios involving data development, when a development task goes live, a data development task entity is constructed based on the input and output table information, and the relationships between the input table and the development task, as well as between the development task and the output table, are established.
3. A data management platform with metadata as its core, employing the method described in claim 1, characterized in that, At least including: The data governance module is configured to perform metadata management using the data management platform implementation method with metadata as the core as described in any one of claims 1-2.
4. The data management platform with metadata as its core as described in claim 3, characterized in that, It also includes a data planning module, which is configured to: store data assets in an orderly manner through unified management of data architecture, storage architecture, and data warehouse data sources.
5. The data management platform with metadata as its core as described in claim 3, characterized in that, It also includes a data integration module, which is configured to: aggregate multi-source heterogeneous data and build a data synchronization channel to enable business data to be quickly put into the warehouse.
6. The data management platform with metadata as its core as described in claim 3, characterized in that, It also includes a data development module, configured to perform model design, self-service development, and API development tasks.
7. The data management platform with metadata as its core as described in claim 3, characterized in that, It also includes a data service module, which is configured to: perform integrated data management across the entire data domain, and enable data sharing and circulation across departments, regions, and levels.
Citation Information
Patent Citations
Data management system and working method
CN113138973A
Batch data fault-tolerant acquisition method based on capture metadata change
CN115712623A
Metadata blood relationship analysis method, system and equipment and storage medium
CN116226159A