Rail transit metadata management system based on multi-source data

By designing a rail transit metadata management system based on multi-source data, the insufficient timeliness and data quality of metadata acquisition in the existing technology is solved, real-time, comprehensive and high-quality acquisition of metadata is achieved, and data management efficiency and quality are improved.

CN120086201APending Publication Date: 2025-06-03BEIJING METRO NETWORK ADMINISTRATION CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510137824.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-07
Publication Date
2025-06-03

AI Technical Summary

Technical Problem

The existing technology metadata collection in the field of rail transit has problems such as insufficient timeliness, easy to generate garbage data, repeated collection and inability to monitor in real time, which has affected the efficiency and quality of data management.

Method used

Design a rail transit metadata management system based on multi-source data, and realize metadata acquisition and integration across multi-source databases by building a metadata acquisition interface that supports multi-source data. The system includes a collection module, a monitoring module, a recording module, a view management module and a comparison module. It uses the open source job scheduling tool XXL-Job technology to perform parallel collection, monitors log file changes in real time, and filters abnormal data through data cleaning rules.

Benefits of technology

Real-time acquisition and high-quality management of metadata are realized, the efficiency and accuracy of metadata acquisition are improved, and the rapid acquisition and management needs of rail transit operation data are met.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120086201A_ABST
    Figure CN120086201A_ABST
Patent Text Reader

Abstract

The invention relates to a rail transit metadata management system based on multi-source data. The system comprises a construction module, an acquisition module, a monitoring module, a recording module, a view management module and a comparison module. According to the system and the method, a unified interface is realized, adapters of all data sources are developed, various data sources such as Hadoop, Mysql, Postgres and GaussDB are supported, different data sources can be adapted to carry out metadata acquisition operation, diversified data acquisition requirements are met, so that comparison of metadata of multiple platforms is realized, log file changes are monitored in real time through a monitoring module, and the data acquisition efficiency is improved. The timeliness of metadata acquisition is solved, and metadata information can be monitored and acquired in real time; abnormal data is filtered through a data cleaning rule, and high quality of metadata information is guaranteed; meanwhile, records can be generated for addition, deletion and modification of the table, and tracking is facilitated; and a table key field marking monitoring mode created through the view is helpful to serve as a basis for key governance content of subsequent data governance.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of data processing and management, and in particular to a rail transit metadata management system based on multi-source data. Background Art

[0002] In today's society, the amount of data is expanding rapidly, and data is becoming the core competitiveness of governments and enterprises. People use data analysis to explore the value of data and provide accurate judgment basis for management decision makers.

[0003] The metadata management system is an important tool for improving the level of sharing, retrieving and understanding enterprise information assets, and is the lubricant of enterprise information management. If an enterprise does not manage metadata or manages it improperly, the information will be lost or hidden and difficult to be used by users, and data integration will be very expensive and unable to effectively support the business. Among them, metadata collection is the core of the metadata management system and the foundation of the entire system. Good metadata management is the basis for effective governance of all data and is helpful for lineage analysis.

[0004] The current metadata collection is mainly divided into two types: manual collection and scheduled collection. Manual collection is to collect the attribute information of the library, table, field, etc. in the specified data source based on the data source information using the corresponding collector. Scheduled collection is to collect the above information at a specified time.

[0005] In the field of rail transit technology, there is not only the basic characteristic that the source data of other industries is very large, but also the data in the field of rail transit technology needs to be collected from multiple heterogeneous systems, such as ticketing systems, monitoring systems, train dispatching systems, equipment operation and maintenance systems, environmental monitoring equipment, etc. When building various systems, in order to meet their own business needs, there are also differences in the selection of database products. In addition, due to the huge amount of rail transit data, in order to meet the real-time nature of the data and the accuracy of the indicator calculation at the same time, the table structure will increase the statistical caliber of the indicator in design. The statistical caliber will be designed in the form of final report (calculating the data from seven days ago to the current day once a day to facilitate the calculation of supplementary numbers) and quick report (instant calculation of the current day's data). Of course, the operation of rail transit involves multiple subsystems, each of which will generate corresponding business data. The data structure and analysis requirements of different business modules are different, resulting in a large number of tables. In addition, historical data needs to be retained for a long time to support analysis and prediction. Data accumulates rapidly. Almost all types of business tables are divided by time (such as day, month, year) or by business (such as line, station, region), resulting in a large number of partitions.

[0006] However, in the case of adopting the above two data collection methods, when the amount of metadata is huge, the collection takes a long time; without cleaning the metadata, a large amount of garbage data is easily generated; the same metadata will be collected repeatedly; and the scheduled collection cannot monitor in real time, making it difficult to meet the requirements of scenarios with rapidly changing data, seriously affecting the efficiency and quality of data management. Therefore, a metadata collection and governance method that can support multiple data sources is needed to achieve comprehensive, accurate, and timely collection and effective management of metadata. Summary of the Invention

[0007] To solve the above technical problems existing in the prior art, the object of the present invention is to provide a rail transit metadata management system based on multi-source data, which can realize metadata collection and integration across multiple source databases and achieve real-time collection of metadata.

[0008] To achieve the above object of the invention, the present invention provides a rail transit metadata management system based on multi-source data, including: a construction module for constructing a metadata collection interface that supports multi-source data, and the metadata collection interface includes a historical data interface and an incremental data interface;

[0009] A collection module for obtaining library, table, view, and field information in the data warehouse through the data collection interface, including batch collection of historical metadata and real-time collection of incremental metadata. The collection process is based on the open-source job scheduling tool XXL-Job technology. By reading the specified table of the data source and batch loading the data, multiple jobs are started in parallel according to the size of the data volume to improve efficiency;

[0010] A monitoring module for real-time monitoring of changes in the log file. When a new data processing script is detected, key information in the script is extracted according to predefined parsing rules to obtain the changed metadata information;

[0011] A recording module for recording the original record of metadata pulling for each table and the table structure change information;

[0012] A view management module for establishing a mapping relationship between the physical table and the view field based on the DDL of the view;

[0013] A comparison module for providing a basis for script modification after data migration.

[0014] According to a technical solution of the present invention, a unified metadata collection interface is designed by using the construction module, and the metadata collection interface is at least adapted to Hadoop, Mysql, Postgres, and GaussDB to implement metadata collection jobs.

[0015] According to a technical solution of the present invention, the acquisition module completes the acquisition of historical metadata through the historical data interface, and the acquisition of historical metadata starts multiple jobs in parallel for data acquisition according to the data volume;

[0016] The acquisition module completes the incremental metadata through the incremental data interface. The incremental metadata acquisition adopts the full-pull method in the early stage of acquisition, and multiple parallel pull jobs can be configured for one data source. In the later stage of acquisition, the method of pulling incremental modified data is adopted, and one pull job is configured for one data source.

[0017] According to a technical solution of the present invention, the monitoring module monitors the metadata addition and modification information by different methods according to different data sources.

[0018] For Mysql data and Hive data, the binlog logs are monitored;

[0019] For Postgres data and GaussDB data, the system tables themselves are monitored to obtain metadata change information.

[0020] According to a technical solution of the present invention, it further includes:

[0021] A generation module, after the monitoring module obtains a new data processing script by monitoring the log file of the data warehouse, generates the lineage of the data and the record information of the data being accessed;

[0022] A storage module, used to transmit the data processing script to the topic of Kafka.

[0023] According to a technical solution of the present invention, it further includes:

[0024] A data cleaning module, configured with data inspection rules, used to perform a cleaning operation on the metadata collected by the acquisition module based on the data inspection rules, and filter out the core data and store it in the data governance repository.

[0025] According to a technical solution of the present invention, the data cleaning module is also used to perform a cleaning operation on the metadata after abnormal data after the monitoring module transmits the data processing script to the topic of Kafka;

[0026] The storage module is also used to store the data processing script after filtering out abnormal data in the data governance repository.

[0027] According to a technical solution of the present invention, the data inspection rules at least include:

[0028] Data integrity rules, used to ensure that the collected metadata information is complete and without defects;

[0029] Data consistency rules are used to determine whether the metadata of different data sources or different parts of the same data source is consistent;

[0030] Data accuracy rules verify whether the values of metadata conform to the expected range and format.

[0031] According to one technical solution of the present invention, the view management module establishes a mapping relationship between physical tables and view fields based on the data definition language of the view. The specific method is as follows:

[0032] Parse the data definition language statements of the view and extract the corresponding relationships between physical tables and view fields;

[0033] Store the corresponding relationships in a mapping table or data structure, and determine the fields of the physical table used to construct the view and the physical table fields corresponding to each field in the view by querying the mapping table or data structure;

[0034] Based on the usage frequency of the corresponding fields, determine the fields that need to be focused on during data governance.

[0035] According to one technical solution of the present invention, before and after data migration, the comparison module respectively collects metadata information of the same table or related tables in the source database and the target database, including table structure, field type, storage method, and performs comparison and analysis on multi-source metadata through a preset comparison algorithm, and re-adapts the script syntax and modifies the database field type after migration.

[0036] Compared with the prior art, the present invention has the following beneficial effects:

[0037] The present invention proposes a rail transit metadata management system based on multi-source data. By implementing a unified interface and developing adapters for each data source, it supports multiple data sources such as Hadoop, Mysql, Postgres, and GaussDB, can adapt to different data sources to carry out metadata collection operations, meet diverse data collection requirements, and thus realize the comparison of metadata on multiple platforms.

[0038] The collection module uses the open-source job scheduling tool XXL-Job technology to collect various types of information such as databases, tables, views, and fields in the data warehouse based on the constructed interface. Whether it is batch collection of historical metadata or real-time collection of incremental metadata, multiple jobs can be flexibly started in parallel according to the data volume, greatly improving the efficiency of metadata collection, ensuring that relevant data such as rail transit operations can be quickly and comprehensively obtained, and providing a sufficient data basis for subsequent management and analysis.

[0039] In the present invention, by monitoring the changes of the log file in real time through the monitoring module, the timeliness of metadata collection is solved, and real-time monitoring and collection of metadata information can be achieved; abnormal data is filtered through data cleaning rules to ensure the high quality of metadata information; at the same time, records are generated for any addition, deletion, or modification of tables, facilitating tracking; and the method of marking and monitoring key fields of the tables created through views helps to serve as the basis for the key governance content in subsequent data governance. BRIEF DESCRIPTION OF THE DRAWINGS

[0040] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following will briefly introduce the drawings required for the embodiments. Obviously, the drawings described below are only some embodiments of the present invention, and those of ordinary skill in the art can obtain other drawings based on these drawings without creative efforts.

[0041] Figure 1 Schematically showing the structural diagram of the rail transit metadata management system based on multi-source data in an embodiment of the present invention;

[0042] Figure 2 Schematically showing the collection process diagram of the collection module in an embodiment of the present invention;

[0043] Figure 3 Schematically showing the monitoring process diagram of the monitoring module in an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0044] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the scope of protection of the present invention.

[0045] As Figure 1As shown in the figure, the present invention provides a rail transit metadata management system based on multi-source data, including a construction module, a collection module, a monitoring module, a recording module, a view management module, and a comparison module, which can solve the timeliness of metadata collection and can monitor and collect metadata information in real time. The present invention solves the problem of abnormal metadata quality, filters abnormal data through data cleaning rules, and ensures that the final metadata information is of high quality. The present invention solves the problem of repeated metadata collection, batch collects historical metadata, and real-time updates incremental metadata. The present invention solves the problem of collecting a large amount of metadata, and solves the efficiency problem of collecting a large amount of data by configuring the number of jobs running simultaneously. Further, when the number of partitions of the table is too large, multiple threads need to be opened for pulling. Version management is performed on the table, and records are generated for any addition, deletion, or modification of the table, which is convenient for tracking. The method of marking key fields of the table created through the view helps to serve as the basis for the key governance content of subsequent data governance.

[0046] Among them, the construction module is used to construct a metadata collection interface that supports multi-source data and is used to implement metadata collection jobs. The metadata collection interface includes a historical data interface and an incremental data interface, and the metadata collection interface is at least adapted to Hadoop, Mysql, Postgres, and GaussDB.

[0047] In order to achieve unified support for multiple data sources, the present invention designs a set of general metadata collection interfaces, which are divided into a historical data interface and an incremental data interface. The historical data interface is used to batch pull the existing metadata in the data source, while the incremental data interface is used to monitor the addition, deletion, and modification changes of the metadata in real time, which can meet the different levels of metadata collection requirements of rail transit, and can not only obtain complete historical metadata, but also capture the dynamic changes of the metadata in a timely manner. At the same time, the system breaks through the information barriers of different rail transit systems by accessing multi-system data sources. Through the collection and governance of the source metadata, users can quickly master the data types involved in each system and the data operation logic of each system, which is convenient for subsequent data fusion development quickly and accurately.

[0048] For different data sources such as Hadoop, Mysql, Postgres, and GaussDB, corresponding adapters are developed to facilitate metadata collection jobs that can adapt to various data sources. These adapters serve as the bridge between the interface and the data source, and are responsible for implementing the specific connection and data interaction operations with each data source. For example, for the Mysql data source, the adapter establishes a connection through technologies such as JDBC, and executes corresponding queries and data extraction operations according to the specifications defined by the interface, so as to ensure that the metadata collection job can accurately obtain the required metadata information from various data sources.

[0049] Among them, JDBC (Java Database Connectivity) is a standard interface provided by Java for connecting to, querying, updating, and managing data in a Java application. It defines a set of APIs that enable Java programs to interact with relational databases.

[0050] Taking the Hadoop metadata collection as an example:

[0051] Data structure metadata collection: used to obtain all library, table, view, and field information in the data warehouse.

[0052] Data execution metadata collection: used to obtain SQL commands executed in the data warehouse, including insert, update, delete, and select.

[0053] Such as Figure 2 As shown, the collection module is used to obtain library, table, view, and field information in the data warehouse through the data collection interface, including batch collection of historical metadata and real-time collection of incremental metadata. The collection process is based on the open-source job scheduling tool XXL-Job technology. By reading the specified table in the data source and batch loading the data, multiple jobs are started in parallel according to the data volume size to improve efficiency.

[0054] The collection module completes the historical metadata collection through the historical data interface. The historical metadata collection starts multiple jobs in parallel for data collection according to the data volume size.

[0055] The present invention develops and designs a metadata collection module based on the open-source job scheduling tool XXL-Job. By reading the specified table in the data source through the data collection interface, all metadata information such as libraries, tables, views, and fields in the data warehouse can be batch loaded into the collection system. According to the data volume size in the data source, multiple jobs can be flexibly configured to start in parallel for data collection, which can greatly improve the data collection efficiency, is conducive to completing the large-scale metadata collection task in a short time, and provides strong support for quickly establishing a comprehensive metadata warehouse for the rail transit system.

[0056] Among them, XXL-Job is a lightweight distributed task scheduling platform mainly used for the timing scheduling and execution of tasks, applicable to various scenarios requiring timing task scheduling, and capable of realizing the distributed management of tasks. XXL-Job supports Web interface operations and provides a simple and intuitive configuration method, and is widely used in scenarios such as microservice architecture, big data processing, and distributed systems.

[0057] The acquisition module completes incremental metadata through the incremental data interface. In the early stage of acquisition, full - volume pulling is adopted. Multiple parallel pulling jobs can be configured for one data source. In the later stage of acquisition, the method of pulling incrementally added, modified, and deleted data is adopted, and one pulling job is configured for one data source. By adopting the above method, full - volume pulling in the early stage can ensure data integrity and improve acquisition efficiency. Incremental pulling in the later stage is conducive to implementing monitoring of metadata changes, reducing system resource consumption, and ensuring system stability.

[0058] Full - volume pulling in the early stage of acquisition can obtain all the existing metadata information in the current data source at one time, laying a foundation for subsequent metadata management and analysis. Through full - volume pulling, a complete snapshot of the data source status can be obtained, including various information such as databases, tables, views, and fields, avoiding omission of important metadata and enabling subsequent data processing and analysis to be based on a complete data set.

[0059] In the later stage of acquisition, the focus is on the added, deleted, and modified information of metadata. By pulling incrementally added, modified, and deleted data, the metadata repository can be updated in real - time or regularly, ensuring that the information in the metadata repository is always synchronized with the latest status of the data source. This is crucial for the continuous and stable operation and management of data warehouses or databases, avoiding data management decision - making mistakes or data processing errors caused by outdated metadata.

[0060] Compared with full - volume pulling, the amount of change in incremental metadata is usually smaller, and one job is sufficient to monitor and acquire subsequent metadata change information. This configuration avoids excessive resource occupation, ensures reasonable utilization of system resources, and at the same time ensures timely access to metadata change information.

[0061] As Figure 3 shown, the listening module is used to monitor changes in the log file in real - time. When a new data - processing script is detected, key information in the script is extracted according to predefined parsing rules to obtain the changed metadata information;

[0062] Based on XXL - Job technology, a connection with the data source is established through the metadata acquisition interface. Once a new metadata change record is found, the changed metadata information is immediately obtained and passed to the subsequent metadata cleaning job for processing to ensure the timeliness and accuracy of metadata.

[0063] For the Mysql data source, by monitoring its binlog logs, the change operations of the database are captured, so as to obtain the information of new addition, modification and deletion of metadata in a timely manner. As a data warehouse tool based on Hadoop, the metadata of Hive is stored in the underlying Mysql. Therefore, the incremental metadata collection is also achieved by monitoring the Binlog log files of the metadata database. For the Postgres data source, since its metadata is stored in its own system tables, the change information of the metadata can be obtained by directly reading the system tables. GaussDB has the same origin as Postgres, so the same reading method as Postgres is adopted to obtain incremental metadata.

[0064] Among them, the Binlog (Binary Log) log is a log file used to record all database modification operations in the MySQL database. It records all data modification events (such as insert, update, delete, etc.), and these operations will be recorded in the binlog for subsequent data recovery, replication and auditing.

[0065] In some embodiments of the present invention, the rail transit metadata management system based on multi-source data further includes:

[0066] A generation module, which generates the lineage of data and the record information of data access after the listening module obtains a new data processing script by listening to the log files of the data warehouse;

[0067] A storage module, which is used to transmit the data processing script to the topic of Kafka.

[0068] A data cleaning module, which is configured with data inspection rules, and is used to perform a cleaning operation on the metadata collected by the collection module based on the data inspection rules, and filter out the core data and store it in the data governance repository.

[0069] For example, when a new data processing script is executed in Hive, the job program can capture the script in real time and transmit it to the topic of Kafka. As a high-performance distributed message queue, Kafka can reliably store and transfer data processing scripts. Subsequently, through the data cleaning operation, the abnormal data in the script is filtered, and the invalid or incorrect information is removed. Finally, the cleaned data is used to perform metadata storage in the data governance repository. In this way, not only can the lineage of data be generated, clearly showing the source and destination of the data, but also the detailed record information of data access can be recorded.

[0070] In addition, the data cleaning module is also used to perform a cleaning operation on the metadata after abnormal data after the listening module transmits the data processing script to the topic of Kafka;

[0071] The storage module is also used to store the data processing script after filtering abnormal data in the data governance repository.

[0072] The collected metadata often contains various noises and abnormal data. In order to ensure the quality of metadata, the present invention introduces a metadata cleaning operation. By configuring detailed data inspection rules, the metadata cleaning operation can comprehensively inspect and filter the collected metadata.

[0073] For example, based on predefined field formats, data types, value ranges and other rules, abnormal data that does not meet the requirements can be located and removed from the collection results. At the same time, data cleaning operations can also standardize metadata, unify metadata formats in different data sources, ensure consistency and standardization of metadata stored in the data governance repository, and provide a high-quality data foundation for subsequent data analysis and applications.

[0074] The data inspection rules at least include:

[0075] Data integrity rules to ensure that the collected metadata information is complete;

[0076] Data consistency rules are used to determine whether the metadata of different data sources or different parts of the same data source are consistent;

[0077] Data accuracy rules verify that metadata values ​​are within expected ranges and formats.

[0078] The recording module is used to record the metadata of each table, pull the original records and table structure change information.

[0079] During the metadata collection process, for each metadata pulling operation of a table, the system of the present invention generates a detailed original record, which at least includes: the initial pulling time of the table, the content and modification time of each subsequent modification to the table, the modifier and other important information.

[0080] By recording this information, the evolution history of the table can be fully traced, providing data managers with a clear timeline to help them understand the changes in the data structure. Whenever the table structure is modified, the system automatically generates a new version of the table and associates it with the previous version for storage. In this way, you can easily trace back to any historical version of the table structure when needed, which plays an important auxiliary role in data consistency checks, troubleshooting, and data recovery.

[0081] The view management module is used to establish the mapping relationship between the physical table and the view field based on the DDL of the view. The view management module establishes the mapping relationship between the physical table and the view field based on the data definition language of the view. The specific method is:

[0082] Parse the data definition language statements of the view, and extract the corresponding relationship between the physical tables and the view fields;

[0083] Store the corresponding relationship in a mapping table or data structure, and determine the fields of the physical table used to construct the view and the physical table fields corresponding to each field in the view by querying the mapping table or data structure;

[0084] Based on the usage frequency of the corresponding fields, determine the fields that need to be focused on during data governance.

[0085] First of all, the analysis requirements of rail transit usually involve complex SQL queries, such as multi-table association, aggregation calculation, grouping statistics, etc. Using views can encapsulate these complex query logics. Secondly, the rail transit industry involves multiple departments and roles (such as operation management, equipment operation and maintenance, dispatching command, financial analysis, etc.). Different roles focus on different indicators and data dimensions. Views can provide customized data display forms for different roles to meet personalized needs. Thirdly, there are sensitive data in the rail transit system (such as passenger information, payment records, equipment status). Directly opening the underlying tables may pose a risk of data leakage. The underlying tables can be filtered and desensitized through views, and only the necessary data is exposed to specific users or departments. Fourthly, the rail transit system involves multiple subsystems (such as passenger flow analysis, equipment operation and maintenance, train dispatching, ticket management, etc.). The data of these systems is usually scattered and stored in different tables. Through views, the data in multiple tables can be integrated to form a logically "unified data source" for business personnel or application programs to use without directly operating the underlying tables.

[0086] For the above reasons, a large number of views will be used in the system. Subsequent ETL (data extraction, transformation, and loading to the destination, etc.) processing reads data through views, so view management is equally important. The DDL (data definition language) of the view is the mapping relationship from the physical table to the view, and usually some special processing is done on the key fields to be used. Therefore, the mapping relationship between some fields of the physical table and the view fields can be established through the DDL of the view. Through this mapping relationship, it can be clearly known which fields have a higher usage frequency, and these fields can be used as the fields that need to be focused on during subsequent data governance. For example, in data quality monitoring, data optimization and other work, these frequently used fields can be preferentially checked and optimized to improve the pertinence and effectiveness of data governance.

[0087] At the same time, by establishing a mapping table in this system and realizing the timed collection of different data sources, the unified governance of multi-source data is ensured, thus solving the difficult problem caused by the data collection in the rail transit field from multiple heterogeneous systems.

[0088] Among them, ETL is the abbreviation of Extract, Transform, Load (extraction, transformation, loading), which is a core concept in data processing and data integration and is widely used in fields such as data warehouses, data marts, and business intelligence (BI) systems. The purpose of the ETL process is to extract data scattered in multiple data sources, perform formatting conversions, and finally load it into a target system (such as a database or data warehouse) for analysis and reporting.

[0089] Of course, the view is the middle part between the data table and the final metrics. Based on rail transit data, line merging is usually required to store some common intermediate calculation results. The extensive use of views simplifies the calculation logic of the final metrics and reduces the time consumption of data queries.

[0090] Systems with large amounts of data often have the need for data sharing, so data migration is common. The comparison module is mainly applied in the data migration process to provide a basis for modifying the scripts after data migration. Rail transit data is widely used in the analysis of urban construction and smart city-related businesses. Therefore, data is often migrated to other systems for use, that is, all data and scripts are migrated from one database system to another. After migration, script syntax re-adaptation and database field type modification are usually carried out. Through the metadata comparison function, the changes in data fields, field types, and storage methods of the same table before and after migration can be intuitively seen, and this is used as the basis for assisting in modifying the scripts after migration.

[0091] Through the governance of metadata by this system, it can ensure that the table structures at each end are consistent during data migration and reduce the occurrence of structural errors.

[0092] Specifically, before and after data migration, the comparison module collects the metadata information of the same table or related tables in the source database and the target database respectively, including table structure, field type, and storage method. Through a preset comparison algorithm, it compares and analyzes multi-source metadata, and re-adapts the script syntax and modifies the database field type after migration, which can intuitively display the changes in data fields, field types, storage methods, etc. of the same table before and after migration. By comparing the metadata differences before and after migration, data managers can quickly discover potential problems and accordingly modify and adjust the scripts after migration, thus ensuring the smooth progress of data migration and the integrity and consistency of data.

[0093] A rail transit metadata management system based on multi-source data of the present invention supports multiple data sources such as Hadoop, Mysql, Postgres, and GaussDB by implementing a unified interface and developing adapters for each data source. It can adapt to different data sources to carry out metadata collection operations, meet diverse data collection requirements, and thus realize the comparison of metadata on multiple platforms.

[0094] The collection module uses the open-source job scheduling tool XXL-Job technology to collect various types of information such as databases, tables, views, and fields in the data warehouse based on the constructed interface. Whether it is the batch collection of historical metadata or the real-time collection of incremental metadata, multiple jobs can be flexibly started in parallel according to the data volume, greatly improving the efficiency of metadata collection, ensuring that relevant data such as rail transit operations can be obtained quickly and comprehensively, and providing an adequate data basis for subsequent management and analysis.

[0095] In the present invention, by monitoring the change of the log file in real time through the monitoring module, the timeliness of metadata collection is solved, and metadata information can be monitored and collected in real time; abnormal data is filtered through data cleaning rules to ensure the high quality of metadata information; at the same time, records will be generated for any addition, deletion, or modification of the table, which is convenient for tracking; and the method of marking and monitoring key fields of the table created by the view helps to serve as the basis for the key governance content of subsequent data governance.

[0096] The parts not elaborated in detail in the present invention belong to the well-known technology in the art.

[0097] Finally, it should also be noted that in this article, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements not only includes those elements, but also includes other elements not expressly listed, or elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in the process, method, article or device comprising the element.

[0098] The various embodiments in this specification are described in a progressive manner. Each embodiment focuses on the differences from other embodiments. The various embodiments can be combined as needed, and the same or similar parts can be referred to each other.

[0099] The foregoing description of the disclosed embodiments enables those skilled in the art to practice or use the present application. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present application. Thus, the present application is not intended to be limited to the embodiments shown herein but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A rail transit metadata management system based on multi-source data, characterized in that: include: A construction module is used to construct a metadata collection interface that supports multi-source data, wherein the metadata collection interface includes a historical data interface and an incremental data interface; The acquisition module is used to obtain the library, table, view, and field information in the data warehouse through the data acquisition interface, including batch acquisition of historical metadata and real-time acquisition of incremental metadata. The acquisition process is based on the open source job scheduling tool XXL-Job technology, which reads the specified table of the data source and loads data in batches. At the same time, multiple jobs are started in parallel to improve efficiency according to the amount of data; The monitoring module is used to monitor the changes of log files in real time. When a new data processing script is detected, the key information in the script is extracted according to the predefined parsing rules to obtain the changed metadata information; The recording module is used to record the metadata of each table, pull the original records and table structure change information; The view management module is used to establish the mapping relationship between the physical table and the view field based on the view DDL; The comparison module is used to provide a basis for script modification after data migration.

2. The rail transit metadata management system based on multi-source data according to claim 1, characterized in that: A unified metadata collection interface is designed using the construction module, and the metadata collection interface is at least compatible with Hadoop, Mysql, Postgres, and GaussDB to implement metadata collection operations.

3. The rail transit metadata management system based on multi-source data according to claim 2 is characterized in that: The acquisition module completes the historical metadata acquisition through the historical data interface. The historical metadata acquisition starts multiple jobs for parallel data acquisition according to the amount of data. The acquisition module completes incremental metadata through the incremental data interface. The incremental metadata acquisition adopts a full-volume pull method in the early stage of acquisition, and one data source can be configured with multiple parallel pull jobs. In the later stage of acquisition, the incremental metadata acquisition adopts a method of pulling added and modified data, and one data source is configured with one pull job.

4. The rail transit metadata management system based on multi-source data according to claim 3 is characterized in that: The monitoring module uses different methods to monitor metadata addition and modification information according to different data sources. For MySQL data and Hive data, monitor the binlog log; For Postgres data and GaussDB data, monitor their own system tables to obtain metadata change information.

5. The rail transit metadata management system based on multi-source data according to claim 1, characterized in that: Also includes: A generation module, after the monitoring module obtains a new data processing script by monitoring the log file of the data warehouse, generates data lineage relationship and data access record information; The storage module is used to transfer the data processing script to the Kafka topic.

6. The rail transit metadata management system based on multi-source data according to claim 5, characterized in that: Also includes: The data cleaning module is configured with data checking rules and is used to clean the metadata collected by the collection module based on the data checking rules, and filter out the core data to be stored in the data governance repository.

7. The rail transit metadata management system based on multi-source data according to claim 6, characterized in that: The data cleaning module is also used to clean the metadata for abnormal data after the monitoring module transmits the data processing script to the topic of Kafka; The storage module is also used to store the data processing script after filtering abnormal data in the data governance repository.

8. The rail transit metadata management system based on multi-source data according to claim 6, characterized in that: The data inspection rules at least include: Data integrity rules to ensure that the collected metadata information is complete; Data consistency rules are used to determine whether the metadata of different data sources or different parts of the same data source are consistent; Data accuracy rules verify that metadata values ​​are within expected ranges and formats.

9. The rail transit metadata management system based on multi-source data according to claim 1, characterized in that: The view management module establishes a mapping relationship between the physical table and the view field based on the data definition language of the view. The specific method is as follows: Parse the data definition language statements of the view and extract the correspondence between the physical table and the view fields; The corresponding relationship is stored in a mapping table or a data structure, and the fields of the physical table used to construct the view and the physical table fields corresponding to each field in the view are determined by querying the mapping table or the data structure; Determine the fields that are important to focus on during data governance based on the usage frequency of the corresponding fields.

10. The rail transit metadata management system based on multi-source data according to claim 1, characterized in that: Before and after data migration, the comparison module collects metadata information of the same table or related tables in the source database and the target database, including table structure, field type, and storage method, compares and analyzes multi-source metadata through a preset comparison algorithm, and re-adapts the script syntax and modifies the database field type after migration.

Citation Information

Patent Citations

  • Metadata-based management and analysis system

    CN112699100A

  • Multi-type database table structure comparison method and system, equipment and storage medium

    CN113672639A

  • Data synchronization method and system based on XXL-JOB and FlinkSQL

    CN114048492A

  • Method for solving change response problem of multi-bin model in real time

    CN115470217A

  • Adaptive metadata acquisition and change tracking system

    CN118689863A