A multi-source data fusion system based on intelligent lake warehouse architecture and a method thereof

The multi-source data fusion system based on the intelligent lake warehouse architecture solves the problem of managing multi-source heterogeneous data, realizes unified data access, standardized processing and efficient storage, reduces costs and risks, and improves data aggregation efficiency and quality.

CN120705234BActive Publication Date: 2025-11-11FUZHOU UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511156350.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-08-19
Publication Date
2025-11-11
Estimated Expiration
2045-08-19

AI Technical Summary

Technical Problem

Existing technologies are insufficient for effectively managing multi-source heterogeneous data, resulting in inconsistent data quality and storage and computing pressures, leading to difficulties in data aggregation and high costs.

Method used

A multi-source data fusion system based on an intelligent lake warehouse architecture is adopted, including data catalog management, access management, lake entry module, aggregation quality inspection and storage management module. Through metadata management, unified access standards, data verification rules and partitioned storage, the system realizes standardized data processing and storage.

Benefits of technology

It improved the efficiency and quality of data aggregation, reduced the risks and development costs of data integration, and enabled centralized data management of multiple business systems.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120705234B_ABST
    Figure CN120705234B_ABST
Patent Text Reader

Abstract

The application provides a multi-source data fusion system and method based on an intelligent lake warehouse architecture, comprising a data directory management module, a data access management module, a data entry lake module, a data aggregation quality inspection module, a data reconciliation module and a data storage management module. The technical solution can improve data aggregation efficiency and data quality, reduce the risk of data integration and the technical threshold of data aggregation, and reduce the development cost of data aggregation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of data management technology, and in particular to a multi-source data fusion system and method based on an intelligent lake warehouse architecture. Background Technology

[0002] With the continuous deepening of information technology construction, massive amounts of data have accumulated, and the sources of data have become increasingly complex. The technical platforms and requirements for data storage and aggregation are becoming increasingly sophisticated. At the same time, the application scenarios, data types, and data scale of data governance and analysis are increasing, leading to higher demands for real-time collection and governance of massive amounts of data. Therefore, technologies for accessing, aggregating, processing, and storing full-volume, multi-source, heterogeneous data all face a series of technical challenges. The following are the problems and shortcomings of existing technologies:

[0003] 1. The data sources and types are diverse, making unified integration and storage management difficult.

[0004] Data sources may include structured data (such as database tables), semi-structured data (such as JSON and XML), unstructured data (such as text, images, and videos), and even real-time streaming data (such as IoT device logs). Different source systems may choose different integration methods due to varying production environment factors (e.g., interface integration for data security, or front-end database exchange for business systems that cannot be modified). Furthermore, different data formats exhibit significant differences in storage structure, field definitions, and encoding methods. Therefore, the coexistence of database tables, log files, streaming data, and text formats necessitates adaptation to different parsing engines, making direct integration and management difficult.

[0005] 2. The quality of data from multiple sources varies.

[0006] Data from multiple sources comes from different systems or production environments. Data sources may have missing key fields, or different systems may store duplicate records of the same entity, which can cause data aggregation to be unable to be shared and applied effectively.

[0007] 3. Storage and computing pressure brought about by the growth in the scale of multi-source data

[0008] As the number and volume of data sources increase, lake warehouse architectures face challenges in terms of storage costs and computing resources. The direct ingestion of heterogeneous data from multiple sources into the lake warehouse architecture, without standardized processing, leads to a large accumulation of duplicate and invalid data, causing storage costs to increase exponentially. Summary of the Invention

[0009] In view of this, the purpose of this invention is to provide a multi-source data fusion system and method based on an intelligent lake warehouse architecture, which improves data aggregation efficiency and data quality, reduces the risk of data integration and the technical threshold of data aggregation, and reduces the development cost of data aggregation.

[0010] To achieve the above objectives, the present invention adopts the following technical solution: a multi-source data fusion system based on an intelligent lake warehouse architecture, including a data catalog management module, a data access management module, a data entry module, a data aggregation and quality inspection module, a data reconciliation module, and a data storage management module;

[0011] The data directory management module includes metadata management, directory registration, directory review, directory retrieval, directory download, directory decommissioning, and data mounting;

[0012] The data access management module includes data access configuration and data aggregation monitoring;

[0013] The data import module includes account creation, data import operation, diverse data processing, data table management, and automatic collection and synchronization of metadata;

[0014] The data aggregation and quality inspection module includes standardized management of data table names, management of data verification rules, import of data verification rules, configuration of table / field verification rules, configuration of verification task priority, and verification task logs.

[0015] The data reconciliation module includes reconciliation time configuration, reconciliation statement management, reconciliation rule configuration, data re-push, and reconciliation result maintenance;

[0016] The data storage management module includes a lake-warehouse integrated data management architecture, which is used to partition and manage the aggregated data according to business attributes and data characteristics.

[0017] In a preferred embodiment, the metadata management module in the data catalog management module is used for the registration and management of the interpretation and definition of accessed data. After the system data is accessed, it automatically analyzes and obtains business metadata information and creates a global-level metadata file to interpret and define the data, forming a consistent and unified data definition within a certain range. For unstructured data or files without built-in metadata, the metadata information is manually sorted and entered into the metadata management module. The metadata management module includes metadata maintenance, metadata generation, metadata acquisition, and metadata query.

[0018] The directory registration is a structured presentation of metadata; the directory registration is used to register and manage system information, database information, basic data information, and data item information;

[0019] The directory review is used to review the compilation and publication of the directory and to fill in comments, including the directory registration review and the directory change review. Once the review is passed, the directory is published and used in the form of a collection of directories.

[0020] The directory retrieval method allows users to search for directory information by directory name, directory status, and hierarchy.

[0021] The directory download is used to download registered directories;

[0022] The directory decommissioning is used to decommission directories that are no longer in use or have been suspended from use;

[0023] The data mounting is used to mount the aggregated data of the core data area to the corresponding data directory; to associate the physically stored dataset with the metadata in the data directory, and to locate and access the actual data through the data directory.

[0024] In a preferred embodiment, the data access management module provides data access methods for different data sources and different data formats, and each external access system accesses data to this system in accordance with a unified data access standard specification; including configuration management and access monitoring for streaming data access, front-end exchange, data import, file transfer, and data entry methods.

[0025] In a preferred embodiment, the data ingestion module provides open data lake storage capabilities, including storage media, storage format, database and table data organization methods, data read and write, snapshot management, metadata management, and table services; raw data from heterogeneous data sources are uniformly merged into a centralized data lake storage based on object storage architecture.

[0026] In a preferred embodiment, the data aggregation quality inspection module provides a custom governance rule configuration module, which is used to aggregate the data accessed from the original data area to the core data area and perform data table and data field verification according to the configured data verification rules.

[0027] In a preferred embodiment, the data aggregation quality inspection module specifically includes:

[0028] 1) Standardized Management of Data Table Names: Standardizes the naming of data tables accessed by the system; supports displaying data tables by organization / database tree, querying data tables by organization / database, classifying data tables by status, querying data tables by name, querying data tables by database name, querying data tables by the most recent modifier, data table editing, field information querying, paginated display of data tables, and paginated navigation of data tables; data table information includes table name, database, affiliated organization, most recent modifier, and most recent modification time;

[0029] 2) Data Validation Rule Management: Configure various data validation rules to validate accessed data. Data validation rules include null value validation, format validation, relationship validation, duplicate validation, comparison validation, logical validation, and sensitive data validation. Supports features such as data validation rule tree display, rule query by rule name, validation rule grouping and categorization, enabling validation rules, viewing validation rule details, querying rules by validation rule name, querying rules by validation rule code, resetting validation rules, displaying the latest version, displaying historical rule versions, enabling historical rule versions, viewing historical rule version details, comparing historical rule versions, paginated display of validation rules, paginated navigation of validation rules, and historical version selection. Data validation rule information includes rule code, rule name, rule description, rule grouping, rule issue level, most recently modified person, most recently modified time, and the latest version.

[0030] 3) Importing data validation rules: Data validation rules can be imported by uploading an Excel file, and Excel templates can be downloaded.

[0031] 4) Table / Field Validation Rule Configuration: Configure table / field validation rules to validate the data tables and fields in the access data tables; supports functions such as displaying the validation field tree, searching for validation fields by field, adding validation rules to data tables, querying rules by rule name, querying validation fields, resetting validation field queries, adding validation rules to validation fields, and setting validation conditions; rule configuration information includes validation fields, validation rules, and validation field information;

[0032] 5) Validation task priority configuration: Configure the data table validation task level to determine the order in which validation tasks are executed; there are five priority levels in total.

[0033] 6) Verification Task Log: Records the status of data table verification tasks, supports data table verification statistics and queries; supports viewing verification task log details, overall task volume statistics, verified data volume statistics, and problematic data volume statistics; verification task log information includes status, error reason, verified data volume, problematic data volume, unverified data volume, task priority, start time, and end time.

[0034] In a preferred embodiment, the data reconciliation module is used to verify and monitor the consistency of the aggregated data;

[0035] By selecting a reconciliation date, the data aggregation count is compared with the data access count. When the data aggregation count is inconsistent with the data access count, the unique identifier rowkey is checked again. The unique identifier that does not exist in the reconciliation list within the reconciliation date is found and returned to the data access party, which then resends the missing data.

[0036] In a preferred embodiment, the data reconciliation module includes the following reconciliation modes:

[0037] (1) Read the business database to reconcile data;

[0038] The data access party allocates business database permissions to the system; customizes reconciliation rules, and compares the data volume of the data access party's business database table and the access database table according to the rules.

[0039] (2) Manual reconciliation;

[0040] When a reconciliation work order is initiated, the aggregation platform's operations and maintenance team negotiates and determines the statistical criteria with the data access party. The data access party calculates the incremental data volume pushed to the business database the previous day according to the statistical criteria and delivers it to the aggregation platform's operations and maintenance team. The aggregation platform's operations and maintenance team calculates the incremental data volume received in the platform's database the previous day according to the statistical criteria and compares it with the data volume provided by the access party. If there is a discrepancy, the data access party provides the push logs from the previous day, and the operations and maintenance team analyzes and provides feedback.

[0041] (3) Real-time data push and reconciliation;

[0042] Modify the data push interface so that before pushing the actual data, the total amount of data to be accessed is sent to the platform. After the platform receives the data, it pushes the actual data. Before the data is stored in the database, the number of received records is counted and compared with the total amount of data to be accessed. If they are inconsistent, the access of this batch of data will fail.

[0043] In a preferred embodiment, the data storage management module of the smart lake warehouse uses a set of distributed cluster servers or nodes to collect, store and manage data from different sources and types; partitioned spaces are designed according to the storage type and usage scenario of the data; computing resources, data resources and component resources of each user are logically isolated; and multi-level permission management is supported within the tenant.

[0044] Users initially set up structured data areas according to different data storage types; based on the data access and storage usage scenarios, raw data areas are created on the basis of the data lake to store raw data from various sources. This is the logical storage location for unprocessed and unintegrated business data. It supports unified control of raw data through centralized management in a virtual form and provides a data source for upper-layer data value-added operations.

[0045] Based on the data aggregation and storage usage scenarios, a core data area and a shared database are created on the basis of the data warehouse. The core data area extracts and integrates multi-source access data and serves as the logical storage location for data aggregation processing. By establishing special topic or theme libraries, aggregated data is classified, stored, and managed, and easy-to-use integration interfaces are provided for access needs of other systems. The shared data area is a separate area for storing shared data, which is logically isolated from the actual data area and stores copies of all data resources that need to be shared.

[0046] This invention also provides a multi-source data fusion method based on an intelligent lakehouse architecture, and a multi-source data fusion system based on an intelligent lakehouse architecture, comprising the following steps:

[0047] Step 1: The backend management user initializes and creates the raw data area, core data area, and shared data area in the data storage management module, and configures the corresponding data storage path for each data area, including object storage, relational database, and document database; after successful configuration, a data area directory and ledger list are generated.

[0048] Step 2: Backend management users configure data aggregation verification rules in the data aggregation quality inspection module;

[0049] Step 3: Backend management users create original data area data entry accounts for data access users in the data entry module;

[0050] Step 4: Data access users register the data resources they need to access in the data catalog management module, fill in the data catalog information, and submit it for backend review;

[0051] Step 5: Backend management users review the data directory information submitted by the data access user in the data directory management module and fill in the review comments. Once approved, the data is published for use in the form of an aggregated data directory.

[0052] Step 6: Data access users select a data access method in the data access management module, configure data access parameters, and implement data connection according to the corresponding access technical specifications;

[0053] Step 7: Data access users configure the parameters for data aggregation and lake entry in the data lake module, and store the accessed data into the data lake;

[0054] Step 8: When the backend management user uses the data warehouse in the data storage management module, the data resources loaded into the lake will be automatically extracted and loaded into the core data area or shared data area, and the data tables and data fields of the data loaded into the lake will be verified according to the configured data verification rules; if the verification is successful, the registration information will be automatically added to the aggregated data directory.

[0055] Step 9: The data reconciliation module regularly performs end-to-end reconciliation of data access data, data entering the lake data, and data aggregation data; in case of data anomalies, the data sender will push the data back to the system.

[0056] Compared with existing technologies, this invention has the following advantages: It is efficient, open, and compatible. It provides centralized support for heterogeneous system data interface and integration management for organizations with multiple business systems. It standardizes the data access process, improves data aggregation efficiency and quality, reduces the risks and technical barriers to data aggregation, and lowers the development costs of data aggregation. Attached Figure Description

[0057] Figure 1 This is a flowchart of the data catalog management process according to a preferred embodiment of the present invention;

[0058] Figure 2 This is a flowchart illustrating the preferred embodiment of the present invention for reading data from the business database for data reconciliation.

[0059] Figure 3 This is a flowchart of the manual reconciliation method according to a preferred embodiment of the present invention;

[0060] Figure 4 This is a flowchart of the data push instant reconciliation method according to a preferred embodiment of the present invention;

[0061] Figure 5 This is a functional relationship diagram of a multi-source data fusion system based on an intelligent lake warehouse architecture, according to a preferred embodiment of the present invention.

[0062] Figure 6 This is a flowchart of a multi-source data fusion method based on an intelligent lake warehouse architecture, which is a preferred embodiment of the present invention. Detailed Implementation

[0063] The present invention will be further described below with reference to the accompanying drawings and embodiments.

[0064] It should be noted that the following detailed descriptions are illustrative and intended to provide further explanation of this application. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application pertains.

[0065] It should be noted that the terminology used herein is for the purpose of describing particular implementations only and is not intended to limit the exemplary implementations according to this application; as used herein, the singular form is intended to include the plural form as well, unless the context clearly indicates otherwise; furthermore, it should be understood that when the terms “comprising” and / or “including” are used in this specification, they indicate the presence of features, steps, operations, devices, components and / or combinations thereof.

[0066] A multi-source data fusion system based on an intelligent lake warehouse architecture, reference Figure 1-6 It includes a data catalog management module, a data access management module, a data lake entry module, a data aggregation and quality inspection module, a data reconciliation module, and a data storage management module.

[0067] The following is a detailed introduction to each module:

[0068] I. Data Directory Management Module

[0069] refer to Figure 1 The data catalog management module is used to manage the catalog of fused data after it has been aggregated and inspected, enabling unified view management of data resources. This includes metadata management, registration, review, retrieval, downloading, decommissioning, and data mounting for the data catalog.

[0070] (1) Metadata Management

[0071] This module is used for the registration and management of accessed data, including its interpretation and definition. After data is accessed, the system automatically analyzes and obtains business metadata information, creating a global metadata file to interpret and define the data, forming a consistent and unified data definition within a certain scope. For unstructured data (such as scanned copies of paper documents) or files without built-in metadata, metadata information (such as creation time, responsible person, and content summary) needs to be manually compiled and entered into the metadata management module. Metadata management mainly includes metadata maintenance, metadata generation, metadata acquisition, and metadata querying.

[0072] 1) Metadata maintenance

[0073] Used for registering, modifying, and viewing business metadata and technical metadata.

[0074] 2) Metadata generation

[0075] XML templates are provided for different categories of metadata.

[0076] 3) Metadata Acquisition

[0077] When creating a data directory, it automatically matches the metadata files from the data access point, and can automatically reference them if a match is found.

[0078] 4) Metadata Query

[0079] Used for quickly locating and finding metadata information.

[0080] (2) Catalog registration

[0081] A data catalog is a structured representation of metadata, used to facilitate users in finding, understanding, and using data. Catalog registration is used to register and manage system information, database information, basic data information, and data item information.

[0082] 1) System Information Registration and Management

[0083] This system allows users to add, modify, and delete application system information, and supports searching for application system information by keyword, application system category, and system status. Registration information includes: system name, affiliated organization name, operating network, scope of use, database creation time, application system category, system status, system classification, remarks, system description, contact person's department, contact person's name, and contact person's mobile phone number.

[0084] 2) Database information registration and management

[0085] It enables the addition, modification, and deletion of database information, and supports searching business databases by data resource name. Registration information includes: database name, database type, and character set encoding.

[0086] 3) Basic Data Information Registration and Management

[0087] This allows for the addition and modification of basic information about data resources. Registration information includes: database, system name, data provider, data name, unit name, directory code, data provider code, associated data code, data name (in English), update cycle, data scope, sharing type, update status, data format, data update cycle, directory category, and data resource summary.

[0088] 4) Data Item Information Registration and Management

[0089] This feature enables the addition, deletion, and batch deletion of data items, and supports Excel upload and download. Registration information includes: data item name, English name of data item, data type, unit of measurement, data length, whether it is a primary key, whether it is nullable, time format, data sensitivity registration, field value range, sharing type, sharing conditions, sharing category, sharing method type, desensitization rules, data precision, and description.

[0090] (3) Catalog review

[0091] The catalog review process involves approving the compilation and publication of the catalog and providing feedback to ensure the accuracy, completeness, and consistency between the registered data catalog and the actual metadata. It includes catalog registration review and catalog change review; once approved, the catalog is published for use in the form of a aggregated catalog.

[0092] 1) Catalog registration and review

[0093] The directory registration information submitted by users is reviewed. The review includes ensuring that the field descriptions, data formats, and data sources do not deviate from the actual data; whether the metadata format, naming rules, and classification logic are consistent; and whether key metadata information is missing.

[0094] 2) Directory Change Review

[0095] Users can update the data catalog online and submit it for review. Once approved, the updated catalog will be published.

[0096] 3) Catalog Release

[0097] Once the registration or change of the directory is approved, it will be released to the public in the form of a departmental aggregated directory.

[0098] (4) Catalog Search

[0099] It is used to query directory information by directory name, directory status, level, etc.

[0100] (5) Download the catalog

[0101] Used for downloading from registered directories.

[0102] (6) Directory offline

[0103] Used to take offline directories that are no longer in use or have been suspended.

[0104] (7) Data mounting

[0105] Used to attach aggregated data from the core data area to the corresponding data directory. It associates physically stored datasets (such as files, database tables, and API interfaces) with metadata in the data directory, allowing users to locate and access the actual data through the data directory.

[0106] II. Data Access Management Module

[0107] The data access management module provides data access methods for different data sources and formats (including structured data, semi-structured data, unstructured data, and streaming data). Various external access systems access data to this system in accordance with unified data access standards and specifications. It mainly includes configuration management and access monitoring for streaming data access, front-end exchange, data import, file transfer, and data entry methods.

[0108] (1) Data access configuration

[0109] Streaming data access: By configuring Kafka streaming data processing parameters, including the number of partitions, the number of replicas, and the data retention time, streaming data can be collected, and the consumer programs for streaming data access and the running status of the Kafka cluster can be monitored.

[0110] The system access is achieved by configuring the collection interface parameters, including the access party name, access database name, associated data table, collection type, directory identifier, network environment, etc., to realize automatic data collection and monitor the system access task and service call status.

[0111] Front-end switching: By configuring front-end switching program parameters, including access party name, target database name, database type, database IP, port, database instance, username, password, etc., it can automatically collect front-end database table data and monitor data exchange tasks, task channels, data connection status, etc.

[0112] Data import: Semi-structured data collection is achieved by uploading data files in EXCEL format.

[0113] File transfer: Upload data files in formats such as text, images, and videos to achieve unstructured data collection, and monitor file transfer tasks and FTP server operation status.

[0114] Data entry: By loading directory metadata information, a data entry interface is automatically generated for data entry, enabling data collection that is not supported by the system.

[0115] (2) Data aggregation and monitoring

[0116] It provides data aggregation and execution monitoring functions, including monitoring lists, monitoring retrieval, and monitoring configuration.

[0117] The monitoring list displays the monitoring status of each access system. The monitoring content includes: the total number of access data items, the total number of data items to be accessed, the number of normal items, and the number of abnormal items.

[0118] The monitoring retrieval supports keyword searches by access system name and business information name.

[0119] The monitoring configuration is used to configure the actual business update cycle, the submission cycle, and the advance reminder time for each directory.

[0120] III. Data Ingestion Module

[0121] Data lake module: Provides open data lake storage capabilities, including storage media, storage format, database and table data organization, data read and write, snapshot management, metadata management, table services (asynchronous indexes, data rearrangement, small file merging), etc. Raw data from heterogeneous data sources is uniformly merged into a centralized data lake storage based on object storage architecture.

[0122] (1) Account creation

[0123] In multi-source data integration, stream and batch processing, or data lake architecture, data with different structures (such as field types, names, and levels) are unified into a consistent schema to create a data lake and authorized accounts to access data lake resources.

[0124] (2) Data entry into the lake

[0125] It provides integrated real-time and offline stream / batch jobs for data ingestion into the data lake, while also offering independent data ingestion job configuration. Real-time ingestion jobs support data sources such as database CDC (MySQL binlog, Oracle redo log), message queues (Kafka, RabbitMQ), and IoT device streams (MQTT protocol); offline ingestion jobs support data sources such as files (JSON), full database export, and API batch retrieval.

[0126] (3) Diverse data processing

[0127] It provides transaction capabilities, database operation capabilities, data compression and reordering capabilities, multimodal indexing, and data snapshot capabilities on top of storage. It supports ACID transactions, enabling atomic operations on data lake tables (such as batch insert, update, and delete). It automatically selects compression algorithms based on data type to reduce storage costs. It supports sorting and storing data by specified fields (such as time or user ID) to improve query efficiency. It supports index rebuilding and deletion, as well as recording metadata such as index size and creation time for easy performance analysis. It supports manually triggered or scheduled automatic snapshot creation (e.g., daily creation) to record the database and table status (including data files, metadata, and indexes) at a specific point in time; when data is accidentally deleted or corrupted, it can be rolled back to a specified version through snapshots.

[0128] (4) Data table management

[0129] Data tables are the core organizational unit of a data lake, supporting operations such as creating tables for structured / semi-structured data, modifying fields, and managing partitions, while being compatible with the read and write requirements of mainstream data processing frameworks. It supports defining table structures using DDL statements, or visually configuring table names, fields (names, types, comments), partition fields, and storage formats through the console, automatically generating and executing DDL statements.

[0130] (5) Automatic collection and synchronization of metadata

[0131] When table fields, partitions, or storage formats are modified, the metadata database is automatically updated to ensure that the metadata is consistent with the actual data.

[0132] IV. Data Aggregation Quality Inspection Module

[0133] The data aggregation and quality inspection module provides a custom governance rule configuration module, which is used to aggregate data from the original data area to the core data area and perform data table and data field verification according to the configured data verification rules to improve the data quality of multi-source data fusion.

[0134] (1) Standardized Management of Data Table Names: Standardize the naming of data tables accessed by the system. Supports functions such as displaying data tables by organization / database tree, querying data tables by organization / database, classifying data tables by status, querying data tables by name, querying data tables by database name, querying data tables by the last modifier, data table editing, field information querying, paginated display of data tables, and paginated navigation of data tables. Data table information includes table name, database, affiliated organization, last modifier, and last modification time.

[0135] (2) Data Validation Rule Management: Multiple data validation rules can be configured to validate accessed data. These rules include, but are not limited to, null value validation, format validation, relationship validation, duplicate validation, comparison validation, logical validation, and sensitive data validation. It supports features such as displaying a data validation rule tree, querying rules by rule name, grouping and categorizing validation rules, enabling validation rules, viewing validation rule details, querying rules by validation rule name, querying rules by validation rule code, resetting validation rules, displaying the latest version, displaying historical rule versions, enabling historical rule versions, viewing historical rule version details, comparing historical rule versions, paginated display of validation rules, paginated navigation of validation rules, and selecting historical versions. Data validation rule information includes rule code, rule name, rule description, rule group, rule issue level, most recently modified person, most recently modified time, and latest version.

[0136] (3) Importing data validation rules: Data validation rules can be imported by uploading an Excel file, and Excel templates can be downloaded.

[0137] Table / Field Validation Rule Configuration: This feature allows you to configure table / field validation rules to validate data tables and their fields. It supports functions such as displaying a validation field tree, searching for validation fields by field, adding validation rules to data tables, querying rules by rule name, querying validation fields, resetting validation field queries, adding validation rules to validation fields, and setting validation conditions. Rule configuration information includes the validation field, validation rule, and validation field information.

[0138] (4) Validation task priority configuration: The data table validation task level can be configured to determine the order in which validation data tasks are executed. There are five priority levels in total.

[0139] (5) Verification Task Log: Records the status of data table verification tasks and supports data table verification statistics and queries. It supports viewing verification task log details, overall task volume statistics, verified data volume statistics, and problematic data volume statistics. Verification task log information includes status, error reason, verified data volume, problematic data volume, unverified data volume, task priority, start time, and end time.

[0140] V. Data Reconciliation Module

[0141] The data reconciliation module is used to verify and monitor the consistency of aggregated data.

[0142] By selecting a reconciliation date, the aggregated data number is compared with the data access number. When the successfully aggregated data number differs from the data access number, the unique identifier (rowkey) is then verified. The unique identifier that does not exist in the verification list within the verification date is found and returned to the data access party, who then resends the missing data.

[0143] (1) Reconciliation time configuration

[0144] Configurable data reconciliation cycle time.

[0145] (2) Billing Management

[0146] Manage the basic information elements of the reconciliation list, including transaction number, reconciliation cycle, reconciliation method, reconciliation date, business item name, business information name, data item, data volume, etc.

[0147] (3) Configuration of reconciliation rules

[0148] Configure data comparison rules for both parties involved in the reconciliation, and determine whether there is any missing data based on the returned comparison results.

[0149] (4) Data re-push

[0150] For data that fails to be pushed or is missing, re-push the data table. You can choose to re-push the data table to which the access provider belongs, or you can re-push it by transaction number.

[0151] (5) Maintenance of reconciliation results

[0152] It allows users to check reconciliation results and automatically sends the results back to both parties involved in the reconciliation.

[0153] refer to Figure 2-4 Data access and reconciliation mode:

[0154] (I) Read the business database to reconcile data.

[0155] The data access party assigns business database permissions to the system (at least query permissions). Custom reconciliation rules are established, and the data volume in the access party's business database tables and the access database tables are compared according to these rules.

[0156] (II) Manual reconciliation

[0157] When a reconciliation work order is initiated, the aggregation platform's operations and maintenance team negotiates and determines the statistical scope with the data access party. The data access party calculates the amount of incremental data pushed to the business database the previous day according to the statistical scope and delivers it to the aggregation platform's operations and maintenance team. The aggregation platform's operations and maintenance team calculates the amount of incremental data received in the platform's database the previous day according to the statistical scope and compares it with the amount of data provided by the access party. If there is a discrepancy, the data access party provides the previous day's push logs (including at least the primary keys of the pushed data), and the operations and maintenance team analyzes and provides feedback.

[0158] (III) Real-time data push and reconciliation

[0159] Modify the data push interface so that before pushing the actual data, the total amount of data to be accessed is sent to the platform. After the platform receives the data, it pushes the actual data. Before the data is stored in the database, the number of received records is counted and compared with the total amount of data to be accessed. If they are inconsistent, the access of this batch of data will fail.

[0160] VI. Data Storage Management Module

[0161] The data storage management module adopts a new integrated lake warehouse data management architecture to manage the aggregated data by partitioning it according to business attributes and data characteristics.

[0162] Intelligent lake warehouses use a distributed cluster of servers or nodes to collect, store, and manage data from different sources and of different types. Partitioned spaces can be designed based on data storage type and usage scenarios to optimize storage efficiency. Computational resources, data resources, and component resources are logically isolated for each user. Multi-level access control is supported within each tenant's premises.

[0163] Users initially set up structured data areas (by business category, confidentiality level, etc.), unstructured areas (image data), and metadata areas (management metadata, technical metadata, business metadata) according to different data storage types. Based on the data access and storage usage scenario, a raw data area is created on the basis of the data lake to store the raw data accessed from various sources. This is the logical storage location for unprocessed and unintegrated business data. It can support unified control of raw data through centralized virtual management and provide a data source for upper-layer data value-added operations.

[0164] Based on the data aggregation and storage usage scenarios, a core data area and a shared database are created based on the data warehouse. The core data area extracts and integrates data from multiple sources, serving as the logical storage location for data aggregation processing. By establishing thematic or topical libraries, aggregated data is categorized, stored, and managed, and easy-to-use integration interfaces are provided for access needs of other systems. The shared data area is a separate area for storing shared data, logically isolated from the actual data area, and stores copies of all data resources that need to be shared.

[0165] This invention also provides a multi-source data fusion management method based on an intelligent lake warehouse architecture, applied to the system, referenced. Figure 6 The method includes the following steps:

[0166] Step 1. In the backend management module, the user initializes and creates partitions such as the raw data area, core data area, and shared data area, and configures the corresponding data storage path for each data partition, including data storage systems such as object storage, relational databases, and document databases. Successful configuration generates a data partition directory and ledger list.

[0167] Step 2. Backend management users configure data aggregation verification rules in the data aggregation quality inspection module;

[0168] Step 3. The backend management user creates a raw data area data entry account for the data access user in the data entry module;

[0169] Step 4. Data access users register the data resources they need to access in the data catalog management module, fill in the data catalog information, and submit it for backend review;

[0170] Step 5. Backend management users review the data directory information submitted by the data access user in the data directory management module and fill in the review comments. Once approved, the data is published for use in the form of an aggregated data directory.

[0171] Step 6. Data access users select a data access method in the data access management module, configure data access parameters, and implement data connection according to the corresponding access technical specifications;

[0172] Step 7. Data access: Users configure the parameters for data aggregation and lake entry in the data lake module, and store the accessed data in the data lake;

[0173] Step 8. When users in the backend management module use the data warehouse, the system will automatically extract and load the data resources into the core data area or shared data area, and perform data table and data field validation on the data entering the lake according to the configured data validation rules. If the validation is successful, the registration information will be automatically added to the aggregated data directory.

[0174] Step 9. The data reconciliation module periodically performs end-to-end reconciliation of data access, data entering the data lake, and data aggregation. In case of data anomalies, the data sender will push the data back to the system.

[0175] This invention provides a multi-source data fusion system and method based on an intelligent lakehouse architecture, integrating multiple data access methods to meet the needs of different production environment systems and different data formats accessing a unified system. It standardizes data access from the source by configuring data quality inspection rules to verify the quality of aggregated data. For systems with problematic data, access must be re-established only after meeting the system's access specifications. Storage is tiered according to data quality levels, including raw data areas, core data areas, and shared data areas. This addresses data quality issues during the data aggregation process and the storage and computational pressure brought about by the growth in the scale of multi-source data, ensuring the timeliness, standardization, completeness, and reliability of the integrated data. It achieves end-to-end data fusion management of multi-source data from registration, access, processing, and storage.

Claims

1. A multi-source data fusion system based on an intelligent lakehouse architecture, characterized in that, It includes a data catalog management module, a data access management module, a data lake module, a data aggregation and quality inspection module, a data reconciliation module, and a data storage management module; The data directory management module includes metadata management, directory registration, directory review, directory retrieval, directory download, directory decommissioning, and data mounting; The data access management module includes data access configuration and data aggregation monitoring; The data import module includes account creation, data import operation, diverse data processing, data table management, and automatic collection and synchronization of metadata; The data aggregation and quality inspection module includes standardized management of data table names, management of data verification rules, import of data verification rules, configuration of table / field verification rules, configuration of verification task priority, and verification task logs. The data reconciliation module includes reconciliation time configuration, reconciliation statement management, reconciliation rule configuration, data re-push, and reconciliation result maintenance; The data storage management module includes a data management architecture that integrates lake warehouses, used to manage the aggregated data by partitioning it according to business attributes and data characteristics; The data access management module provides data access methods for different data sources and formats, and each external access system accesses data into this system in accordance with a unified data access standard and specification. This includes configuration management and access monitoring for streaming data access, front-end switching, data import, file transfer, and data entry methods; The data aggregation quality inspection module provides a custom governance rule configuration module, which is used to aggregate the data accessed from the original data area to the core data area and perform data table and data field verification according to the configured data verification rules. The data reconciliation module is used to verify and monitor the consistency of the aggregated data; By selecting a reconciliation date, the data aggregation count is compared with the data access count. When the data aggregation count is inconsistent with the data access count, the unique identifier rowkey is checked again. The unique identifier that does not exist in the reconciliation list within the reconciliation date is found and returned to the data access party, which then resends the missing data.

2. The multi-source data fusion system based on an intelligent lakehouse architecture according to claim 1, characterized in that, The metadata management module in the data catalog management module is used for the registration and management of the interpretation and definition of the accessed data. After the system data is accessed, it automatically analyzes and obtains business metadata information, and creates a global-level metadata file to interpret and define the data, forming a consistent and unified data definition. For unstructured data or files without built-in metadata, the metadata information is manually sorted out and entered into the metadata management system. Metadata management includes metadata maintenance, metadata generation, metadata retrieval, and metadata querying; The directory registration is a structured representation of metadata; The directory registration is used for the registration and management of system information, database information, basic data information, and data item information; The directory review is used to review the compilation and publication of the directory and to fill in comments, including the directory registration review and the directory change review. Once the review is passed, the directory is published and used in the form of a collection of directories. The directory retrieval method allows users to search for directory information by directory name, directory status, and hierarchy. The directory download is used to download registered directories; The directory decommissioning is used to decommission directories that are no longer in use or have been suspended from use; The data mounting is used to mount the aggregated data of the core data area to the corresponding data directory; to associate the physically stored dataset with the metadata in the data directory, and to locate and access the actual data through the data directory.

3. The multi-source data fusion system based on an intelligent lakehouse architecture according to claim 1, characterized in that, The data ingestion module provides open data lake storage capabilities, including storage media, storage format, database and table data organization methods, data read and write, snapshot management, metadata management, and table services; it is used to uniformly merge raw data from heterogeneous data sources into a centralized data lake storage based on object storage architecture.

4. A multi-source data fusion system based on an intelligent lakehouse architecture according to claim 1, characterized in that, The data aggregation and quality inspection module specifically includes: 1) Standardized Management of Data Table Names: Standardizes the naming of data tables accessed by the system; supports displaying data tables by organization / database tree, querying data tables by organization / database, classifying data tables by status, querying data tables by name, querying data tables by database name, querying data tables by the most recent modifier, data table editing, field information querying, paginated display of data tables, and paginated navigation of data tables; data table information includes table name, database, affiliated organization, most recent modifier, and most recent modification time; 2) Data Validation Rule Management: Configure various data validation rules to validate accessed data. Data validation rules include null value validation, format validation, relationship validation, duplicate validation, comparison validation, logical validation, and sensitive data validation. Supports features such as data validation rule tree display, rule query by rule name, validation rule grouping and categorization, enabling validation rules, viewing validation rule details, querying rules by validation rule name, querying rules by validation rule code, resetting validation rules, displaying the latest version, displaying historical rule versions, enabling historical rule versions, viewing historical rule version details, comparing historical rule versions, paginated display of validation rules, paginated navigation of validation rules, and historical version selection. Data validation rule information includes rule code, rule name, rule description, rule grouping, rule issue level, most recently modified person, most recently modified time, and the latest version. 3) Importing data validation rules: Data validation rules can be imported by uploading an Excel file, and Excel templates can be downloaded. 4) Table / Field Validation Rule Configuration: Configure table / field validation rules to validate the data tables and fields in the access data tables; supports functions such as displaying the validation field tree, searching for validation fields by field, adding validation rules to data tables, querying rules by rule name, querying validation fields, resetting validation field queries, adding validation rules to validation fields, and setting validation conditions; rule configuration information includes validation fields, validation rules, and validation field information; 5) Validation task priority configuration: Configure the data table validation task level to determine the order in which validation tasks are executed; there are five priority levels in total. 6) Verification Task Log: Records the status of data table verification tasks, supports data table verification statistics and queries; supports viewing verification task log details, overall task volume statistics, verified data volume statistics, and problematic data volume statistics; verification task log information includes status, error reason, verified data volume, problematic data volume, unverified data volume, task priority, start time, and end time.

5. A multi-source data fusion system based on an intelligent lakehouse architecture according to claim 1, characterized in that, The data reconciliation module includes the following reconciliation modes: (1) Read the business database to reconcile data; The data access party allocates business database permissions to the system; customizes reconciliation rules, and compares the data volume of the data access party's business database table and the access database table according to the rules. (2) Manual reconciliation; When a reconciliation work order is initiated, the aggregation platform's operations and maintenance team negotiates and determines the statistical criteria with the data access party. The data access party calculates the incremental data volume pushed to the business database the previous day according to the statistical criteria and delivers it to the aggregation platform's operations and maintenance team. The aggregation platform's operations and maintenance team calculates the incremental data volume received in the platform's database the previous day according to the statistical criteria and compares it with the data volume provided by the access party. If there is a discrepancy, the data access party provides the push logs from the previous day, and the operations and maintenance team analyzes and provides feedback. (3) Real-time data push and reconciliation; Modify the data push interface so that before pushing the actual data, the total amount of data to be accessed is sent to the platform. After the platform receives the data, it pushes the actual data. Before the data is stored in the database, the number of received records is counted and compared with the total amount of data to be accessed. If they are inconsistent, the access of this batch of data will fail.

6. A multi-source data fusion system based on an intelligent lakehouse architecture according to claim 1, characterized in that, In the data storage management module, the smart lake warehouse uses a set of distributed cluster servers or nodes to collect, store and manage data from different sources and of different types. Design partition spaces based on the data storage type and usage scenario; Each user's computing resources, data resources, and component resources are logically isolated; The tenant's internal access control supports multi-level permission management; Users initially set up structured data areas according to different data storage types; based on the data access and storage usage scenarios, raw data areas are created on the basis of the data lake to store raw data from various sources. This is the logical storage location for unprocessed and unintegrated business data. It supports unified control of raw data through centralized management in a virtual form and provides a data source for upper-layer data value-added operations. Based on the data aggregation and storage usage scenario, a core data area and a shared database are created based on the data warehouse; the core data area extracts and integrates multi-source access data, which serves as the logical storage location for data aggregation and processing. By establishing specialized or thematic libraries, aggregated data is categorized, stored, managed, and provided with easy-to-use integration interfaces to meet the access needs of other systems; the shared data area is a separate area for storing shared data, logically isolated from the actual data area, and stores copies of all data resources that need to be shared.

7. A multi-source data fusion method based on an intelligent lakehouse architecture, characterized in that, A multi-source data fusion system based on an intelligent lakehouse architecture, as described in any one of claims 1-6, includes the following steps: Step 1: The backend management user initializes and creates the raw data area, core data area, and shared data area in the data storage management module, and configures the corresponding data storage path for each data area, including object storage, relational database, and document database; after successful configuration, a data area directory and ledger list are generated. Step 2: Backend management users configure data aggregation verification rules in the data aggregation quality inspection module; Step 3: Backend management users create original data area data entry accounts for data access users in the data entry module; Step 4: Data access users register the data resources they need to access in the data catalog management module, fill in the data catalog information, and submit it for backend review; Step 5: Backend management users review the data directory information submitted by the data access user in the data directory management module and fill in the review comments. Once approved, the data is published for use in the form of an aggregated data directory. Step 6: Data access users select a data access method in the data access management module, configure data access parameters, and implement data connection according to the corresponding access technical specifications; Step 7: Data access users configure the parameters for data aggregation and lake entry in the data lake module, and store the accessed data into the data lake; Step 8: When the backend management user uses the data warehouse in the data storage management module, the data resources loaded into the lake will be automatically extracted and loaded into the core data area or shared data area, and the data tables and data fields of the data loaded into the lake will be verified according to the configured data verification rules; if the verification is successful, the registration information will be automatically added to the aggregated data directory. Step 9: The data reconciliation module regularly performs end-to-end reconciliation of data access data, data entering the lake data, and data aggregation data; in case of data anomalies, the data sender will push the data back to the system.

Citation Information

Patent Citations

  • Multi-source heterogeneous data-oriented Hudi data uptake method and system

    CN118503229A

  • Petroleum data governance method based on business resource model

    CN118643030A