Lake and cabin integration implementation method and device based on dynamic change table structure
By using MinIO, Hudi and Flink CDC in the Hucang integrated architecture, the table structure is updated dynamically, which solves the problems of complex data flow and the library table structure in the existing technology that cannot be changed dynamically, and achieves more efficient data management and utilization.
Patent Information
- Application Number
- CN202510311064.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-17
- Publication Date
- 2025-06-13
AI Technical Summary
The existing Hucang integrated architecture has complexity and high maintenance costs in data flow and library table structure management, and cannot dynamically change the library table structure.
MinIO is used as storage, Hudi is used as the data lake framework, data synchronization is used using Flink CDC, and class dynamic update table structure is implemented through the Flink DataStream API and custom DebeziumDeserializationSchema.
It realizes the reduction of data flow times, reduces system complexity and maintenance costs, supports dynamic changes in the library table structure, and improves the management and utilization efficiency of data in the lake warehouse.
Smart Images

Figure CN120144680A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of big data, and specifically provides a method and device for realizing lakehouse integration based on dynamically changing table structures. Background Art
[0002] With the popularization of the Internet and mobile devices, the amount of data has grown exponentially, and most industries have entered the big data era. In order to meet both offline and real-time computing requirements, current big data processing systems are gradually adopting the lakehouse integration architecture. However, the current lakehouse integration architecture has the following problems:
[0003] (1) Most use Kafka as middleware. Data needs to be synchronized to Kafka first and then from Kafka to the data lake. Although this method decouples the association between data collection and the data warehouse, the introduction of Kafka increases the number of data transfers and the complexity of the system.
[0004] (2) Since the table structure in the original data layer of the data warehouse is determined before data synchronization, when the table structure of the data source changes, the data warehouse synchronization task needs to be stopped first. After modifying the table structure in the original data layer, the synchronization task needs to be restarted. This way of being unable to dynamically change the table structure of the database obviously increases the difficulty and cost of data warehouse maintenance. Summary of the Invention
[0005] The present invention aims at the above-mentioned deficiencies of the prior art and provides a practical method for realizing lakehouse integration based on dynamically changing table structures.
[0006] The further technical task of the present invention is to provide a device for realizing lakehouse integration based on dynamically changing table structures that is reasonably designed, safe and applicable.
[0007] The technical solution adopted by the present invention to solve its technical problems is as follows:
[0008] A method for realizing lakehouse integration based on dynamically changing table structures uses MinIO as storage and Hudi as a data lake framework, and uses Flink CDC to store the original data in the Hudi table according to the original table structure, including the characteristics of building a warehouse on the lake, dynamically changing the table structure, and metadata management. It directly uses Flink CDC for data synchronization and interfaces with the data calculation layer at the code level;
[0009] When storing the original data in the Hudi table, the Hudi table corresponds to the original data layer in the data warehouse. When the structure of the business data table changes, the corresponding table structure in the ODS layer will be automatically changed during the data synchronization process;
[0010] Develop data synchronization using the Flink DataStream API. Obtain changes to the table structure by customizing the implementation class of DebeziumDeserializationSchema, and then modify the corresponding table structure in the ODS layer.
[0011] Furthermore, when modifying the corresponding table structure in the ODS layer, the following steps are involved:
[0012] (1) Implement the getProducedType method, and use a quadruple for the returned type information;
[0013] (2) Implement the deserialize method, and implement the logic for converting data in the deserialize method;
[0014] (3) When synchronizing data to the lakehouse, determine whether there are table structure change operations.
[0015] Furthermore, in step (1), the first element describes the table structure change situation: N indicates no change, D indicates deleting a field, A indicates adding a field, and M indicates modifying a field;
[0016] The second element describes the data operation: A indicates adding, U indicates modifying, and D indicates deleting;
[0017] The third element uses a JSON string to describe the field information that has changed, and the fourth element also describes the data content read in the form of a JSON string, including the field name and the corresponding value.
[0018] Furthermore, in step (2), it includes obtaining the field change information and data content from the SourceRecord, converting them into the quadruple defined in getProducedType, and finally collecting them by the Collector.
[0019] Furthermore, in step (3), when synchronizing data to the lakehouse, determine whether there are table structure change operations. If so, update the table structure; otherwise, synchronize the changed data.
[0020] Furthermore, in the metadata management, metadata provides basic information about the data, and metadata management catalogs and indexes all the data in the data lake correctly to enable its location and use.
[0021] Furthermore, relying on the metadata management platform, establish a metadata collection task to regularly extract the latest metadata information from the data lake, store or update it in MySQL, and display it in the form of a Web UI.
[0022] A lakehouse integration implementation device based on a dynamically changing table structure, comprising: at least one memory and at least one processor;
[0023] The at least one memory is used for storing machine-readable programs;
[0024] The at least one processor is used for calling the machine-readable programs and executing a lakehouse integration implementation method based on a dynamically changing table structure.
[0025] Compared with the prior art, the lakehouse integration implementation method and device based on a dynamically changing table structure of the present invention have the following prominent beneficial effects:
[0026] The present invention solves the problems of a large number of data transfer times, high maintenance costs, and inability to dynamically change the database table structure existing in the current lakehouse integration system, and improves the management and utilization of data in the lakehouse by regularly extracting metadata information, effectively reducing the possibility of the data lake degenerating into a data swamp. BRIEF DESCRIPTION OF THE DRAWINGS
[0027] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0028] Att Figure 1 is a flowchart of a lakehouse integration implementation method based on a dynamically changing table structure. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0029] In order to enable those skilled in the art to better understand the solutions of the present invention, the following will further elaborate on the present invention in conjunction with specific embodiments. Obviously, the described embodiments are only some embodiments of the present invention, rather than all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the scope of protection of the present invention.
[0030] The following gives a best embodiment:
[0031] Such as Figure 1As shown in the figure, in the lakehouse integration implementation method based on the dynamically changing table structure in this embodiment, the data lake technology Hudi and the real-time data synchronization technology Flink CDC are introduced. The business data is directly synchronized to the data lake in real time, and a data warehouse is built on the lake, so as to realize that one architecture can meet both offline and real-time processing requirements. Based on the real-time data synchronization of Flink CDC, the acquisition of the library table structure is extended. Once the table structure change is detected, the corresponding table structure in the data lake is dynamically updated. Considering the limitations of HDFS in terms of storage, operation and maintenance costs, and scalability, MinIO object storage compatible with the S3 protocol is used as the storage medium of the data lake.
[0032] Lakehouse integration architecture design. In the present invention, MinIO is used as the storage, Hudi is used as the data lake framework, and FlinkCDC is used to store the original data in the Hudi table according to the original table structure, including the characteristics of building a warehouse on the lake, dynamically changing the table structure, and metadata management. The architecture is as Figure 1 shown. Using Flink CDC for direct data synchronization can reduce the data transmission link, and at the same time seamlessly connect to the data calculation layer at the code level, reducing the R & D cost. Different from the offline data warehouse that uses Hive tables to build each layer of the data warehouse, Hudi tables are used to implement the layering of the data warehouse to realize building a warehouse on the lake.
[0033] Dynamically changing the table structure. For the business system, due to the change of requirements, the library table structure may also need to be changed accordingly. If the code needs to be modified and the Flink task needs to be restarted every time there is a change, it will bring a lot of unnecessary maintenance costs. Therefore, it is necessary for the library table structure to be able to change dynamically.
[0034] When storing the original data in the Hudi table, the corresponding Hudi table is the original data layer (Operation Data Store, ODS) in the data warehouse. Therefore, when the business data table structure changes, the corresponding table structure in the ODS layer will be automatically changed during the data synchronization process.
[0035] Since the process of Flink CDC reading the binlog and synchronizing the data to the target table obtains the table structure information of the target table from Flink SQL, even if there is table structure change information in the binlog data, the old table structure is still used by Flink CDC, resulting in the inability to dynamically update the table structure information of the target table. Therefore, in this application, Flink DataStream API is used for the development of data synchronization, and the table structure change is obtained by customizing the DebeziumDeserializationSchema implementation class, and then the corresponding table structure in the ODS layer is modified.
[0036] The implementation steps are as follows:
[0037] (1) Implement the getProducedType method. The returned type information uses a quadruple. The first element describes the change situation of the table structure: N means no change, D means deleting a field, A means adding a field, and M means modifying a field.
[0038] The second element describes the data operation: A means adding, U means modifying, and D means deleting.
[0039] The third element uses a JSON string to describe the field information that has changed.
[0040] The fourth element also describes the data content read in the form of a JSON string, including the field name and the corresponding value.
[0041] (2) Implement the deserialize method. In this method, implement the logic for converting data, including obtaining the field change information and data content from the SourceRecord, and converting them into the quadruple defined in getProducedType, and finally collected by the Collector.
[0042] (3) When synchronizing data to the lakehouse, judge whether there is an operation to change the table structure. If so, update the table structure; otherwise, synchronize the changed data.
[0043] Metadata management. Metadata provides basic information about data, such as the source, structure, content, and format of the data. Metadata management ensures that all data in the data lake is correctly cataloged and indexed, making it easier to locate and use. Without reasonable metadata management, it is almost impossible to effectively understand and use data, leading to a data swamp. Therefore, metadata management is indispensable in the construction of lakehouse integration.
[0044] Taking the data warehouse layering as the dimension, manage the metadata information of each layer by adopting a hierarchical meta-model of data catalog, database, table, and field. By establishing a metadata collection task, extract the latest metadata information from the data lake regularly, store or update it in MySQL, and display it in the form of a Web UI.
[0045] This application uses the data lake framework Hudi and the object storage engine Minio as the data management and storage foundation for the lakehouse. It uses Flink CDC to synchronize business data from multiple heterogeneous data sources into the data lake as the ODS layer, and obtains the change information of the table structure by customizing the DebeziumDeserializationSchema to synchronize and modify the table structure of the ODS layer. For the data layering DWD and DWS in the lakehouse, Flink is used for stream-batch integrated computing, and finally the processed data is provided for data services to use. Relying on the metadata management platform, the metadata information in the lakehouse is updated regularly.
[0046] Using MinIO as the storage medium, integrating Hudi in the Flink application, using Flink as the computing engine for data synchronization and stream-batch processing, and relying on the metadata extraction scheduled task, the table and library metadata in the lakehouse are extracted into MySQL and visualized in the form of a Web UI. For the dynamically changing table structure during data synchronization, taking the synchronization of user basic information table data as an example, when a new address field "address" is added to the table, the deserialize method will convert the change information into a quadruple ("A","",[{"columnName":"address","type":"varchar","length":255}],{}). When synchronizing data into the lakehouse, the DDL statement for adding fields is executed. If data is added after the new field is added, a corresponding quadruple example is ("N","A",[],{"id":123,"address":"Beijing"}), and at this time, a new piece of data is added without the need to update the table structure.
[0047] Based on the above method, the lakehouse integration implementation device based on the dynamically changing table structure in this embodiment includes: at least one memory and at least one processor;
[0048] The at least one memory is used to store machine-readable programs;
[0049] The at least one processor is used to call the machine-readable program and execute the lakehouse integration implementation method based on the dynamically changing table structure.
[0050] The above specific implementation manners are only specific cases of the present invention. The patent protection scope of the present invention includes but is not limited to the above specific implementation manners. Any technical solution that conforms to the above specific implementation manners of the present invention and any appropriate changes or substitutions made by those of ordinary skill in the art shall fall within the patent protection scope of the present invention.
[0051] Although embodiments of the present invention have been shown and described, it will be understood by those of ordinary skill in the art that various changes, modifications, substitutions and variations can be made to these embodiments without departing from the principles and spirit of the present invention, and the scope of the present invention is defined by the appended claims and their equivalents.
Claims
1. A lake-warehouse integration implementation method based on a dynamic change table structure is characterized by: MinIO is used as storage, Hudi is used as the data lake framework, and Flink CDC is used to store the original data in the Hudi table according to the original table structure. It includes the features of building warehouses on the lake, dynamically changing table structures, and metadata management. Flink CDC is used to directly synchronize data and connect to the data computing layer at the code level. When the original data is stored in the Hudi table, the Hudi table corresponds to the original data layer in the data warehouse. When the business data table structure changes, the corresponding table structure in the ODS layer will be automatically changed during the data synchronization process; Use Flink DataStream API for data synchronization development, obtain table structure changes by customizing the DebeziumDeserializationSchema implementation class, and then modify the corresponding table structure in the ODS layer.
2. The lake-warehouse integration implementation method based on the dynamic change table structure according to claim 1 is characterized in that: When modifying the corresponding table structure in the ODS layer, the following steps are required: (1) Implement the getProducedType method and return the type information in a four-tuple format; (2) Implementing a deserialize method, in which logic for converting data is implemented; (3) When synchronizing data to the lake warehouse, determine whether there is any table structure change operation.
3. The lake-warehouse integration implementation method based on the dynamic change table structure according to claim 2 is characterized in that: In step (1), the first element describes the changes in the table structure: N indicates no change, D indicates a deleted field, A indicates a newly added field, and M indicates a modified field; The second element describes the data operation: A for adding, U for modifying, and D for deleting; The third element uses a JSON string to describe the changed field information. The fourth element also describes the read data content in the form of a JSON string, including the field name and the corresponding value.
4. The lake-warehouse integration implementation method based on the dynamic change table structure according to claim 3 is characterized in that: In step (2), field change information and data content are obtained from SourceRecord, converted into the four-tuple defined in getProducedType, and finally collected by Collector.
5. The lake-warehouse integration implementation method based on the dynamic change table structure according to claim 4 is characterized in that: In step (3), when synchronizing data to the lake warehouse, determine whether there is a table structure change operation. If so, update the table structure; otherwise, synchronize the changed data.
6. The lake-warehouse integration implementation method based on the dynamic change table structure according to claim 5 is characterized in that: In the metadata management, metadata provides basic information about the data. Metadata management correctly catalogs and indexes all data in the data lake to enable it to be located and used.
7. The lake-warehouse integration implementation method based on the dynamic change table structure according to claim 6 is characterized in that: Relying on the metadata management platform, a metadata collection task is established to extract the latest metadata information from the data lake on a regular basis, store or update it in MySQL, and display it in the form of a Web UI.
8. A lake-warehouse integration implementation device based on a dynamic change table structure, characterized in that: include: at least one memory and at least one processor; The at least one memory is used to store a machine-readable program; The at least one processor is configured to call the machine-readable program to execute the method according to any one of claims 1 to 7.