A Hybrid Distributed Data Lakehouse System and Its Operation Method

By designing a hybrid distributed data lake warehouse system, the contradiction between data decentralization and distributed storage in the big data platform is solved, and efficient data processing and system stability are achieved.

CN116842105BActive Publication Date: 2025-06-17YUNNAN QUANYAN TECH INFORMATION CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310819463.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-07-06
Publication Date
2025-06-17
Estimated Expiration
2043-07-06

AI Technical Summary

Technical Problem

In the big data platform, data decentralization under the microservice architecture and data centralization under distributed storage are contradictory, resulting in inefficient data management and processing.

Method used

Design a hybrid distributed data lake warehouse system, including the underlying data lake warehouse, data engine, semantic interpretation engine, multi-source data target distribution module and data model engine, to realize multi-source heterogeneous management and efficient processing of data.

Benefits of technology

It solves the contradiction between data decentralization and distributed storage in the big data platform, and improves the efficiency of data processing and the high availability and stability of the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116842105B_ABST
    Figure CN116842105B_ABST
Patent Text Reader

Abstract

The present invention discloses a hybrid distributed data lakehouse system and its operation method, including: the underlying data lakehouse supports coexistence of multiple data types and realizes mutual sharing of data. The parameters input by the data engine obtain or operate on the data in the data lakehouse. The data engine is responsible for organizing the original data and cooperating with the data product engine and the semantic interpretation engine. The data product engine is responsible for secondary processing, computing and analyzing the data in the data lakehouse, and finally providing data support for upper-layer applications. The semantic interpretation engine stores the standardized templates and parses them into data warehouse execution commands. The data model engine is for user-oriented visual operations. The multi-source data target distribution module is the middleware in the hybrid distributed data lakehouse system. The advantages of the present invention are: it can process various data forms, can seamlessly call a large amount of metadata, perform on-demand model operations with high efficiency and low consumption, and improve the high availability and stability of the system.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of data processing, and particularly relates to a hybrid distributed data lake warehouse system and an operation method thereof. Background Art

[0002] The heterogeneous data architecture is a data processing architecture that allows enterprises to integrate data from different sources and in different formats for better data analysis and decision-making. In this architecture, data is usually stored in different databases or data warehouses and uses different data structures and standards.

[0003] In the era of big data, the demand for heterogeneous data architectures is very prominent, and massive amounts of data of different types and from different data sources need to be centrally collected and managed.

[0004] The lakehouse technology is an emerging technology system that emerged in the era of big data. The data lakehouse integrates two systems: the data warehouse and the data lake:

[0005] Data warehouse: suitable for storing structured, high-information-density, processed data. For example, the association information, portrait information, etc. obtained through big data analysis can be stored in the data warehouse;

[0006] Data lake: suitable for storing unstructured, low-information-density, un-cleaned data. For example, the log information, long text information, IoT raw message information, etc. obtained in production can be directly stored in the data lake.

[0007] The advantage of the lakehouse technology is that it inherits the flexibility and scalability of the data lake and also has the characteristics of data quality and security of the data warehouse. This architecture can increase the speed of data processing to a level comparable to that of the data warehouse, and at the same time can process various types of data, including structured, semi-structured, and unstructured data. In addition, the lakehouse technology also supports emerging data processing technologies such as real-time data processing and machine learning.

[0008] However, in the microservices architecture of the big data platform, there is a contradiction between data decentralization and data centralization in distributed storage.

[0009] Abbreviation and Definition of Key Terms

[0010] DBMS: Database Management System, database management system;

[0011] DLW: Data Lake Warehouse, data lake warehouse;

[0012] DM:Data Model,data model;

[0013] BI:Business Intelligence, business intelligence;

[0014] ML: Machine Learning;

[0015] NLP: Natural Language Processing, natural language processing;

[0016] JDBC: Java Data Base Connectivityjava, database connector;

[0017] SDK: Software Development Kit, software development kit;

[0018] API: Application Programming Interface, application programming interface;

[0019] XML: Extensible Markup Language, extensible markup language;

[0020] JSON: JavaScript Object Notation JavaScript, object notation. Summary of the invention

[0021] In view of the defects of the prior art, the present invention provides a hybrid distributed data lake warehouse system and an operation method thereof.

[0022] In order to achieve the above invention object, the technical solution adopted by the present invention is as follows:

[0023] A hybrid distributed data lake warehouse system, including: an underlying data lake warehouse, a data engine, a semantic interpretation engine, a multi-source data target distribution module and a data model engine;

[0024] Bottom-layer data lake warehouse: includes data warehouse and data lake. The bottom layer supports the coexistence of multiple data types and enables mutual sharing of data.

[0025] The data warehouse includes MySQL and Redis, and the data lake includes: HDFS, ClickHouse, ElastiSearch, and HBase.

[0026] The data engine consists of a multi-source data management engine, a data product engine, and a data calculation engine.

[0027] The multi-source data management engine is responsible for parsing the input parameters, which include SQL, connection method, database code, etc., and directly interacts with the underlying data lake warehouse using native JDBC or native connectors to obtain or operate on the data in the data lake warehouse.

[0028] The data calculation engine integrates the basic big data calculation functions and is responsible for two system operations: ① It is responsible for sorting out the raw data in the data lake warehouse and delivering it to the data product engine for processing. ② It is responsible for cooperating with the data product engine to provide computing power support for it.

[0029] The data calculation engine integrates the following: BI module, ML module, NLP module, and ScriptLanguage module;

[0030] The data product engine is responsible for secondary processing, computing, and analyzing the data in the data lake warehouse, and finally providing data support for the upper-layer applications.

[0031] The data product engine integrates a database configuration module, a database driver module, DBSession, and a database connection pool;

[0032] The database configuration module stores the database configurations configured by users in the data model engine. When it is necessary to interact with the database, the corresponding database configurations will be obtained from this module first. The configurations include database connection information and database driver information.

[0033] The function of the database driver module is to allocate different driver versions for each database request for connecting to the database. The database driver information retrieved from the database configuration module will be passed into this module, and this module will then call DBSession to establish a connection to the corresponding database.

[0034] DBSession is used to create a connection and session between the program and the database and manage the database transactions therein. All database requests need to be managed through DBSession.

[0035] The database connection pool is responsible for allocating, managing, and releasing database connections. It allows the application to reuse an existing database connection instead of establishing a new one; it releases database connections whose idle time exceeds the maximum idle time to avoid database connection leaks caused by not releasing database connections.

[0036] The semantic interpretation engine, whose function is to parse the standardized templates of class SQL templates, XML templates, and JSON parameter templates stored in the data model engine into SQL statements for the data warehouse, operations on the data lake, or execution commands.

[0037] The data model engine is for users' visual operations. The data model engine provides a web - side entry to facilitate users in managing database information and SQL templates. It supports a visual editing window where users can write SQL statements or configure SQL templates according to standard SQL specifications or template configuration specifications.

[0038] The multi - source data target distribution module is a middleware in the hybrid distributed data lake - warehouse system, playing a connecting role between the upper and lower layers, and is used to connect the underlying data lake - warehouse, data engine, semantic interpretation engine, and data model engine.

[0039] Furthermore, the semantic interpretation engine integrates: an XML parser, a JSON parser, a Script parser, a DBSession reflector, a dynamic parameter interpreter, a class - SQL mapper, a semantic node manager, a Hybrid hybrid queryer, a semantic encapsulator, and a semantic executor;

[0040] The XML parser is used to convert XML into an XML DOM object;

[0041] The JSON parser is used to parse JSON - formatted data according to the semantic template.

[0042] The Script parser is used to parse the script language into executable statements.

[0043] The DBSession reflector is used to create a session between the program and the database.

[0044] The dynamic parameter interpreter is used to parse dynamic parameters according to the pre - edited parameter template.

[0045] The class - SQL mapper is used to parse and convert general SQL language according to the specific database characteristics.

[0046] The role of the semantic node manager is to cooperate with various parsers to manage the nodes of the parsed SQL statements. The management content includes query conditions, sorting conditions, linking conditions, etc. in the SQL statements.

[0047] The Hybrid hybrid queryer is used for the hybrid query of multiple requests in a unified transaction, enabling the H2DLW service to support multiple database connections in one transaction. Specifically, it combines multiple different - type database requests into one transaction, and when performing transaction rollback, it can roll back different databases at once.

[0048] The semantic encapsulator is used to generate a semantic matching template from the pre - configured and stored data semantic template for quick matching when various parsers are parsing.

[0049] The semantic executor is used to execute the parsed data semantic template, such as a properly parsed XML template, and execute the SQL statement.

[0050] Furthermore, in the data model engine, users can configure the database information they want to connect to in the database operation interface, including: database connection address, database type, and account password information for logging in to the database. Among them, the data source management module stores information about the database types that the H2DLW service can connect to. These information are not allowed to be configured by users, and users can only select the corresponding data source type when configuring the database connection information.

[0051] Users can also configure SQL statements, SQL templates, and XML templates in the visual operation interface. Before configuration, users need to manually select the SQL type. After configuration, the system will add a header modification tag to the template configured by the user according to the selected SQL type, and finally generate a standard XML template.

[0052] During the process of configuring the template, users can debug in the visual interface to view the SQL statement generated by the template and the result after the execution of the SQL.

[0053] After configuration, this template will generate a unique ID. Users only need to use this ID in their own programs without restarting the H2DLW service. That is, the hot deployment function of the data model.

[0054] Users can also modify the already configured template in the visual interface. After modification, the unique ID of this template will not change. That is, the hot update function of the data model.

[0055] Furthermore, the data model engine also includes a set of Java SDK toolkits. The Java SDK toolkits contain a small data model engine and a network request module. Users can use Maven to introduce the toolkits and call the API methods in them to generate a temporary JSON template for database interaction. After generating the JSON template, users need to call the API method of the request module to send a request to the H2DLW service. Then the request module will wrap the obtained result data into the Java object required by the user through the reflection mechanism of Java.

[0056] The present invention also discloses an operation method for a hybrid distributed data lake warehouse system, including: after the multi-source data target distribution module receives a request initiated by a user, it obtains a template in the data model engine, parses the template and parameters in the semantic interpretation engine, returns SQL or commands, and then sends the SQL or commands to the data model engine. The data model engine sends a request to the underlying data lake warehouse, and the underlying data lake warehouse returns results or data to the multi-source data target distribution module, which presents them to the user.

[0057] Compared with the prior art, the advantages of the present invention are as follows:

[0058] It solves the contradiction between data decentralization in the microservices architecture and data centralization in distributed storage in the big data platform.

[0059] BI, ML, and NLP can quickly and directly apply the processed data products of the hybrid distributed data lake warehouse system, and can also seamlessly call the massive metadata in the data lake warehouse of the entire hybrid distributed data lake warehouse system to perform on-demand model operations with high efficiency and low consumption.

[0060] The hybrid distributed data lake warehouse system realizes the underlying management of multi-source heterogeneous data internally, reducing the construction and management of the data underlying layer by other microservice modules; with one configuration management, multiple services benefit, improving the high availability and stability of the system. Description of the Drawings

[0061] Figure 1 It is a schematic structural diagram of the hybrid distributed data lake warehouse system according to an embodiment of the present invention.

[0062] Figure 2 It is a flowchart of the operation of the multi-source data target distribution module according to an embodiment of the present invention. Detailed Embodiments

[0063] To make the objectives, technical solutions, and advantages of the present invention clearer and more understandable, the following further elaborates on the present invention in detail with reference to the drawings and by listing embodiments.

[0064] The full name of H2DLW: Hybrid Distributed Data Lake Warehouse. The design of H2DLW integrates the distributed service design of microservices, the HDFS distributed file system, and the distributed storage system for multi-source heterogeneous data, forming a hybrid distributed application model. In H2DLW, a unified microservice module dynamically manages relational databases, columnar databases, in-memory databases, file system libraries, etc.

[0065] H2DLW provides unified data slice management, data directory management, and data permission hierarchical management for other microservice modules or external applications, and provides data access capabilities in the form of API, SQL, language scripts, etc. It supports dynamic update and loading of new data sources and new data access requirements in a non-stop high-availability mode.

[0066] H2DLW integrates BI, ML, NLP and other model capabilities internally. BI, ML and NLP can quickly and directly apply data products processed by H2DLW, and can also seamlessly call the massive metadata in the entire H2DLW data lake warehouse to perform on-demand model operations with high efficiency and low consumption.

[0067] H2DLW implements the underlying management of multi-source heterogeneous data internally. Other microservice modules no longer need to build data drivers themselves, do not need to manage thread scheduling during data storage, and do not need to worry about connection pool overflow. Public service data needs are configured once in H2DLW and shared by multiple services. No redundant configuration information is generated in each microservice module, and streamlined configuration improves security. When the data source or data set changes, all related application microservice modules no longer need to adjust their configurations and restart, greatly improving the high availability and reliability of the entire system.

[0068] like Figure 1 As shown, the present invention provides a hybrid distributed data lake warehouse system and an operation method thereof, including: an underlying data lake warehouse, a data engine, a semantic interpretation engine, a multi-source data target distribution module and a data model engine;

[0069] The underlying data lake warehouse: connects the data warehouse and the data lake, and integrates the advantages of the two architectures. Its underlying layer supports the coexistence of multiple data types and enables mutual sharing of data.

[0070] The data warehouse includes MySQL and Redis, and the data lake includes: HDFS, ClickHouse, ElastiSearch, and HBase.

[0071] The data engine consists of a multi-source data management engine, a data product engine, and a data calculation engine.

[0072] The multi-source data management engine is responsible for parsing the input parameters, including SQL, connection method, database code, etc. It uses native JDBC or native connectors to directly interact with the underlying data lake warehouse to obtain or operate the data in the data lake warehouse.

[0073] The data computing engine integrates the basic computing functions of big data. It is mainly responsible for two system businesses: 1. Organizing the original data in the data lake warehouse and transmitting it to the data product engine for processing. 2. Cooperating with the data product engine to provide computing power support.

[0074] The data calculation engine integrates: a BI module, an ML module, an NLP module, and a ScriptLanguage module;

[0075] The BI module is used to provide a calculation engine for business analysis and data mining for business intelligence (BI).

[0076] The ML module is used to provide a calculation engine for machine learning (ML) modeling.

[0077] The NLP module is used to provide a low-level calculation engine for natural language processing (NLP).

[0078] The ScriptLanguage module is used to provide basic engine support for the automated execution of script languages.

[0079] The data product engine is mainly responsible for secondary processing, operation, and analysis of the data in the data lake warehouse, and finally provides data support for upper-layer applications.

[0080] The data product engine integrates a database configuration module, a database driver module, a DBSession, and a database connection pool;

[0081] The database configuration module stores the database configurations configured by users in the data model engine. When interacting with the database, the corresponding database configurations will be obtained from this module first. The configurations mainly include database connection information and database driver information.

[0082] The function of the database driver module is to allocate different driver versions for each database request for connecting to the database. The database driver information retrieved from the database configuration module will be passed into this module, and then this module will call the DBSession to establish a connection to the corresponding database.

[0083] The DBSession is used to create a connection and session between the program and the database and manage the database transactions therein. All database requests need to be managed through the DBSession.

[0084] The database connection pool is responsible for allocating, managing, and releasing database connections. It allows the application program to reuse an existing database connection instead of establishing a new one; it releases database connections whose idle time exceeds the maximum idle time to avoid database connection leakage caused by not releasing database connections.

[0085] The semantic interpretation engine, whose main function is to parse the standardized templates such as SQL templates, XML templates, and JSON parameter templates stored in the data model engine into SQL statements for the data warehouse, operations for the data lake, or execution commands.

[0086] The semantic interpretation engine integrates the following components: XML parser, JSON parser, Script parser, DBSession reflector, dynamic parameter interpreter, class SQL mapper, semantic node manager, Hybrid hybrid query engine, semantic encapsulator, semantic executor;

[0087] The XML parser is used to convert XML into an XML DOM object;

[0088] The JSON parser is used to parse JSON-formatted data according to the semantic template.

[0089] The Script parser is used to parse the script language into executable statements.

[0090] The DBSession reflector is used to create a session between the program and the database.

[0091] The dynamic parameter interpreter is used to parse dynamic parameters according to the pre-edited parameter template.

[0092] The class SQL mapper is used to parse and convert the general SQL language according to the specific database characteristics.

[0093] The main function of the semantic node manager is to cooperate with various parsers to manage the nodes of the parsed SQL statements. The management content includes query conditions, sorting conditions, linking conditions, etc. in the SQL statements.

[0094] The Hybrid hybrid query engine is used for the hybrid query of multiple requests in a unified transaction, enabling the H2DLW service to support multiple database connections in one transaction. Specifically, it combines multiple different types of database requests into one transaction, and when rolling back the transaction, it can roll back different databases at once.

[0095] The semantic encapsulator is used to generate a semantic matching template from the pre-configured and stored data semantic template for quick matching when various parsers are parsing.

[0096] The semantic executor is used to execute the parsed data semantic template, such as a normally parsed XML template, and execute the SQL statement.

[0097] The data model engine is for user visual operations. The data model engine provides a web-based entry to facilitate users to manage database information and SQL templates. It supports a visual editing window where users can write SQL statements or configure SQL templates according to the standard SQL specification or template configuration specification.

[0098] The main function of the data model engine is to provide users with an operating space for H2DLW, including a visual operation interface and a Java SDK toolkit.

[0099] The functional modules of the data model engine include: data source management module, database management module, data model modification module, data model generation module, visual statement configuration module, visual statement debugging module, data model hot deployment module, data model hot update module;

[0100] The data model engine provides users with a visual operation interface.

[0101] Users can configure the database information they want to connect to in the database operation interface, such as the database connection address, database type, account password for logging in to the database, etc. Among them, the data source management module stores the information of the database types that the H2DLW service can connect to. These information are not allowed to be configured by users, and users can only select the corresponding data source type when configuring the database connection information.

[0102] Users can also configure SQL statements, SQL templates, and XML templates in the visual operation interface. Before configuration, users need to manually select the SQL type. After configuration, the system will add header modification tags to the templates configured by users according to the selected SQL type, and finally generate a standard XML template.

[0103] During the process of configuring the template, users can debug in the visual interface to view the SQL statements generated by the template and the results after the execution of the SQL.

[0104] After configuration, this template will generate a unique ID. Users only need to use this ID in their own programs without restarting the H2DLW service. That is, the hot deployment function of the data model.

[0105] Users can also modify the already configured template in the visual interface. After modification, the unique ID of this template will not change. That is, the hot update function of the data model.

[0106] The multi-source data target distribution module is the middleware in the hybrid distributed data lake warehouse system, playing a connecting role, and is used to connect the underlying data lake warehouse, data engine, semantic interpretation engine, and data model engine;

[0107] such as Figure 2As shown, after the multi-source data target distribution module receives a request initiated by the user, it obtains a template in the data model engine, parses the template and parameters in the semantic interpretation engine, returns SQL or commands, and then sends the SQL or commands to the data model engine. The data model engine sends a request to the underlying data lake repository, and the underlying data lake repository returns the results or data to the multi-source data target distribution module, which presents them to the user.

[0108] The hybrid distributed data lake repository system also provides a set of Java SDK toolkits, which includes a small data model engine and a network request module. Users can use Maven to introduce the toolkit and call the API methods therein to generate a temporary JSON template for database interaction. After generating the JSON template, the user needs to call the API method of the request module to send a request to the H2DLW service. Then the request module will wrap the obtained result data into the Java object required by the user through the reflection mechanism of Java.

[0109] Those of ordinary skill in the art will realize that the embodiments described herein are for helping the reader understand the implementation methods of the present invention, and it should be understood that the protection scope of the present invention is not limited to such specific statements and embodiments. Those of ordinary skill in the art can make various other specific deformations and combinations that do not depart from the essence of the present invention based on the technical revelations disclosed in the present invention, and these deformations and combinations are still within the protection scope of the present invention.

Claims

1. A hybrid distributed data lakehouse system, comprising: The underlying data lake warehouse, data engine, semantic interpretation engine, multi-source data target distribution module and data model engine; Bottom-layer data lake warehouse: includes data warehouse and data lake. The bottom layer supports the coexistence of multiple data types and enables mutual sharing of data. The data warehouse includes MySQL and Redis, and the data lake includes: HDFS, ClickHouse, ElastiSearch, and HBase; The data engine consists of a multi-source data management engine, a data product engine, and a data calculation engine; The multi-source data management engine is responsible for parsing the input parameters, including SQL, connection mode, and database code. It uses native JDBC or native connectors to directly interact with the underlying data lake warehouse to obtain or operate the data in the data lake warehouse. The data computing engine integrates the basic computing functions of big data. It is responsible for two system businesses:

1. Organizing the raw data in the data lake warehouse and transmitting it to the data product engine for processing; 2. Cooperating with the data product engine to provide computing power support for it; The data computing engine integrates: BI module, ML module, NLP module and ScriptLanguage module; The BI module is used to provide a computing engine for business analysis and data mining for business intelligence; The ML module is used to provide a computing engine for machine learning modeling; The NLP module is used to provide the underlying computing engine for natural language processing; The ScriptLanguage module provides basic engine support for the automated execution of script languages; The data product engine is responsible for secondary processing, calculation and analysis of the data in the data lake warehouse, and ultimately provides data support for upper-level applications; The multi-source data management engine integrates a database configuration module, a database driver module, a DBSession, and a database connection pool; The database configuration module stores the database configuration configured by the user in the data model engine. When interaction with the database is required, the corresponding database configuration will be obtained from this module first. The configuration includes database connection information and database driver information; The function of the database driver module is to assign a different driver version to each database request for connecting to the database; the database driver information taken out from the database configuration module will be passed to this module, and this module will then call DBSession to establish a connection to the corresponding database; The DBSession is used to create a connection and session between the program and the database, and manage database transactions therein; all database requests need to be managed through the DBSession; The database connection pool is responsible for allocating, managing, and releasing database connections. It allows applications to reuse an existing database connection instead of re-establishing one. It releases database connections whose idle time exceeds the maximum idle time to avoid missing database connections due to failure to release database connections. The semantic interpretation engine is used to parse the standardized templates of SQL-like templates, XML templates, and JSON parameter templates stored in the data model engine into SQL statements for the data warehouse, operations for the data lake, or execution commands; The data model engine is for users' visual operations; the data model engine provides a web - side entry to facilitate users' management of database information and SQL templates; it supports a visual editing window where users can write SQL statements or configure SQL templates according to standard SQL specifications or template configuration specifications; After receiving a request initiated by the user, the multi - source data target distribution module obtains a template in the data model engine, parses the template and parameters in the semantic interpretation engine, returns SQL or commands, and then sends the SQL or commands to the data model engine. The data model engine sends a request to the underlying data lake repository, and the underlying data lake repository returns results or data to the multi - source data target distribution module, which presents them to the user.

2. The hybrid distributed data lakehouse system according to claim 1, characterized in that: The semantic interpretation engine integrates the following: XML parser, JSON parser, Script parser, DBSession reflector, dynamic parameter interpreter, class SQL mapper, semantic node manager, Hybrid hybrid query engine, semantic encapsulator, semantic executor; The XML parser is used to convert XML into an XML DOM object; The JSON parser is used to parse JSON - formatted data according to the semantic template; The Script parser is used to parse script languages into executable statements; The DBSession reflector is used to create a session between the program and the database; The dynamic parameter interpreter is used to parse dynamic parameters according to a pre - edited parameter template; The class SQL mapper is used to parse and convert general SQL language according to the specific database characteristics; The role of the semantic node manager is to cooperate with various parsers to manage the nodes of the parsed SQL statements. The management content includes query conditions, sorting conditions, and linking conditions in the SQL statements; The Hybrid hybrid query engine is used for the hybrid query of multiple requests in a transaction, enabling the H2DLW service to support multiple database connections in one transaction; specifically, it combines multiple different types of database requests into one transaction, and when performing transaction rollback, it can roll back different databases at once; The semantic encapsulator is used to generate a semantic matching template from pre - configured and stored data semantic templates for quick matching by various parsers during parsing; The semantic executor is used to execute the parsed data semantic template, such as a normally parsed XML template, and execute the SQL statement.

3. The hybrid distributed data lakehouse system according to claim 1, characterized in that: In the data model engine, users can configure the database information they want to connect to in the database operation interface, including: database connection address, database type, account password information for logging in to the database; among them, the data source management module stores information about the database types that the H2DLW service can connect to. These information are not allowed to be configured by users, and users can only select the corresponding data source type when configuring database connection information; Users can also configure SQL statements, SQL templates, and XML templates in the visual operation interface. Before configuration, users need to manually select the SQL type first. After the configuration is completed, the system will add header modification tags to the templates configured by the user according to the selected SQL type, and finally generate a standard XML template; During the process of configuring the template, users can debug in the visual interface to view the SQL statements generated by the template and the results after the execution of the SQL; After the configuration is completed, this template will generate a unique ID. Users only need to use this ID in their own programs without restarting the H2DLW service; that is, the hot deployment function of the data model; Users can also modify the already configured templates in the visual interface. After the modification is completed, the unique ID of this template will not change, that is, the hot update function of the data model.

4. A hybrid distributed data lakehouse system according to claim 3, wherein: The data model engine also includes a set of Java SDK toolkits. The Java SDK toolkits contain a small data model engine and a network request module; users use Maven to introduce the toolkits and call the API methods in them to generate temporary JSON templates for database interaction; after generating the JSON templates, users need to call the API methods of the request module to send requests to the H2DLW service; After that, the request module will wrap the obtained result data into the Java objects required by the user through the reflection mechanism of Java.

Citation Information

Patent Citations

  • Automated development platform based on model configuration

    CN105549982A

  • Multi-source data management method and system and data management apparatus

    CN109542871A