Method for reading data of multiple databases

By automatically matching the read protocol of database type, data is stored in a ring queue in Arrow format, solving the high-performance computing problem in federated queries and achieving efficient data computing and storage.

CN120216530APending Publication Date: 2025-06-27HARVEST VISION TECH (BEIJING) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510166904.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-14
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

In the prior art, it is difficult to realize high-performance data calculations when conducting federal queries, and it is necessary to synchronize data to the same column-store MPP database, resulting in inefficiency.

Method used

High-performance computing is achieved by automatically matching the read protocol of database types, reading data directly from different types of databases and writing them to a ring queue and storing them in an Arrow format.

Benefits of technology

It realizes the ability to directly read the general column storage Arrow format from different types of databases, avoids the data synchronization process, improves data computing performance, and supports vectorized computing of the latest CPUs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120216530A_ABST
    Figure CN120216530A_ABST
Patent Text Reader

Abstract

The invention discloses a method for reading data of a plurality of databases. The method comprises the following steps of: 1, automatically matching a reading protocol according to the types of the databases; step 2, reading the corresponding database through the automatically matched reading protocol at the same time; and step 3, writing the read data into the annular queue in such a manner, wherein the manner enables the data in the annular queue to be in an Arrow format.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to database languages, and particularly to a method for reading data from multiple databases. Background Art

[0002] Federated queries can solve the feasibility problem of multi-source data, but high performance has always been an important factor restricting the development of federated queries. Therefore, in order to obtain high-performance data computing, the current practice usually synchronizes (ETL) data of various types of databases to the same columnar MPP database, and then the capabilities of the MPP database can be utilized for high-performance computing.

[0003] Arrow is an advanced in-memory columnar data format, and the Arrow format is ready for high-performance data computing. If there is a software implementation that can solve the federation problem while outputting a general columnar structure (Arrow) in memory, it can not only avoid ETL but also be ready for high-performance computing. Summary of the Invention

[0004] A method for reading data from multiple databases, comprising the following steps: Step 1, automatically matching a reading protocol according to the type of the database; Step 2, simultaneously reading the corresponding databases through the automatically matched reading protocol; Step 3, writing the read data into a circular queue in such a way that the data in the circular queue is in Arrow format.

[0005] Preferably, the reading protocol is a correspondence table between the data types of different types of databases and the Arrow type set in advance.

[0006] Preferably, the types include at least one of Oracle, Mysql, ClickHouse, Sqlserver, duckDB, and postgresql.

[0007] Preferably, for the duckDB type, it is read directly.

[0008] Preferably, after Step 1 and before Step 2, it includes: classifying into ODBC operations and ADBC operations according to the type of the database. For ODBC operations, it normally enters Step 2, and for ADBC operations, it directly jumps to Step 3.

[0009] Preferably, the ODBC operations correspond to relational databases, including mysql, oracle, pg, ClickHouse, and doris, and the ADBC operations correspond to databases where the read data is in Arrow format.

[0010] Preferably, the Rust language is used.

[0011] Preferably, it further includes step 4, and the data in Arrow format is provided for use by a high-performance computing engine.

[0012] Preferably, the Arrow format is a tabular data format for facilitating data analysis operations.

[0013] Preferably, the Arrow format is suitable for vectorized computing of the latest CPUs and is independent of database languages.

[0014] In the implementation of federated queries such as Presto and Spark, they already have the ability to shield the problem of a large variety of field types brought about by a large variety of databases and do not require data synchronization. However, the implementation of these components for the data type format read from the database is based on their own type format (row storage). If columnar storage high-performance computing is to be performed, it is often necessary to perform serialization conversion on the read data (serialize the self row storage type into a columnar storage format), and then use an engine that supports Arrow for computing. The time consumed by big data during serialization is huge. This is also an important reason why big data computing performance of Presto, Spark, etc. is usually inferior to MPP databases such as ClickHouse and Doris.

[0015] For columnar storage databases such as ClickHouse and Doris, the expression of the data format in memory is neither the general Arrow format (both are their own columnar storage formats), nor do they have complete federated query capabilities.

[0016] The implementation of the method of the present invention has the ability to directly read the general columnar storage Arrow format from different types of databases. Any computing engine that supports the Arrow format can obtain the data read by this software without serialization and deserialization and directly perform high-performance computing. In the present invention, the logic and implementation method for the data federation to uniformly read as the Arrow format; the correspondence between the data federation reading as the Arrow format and each field of each database. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] Figure 1 It is a simplified view of the method of the present invention; Figure 2 It shows the conversion relationship between common types in Oracle and Arrow types; Figure 3 It shows the conversion relationship between common types in Mysql and ClickHouse and Arrow types; Figure 4Shows the conversion relationship between common types in SqlServer and Arrow types; Figure 5 Shows the conversion relationship between common types in duckDB and Arrow types; Figure 6 Shows the conversion relationship between common types in postgresql and Arrow types. Detailed implementation

[0018] As Figure 1 shown, a method for reading data from multiple databases includes the following steps: Step 1, automatically match the reading protocol according to the type of the database; Step 2, simultaneously read the corresponding database through the automatically matched reading protocol; Step 3, write the read data into a circular queue in such a way that the data in the circular queue is in Arrow format.

[0019] As Figures 2 - 6 shown, preferably, the reading protocol is a correspondence table between data types of different types of databases and Arrow types set in advance.

[0020] Preferably, the types include at least one of Oracle, Mysql, ClickHouse, Sqlserver, duckDB, and postgresql.

[0021] Preferably, directly read for the duckDB type.

[0022] Preferably, after step 1 and before step 2, it includes: classifying into ODBC operations and ADBC operations according to the type of the database. For ODBC operations, normally enter step 2, and for ADBC operations, directly jump to step 3.

[0023] Preferably, the ODBC operations correspond to relational databases, including mysql, oracle, pg, ClickHouse, and doris, and the ADBC operations correspond to databases where the read data is in Arrow format.

[0024] Preferably, use the Rust language.

[0025] Preferably, it further includes step 4, providing the data in Arrow format for use by a high-performance computing engine.

[0026] Preferably, the Arrow format is a tabular data format for easy data analysis operations.

[0027] Preferably, the Arrow format is suitable for vectorized computing of the latest CPUs and is independent of database languages.

[0028] Term Definition: Arrow: Apache Arrow defines a format for representing tabular data in memory. This format is particularly optimized for analytical operations. For example, the columnar format can make full use of the advantages of modern CPUs for vectorized computing. Moreover, Arrow also defines the IPC format to serialize the data in memory for network transmission or to persist the data in the form of files. The format defined by Arrow is language-independent, so any language can implement the format defined by Arrow. The Arrow project provides SDKs for almost all mainstream programming languages.

[0029] Data Federation: Data federation eliminates the need to create another database or data warehouse and manage the integration with the central data store.

[0030] ODBC&ADBC: A driver for reading database data. Through ODBC, data can be read from relational databases (such as mysql, oracle, pg, ClickHouse, doris, etc.). If the database itself supports the ADBC protocol, data can also be directly read from the database using ADBC (such as duckDB), and the data read out is directly in the Arrow format.

[0031] The present invention is developed using the Rust language and uses ODBC / ADBC / other methods to concurrently read data from the database. The reading method is as follows: According to the database table partitioning method (obtained from the database metadata), data is concurrently read from different databases at the same time (the number of concurrent reads depends on the number of database table partitions), and the read data is circularly written into a circular queue. And try to enter the queue in batches through the method of avoiding memory copying and serialization (if it is ODBC reading, serialization is required). At this time, the in-memory data in the queue is already in the Arrow format and can be directly used by the outside.

[0032] Therefore, while the present invention realizes the functions that data federation should have, by defining a unified Arrow in-memory format for output, it lays a foundation for solving the high-performance computing of data federation.

[0033] The preferred embodiments of the present invention are exemplarily given above with reference to the accompanying drawings. However, those skilled in the art should understand that without exceeding the scope of protection of the appended claims, those skilled in the art can form new technical solutions by adding features, combining features, etc.

Claims

1. A method for reading data from multiple databases, characterized in that: The steps include: Step 1, automatically matching the reading protocol according to the type of the database; Step 2, simultaneously reading the corresponding database through the automatic matching reading protocol; Step 3: write the read data into the circular queue in such a way that the data in the circular queue is in Arrow format.

2. The method for reading data from multiple databases according to claim 1, characterized in that: The reading protocol is a preset correspondence table between data types of different types of databases and Arrow types.

3. The method for reading data from multiple databases according to claim 2, characterized in that: The types include at least one of Oracle, Mysql and ClickHouse, Sqlserver, duckDB, postgresql.

4. The method for reading data from multiple databases according to claim 3, characterized in that: For duckDB type, read directly.

5. The method for reading data from multiple databases according to claim 1, characterized in that: After step 1 and before step 2, the method includes: dividing the operation into ODBC operation and ADBC ​​operation according to the type of database, and normally proceeding to step 2 for ODBC operation, and directly jumping to step 3 for ADBC ​​operation.

6. The method for reading data from multiple databases according to claim 5, characterized in that: The ODBC operation corresponds to relational databases, including mysql, oracle, pg, ClickHouse and doris, and the ADBC ​​operation corresponds to a database whose data is read in Arrow format.

7. The method for reading data from multiple databases according to any one of claims 1 to 6, characterized in that: Use Rust language.

8. The method for reading data from multiple databases according to claim 7, characterized in that: The method further includes step 4, wherein the data in the Arrow format is provided to a high performance computing engine for use.

9. The method for reading data from multiple databases according to claim 8, characterized in that: The Arrow format is a tabular data format to facilitate data analysis operations.

10. The method for reading data from multiple databases according to claim 9, characterized in that: The Arrow format is suitable for vectorized computing on the latest CPUs and is independent of the database language.