Database data source parallel reading and predicate push-down method and device
By specifying the grouping fields and number, using jsqlParser to parse the filtering conditions and inject the grouping filter, combined with parallelism control and grouping expression templates, the problems of limited parallelism and lack of support for predicate pushdown when Spark reads JDBC data sources are solved, thus improving query performance and resource utilization efficiency.
Patent Information
- Application Number
- CN202510757578.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-09
- Publication Date
- 2025-11-07
AI Technical Summary
Spark's parallelism is limited when reading JDBC data sources. Data skew and predicate pushdown are not supported, leading to performance degradation, especially in non-aggregate SQL queries and complex SQL queries.
By specifying the grouping fields and number, the filter conditions are parsed using jsqlParser, and the grouping filter is dynamically injected. Parallelism control and grouping expression templates are introduced to optimize the SQL statement to achieve predicate pushdown and parallel reading.
It improves the parallel reading capability of JDBC data sources, enhances query performance through dynamic predicate pushdown, avoids data skew and excessive resource consumption, adapts to the characteristics of different databases, and enhances versatility.
Smart Images

Figure CN120910090A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The application relates to the field of JDBC data source reading, and in particular to a database data source parallel reading and predicate pushdown method and device. BACKGROUND
[0002] 1. Parallelism problem of Spark reading JDBC data source
[0003] When Spark reads a JDBC data source, the default parallelism is 1. Although parameters can be specified to support parallel reading by range partitioning, this method has many limitations. Specifically, the field used for data splitting must be specified, and the data type of the field must be numerical or date, otherwise it is not applicable. If the table does not have a numerical or date type field, the data cannot be read in parallel. In addition, the splitting method is based on range, and the amount of data read in parallel may be uneven, leading to data skew and affecting query performance. Therefore, although the parallelism can be improved by range splitting, there are many limitations.
[0004] Suppose we have a table named users that contains an id field (integer) and some other fields. We want to use the id field as the partition field and partition by the range of the id field to read data in parallel in multiple tasks.
[0005]
[0006]
[0007] Working principle
[0008] Partition basis: partitionColumn divides the data into multiple partitions by the specified range (defined by lowerBound and upperBound).
[0009] Parallel reading: Spark creates multiple parallel JDBC query tasks according to the specified numPartitions and the range divided by partitionColumn.
[0010] For example, suppose id is divided into 3 partitions, and Spark generates the following query:
[0011] SELECT*(SELECT id,name,area FROM users)SPARKGEN WHERE id BETWEEN 1AND10000;
[0012] SELECT * FROM (SELECT id, name, area FROM users) SPARKGEN WHERE id BETWEEN 10001 AND 20000;
[0013] SELECT * FROM (SELECT id, name, area FROM users) SPARKGEN WHERE id BETWEEN 20001 AND 30000;
[0014] When using range partitioning, ensure that the data distribution of partitionColumn is uniform. If the data volume of some ranges in the dataset is very large, while the data volume of other ranges is small, it may cause some tasks to be overloaded (data skew).
[0015] 2. Non-aggregated SQL query predicate pushdown problem
[0016] When Spark executes a SQL query, if the optimizer of the underlying database does not support predicate pushdown, it will cause performance degradation. Take the following SQL as an example:
[0017] select id, name, area from tb;
[0018] Assuming we divide the query into 3 partitions, Spark will generate the following SQL query for each partition, and each partition corresponds to a concurrent query:
[0019] SELECT * FROM (SELECT id, name, area FROM users) SPARKGEN WHERE id BETWEEN 1 AND 10000;
[0020] SELECT * FROM (SELECT id, name, area FROM users) SPARKGEN WHERE id BETWEEN 10001 AND 20000;
[0021] SELECT * FROM (SELECT id, name, area FROM users) SPARKGEN WHERE id BETWEEN 20001 AND 30000;
[0022] For databases that support predicate pushdown (such as MySQL, Oracle, etc.), the database optimizer will automatically push the where condition to the subquery, and perform data filtering in advance, thereby improving performance. The actual SQL executed by the database is as follows:
[0023] select * from (select id, name, area from users where id between 1 and 10000) SPARKGEN;
[0024] select * from (select id, name, area from users where id between 10001 and 20000) SPARKGEN;
[0025] select * from (select id, name, area from users where id between 20001 and 30000) SPARKGEN;
[0026] However, for databases that do not support predicate pushdown (such as ClickHouse), the database will first execute the subquery, return all results, and then filter the where condition, which will cause significant performance problems because the amount of data to be processed is very large at this time.
[0027] 3. Predicate pushdown problem in aggregate SQL and complex SQL
[0028] For aggregate operations, various databases cannot implement predicate pushdown. For complex SQL, the optimizer cannot always implement predicate pushdown.
[0029] For example, consider the following SQL query:
[0030]
[0031]
[0032] Assume we divide the query into 3 partitions, and Spark will generate SQL queries similar to the following for each partition.
[0033] Statement 1:
[0034]
[0035] Statement 2:
[0036]
[0037] Statement 3:
[0038]
[0039] In this case, since the where condition filter cannot be applied before performing the aggregation operation, the database first performs a complete aggregation calculation, and only after the calculation is completed, the where condition filter is performed. This operation results in a large amount of unnecessary data processing, greatly affecting the performance. SUMMARY
[0040] To solve the above-mentioned problems existing in the JDBC data source reading data, the application provides a database data source parallel reading and predicate pushdown method and device, mainly including improving parallel reading capability, implementing predicate pushdown, introducing parallelism control mechanism, and providing flexible grouping expression template and the like. Through these innovative methods, the efficiency of accessing the JDBC data source is greatly improved, which helps to speed up the performance of big data analysis applications.
[0041] To achieve the above object, the application adopts the following technical solutions:
[0042] In an embodiment of the application, a database data source parallel reading and predicate pushdown method is provided, which comprises:
[0043] For a database table, the fields used for grouping, the required number of groups, and the parallelism threshold are specified;
[0044] The SQL statement of the JDBC data source is split into multiple groups for execution, ensuring that the number of groups executed simultaneously does not exceed the set parallelism threshold;
[0045] The jsqlParser component is used to parse the filter conditions of the table and dynamically injected into the filter conditions of the groups;
[0046] For different types of databases, the corresponding grouping expression template is used, allowing users to customize the grouping rules.
[0047] Further, according to the characteristics of different types of databases, the grouping field position is marked by a specific placeholder to generate the grouping expression template.
[0048] Further, the selection of the placeholder is as follows:
[0049] The primary key field of the database table is preferentially selected as the placeholder;
[0050] If the table has no primary key or the database does not support the primary key concept, then a column with high selectivity is selected as the placeholder;
[0051] The columns suitable for being the placeholder are identified by querying the system table of the database;
[0052] For databases that do not support statistical information, the selectivity of the column is calculated using the algorithm function provided by the database.
[0053] Further, the SQL statements of the JDBC data source are grouped, and a filter condition is added in the WHERE clause of the original SQL statement; the filter condition is the result of the number calculated by the grouping expression modulo the grouping number, and the value range is from 0 to the grouping number minus 1.
[0054] In an embodiment of the present application, a device for parallel reading and predicate pushdown of a database data source is also provided, and the device comprises:
[0055] A grouping field and number specifying module is configured to specify, for a database table, a field used for grouping, a required grouping number, and a parallelism threshold value;
[0056] A parallelism control module is configured to split the SQL statement of the JDBC data source into multiple groups for execution, and ensure that the number of groups executed simultaneously does not exceed the set parallelism threshold value;
[0057] A predicate pushdown implementation module is configured to use a jsqlParser component to parse the filter condition of a table, and dynamically inject the filter condition into the grouping condition;
[0058] A database type adaptation module is configured to use a corresponding grouping expression template for different types of databases, and allow a user to customize a grouping rule.
[0059] Further, according to the characteristics of different types of databases, a grouping field position is marked by a specific placeholder to generate a grouping expression template.
[0060] Further, the selection of the placeholder is as follows:
[0061] The primary key field of the database table is preferentially selected as the placeholder;
[0062] If the table has no primary key or the database does not support the primary key concept, a column with high selectivity is selected as the placeholder;
[0063] The columns suitable for being the placeholder are identified by querying the system table of the database;
[0064] For a database that does not support statistical information, an algorithm function provided by the database is used to count the selectivity of the column.
[0065] Further, the SQL statements of the JDBC data source are grouped, and a filter condition is added in the WHERE clause of the original SQL statement; the filter condition is the result of the number calculated by the grouping expression modulo the grouping number, and the value range is from 0 to the grouping number minus 1.
[0066] In an embodiment of the present application, a computer device is also provided, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the aforementioned parallel reading of database data source and predicate pushdown when executing the computer program.
[0067] In an embodiment of the present application, a computer readable storage medium is also provided, which stores a computer program for executing the parallel reading of database data source and predicate pushdown.
[0068] Advantages:
[0069] 1. The present application supports custom partition fields and partition ranges, improving the parallel reading capability from JDBC data sources.
[0070] 2. The present application realizes dynamic predicate pushdown, automatically injecting query conditions into subqueries to improve the query performance of different types of databases.
[0071] 3. The present application introduces a parallelism control mechanism to ensure that query load does not excessively squeeze underlying database and cluster resources.
[0072] 4. The present application provides flexible grouping expression templates to adapt to the characteristics of various databases, enhancing the universality of the solution. BRIEF DESCRIPTION OF DRAWINGS
[0073] Figure 1 is a flow chart of the method for parallel reading of database data source and predicate pushdown of the present application;
[0074] Figure 2 is a schematic diagram of the structure of the apparatus for parallel reading of database data source and predicate pushdown of the present application;
[0075] Figure 3 is a schematic diagram of the structure of the computer device of the present application. DETAILED DESCRIPTION
[0076] The principles and spirits of the present application will be described below with reference to a number of exemplary embodiments, and it should be understood that these embodiments are given only to enable those skilled in the art to better understand and implement the present application, and do not limit the scope of the present application in any way. On the contrary, these embodiments are provided to make the present disclosure more thorough and complete, and to fully convey the scope of the present disclosure to those skilled in the art.
[0077] Those skilled in the art know that the embodiments of the present application can be implemented as an apparatus, a device, a equipment, a method or a computer program product. Therefore, the present disclosure can be embodied in the form of a complete hardware, a complete software (including firmware, resident software, microcode, etc.), or a combination of hardware and software.
[0078] According to an embodiment of the present application, a method for database data source parallel reading and predicate pushdown is proposed, mainly including improving parallel reading capability, implementing predicate pushdown, introducing parallelism control mechanism, and providing flexible grouping expression template, etc. Through these innovative methods, the efficiency of accessing JDBC data source is greatly improved, which helps to accelerate the performance of big data analysis applications.
[0079] The principles and spirits of the present application will be explained in detail below with reference to several representative embodiments of the present application.
[0080] In order to improve the performance and security of database query, the JDBC data source of Spark is optimized as follows Figure 1 The specific steps are as follows:
[0081] 1. Specify the grouping field and quantity: For the database table, specify the field used for grouping, the required number of groups, and the parallelism threshold.
[0082] 2. Parallelism control: The SQL statement of the JDBC data source is split into multiple groups for execution, ensuring that the number of groups executed simultaneously does not exceed the set parallelism threshold. If the parallelism is not controlled, it may cause too many groups to be executed simultaneously, causing great performance pressure on the database and the host running the Spark task.
[0083] 3. Predicate pushdown implementation: To implement predicate pushdown, we can use the jsqlParser component to parse the filter conditions of the table and dynamically inject them into the filter conditions of the groups. In this way, we can filter data in advance in the subquery stage, avoiding the performance pressure caused by too large amount of data read at one time. Dynamic injection of filter conditions avoids errors that may be caused by parsing the complete SQL statement, thereby significantly improving performance and security. JSQLParser is an open source Java library for parsing SQL statements. It can parse SQL statements into Abstract Syntax Tree (AST), allowing developers to analyze and manipulate SQL statements programmatically.
[0084] 4. Database type adaptation: For different types of databases such as MySQL and ClickHouse, the grouping methods are different. By using the grouping expression template, users can customize the grouping rules to achieve better scalability.
[0085] Through the above steps, the database query can be effectively optimized to ensure the query efficiency and system stability.
[0086] Detailed implementation
[0087] 1. Added parameters
[0088] On the basis of spark Jdbc data source, the following parameters are added:
[0089] (1) Grouping expression template: This parameter allows users to select one or more fields from a database table and process them according to a custom SQL expression to generate a set of numbers. By setting the grouping number, the system will decompose the original query SQL into multiple subqueries to execute in a grouped manner.
[0090] (2) Grouping number: This parameter controls the total number of generated SQL groups. Increasing the grouping number will reduce the amount of data processed by each SQL query, but at the same time, it will increase the number of queries executed. Therefore, users should adjust the grouping number flexibly according to the data size and business requirements to achieve the best balance between performance and resource consumption.
[0091] (3) Parallelism threshold: This parameter defines the maximum number of SQL groups that can be executed simultaneously. For example, if the grouping number is set to 10 and the parallelism threshold is set to 4, the system will execute a maximum of 4 SQL queries for each group at the same time. By adjusting the parallelism threshold, resource utilization and execution efficiency can be optimized, especially when dealing with large-scale data.
[0092] 2、How to group
[0093] The grouping expression template is implemented according to the characteristics of different databases, which uses specific placeholders to mark field positions so that these placeholders can be replaced with actual database table fields during task execution. In MySQL, the crc32 function can be used to convert data into a 32-bit unsigned cyclic redundancy check value, which is a numerical value. For example, if the id field is used as the grouping basis, the grouping expression template can be set to crc32(${placeHolder}), and after replacing the placeholder, the expression becomes crc32(id), each id record will correspond to a number. If there are multiple grouping fields, you can use the concat_ws function (string concatenation function) to connect them with commas, such as crc32(concat_ws(',',id,name)).
[0094] For ClickHouse database, the grouping expression template can be configured as CityHash64(${placeHolder}), where CityHash64 is a high-performance non-encryption hash function used to calculate the 64-bit hash value of a string, and the return value is also a number. Different databases may have different implementations.
[0095] Once the grouping expression template and the number of groups are determined, the SQL statement can be grouped by adding a filter condition in the WHERE clause of the original SQL statement. The filter condition is the result of the grouping expression modulo the number of groups, which ranges from 0 to the number of groups minus 1.
[0096] Let's take a simple SQL statement as an example to illustrate the principle of grouping:
[0097] SELECT devid, devname, vendor FROM cm_res_device;
[0098] Suppose we use devid as the grouping field and choose crc32(${placeHolder}) as the template, with the number of groups set to 3. Then:
[0099] The grouping expression is: crc32(devid).
[0100] The modulo expression is: crc32(devid) % 3, which results in 0, 1, or 2.
[0101] Accordingly, we can generate three SQL query statements for the three groups:
[0102] SELECT devid, devname, vendor FROM cm_res_device WHERE crc32(devid) % 3 = 0;
[0103] SELECT devid, devname, vendor FROM cm_res_device WHERE crc32(devid) % 3 = 1;
[0104] SELECT devid, devname, vendor FROM cm_res_device WHERE crc32(devid) % 3 = 2;
[0105] In this way, the original query is split into three queries, each of which processes approximately 1 / 3 of the data in the table, thereby improving query performance.
[0106] 3. Selection of placeholders
[0107] To ensure the balance of data after grouping, we should prefer to choose fields with uniform data distribution as placeholders. In general, the primary key of a database table is the best choice, as it is usually unique and uniformly distributed. If the table does not have a primary key or the database does not support the primary key concept, the most suitable placeholder needs to be selected according to the characteristics of the database. The following is the specific selection process:
[0108] (1) Obtain primary key field
[0109] Obtain the primary key field from the metadata of the database data source and replace the placeholder with it.
[0110] If the primary key contains multiple fields, use database-specific functions or operators to concatenate them to replace the placeholder. For example, in MySQL, you can use the concat_ws function, while in Oracle, you can use the "||" concatenator.
[0111] (2) Select columns with high selectivity
[0112] If the table does not have a primary key, select a column with high selectivity as the placeholder field.
[0113] The formula for calculating selectivity is: the number of unique values of the column divided by the total number of records.
[0114] The closer the selectivity is to 1, the more uniform the value distribution of the column, and the more suitable it is as a placeholder.
[0115] (3) Use database statistics
[0116] To avoid affecting the performance of the database, we do not directly count the selectivity of each column, but instead look at the database statistics.
[0117] Mainstream databases (such as Oracle, MySQL) support periodic collection of statistics in the background, and we can query the system table of the database to identify columns suitable for placeholders.
[0118] Example: Query selectivity of Oracle and MySQL
[0119] For Oracle databases, you can query the dba_tab_col_statistics view to view the cardinality and selectivity of each column. The dba_tab_col_statistics view in MySQL is a metadata view that displays the statistical information of tables and columns in the database. These statistical information includes the minimum value, maximum value, number of null values, number of unique values, etc., mainly used by the query optimizer to generate efficient execution plans. The following is a SQL script to query the cardinality and selectivity of each column in the TEST table:
[0120]
[0121]
[0122] The INFORMATION_SCHEMA.STATISTICS table of MySQL provides statistics of indexes and columns, including Cardinality. The selectivity of a column can be obtained through the following SQL query:
[0123]
[0124]
[0125] (4) For databases that do not support statistics
[0126] For databases that do not support statistics, the selectivity of a column can be calculated using the algorithmic functions provided by the database. For example, ClickHouse can use the HyperLogLog function to calculate the number of unique values, while ClickHouse stores the total number of records in the files of each data segment, which is also very efficient. HyperLogLog (HLL for short) is a high-efficiency probabilistic algorithm for estimating the cardinality (i.e., counting the number of unique values) in ClickHouse, which can handle millions of data entries per second.
[0127] (5) User manually specifies grouping fields
[0128] To compensate for the scenarios not supported by the above methods, users can manually specify grouping fields.
[0129] Through these optimization measures, we can more efficiently and accurately select columns suitable for placeholders, thereby improving the balance and efficiency of data processing.
[0130] 4. Control of parallelism
[0131] A parallelism control mechanism is introduced. In local mode, the parallelism size is controlled by the local[N] parameter, and in Yarn mode, the parallelism is set by the spark.default.parallelism parameter (spark's control task parallelism parameter) to ensure that the query load does not excessively squeeze the underlying database and cluster resources. The local[N] parameter is a parameter for submitting tasks in local mode, where N represents the number of available CPU cores. For example, if N is 3, it means using 3 CPU cores for calculation.
[0132] 5. Predicate pushdown
[0133] In the SQL statement, the placeholder #<dame_pushdown_holder> is used to mark the position where grouping filtering needs to be performed. When generating the grouping SQL expression, the system will replace this placeholder with the actual grouping condition expression, thereby dynamically constructing the final query statement.
[0134] For SQL query statement:
[0135]
[0136] Use #<dame_pushdown_holder> to specify the placeholder position:
[0137]
[0138] The above example specifies the position of the group filter expression replacement through the placeholder #<dame_pushdown_holder>.
[0139] In the execution of data reading, the SQL statement will be grouped into 3 SQL statements, respectively executed, and the generated SQL grouping statement is as follows:
[0140] SQL group 1:
[0141]
[0142] SQL group 2:
[0143]
[0144]
[0145] SQL group 3:
[0146]
[0147] It should be noted that although the operations of the method of the present application are described in a specific order in the above embodiments and drawings, this does not require or imply that the operations must be performed in this specific order, or that all the shown operations must be performed to achieve the desired results. Additionally or alternatively, certain steps can be omitted, multiple steps can be combined into one step, and / or one step can be divided into multiple steps.
[0148] In order to more clearly explain the method of parallel reading and predicate pushdown of the above database data source, a specific embodiment will be described below, however, it should be noted that this embodiment is only for better illustration of the present application and does not constitute an improper limitation on the present application.
[0149] Implementation case:
[0150] The following code specifies three parameters: the number of groups, the group expression and the parallelism threshold, and uses the placeholder #<dame_pushdown_holder> to mark the position of the filter group expression replacement. Load the table of the database through the JDBC data source and read in parallel.
[0151]
[0152]
[0153] Based on the same inventive concept, the present application further provides a device for parallel reading of database data source and predicate pushdown. The implementation of the device can refer to the implementation of the above-mentioned method, and the repeated parts will not be described herein. The term "module" used below can be a combination of software and / or hardware that realizes a predetermined function. Although the device described in the following embodiments is preferably realized in software, the implementation of hardware or a combination of software and hardware is also possible and is conceived.
[0154] Figure 2 is a structural schematic diagram of the device for parallel reading of database data source and predicate pushdown of the present application. As shown in Figure 2 , the device comprises:
[0155] A grouping field and number specifying module 101 is configured to specify, for a database table, a field used for grouping, a required grouping number, and a parallelism threshold.
[0156] According to the characteristics of different types of databases, the grouping field position is marked by a specific placeholder, and a grouping expression template is generated;
[0157] The selection of the placeholder is as follows:
[0158] The primary key field of the database table is preferentially selected as the placeholder;
[0159] If the table has no primary key or the database does not support the primary key concept, a column with high selectivity is selected as the placeholder;
[0160] The columns suitable for being the placeholder are identified by querying the system table of the database;
[0161] For databases that do not support statistical information, the selectivity of the column is calculated by using an algorithm function provided by the database.
[0162] The SQL statement of the JDBC data source is subjected to grouping processing, and a filtering condition is added in the WHERE clause of the original SQL statement; the filtering condition is the result of taking the modulus of the number calculated by the grouping expression with the grouping number, and the value range is from 0 to grouping number minus 1.
[0163] A parallelism control module 102 is configured to split the SQL statement of the JDBC data source into multiple groups for execution, and ensure that the number of groups executed simultaneously does not exceed the set parallelism threshold.
[0164] The predicate pushdown implementation module 103 is configured to parse the filter condition of the table using the jsqlParser component and dynamically inject into the filter condition of the grouping.
[0165] The database type adaptation module 104 is configured to use corresponding grouping expression templates for different types of databases to allow user to customize grouping rules.
[0166] It should be noted that although several modules of the apparatus for parallel reading of database data sources and predicate pushdown are mentioned in the foregoing detailed description, such division is merely exemplary and not mandatory. In fact, according to the embodiments of the present application, the features and functions of two or more modules described above can be embodied in one module. Conversely, the features and functions of one module described above can be further divided into several modules.
[0167] Based on the foregoing inventive concept, as shown in Figure 3 The present application further proposes a computer device 200, which comprises a memory 210, a processor 220, and a computer program 230 stored in the memory 210 and executable on the processor 220, wherein the processor 220 implements the foregoing method for parallel reading of database data sources and predicate pushdown when executing the computer program 230.
[0168] Based on the foregoing inventive concept, the present application further proposes a computer-readable storage medium, which stores a computer program for executing the foregoing method for parallel reading of database data sources and predicate pushdown.
[0169] The method and apparatus for parallel reading of database data sources and predicate pushdown proposed by the present application have the following highlights:
[0170] 1. Support for customizing partition fields and partition ranges to improve the parallel reading capability from JDBC data sources.
[0171] 2. Dynamic predicate pushdown is implemented to automatically inject query conditions into subqueries to improve the query performance of different types of databases.
[0172] 3. A parallelism control mechanism is introduced to ensure that the query load does not excessively squeeze the underlying database and cluster resources.
[0173] 4. Flexible grouping expression templates are provided to adapt to the characteristics of various databases and enhance the universality of the solution.
[0174] While the principles and spirit of the application have been described with reference to several specific embodiments, it is to be understood that the application is not limited to the specific embodiments disclosed, and that the division of the aspects is not meant to imply that features from these aspects cannot be combined to benefit, but is merely for ease of presentation. The application is intended to cover various modifications and equivalent arrangements included within the spirit and scope of the appended claims.
[0175] The scope of the protection of the application is shown by the appended claims, and it is understood that all modifications and equivalent arrangements within the scope of the claims and variations thereof are to be included in the scope of the application.
Claims
1. A method for database data source parallel read with predicate pushdown, the method comprising: The method comprises: For a database table, a field used for grouping, a required grouping number, and a parallelism threshold are specified; SQL statements of a JDBC data source are split into multiple grouping executions, ensuring that the number of simultaneously executed groupings does not exceed the set parallelism threshold; A jsqlParser component is used to parse filter conditions of a table and dynamically inject into filter conditions of groupings; For different types of databases, corresponding grouping expression templates are used to allow users to customize grouping rules.
2. The method of database data source parallel read with predicate pushdown according to claim 1, characterized in that, According to characteristics of different types of databases, grouping field positions are marked by specific placeholders to generate grouping expression templates.
3. The method of database data source parallel read with predicate pushdown according to claim 2, characterized in that, The selection of the placeholders is as follows: A primary key field of a database table is preferentially selected as a placeholder; If the table has no primary key or the database does not support the primary key concept, a column with high selectivity is selected as a placeholder; Suitable columns as placeholders are identified by querying system tables of the database; For databases that do not support statistical information, an algorithm function provided by the database is used to count selectivity of the column.
4. The method of database data source parallel read with predicate pushdown according to claim 1, characterized in that, Grouping processing is performed on SQL statements of a JDBC data source, and filter conditions are added in a WHERE clause of an original SQL statement; The filter conditions are results of modulus operation of a number calculated by a grouping expression on a grouping number, and the value range is from 0 to grouping number minus 1.
5. An apparatus for database data source parallel read with predicate pushdown, the apparatus comprising: The device comprises: A grouping field and number specifying module for specifying, for a database table, a field used for grouping, a required grouping number, and a parallelism threshold; A parallelism control module for splitting SQL statements of a JDBC data source into multiple grouping executions, ensuring that the number of simultaneously executed groupings does not exceed the set parallelism threshold; A predicate pushdown implementation module for using a jsqlParser component to parse filter conditions of a table and dynamically injecting into filter conditions of groupings; A database type adaptation module for using, for different types of databases, corresponding grouping expression templates to allow users to customize grouping rules.
6. The apparatus according to claim 5, wherein, According to characteristics of different types of databases, grouping field positions are marked by specific placeholders to generate grouping expression templates.
7. The apparatus according to claim 6, wherein, The selection of the placeholders is as follows: A primary key field of a database table is preferentially selected as a placeholder; If the table has no primary key or the database does not support the primary key concept, a column with high selectivity is selected as a placeholder; Suitable columns as placeholders are identified by querying system tables of the database; For databases that do not support statistical information, an algorithm function provided by the database is used to count selectivity of the column.
8. The apparatus of claim 5, wherein, Grouping processing is performed on SQL statements of a JDBC data source, and filter conditions are added in a WHERE clause of an original SQL statement; The filter conditions are results of modulus operation of a number calculated by a grouping expression on a grouping number, and the value range is from 0 to grouping number minus 1.
9. A computer device comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, The processor implements the method of any one of claims 1-4 when executing the computer program.
10. A computer-readable storage medium, characterized in that, The computer readable storage medium stores a computer program for executing the method of any one of claims 1-4.