Data processing method and device, equipment, medium and program product

By obtaining the execution plan data in the distributed database system, determining the estimated output data unit and number of execution instances of the operation operator, the problem of inaccurate parallelism setting in the prior art is solved, and query execution efficiency is improved.

CN120011399APending Publication Date: 2025-05-16CHINA MOBILE INFORMATION TECHNOLOGY CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510104267.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-22
Publication Date
2025-05-16

AI Technical Summary

Technical Problem

When setting the parallelism of a distributed database system, the accuracy rate is low, which can easily cause resource utilization imbalances and affect the execution efficiency of query.

Method used

By obtaining the execution plan data in the query instruction, including the operation operator and the number of input data units, the estimated output data units of the operation operator are determined, and the number of execution instances is determined based on the input data, the estimated output data and the number of processing data units of a single instance.

Benefits of technology

It improves the accuracy of the number of execution instances in the query statement, fully considers the characteristics of the operation operator and the data processing volume, effectively avoids the problem of inaccurate parallelism setting, and improves the execution efficiency of query instructions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120011399A_ABST
    Figure CN120011399A_ABST
Patent Text Reader

Abstract

The invention discloses a data processing method and device, equipment, a medium and a program product, and is applied to the technical field of data processing. The data processing method comprises the steps of obtaining execution plan data corresponding to each level of query statement in a query instruction, wherein the execution plan data comprises an operator and a first input data unit number; according to the operator and the first input data unit number, determining an estimated output data unit number of the operator; according to the first input data unit number, the estimated output data unit number and the single instance processing data unit number corresponding to the operator, the number of execution instances used for executing the query statement is determined. According to the data processing method, the accuracy of the number of the execution examples determined for the whole query statement is improved by considering the data processing amount and the operator characteristics involved in the query statement, and the problem of inaccurate parallelism degree setting caused by neglecting the characteristics of a superior sub-plan is effectively avoided.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application belongs to the field of data processing technology, and in particular, relates to a data processing method, device, equipment, medium and program product. Background Art

[0002] A distributed database system is a database system that stores and manages data on multiple computers, which can be located in different geographical locations and connected through a network. When a distributed database system processes a query instruction input by a user, it first builds an execution plan with a hierarchical structure corresponding to the query instruction. The execution plan includes multiple ordered sub-plans, which record how the distributed database system processes each part of the query instruction to finally obtain the query result.

[0003] In the related art, when determining the concurrency of the upper-level sub-plan, the maximum value of the concurrency in the lower-level sub-plan is usually determined as the concurrency of the upper-level sub-plan. The accuracy of the parallelism set in this way is low, which is easy to cause imbalance in resource utilization and affect the execution efficiency of the query. Summary of the invention

[0004] The embodiments of the present application provide a data processing method, apparatus, device, medium and program product, which can solve the problem of inaccurate parallelism setting.

[0005] In a first aspect, an embodiment of the present application provides a data processing method, the data processing method comprising:

[0006] Obtaining execution plan data corresponding to each level of query statements in the query instruction, the execution plan data including operation operators and the number of first input data units;

[0007] Determining an estimated number of output data units of the operation operator according to the operation operator and the first number of input data units;

[0008] The number of execution instances used to execute the query statement is determined according to the first number of input data units, the estimated number of output data units, and the number of single-instance processed data units corresponding to the operation operator.

[0009] In some possible implementations of the embodiments of the present application, the number of operators is at least two, and the at least two operators include a first operator and a second operator for processing output data of the first operator; the estimated number of output data units includes a first estimated number of output data units of the first operator and a second estimated number of output data units of the second operator. Based on this, determining the estimated number of output data units of the operator according to the operator and the first number of input data units may also include:

[0010] Determine a first estimated number of output data units according to a first output data unit number statistical algorithm corresponding to the first operation operator and the first input data unit number;

[0011] determining the first estimated output data unit number as the second input data unit number of the second operation operator;

[0012] A second estimated number of output data units is determined according to a second output data unit number statistical algorithm corresponding to the second operation operator and the second input data unit number.

[0013] In some possible implementations of the embodiments of the present application, the number of data units processed by a single instance includes the number of data units processed by a single instance corresponding to a first operation operator and the number of data units processed by a single instance corresponding to a second operation operator. Based on this, the number of execution instances used to execute the query statement is determined according to the first number of input data units, the estimated number of output data units, and the number of data units processed by a single instance corresponding to the operation operator, including:

[0014] Determine the first execution instance number of the first operation operator according to the first input data unit number, the first estimated output data unit number and the single instance processing data unit number corresponding to the first operation operator;

[0015] Determine the number of second execution instances of the second operation operator according to the second number of input data units, the second estimated number of output data units, and the number of single-instance processed data units corresponding to the second operation operator;

[0016] The maximum number of the first execution instance number and the second execution instance number is determined as the execution instance number used to execute the query statement.

[0017] In some possible implementations of the embodiments of the present application, the number of operators is one, based on which the estimated number of output data units of the operator is determined according to the operator and the first number of input data units, including: determining the estimated number of output data units of the operator according to the output data unit number statistical algorithm corresponding to the operator and the first number of input data units.

[0018] In some possible implementations of the embodiments of the present application, the number of operation operators is one; determining the number of execution instances for executing the query statement according to the first number of input data units, the estimated number of output data units, and the number of single-instance processed data units corresponding to the operation operator includes:

[0019] Determine the total number of data units to be processed by the operation operator according to the sum of the first input data unit number and the estimated output data unit number;

[0020] Determine the number of execution instances of the operation operator according to the total number of data units to be processed and the number of data units processed by a single instance corresponding to the operation operator;

[0021] The number of execution instances of the operation operator is determined as the number of execution instances used to execute the query statement.

[0022] In some possible implementations of the embodiments of the present application, the operation operator includes a first type of operation operator, the first type of operation operator is used to perform a connection operation on the input data of the first type of operation operator, the input data of the first type of operation operator includes at least two input data tables, and before performing the step of determining the number of execution instances for executing the query statement according to the first input data unit number, the estimated output data unit number, and the single instance processing data unit number corresponding to the operation operator, the data processing method further includes:

[0023] Determine the input data table with the largest number of input data units in at least two input data tables as the target input data table;

[0024] Determine a unit number change state of a first-category operation operator according to the number of input data units and the estimated number of output data units of the target input data table, where the unit number change state includes a unit number increase state or a unit number decrease state;

[0025] According to the association information between the preset unit number change state and the first preset instance processing data unit number, the first preset instance processing data unit number associated with the unit number change state is determined as the single instance processing data unit number corresponding to the operation operator.

[0026] In some possible implementations of the embodiments of the present application, the operation operator includes a second type of operation operator, and the second type of operation operator is used to perform an aggregation operation on the input data of the second type of operation operator. Before performing the step of determining the number of execution instances for executing the query statement according to the first input data unit number, the estimated output data unit number, and the single instance processing data unit number corresponding to the operation operator, the data processing method further includes:

[0027] According to the association information between the preset operator type and the second preset instance processing data unit number, the second preset instance processing data unit number associated with the second type of operation operator is determined as the single instance processing data unit number.

[0028] In a second aspect, an embodiment of the present application provides a data processing device, the data processing device comprising:

[0029] An acquisition module, used to acquire execution plan data corresponding to each level of query statements in the query instruction, the execution plan data including an operation operator and a first input data unit number;

[0030] A first determination module, used to determine the estimated number of output data units of the operation operator according to the operation operator and the first number of input data units;

[0031] The second determination module is used to determine the number of execution instances used to execute the query statement according to the first number of input data units, the estimated number of output data units and the number of single-instance processed data units corresponding to the operation operator.

[0032] In a third aspect, an embodiment of the present application provides an electronic device, comprising: a processor and a memory storing computer program instructions; when the processor executes the computer program instructions, the data processing method as described in any one of the first aspects is implemented.

[0033] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium having computer program instructions stored thereon, and when the computer program instructions are executed by a processor, a data processing method as described in any one of the first aspects is implemented.

[0034] In a fifth aspect, an embodiment of the present application provides a computer program product, which includes a computer program or instructions, and when the computer program or instructions are executed by a processor, they implement the data processing method as described in any one of the first aspects.

[0035] The data processing method, apparatus, device, medium and program product of the embodiments of the present application take into account the manner and efficiency of the operator in processing data by determining the estimated number of output data units through the operator and the first number of input data units, so that the estimated number of output data units obtained in this way can accurately reflect the size of the data after being processed by the operator. Furthermore, the number of data units processed by a single instance of the operator takes into account the processing capability of the execution instance of a single operator under normal load conditions. The number of execution instances obtained based on the coordination of the number of data units processed by a single instance, the first number of input data units and the estimated number of output data units is based on the operator. In this way, by considering the characteristics of the operator itself in the query statement, the accuracy of the number of execution instances determined for the entire query statement is greatly improved. BRIEF DESCRIPTION OF THE DRAWINGS

[0036] In order to more clearly illustrate the technical solution of the embodiments of the present application, the following is a brief introduction to the drawings required for use in the embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0037] Figure 1 A schematic diagram showing a flow chart of a data processing method provided by some embodiments of the present application;

[0038] Figure 2 A schematic flow chart of a method for determining the number of data units processed by a single instance corresponding to an operation operator in a data processing method provided in some embodiments of the present application is shown;

[0039] Figure 3 A flowchart showing a specific implementation of step 130 provided in some embodiments of the present application is shown;

[0040] Figure 4 A flowchart showing a specific implementation of step 120 provided in some embodiments of the present application is shown;

[0041] Figure 5 A flowchart showing another specific implementation of step 130 provided in some embodiments of the present application is shown;

[0042] Figure 6 A schematic diagram showing the structure of a data processing device provided in some embodiments of the present application is shown;

[0043] Figure 7 A schematic diagram of the structure of an electronic device provided in some embodiments of the present application is shown. DETAILED DESCRIPTION

[0044] The features and exemplary embodiments of various aspects of the present application will be described in detail below. In order to make the purpose, technical solutions and advantages of the present application clearer, the present application will be further described in detail below in conjunction with the accompanying drawings and specific embodiments. It should be understood that the specific embodiments described herein are only intended to explain the present application, rather than to limit the present application. For those skilled in the art, the present application can be implemented without the need for some of these specific details. The following description of the embodiments is only to provide a better understanding of the present application by illustrating the examples of the present application.

[0045] It should be noted that, in this article, relational terms such as first and second, etc. are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, the elements defined by the sentence "include..." do not exclude the existence of other identical elements in the process, method, article or device including the elements.

[0046] It should be noted that the acquisition, storage, use and processing of data in the embodiments of the present application are in compliance with the relevant provisions of national laws and regulations.

[0047] It should be noted that in the embodiments of the present application, certain software, components, models and other existing solutions in the industry may be mentioned, and they should be regarded as exemplary. Their purpose is only to illustrate the feasibility of implementing the technical solution of the present application, but it does not mean that the applicant has or will necessarily use the solution.

[0048] Before describing the technical solutions provided by the embodiments of the present application, in order to facilitate the understanding of the embodiments of the present application, the present application first specifically describes the related technologies involved:

[0049] In a distributed database system, the concurrency of an upper-level sub-plan refers to the number of instances of processing the upper-level sub-plan started at the same time when the distributed database system executes the parallel query task of the upper-level sub-plan.

[0050] In a specific implementation scenario, if the query instruction is "select my1.c1 from my1, t2wheremy1.c1=t2.c1 group by my1.c1". The execution logic of this query instruction is to filter out data records that meet the condition that my1.c1=t2.c1 from the two data tables to be queried, namely my1 and t2, and group them according to the field my1.c1. It can be understood that when the join key (Join) and the group key (Group Key) are consistent, in the generated execution plan, the join (join) and aggregate (aggregate) operators will be placed in the same sub-plan. Based on this, the execution plan generated by the above query instruction is specifically as follows.

[0051] 1. Calculation sub-plan

[0052] Hash Aggregate

[0053] Group Key:my1.c1

[0054] Hash Join Hash Cond:(my1.c1=t2.c1)

[0055] 2. Scanning sub-plan

[0056] Scanning sub-plan 1:

[0057] Redistribute Motion Hash Key:my1.c

[0058] Scan on my1

[0059] Scanning sub-plan 2:

[0060] Redistribute Motion Hash Key:t2.c1

[0061] Scan on t2

[0062] In view of the above implementation scenario, in related technologies, the following method is usually used to calculate the degree of parallelism:

[0063] The basis for determining the parallelism of scanning sub-plan 1 and scanning sub-plan 2 is the number of data files of the data table to be scanned, that is, the parallelism of scanning sub-plan 1 is equal to the number of instances corresponding to the number of data files to be scanned by scanning sub-plan 1, and the parallelism of scanning sub-plan 2 is equal to the number of instances corresponding to the number of data files to be scanned by scanning sub-plan 2. As for the parallelism of the calculation sub-plan, its value is directly equal to the larger value of the parallelism of scanning sub-plan 1 and the parallelism of scanning sub-plan 2.

[0064] It is understandable that if there is an upper-level computing sub-plan, then the parallelism of each computing sub-plan can be determined one by one according to the lower-level sub-plan in this bottom-up manner.

[0065] However, this method of directly equating the parallelism of the upper-level sub-plan with the parallelism of the lower-level sub-plan does not take into account the unique factors of the upper-level sub-plan itself, such as differences in data processing volume, operation complexity, and data processing type. Such improper parallelism settings are very likely to cause uneven resource utilization, which in turn has a negative impact on the execution efficiency and overall performance of the query. For example, too high a parallelism may cause small data query tasks to occupy too much system resources, resulting in resource waste; while too low a parallelism will make the processing of large data query tasks abnormally slow.

[0066] In order to solve the above problems in the related art, the present application provides a data processing method, device, equipment, medium and program product. Figure 1 To Attachment Figure 5 , the data processing method provided in the embodiment of the present application is described in detail through specific embodiments and their application scenarios.

[0067] Figure 1 The flowchart of the data processing method provided by some embodiments of the present application is shown. The execution subject of the data processing method can be a distributed database system, such as Figure 1 As shown, the data processing method may include steps 110 to 130.

[0068] Step 110, obtaining execution plan data corresponding to each level of query statements in the query instruction, the execution plan data including operation operators and the number of first input data units.

[0069] Step 120, determining the estimated output data unit number of the operation operator according to the operation operator and the first input data unit number.

[0070] Step 130 , determining the number of execution instances used to execute the query statement according to the first number of input data units, the estimated number of output data units, and the number of single-instance processed data units corresponding to the operation operator.

[0071] Thus, the estimated number of output data units determined by the operator and the first number of input data units takes into account the mode and efficiency of the operator in processing data, so that the estimated number of output data units obtained can accurately reflect the size of the data volume after being processed by the operator. Further, the number of data units processed by a single instance of the operator takes into account the processing capacity of the execution instance of a single operator under normal load conditions. The number of execution instances obtained based on the number of data units processed by a single instance, the first number of input data units and the estimated number of output data units is based on the operator. In this way, by considering the characteristics of the operator itself in the query statement, the accuracy of the number of execution instances determined for the entire query statement is greatly improved. Compared with the method of determining the concurrency of the upper sub-plan only based on the maximum value of the concurrency of the lower sub-plan in the aforementioned related technology, this method fully considers the data processing volume and operator characteristics involved in the query statement itself, effectively avoiding the problem of inaccurate parallelism setting caused by ignoring the characteristics of the upper sub-plan itself, thereby greatly improving the execution efficiency of the query instruction.

[0072] The above steps are described in detail below, as shown below.

[0073] First, step 110 is involved. In an embodiment of the present application, when executing a query instruction input by a user, the distributed database system generates a detailed execution plan, which includes a plurality of ordered sub-plans, each of which undertakes a specific data processing task. Among them, the query instruction covers multiple levels of query statements, each level of query statement corresponds to a sub-plan, and has corresponding execution plan data. Among them, the operation operator is the basic operation unit in the sub-plan corresponding to the query statement, which is used to execute the data processing task corresponding to the query statement. The first number of input data units refers to the number of initial input data that the operation operator needs to process, such as the number of input data rows.

[0074] Exemplarily, the operation operator may include any of the following scanning operators, connection operators, and aggregation operators. Correspondingly, the data processing task corresponding to the query statement may include any of the following: scanning a data table, connecting a data table, filtering data, and aggregating data.

[0075] Secondly, with respect to step 120, the estimated number of output data units is based on the type of the operation operator and the number of input data units, and is the estimated number of output data units after being processed by the operator.

[0076] Then, in step 130, the number of data units processed by a single instance refers to the upper limit of the number of data units that can be processed at one time by an execution instance, such as a database connection or a thread. Exemplarily, according to the estimated number of output data units corresponding to each level of query statements, the number of first input data units, and the number of data units processed by a single instance corresponding to the operation operator, the distributed database system calculates the minimum number of execution instances required to execute each level of query statements.

[0077] In some embodiments of the present application, in order to accurately obtain the number of single-instance processing data units corresponding to an operator, the number of single-instance processing data units corresponding to the operator can be obtained according to the operator type of the operator. The operator may include a first type of operator, and the first type of operator is used to perform a join operation on the input data of the first type of operator, and the input data of the first type of operator includes at least two input data tables. Exemplarily, the first type of operator may be a hash join, which is an operator used to perform a join operation, and is used to join and merge two or more input data tables according to specific rules, i.e., based on a hash function to perform hash calculations on the join keys, to generate a new output data table. For example, in a distributed database system, when it is necessary to associate and query data from different data tables but with an associated relationship, hash join can efficiently combine data rows that meet the join conditions. Based on this, if Figure 2 As shown, before the above step 130, the data processing method may further include steps 210 to 230.

[0078] Step 210: determine the input data table with the largest number of input data units among the at least two input data tables as the target input data table.

[0079] Exemplarily, the number of data units in the input data of the first type of operation operator hash join, i.e., at least two input data tables, can be counted through the built-in statistical function or query tool of the distributed database system, and then the number of data units of each of the at least two input data tables can be compared to find the input data table with the largest number of data units and determine it as the target input data table.

[0080] Step 220, determining the unit number change state of the first type of operation operator according to the input data unit number of the target input data table and the estimated output data unit number, the unit number change state including the unit number increase state or the unit number decrease state.

[0081] The unit number change state refers to the change trend of the estimated output data unit number obtained after processing by the first type of operation operator compared with the input data unit number of the target input data table. The unit number increase state means that after processing by the first type of operation operator, the estimated output data unit number is more than the input data unit number of the target input data table; the unit number decrease state means that the estimated output data unit number is less than the input data unit number of the target input data table.

[0082] Exemplarily, after the target input data table is determined, the number of input data units of the target input data table is compared with the estimated number of output data units. If the estimated number of output data units is greater than the number of input data units of the target input data table, then the unit number change state of the first type of operation operator is determined to be a unit number increase state; conversely, if the estimated number of output data units is less than the number of input data units of the target input data table, then its unit number change state is determined to be a unit number decrease state.

[0083] Step 230: According to the association information between the preset unit number change state and the first preset instance processing data unit number, the first preset instance processing data unit number associated with the unit number change state is determined as the single instance processing data unit number corresponding to the operation operator.

[0084] Exemplarily, association information between a preset unit number change state and a first preset instance processing data unit number can be established in advance in a distributed database system. After the unit number change state is determined through step 220, the first preset instance processing data unit number corresponding to the currently determined unit number change state can be searched and obtained from the stored association information based on this association information, and then determined as the single instance processing data unit number corresponding to the first type of operation operator hash join.

[0085] Therefore, the impact trend of the first type of operator on the data volume, that is, the unit number change state, can be determined through the target input data table with the relatively largest data volume. Then, based on the preset information associated with this unit number change state, that is, the number of data units processed by the first preset instance, the number of data units processed by the single instance corresponding to the operator can be accurately determined, which helps to better adapt to the data processing needs in different data volume change scenarios.

[0086] In some embodiments of the present application, before the above step 210, the data processing method may further include determining the number of data units processed by the first preset instance corresponding to the change state of the preset number of units. Specifically, for the first type of operation operator, taking hash join as an example, the following test is carried out:

[0087] Test scenario 1: In a single-node environment, after executing the hash join operation, the number of output data units increases.

[0088] In this test scenario, the purpose is to obtain the total computing amount value corresponding to the time when the performance of a single hash join execution instance shows a significant downward trend when the total computing amount gradually increases, and determine this value as the single instance standard computing amount corresponding to the state of increasing unit number. Among them, the total computing amount refers to the sum of the number of input data units of all first-class operators as input and the number of output data units generated by the hash join operation.

[0089] Test scenario 2: In a single-node environment, after executing the hash join operation, the number of output data units decreases.

[0090] In this test scenario, the test purpose is also to find the total computing value corresponding to the time when the performance of a single hash join execution instance shows a significant downward trend during the continuous increase of the total computing amount, and to determine this value as the single instance standard computing amount corresponding to the state of reduced unit number. The total computing amount refers to the sum of the number of input data units of all first-class operators as input and the number of output data units generated by the hash join operation.

[0091] It is worth noting that the above-mentioned single-node environment specifically refers to the execution environment of a sub-plan corresponding to a query statement containing a hash join operator. In this environment, all data processing operations related to the sub-plan, including but not limited to hash join operations and associated input data reading, intermediate result storage and other processes, are completed on the same computing node, and there is no distributed processing across nodes.

[0092] In some embodiments of the present application, the operator includes a second type of operator, and the second type of operator is used to perform an aggregation operation on the input data of the second type of operator. Exemplarily, the second type of operator can be HashAggregate, which is an operator for performing an aggregation operation, and is used to perform an aggregation task on the input data based on a hash algorithm. Based on this, before the above step 130, the data processing method may also include, according to the association information of the preset operator type and the second preset instance processing data unit number, determining the second preset instance processing data unit number associated with the second type of operator as the single instance processing data unit number.

[0093] Exemplarily, association information between a preset operator type and a second preset instance processing data unit number can be established in advance in the distributed database system. When the second type of operator is identified, the second preset instance processing data unit number associated with the second type of operator is searched, and the obtained second preset instance processing data unit number associated with the second type of operator is determined as the single instance processing data unit number corresponding to the operator.

[0094] Therefore, by determining the number of data units processed by a single instance corresponding to the second type of operator based on the preset association information, the number of data units processed by each instance of the second type of operator can be accurately configured according to the characteristics of the aggregation operation of the second type of operator and the actual data processing environment requirements. In this way, the second type of operator can better adapt to different data volumes when performing aggregation operations, avoiding problems such as memory overflow caused by processing too much data, or waste of resources and low efficiency caused by processing too little data.

[0095] In some embodiments of the present application, the data processing method may further include determining the number of second preset instance processing data units corresponding to the second type of operation operator. Specifically, for the second type of operation operator, taking hash Aggregate as an example, the following test is performed:

[0096] When running a hash Aggregate operation in a single-node environment, by monitoring the performance changes of a single hash Aggregate execution instance, the total computational value corresponding to the moment when its performance shows a significant downward trend is determined as the number of data units processed by the second preset instance corresponding to the second type of operator. The total computational value refers to the sum of the number of input data units of all second type operators as input and the number of output data units generated by the hash Aggregate operation.

[0097] In some embodiments of the present application, the number of operation operators may be one. Based on this, the above step 120 may specifically include determining the estimated number of output data units of the operation operator according to the output data unit number statistical algorithm corresponding to the operation operator and the first input data unit number.

[0098] Among them, the output data unit number statistical algorithm is an algorithm that takes into account multiple factors such as the internal logic of the operator, data distribution characteristics, historical statistical information, etc. It is used to estimate the output data unit number of the operator based on the type of the operator and the number of input data units. Different operators can correspond to different output data unit number statistical algorithms.

[0099] Exemplarily, for a scan operator, when there is no filtering condition in the scan operator, the output data unit number statistical algorithm corresponding to the scan operator can be expressed as estimating that the output data unit number is equal to the first input data unit number.

[0100] Therefore, by using the output data unit number statistics algorithm corresponding to the operator, the distributed database system can more accurately estimate the size of the query statement execution result, that is, the estimated output data unit number. In addition, the output data unit number statistics algorithm corresponding to the operator can adapt to the data distribution characteristics and query mode of the operator, thereby providing a more stable and reliable estimation result.

[0101] When the number of operators is one, such as Figure 3 As shown, the above step 130 may specifically include steps 1301a to 1303a.

[0102] Step 1301a, determining the total number of data units to be processed by the operator according to the sum of the first input data unit number and the estimated output data unit number.

[0103] Exemplarily, when there is only one operator, the data volume of the input data of the operator, i.e., the first input data unit number, and the data volume of the output data estimated to be obtained after processing by the operator, i.e., the estimated output data unit number, are comprehensively considered. The total value obtained by adding the data unit numbers of these two parts represents the total data volume that needs to be processed by this operator.

[0104] Step 1302a, determining the number of execution instances of the operation operator according to the total number of data units to be processed and the number of data units processed by a single instance corresponding to the operation operator.

[0105] Exemplarily, after obtaining the total number of data units to be processed and the number of single-instance processing data units corresponding to the operation operator, a division algorithm is used to calculate the number of execution instances of the operation operator, that is, the total number of data units to be processed is divided by the number of single-instance processing data units corresponding to the operation operator, and the quotient obtained is determined as the number of execution instances of the operation operator.

[0106] Step 1303a, determining the number of execution instances of the operation operator as the number of execution instances used to execute the query statement.

[0107] Therefore, the number of execution instances calculated by the size of the data volume that the operator needs to process and the number of data units processed by the corresponding single instance of the operator can reasonably allocate execution resources according to the data processing capacity of the operator itself and the total amount of data actually to be processed, so that the data processing task can be completed efficiently in parallel through an appropriate number of instances, avoiding the waste of resources due to too many execution instances, or the inefficiency of data processing due to too few execution instances, and the inability to complete the task on time. In this way, the accurate number of execution instances is set for the data processing task of the query statement, ensuring that the data processing flow involved in the entire query instruction can effectively allocate resources and control task execution, and improve the efficiency and accuracy of the entire data processing of the query instruction.

[0108] In some embodiments of the present application, the number of operators is at least two, and the at least two operators include a first operator and a second operator for processing the output data of the first operator. The estimated number of output data units includes a first estimated number of output data units of the first operator and a second estimated number of output data units of the second operator. Among them, there is a sequential data processing association relationship between the first operator and the second operator, that is, the second operator is used to process the data output by the first operator. For example, the first operator can be an operation for filtering the original data, and the second operator can be an operation for further processing such as aggregation statistics on the filtered data. The first estimated number of output data units and the second estimated number of output data units respectively represent the number of units of output data that are expected to be obtained by the first operator and the second operator after processing the corresponding input data. Based on this, if Figure 4 As shown, the above step 120 may specifically include steps 1201 to 1203.

[0109] Step 1201, determining a first estimated number of output data units according to a first output data unit number statistical algorithm corresponding to a first operation operator and a first input data unit number.

[0110] The first output data unit number statistical algorithm refers to an algorithm for calculating the number of output data units of the first operation operator determined based on the operation rules of the first operation operator itself, the characteristics of the input data, and other factors. For example, for the first operation operator that performs a data grouping operation, its first output data unit number statistical algorithm can be to estimate the number of output data units based on the number of groups and the average distribution of data in each group.

[0111] Exemplarily, the first input data unit number is processed by a first output data unit number statistical algorithm to obtain a first estimated output data unit number.

[0112] Step 1202: determine the first estimated number of output data units as the second number of input data units of the second operator.

[0113] Step 1203: Determine a second estimated number of output data units according to a second output data unit number statistical algorithm corresponding to the second operation operator and the second input data unit number.

[0114] The second output data unit number statistical algorithm refers to an algorithm for calculating the output data unit number of the second operator determined based on the operation rules of the second operator itself, the characteristics of the input data, and other factors. For example, for the second operator that performs a sum aggregation operation, its second output data unit number statistical algorithm can be to estimate the output data unit number based on the number of fields participating in the aggregation, the degree of discreteness of the data, and the like.

[0115] Exemplarily, the second input data unit number is processed by a second output data unit number statistical algorithm to obtain a second estimated output data unit number.

[0116] Therefore, by using the output data unit number statistical algorithm corresponding to the operator combined with the actual number of input data units, the data volume after processing by the first operator, that is, the first estimated output data unit number, can be accurately determined. Furthermore, an association in data volume between the two operators, that is, the first operator and the second operator, is established, so that the second input data unit number of the second operator can be determined based on the output of the first operator, and then the data volume after processing by the second operator, that is, the second estimated output data unit number, can be more accurately estimated.

[0117] In some embodiments of the present application, the number of single-instance processed data units includes the number of single-instance processed data units corresponding to the first operation operator and the number of single-instance processed data units corresponding to the second operation operator. Figure 5 As shown, the above step 130 includes steps 1301b to 1303b.

[0118] Step 1301b: Determine the first execution instance number of the first operation operator according to the first input data unit number, the first estimated output data unit number and the single instance processing data unit number corresponding to the first operation operator.

[0119] The first execution instance number refers to the number of specific execution instances that need to be allocated to the first operator in order to complete the data processing task of the first operator. The number of single-instance processing data units corresponding to the first operator can be determined according to the operator type of the first operator and the above step of determining the corresponding single-instance processing data unit number based on the type of the operator.

[0120] Exemplarily, step 1301b may specifically include:

[0121] Obtaining a total number of data units to be processed by the first operation operator by adding the first number of input data units and the first estimated number of output data units;

[0122] The total number of data units to be processed by the first operator is divided by the number of data units processed by a single instance corresponding to the first operator to obtain the number of first execution instances of the first operator.

[0123] Step 1302b: Determine the second execution instance number of the second operation operator according to the second input data unit number, the second estimated output data unit number and the single instance processing data unit number corresponding to the second operation operator.

[0124] The second execution instance number refers to the number of specific execution instances that need to be allocated to the second operator in order to complete the data processing task of the second operator. The number of single-instance processing data units corresponding to the second operator can be determined according to the operator type of the second operator and the above step of determining the corresponding single-instance processing data unit number based on the type of the operator.

[0125] Exemplarily, step 1302b may specifically include:

[0126] Obtaining a total number of data units to be processed by the second operator by adding the second number of input data units and the second estimated number of output data units;

[0127] The total number of data units to be processed by the second operator is divided by the number of data units processed by a single instance corresponding to the second operator to obtain the number of second execution instances of the second operator.

[0128] Step 1303b: determine the maximum number of the first execution instance number and the second execution instance number as the execution instance number used to execute the query statement.

[0129] Therefore, by selecting the maximum value of the first execution instance number and the second execution instance number as the execution instance number for executing the query statement, it can be ensured that in the entire data processing process, both the first operator and the second operator have sufficient execution instances to process their respective data, avoiding data processing jams or incompleteness due to insufficient number of execution instances, and overall ensuring that the data processing tasks involved in the query statement can be executed smoothly and efficiently, and resources can be used reasonably.

[0130] Based on the data processing method provided in the above embodiment, the present application also provides a specific implementation of a data processing device. Please refer to the following embodiment.

[0131] See also Figure 6 , the data processing device 600 provided in the embodiment of the present application includes:

[0132] An acquisition module 610 is used to acquire execution plan data corresponding to each level of query statements in the query instruction, where the execution plan data includes an operation operator and a first input data unit number;

[0133] A first determination module 620, configured to determine an estimated number of output data units of the operation operator according to the operation operator and the first number of input data units;

[0134] The second determination module 630 is used to determine the number of execution instances used to execute the query statement according to the first number of input data units, the estimated number of output data units and the number of single-instance processed data units corresponding to the operation operator.

[0135] Therefore, the estimated number of output data units determined by the first determination module 620 through processing the operator and the first input data unit number obtained by the acquisition module 610 takes into account the way and efficiency of the operator processing data, so that the estimated number of output data units obtained can accurately reflect the size of the data after being processed by the operator. Furthermore, the number of data units processed by a single instance of the operator takes into account the processing capacity of the execution instance of a single operator under normal load conditions. The number of execution instances obtained by the second determination module 630 through the coordinated processing of the single instance processing data unit number, the first input data unit number and the estimated output data unit number is based on the operator. In this way, by considering the characteristics of the operator itself in the query statement, the accuracy of the number of execution instances determined for the entire query statement is greatly improved. Compared with the aforementioned related technology that determines the concurrency of the upper-level sub-plan only based on the maximum concurrency of the lower-level sub-plan, this method fully considers the data processing volume and operation operator characteristics involved in the query statement itself, and effectively avoids the problem of inaccurate parallelism setting caused by ignoring the characteristics of the upper-level sub-plan itself, thereby greatly improving the execution efficiency of the query instructions.

[0136] In some embodiments of the present application, the first determining module 620 may be specifically used for:

[0137] In a case where the number of operators is at least two, the at least two operators include a first operator and a second operator for processing output data of the first operator, and the estimated number of output data units includes a first estimated number of output data units of the first operator and a second estimated number of output data units of the second operator, determining the first estimated number of output data units according to a first output data unit number statistical algorithm corresponding to the first operator and the first input data unit number;

[0138] determining the first estimated output data unit number as the second input data unit number of the second operation operator;

[0139] A second estimated number of output data units is determined according to a second output data unit number statistical algorithm corresponding to the second operation operator and the second input data unit number.

[0140] In some embodiments of the present application, the second determination module 630 may be specifically used to:

[0141] In a case where the number of single-instance processed data units includes the number of single-instance processed data units corresponding to the first operation operator and the number of single-instance processed data units corresponding to the second operation operator, determining the first execution instance number of the first operation operator according to the first input data unit number, the first estimated output data unit number, and the number of single-instance processed data units corresponding to the first operation operator;

[0142] Determine the number of second execution instances of the second operation operator according to the second number of input data units, the second estimated number of output data units, and the number of single-instance processed data units corresponding to the second operation operator;

[0143] The maximum number of the first execution instance number and the second execution instance number is determined as the execution instance number used to execute the query statement.

[0144] In some embodiments of the present application, the first determining module 620 may also be used to:

[0145] When the number of operation operators is one, the estimated number of output data units of the operation operator is determined according to the output data unit number statistical algorithm corresponding to the operation operator and the first input data unit number.

[0146] In some embodiments of the present application, the second determining module 630 may also be used to:

[0147] When the number of the operation operator is one, determining the total number of data units to be processed by the operation operator according to the sum of the first number of input data units and the estimated number of output data units;

[0148] Determine the number of execution instances of the operation operator according to the total number of data units to be processed and the number of data units processed by a single instance corresponding to the operation operator;

[0149] The number of execution instances of the operation operator is determined as the number of execution instances used to execute the query statement.

[0150] In some embodiments of the present application, the data processing device may further include:

[0151] A third determination module is used to determine the input data table with the largest number of input data units in the at least two input data tables as the target input data table before determining the number of execution instances for executing the query statement based on the first input data unit number, the estimated output data unit number, and the single-instance processing data unit number corresponding to the operation operator when the operation operator includes the first type of operation operator, the first type of operation operator is used to perform a connection operation on the input data of the first type of operation operator, the input data of the first type of operation operator includes at least two input data tables, and the number of execution instances for executing the query statement is determined according to the first input data unit number, the estimated output data unit number, and the single-instance processing data unit number corresponding to the operation operator;

[0152] A fourth determination module is used to determine a unit number change state of the first type of operation operator according to the number of input data units and the estimated number of output data units of the target input data table, where the unit number change state includes a unit number increase state or a unit number decrease state;

[0153] The fifth determination module is used to determine the first preset instance processing data unit number associated with the unit number change state as the single instance processing data unit number corresponding to the operation operator according to the association information between the preset unit number change state and the first preset instance processing data unit number.

[0154] In some embodiments of the present application, the data processing device may further include:

[0155] The sixth determination module is used to determine the second preset instance processing data unit number associated with the second type of operation operator as the single instance processing data unit number according to the association information between the preset operator type and the second preset instance processing data unit number before the operation operator includes the second type of operation operator, the second type of operation operator is used to perform aggregation operations on the input data of the second type of operation operator and the number of execution instances used to execute the query statement is determined based on the first input data unit number, the estimated output data unit number and the single instance processing data unit number corresponding to the operation operator.

[0156] Each module of the data processing device 600 provided in the embodiment of the present application can realize Figures 1 to 5 The functions of each step of the data processing method are provided, and the corresponding technical effects can be achieved. For the sake of concise description, they will not be repeated here.

[0157] Figure 7 A schematic diagram of the hardware structure of an electronic device provided in some embodiments of the present application is shown.

[0158] The electronic device may include a processor 701 and a memory 702 storing computer program instructions.

[0159] Specifically, the processor 701 may include a central processing unit (CPU), or an application specific integrated circuit (ASIC), or may be configured to implement one or more integrated circuits of the embodiments of the present application.

[0160] The memory 702 may include a large capacity memory for data or instructions. By way of example and not limitation, the memory 702 may include a hard disk drive (HDD), a floppy disk drive, a flash memory, an optical disk, a magneto-optical disk, a magnetic tape, or a universal serial bus (USB) drive or a combination of two or more of these. In appropriate cases, the memory 702 may include a removable or non-removable (or fixed) medium. In appropriate cases, the memory 702 may be inside or outside the integrated gateway disaster recovery device. In a specific embodiment, the memory 702 is a non-volatile solid-state memory.

[0161] In a particular embodiment, the memory 702 may include a read-only memory (ROM), a random access memory (RAM), a magnetic disk storage medium device, an optical storage medium device, a flash memory device, an electrical, optical or other physical / tangible memory storage device. Therefore, generally, the memory 702 includes one or more tangible (non-transitory) computer-readable storage media (e.g., a memory device) encoded with software including computer-executable instructions, and when the software is executed (e.g., by one or more processors), it is operable to perform the operations described with reference to the data processing method according to the first aspect of the present application.

[0162] The processor 701 implements any one of the data processing methods in the above embodiments by reading and executing computer program instructions stored in the memory 702 .

[0163] In one example, the electronic device may further include a communication interface 703 and a bus 710. Figure 7 As shown, the processor 701, the memory 702, and the communication interface 703 are connected via a bus 710 and communicate with each other.

[0164] The communication interface 703 is mainly used to implement communication between various modules, devices, units and / or equipment in the embodiments of the present application.

[0165] Bus 710 includes hardware, software or both, and the parts of electronic equipment are coupled to each other.For example, but not limitation, bus may include accelerated graphics port (AGP) or other graphics bus, enhanced industrial standard architecture (EISA) bus, front side bus (FSB), hypertransport (HT) interconnection, industrial standard architecture (ISA) bus, infinite bandwidth interconnection, low pin count (LPC) bus, memory bus, micro channel architecture (MCA) bus, peripheral component interconnection (PCI) bus, PCI-Express (PCI-X) bus, serial advanced technology attachment (SATA) bus, video electronics standard association local (VLB) bus or other suitable bus or two or more of these combinations. In appropriate cases, bus 710 may include one or more buses. Although the present application embodiment describes and shows a specific bus, the application considers any suitable bus or interconnection.

[0166] The electronic device can execute the data processing method in the embodiment of the present application, thereby realizing the combination Figures 1 to 6 The data processing method and device described.

[0167] In addition, in combination with the data processing method in the above embodiment, the embodiment of the present application may provide a computer-readable storage medium for implementation. The computer-readable storage medium stores computer program instructions; when the computer program instructions are executed by the processor, any one of the data processing methods in the above embodiment is implemented. Examples of computer-readable storage media include non-transitory computer-readable storage media, such as portable disks, hard disks, random access memories (RAM), read-only memories (ROM), erasable programmable read-only memories (EPROM or flash memory), portable compact disk read-only memories (CD-ROM), optical storage devices, magnetic storage devices, etc.

[0168] In addition, in combination with the data processing method in the above embodiment, the embodiment of the present application may provide a computer program product for implementation. The program product is stored in a storage medium, and may specifically include a computer program or instruction, which implements any of the data processing methods in the above embodiment when executed by a processor. The program product is executed by at least one processor to implement the various processes of the above data processing method embodiment, and can achieve the same technical effect, so it will not be repeated here to avoid repetition.

[0169] It should be clear that the present application is not limited to the specific configuration and processing described above and shown in the figures. For the sake of simplicity, a detailed description of the known method is omitted here. In the above embodiments, several specific steps are described and shown as examples. However, the method process of the present application is not limited to the specific steps described and shown, and those skilled in the art can make various changes, modifications and additions, or change the order between the steps after understanding the spirit of the present application.

[0170] The functional blocks shown in the above-described block diagram can be implemented as hardware, software, firmware or a combination thereof. When implemented in hardware, it can be, for example, an electronic circuit, an application specific integrated circuit (ASIC), appropriate firmware, a plug-in, a function card, etc. When implemented in software, the elements of the present application are programs or code segments that are used to perform the required tasks. The program or code segment can be stored in a machine-readable medium, or transmitted on a transmission medium or a communication link by a data signal carried in a carrier wave. "Machine-readable medium" can include any medium capable of storing or transmitting information. Examples of machine-readable media include electronic circuits, semiconductor memory devices, ROM, flash memory, erasable ROM (EROM), floppy disks, CD-ROMs, optical disks, hard disks, optical fiber media, radio frequency (RF) links, etc. The code segment can be downloaded via a computer network such as the Internet, an intranet, etc.

[0171] It should also be noted that the exemplary embodiments mentioned in this application describe some methods or systems based on a series of steps or devices. However, this application is not limited to the order of the above steps, that is, the steps can be performed in the order mentioned in the embodiment, or in a different order from the embodiment, or several steps can be performed simultaneously.

[0172] Aspects of the present disclosure are described above with reference to the flowchart and / or block diagram of the method, device (system) and computer program product according to the embodiment of the present disclosure. It should be understood that each box in the flowchart and / or block diagram and the combination of each box in the flowchart and / or block diagram can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device to produce a machine so that these instructions executed by the processor of the computer or other programmable data processing device enable the implementation of the function / action specified in one or more boxes of the flowchart and / or block diagram. Such a processor can be, but is not limited to, a general-purpose processor, a special-purpose processor, a special application processor, or a field programmable logic circuit. It can also be understood that each box in the block diagram and / or flowchart and the combination of boxes in the block diagram and / or flowchart can also be implemented by dedicated hardware that performs a specified function or action, or can be implemented by a combination of dedicated hardware and computer instructions.

[0173] The above is only a specific implementation of the present application. Those skilled in the art can clearly understand that for the convenience and simplicity of description, the specific working processes of the systems, modules and units described above can refer to the corresponding processes in the aforementioned method embodiments, and will not be repeated here. It should be understood that the protection scope of the present application is not limited to this. Any technician familiar with the technical field can easily think of various equivalent modifications or replacements within the technical scope disclosed in this application, and these modifications or replacements should be included in the protection scope of this application.

Claims

1. A data processing method, characterized in that: include: Obtaining execution plan data corresponding to each level of query statements in the query instruction, wherein the execution plan data includes an operation operator and a first input data unit number; Determining an estimated number of output data units of the operation operator according to the operation operator and the first number of input data units; The number of execution instances used to execute the query statement is determined according to the first number of input data units, the estimated number of output data units, and the number of single-instance processed data units corresponding to the operation operator.

2. The method according to claim 1, characterized in that The number of the operation operators is at least two, and the at least two operation operators include a first operation operator and a second operation operator for processing output data of the first operation operator; the estimated number of output data units includes a first estimated number of output data units of the first operation operator and a second estimated number of output data units of the second operation operator; The step of determining the estimated number of output data units of the operation operator according to the operation operator and the first number of input data units includes: Determine the first estimated number of output data units according to a first output data unit number statistical algorithm corresponding to the first operation operator and the first number of input data units; Determining the first estimated number of output data units as the second number of input data units of the second operation operator; The second estimated number of output data units is determined according to a second output data unit number statistical algorithm corresponding to the second operation operator and the second input data unit number.

3. The method according to claim 2, characterized in that The number of data units processed by a single instance includes the number of data units processed by a single instance corresponding to the first operation operator and the number of data units processed by a single instance corresponding to the second operation operator; The determining the number of execution instances for executing the query statement according to the first number of input data units, the estimated number of output data units, and the number of single-instance processed data units corresponding to the operation operator includes: Determining a first execution instance number of the first operation operator according to the first input data unit number, the first estimated output data unit number, and the single-instance processing data unit number corresponding to the first operation operator; Determine the number of second execution instances of the second operation operator according to the second number of input data units, the second estimated number of output data units, and the number of single-instance processed data units corresponding to the second operation operator; The maximum number of the first execution instance number and the second execution instance number is determined as the execution instance number used to execute the query statement.

4. The method according to claim 1, characterized in that: The number of the operation operator is one; and determining the estimated number of output data units of the operation operator according to the operation operator and the first number of input data units includes: Determine the estimated number of output data units of the operation operator according to the output data unit number statistical algorithm corresponding to the operation operator and the first number of input data units.

5. The method according to claim 4, characterized in that The number of the operation operators is one; and determining the number of execution instances for executing the query statement according to the first number of input data units, the estimated number of output data units, and the number of single-instance processed data units corresponding to the operation operator includes: Determining the total number of data units to be processed by the operation operator according to the sum of the first number of input data units and the estimated number of output data units; Determine the number of execution instances of the operation operator according to the total number of data units to be processed and the number of single-instance processing data units corresponding to the operation operator; The number of execution instances of the operation operator is determined as the number of execution instances used to execute the query statement.

6. The method according to any one of claims 1, 4 to 5, characterized in that: The operation operator includes a first type of operation operator, the first type of operation operator is used to perform a connection operation on input data of the first type of operation operator, and the input data of the first type of operation operator includes at least two input data tables; Before determining the number of execution instances for executing the query statement according to the first number of input data units, the estimated number of output data units, and the number of single-instance processed data units corresponding to the operation operator, the method further includes: Determine the input data table with the largest number of input data units in the at least two input data tables as the target input data table; Determine a unit number change state of the first type of operation operator according to the number of input data units of the target input data table and the estimated number of output data units, wherein the unit number change state includes a unit number increase state or a unit number decrease state; According to the association information between the preset unit number change state and the first preset instance processing data unit number, the first preset instance processing data unit number associated with the unit number change state is determined as the single instance processing data unit number corresponding to the operation operator.

7. The method according to any one of claims 1, 4 to 5, characterized in that: The operation operators include second-type operation operators, and the second-type operation operators are used to perform aggregation operations on input data of the second-type operation operators; Before determining the number of execution instances for executing the query statement according to the first number of input data units, the estimated number of output data units, and the number of single-instance processed data units corresponding to the operation operator, the method further includes: According to the association information between the preset operator type and the second preset instance processing data unit number, the second preset instance processing data unit number associated with the second type of operation operator is determined as the single instance processing data unit number.

8. A data processing device, characterized in that: include: An acquisition module, used to acquire execution plan data corresponding to each level of query statements in the query instruction, wherein the execution plan data includes an operation operator and a first input data unit number; A first determining module, configured to determine an estimated number of output data units of the operation operator according to the operation operator and the first number of input data units; The second determination module is used to determine the number of execution instances used to execute the query statement according to the first number of input data units, the estimated number of output data units and the number of single-instance processing data units corresponding to the operation operator.

9. An electronic device, characterized in that: The device comprises: a processor and a memory storing computer program instructions; When the processor executes the computer program instructions, the data processing method according to any one of claims 1 to 7 is implemented.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer program instructions, and when the computer program instructions are executed by a processor, the data processing method according to any one of claims 1 to 7 is implemented.

11. A computer program product, characterized in that The invention comprises a computer program, wherein when the computer program is executed, the data processing method according to any one of claims 1 to 7 is implemented.