A data processing method and apparatus
By using adaptive data sampling and sharding, the problem of data bloat in big data processing is solved, achieving more efficient data sharding and computing performance.
Patent Information
- Application Number
- CN202110871021.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-07-30
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2041-07-30
AI Technical Summary
Existing technologies cannot effectively address the data bloat problem in big data processing, especially in distributed parallel data processing. Uneven data sharding causes a few tasks to process at a speed far below the average, slowing down the entire computing process.
An adaptive sampling method is used to sample the data table to obtain the sample data distribution. Based on the sample data distribution, the output data volume corresponding to the join operation of the shard is calculated, thereby selecting the target shard and performing data sharding.
By using adaptive data sampling and sharding, the target shards that need to be sharded again are accurately selected, avoiding data bloat, improving the efficiency and uniformity of data processing, and reducing computation time.
Smart Images

Figure CN113590322B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of big data technology, and in particular, to a data processing method and apparatus. Background Art
[0002] Data skew has always been a major pain point in the field of big data. It refers to the situation in the distributed data parallel processing process that due to uneven data sharding, a large amount of data is concentrated on one or a few computing nodes, resulting in the processing speed of a few tasks being much lower than the average speed, which slows down the entire computing process. One common data skew situation is called data explosion, which is manifested as follows: the size of the input shard is not significantly larger than the average level, but the corresponding data only contains a few different keys. After the Join operation is executed, the amount of data output by the corresponding task is much larger than that of other tasks.
[0003] In the process of implementing the present invention, the inventors found that there are at least the following problems in the prior art:
[0004] The existing computing engines are divided based on the size of the input shards and cannot handle the scenario of data explosion. Moreover, in the actual production environment, the situation is more complex, and multiple data skew scenarios (including data explosion) often occur together, while the existing technical solutions cannot handle the situation including data explosion. Summary of the Invention
[0005] In view of this, embodiments of the present invention provide a data processing method and apparatus to solve the technical problem of being unable to handle data explosion.
[0006] To achieve the above object, according to one aspect of the embodiments of the present invention, a data processing method is provided, including:
[0007] Sampling the data to be processed in the data table in an adaptive manner to obtain the sample data distribution of the data table;
[0008] Dividing the data to be processed into multiple shards, and respectively calculating the amount of output data corresponding to the join operations of the multiple shards according to the sample data distribution of the data table;
[0009] According to the amount of output data corresponding to the join operations of the multiple shards, screening out target shards from the multiple shards, and performing data sharding on the target shards.
[0010] Optionally, sampling the data to be processed in the data table in an adaptive manner to obtain the sample data distribution of the data table, including:
[0011] Obtaining the data to be processed from the data table to obtain the keys in the data to be processed;
[0012] Determine whether the cumulative number of processed items is greater than or equal to the first quantity threshold; if not, write the key in the data to be processed into the count table, and use the key in the count table as the sampling result; if so, use the reservoir sampling algorithm to sample the data to be processed;
[0013] Obtain the sample data distribution of the data table based on the sampling result of the data to be processed.
[0014] Optionally, using the reservoir sampling algorithm to sample the data to be processed includes:
[0015] Determine whether the cumulative number of processed items is greater than or equal to the second quantity threshold; if not, write the key in the data to be processed into the sampling array; if so, use the reservoir sampling algorithm to replace one key in the sampling array with the key in the data to be processed;
[0016] Use the key in the sampling array as the sampling result.
[0017] Optionally, obtaining the sample data distribution of the data table based on the sampling result of the data to be processed includes:
[0018] Calculate the cumulative sampling amount of each sample respectively; wherein, the sampling result of the data to be processed includes multiple samples;
[0019] For each data table, divide the data volume of the data to be processed by the cumulative sampling amount of each sample to obtain the weight of each sample in the data table, thereby obtaining the sample data distribution of the data table.
[0020] Optionally, calculating the output data volume corresponding to the join operation of the multiple shards based on the sample data distribution of the data table includes:
[0021] For any one of the multiple shards, filter out the target samples assigned to the shard from the sampling result; wherein each data table contains the target samples;
[0022] Calculate the output data volume corresponding to the join operation of the shard according to the weights of the respective target samples in the data table.
[0023] Optionally, calculating the output data volume corresponding to the join operation of the shard according to the weights of the respective target samples in the data table includes:
[0024] Multiply the weights of each target sample in each data table and accumulate the products to obtain the output data volume corresponding to the join operation of the shard.
[0025] Optionally, multiply the weights of each target sample in each data table and accumulate the products to obtain the output data volume corresponding to the join operation of the shard, including:
[0026] If the type of the join operation is an inner join, multiply the weights of each target sample in each data table and accumulate the products to obtain the output data volume corresponding to the join operation of the shard;
[0027] If the type of the join operation is a left join, multiply the weights of each target sample in each data table and accumulate the products, and then add the weights of each target sample in the left join data table to obtain the output data volume corresponding to the join operation of the shard;
[0028] If the type of the join operation is a right join, multiply the weights of each target sample in each data table and accumulate the products, and then add the weights of each target sample in the right join data table to obtain the output data volume corresponding to the join operation of the shard;
[0029] If the type of the join operation is a full join, multiply the weights of each target sample in each data table and accumulate the products, and then add the weights of each target sample in the left join data table and the right join data table to obtain the output data volume corresponding to the join operation of the shard.
[0030] Optionally, according to the output data volumes corresponding to the join operations of the multiple shards, screen out target shards from the multiple shards, and perform data sharding on the target shards, including:
[0031] Sort the output data volumes corresponding to the join operations of the multiple shards, and screen out the median shard of the median from the multiple shards;
[0032] Screen out target shards from the multiple shards whose output data volume is greater than or equal to a preset multiple of the output data volume corresponding to the join operation of the median shard;
[0033] Divide each of the target shards into multiple shards.
[0034] In addition, according to another aspect of the embodiments of the present invention, a data processing device is provided, including:
[0035] A sampling module, configured to sample the data to be processed in the data table in an adaptive manner to obtain the sample data distribution of the data table;
[0036] A calculation module, configured to divide the data to be processed into multiple shards, and calculate the output data volumes corresponding to the join operations of the multiple shards respectively according to the sample data distribution of the data table;
[0037] The sharding module is used to screen out target shards from the multiple shards according to the output data volume corresponding to the connection operation of the multiple shards, and perform data sharding on the target shards.
[0038] Optionally, the sampling module is further used for:
[0039] Obtain the data to be processed from the data table to obtain the key in the data to be processed;
[0040] Judge whether the cumulative number of processed items is greater than or equal to the first quantity threshold; if not, write the key in the data to be processed into the counting table, and use the key in the counting table as the sampling result; if so, use the reservoir sampling algorithm to sample the data to be processed;
[0041] Obtain the sample data distribution of the data table according to the sampling result of the data to be processed.
[0042] Optionally, the sampling module is further used for:
[0043] Judge whether the cumulative number of processed items is greater than or equal to the second quantity threshold; if not, write the key in the data to be processed into the sampling array; if so, use the reservoir sampling algorithm to replace one key in the sampling array with the key in the data to be processed;
[0044] Use the key in the sampling array as the sampling result.
[0045] Optionally, the sampling module is further used for:
[0046] Calculate the cumulative sampling volume of each sample respectively; wherein, the sampling result of the data to be processed includes multiple samples;
[0047] For each data table, divide the data volume of the data to be processed by the cumulative sampling volume of each sample to obtain the weight of each sample in the data table, so as to obtain the sample data distribution of the data table.
[0048] Optionally, the calculation module is further used for:
[0049] For any one of the multiple shards, screen out the target samples assigned to the shard from the sampling results; wherein, each data table contains the target samples;
[0050] Calculate the output data volume corresponding to the connection operation of the shard according to the weights of the respective target samples in the data table.
[0051] Optionally, the calculation module is further used for:
[0052] Multiply the weights of each target sample in each data table and accumulate the products to obtain the output data volume corresponding to the join operation of the shard.
[0053] Optionally, the calculation module is further configured to:
[0054] If the type of the join operation is an inner join, multiply the weights of each target sample in each data table and accumulate the products to obtain the output data volume corresponding to the join operation of the shard;
[0055] If the type of the join operation is a left join, multiply the weights of each target sample in each data table and accumulate the products, and then add the weights of each target sample in the left join data table to obtain the output data volume corresponding to the join operation of the shard;
[0056] If the type of the join operation is a right join, multiply the weights of each target sample in each data table and accumulate the products, and then add the weights of each target sample in the right join data table to obtain the output data volume corresponding to the join operation of the shard;
[0057] If the type of the join operation is a full join, multiply the weights of each target sample in each data table and accumulate the products, and then add the weights of each target sample in the left join data table and the right join data table to obtain the output data volume corresponding to the join operation of the shard.
[0058] Optionally, the sharding module is further configured to:
[0059] Sort the output data volumes corresponding to the join operations of the multiple shards, and filter out the median shard with the median value from the multiple shards;
[0060] Filter out target shards from the multiple shards whose output data volume is greater than or equal to a preset multiple of the output data volume corresponding to the join operation of the median shard;
[0061] Divide each of the target shards into multiple shards.
[0062] According to another aspect of the embodiments of the present invention, there is also provided an electronic device, including:
[0063] One or more processors;
[0064] A storage device for storing one or more programs,
[0065] When the one or more programs are executed by the one or more processors, the one or more processors implement the method described in any of the above embodiments.
[0066] According to another aspect of the embodiments of the present invention, there is also provided a computer-readable medium, on which a computer program is stored, and when the program is executed by a processor, the method described in any of the above embodiments is implemented.
[0067] One embodiment of the above invention has the following advantages or beneficial effects: Since the technical means of sampling the data to be processed in the data table in an adaptive manner to obtain the sample data distribution of the data table, calculating the output data volume corresponding to the join operations of multiple shards respectively according to the sample data distribution of the data table, and thus screening out the target shard and performing data sharding on it are adopted, the technical problem of being unable to handle data expansion in the prior art is overcome. The embodiments of the present invention obtain the sample data distribution of the data table through adaptive data sampling to obtain the overall distribution of the data, and then calculate the output data volume of each allocation according to the sample data distribution, accurately screening out the target shard that needs to perform data sharding again, so as to limit the processing of data expansion and avoid data expansion.
[0068] The further effects of the above non-conventional optional manner will be described below in conjunction with the specific embodiments. BRIEF DESCRIPTION OF THE DRAWINGS
[0069] The drawings are used to better understand the present invention and do not constitute an improper limitation to the present invention. Among them:
[0070] Figure 1 is a schematic diagram of the basic principle of Spark AQE;
[0071] Figure 2 is a schematic diagram of the basic principle of the dynamic optimization strategy in Spark AQE;
[0072] Figure 3 is a schematic diagram of the main process of the data processing method according to the embodiments of the present invention;
[0073] Figure 4 is a schematic diagram of the main process of the data processing method according to a reference embodiment of the present invention;
[0074] Figure 5 is a schematic diagram of the main process of the data processing method according to another reference embodiment of the present invention;
[0075] Figure 6a and 6b is a comparison schematic diagram before and after adopting the data processing method of the embodiments of the present invention;
[0076] Figure 7 is a schematic diagram of the main modules of the data processing device according to the embodiments of the present invention;
[0077] Figure 8It is an exemplary system architecture diagram to which the embodiments of the present invention can be applied;
[0078] Figure 9 It is a schematic structural diagram of a computer system of a terminal device or a server suitable for implementing the embodiments of the present invention. Detailed implementation manners
[0079] The following describes exemplary embodiments of the present invention with reference to the accompanying drawings. Various details of the embodiments of the present invention are included to facilitate understanding, and they should be considered merely exemplary. Therefore, those of ordinary skill in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present invention. Similarly, for the sake of clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.
[0080] Distributed computing engines such as Spark and Hive all provide certain processing capabilities for data skew. Taking Spark as an example, since its version 3.0, it has introduced the Adaptive Query Execution (AQE) function during runtime, which can adaptively and dynamically adjust and optimize subsequent computing links based on the statistical information of the completed stages.
[0081] A very important strategy in AQE is: dynamically optimizing the skewed Join connection. As Figure 1 shown, taking the join of table A and table B as an example to introduce its basic principle. Assume that the data volume of the shard A0 of table A is significantly larger than that of other shards. After performing hash shuffle, the data volume of the computing tasks for processing A0 and B0 is significantly larger than that of other shards, resulting in the slowdown of the entire computing link. shuffle: In distributed computing, the process of data distribution over the network; generally, shuffle is based on hash values, and data corresponding to the same hash value will be placed in the same shard.
[0082] The dynamic optimization strategy in Spark AQE will count the size of each shard. When it is found that A0 is significantly larger than the median of each shard, the division of A0 will be triggered. At the same time, to ensure the smooth progress of the calculation, B0 will be replicated, as Figure 2 shown. After such processing, although the total number of tasks increases (from 4 to 5), the data of each task is more uniform, and the calculation time is almost the same, bringing better performance to the entire query calculation.
[0083] In the actual production environment, there are various data skew scenarios. One common data skew situation is called data explosion, which is manifested as follows: the size of the input shard is not significantly larger than the average level, but the corresponding data contains only a few different keys. After the Join operation is executed, the amount of data output by the corresponding task is much larger than that of other tasks. Existing computing engines are divided based on the size of the input shard and cannot handle this data explosion scenario. For example, the sizes of the input shards A3 and B3 do not show skew, but there is a frequent key = Hot_Sku in both A3 and B3. The frequencies of this key on the A3 side and the B3 side are N_A and N_B respectively. Then, in the result after the Join operation, the frequency of this key is N_A * N_B. When N_A and N_B are normal compared to the input shard size, but N_A * N_B is large compared to the output shard size, data explosion will occur.
[0084] To solve the problem of data explosion existing in the prior art, the embodiments of the present invention, in a distributed environment, estimate the amount of output data of the Join operation through adaptive data statistics and analysis, and accordingly divide the data shards to handle data explosion.
[0085] Figure 3 It is a schematic diagram of the main process of the data processing method according to the embodiments of the present invention. As an embodiment of the present invention, as Figure 3 shown, the data processing method may include:
[0086] Step 301, sample the data to be processed in the data table in an adaptive manner to obtain the sample data distribution of the data table.
[0087] First, sample the data to be processed in the data table in an adaptive manner. The sampling logic is executed in a distributed form at the executor. Through adaptive data sampling, sample data can be collected from the data table, and thus the sample data distribution of the data table can be calculated. Executor: In the Spark computing engine, the node that actually executes the specific computing task.
[0088] Optionally, step 301 may include: obtaining the data to be processed from the data table to obtain the key in the data to be processed; determining whether the cumulative number of processed items is greater than or equal to the first quantity threshold; if not, writing the key in the data to be processed into the counting table, and using the key in the counting table as the sampling result; if so, sampling the data to be processed using the reservoir sampling algorithm; obtaining the sample data distribution of the data table according to the sampling result of the data to be processed. The first quantity threshold (MapSize) can be preset in advance and then incrementally updated in the form of a stream. Each time a piece of data to be processed is obtained from the data table, the key is extracted from this piece of data to be processed, and the cumulative number of processed items is recorded in a cumulative manner. If the cumulative number of processed items is less than the first quantity threshold, enter the counting mode (Counting). In this mode, write this key into the counting table (CountMap), and use all the keys in the counting table as the sampling result; if the cumulative number of processed items is greater than or equal to the first quantity threshold, that is, the size of the counting table reaches the preset first quantity threshold, enter the sampling mode (Sampling), that is, use the reservoir sampling algorithm to sample the data to be processed, ensuring that each key in the entire stream is sampled and retained with equal probability. Reservoir Sampling: A streaming sampling algorithm that can randomly extract n data with equal probability in o(n) time.
[0089] It should be noted that the data to be processed contains the key and its corresponding value. In the embodiments of the present invention, only the key is required, so only the key needs to be extracted from the data to be processed.
[0090] Optionally, sampling the data to be processed using the reservoir sampling algorithm includes: determining whether the cumulative number of processed items is greater than or equal to the second quantity threshold; if not, writing the key in the data to be processed into the sampling array; if so, using the reservoir sampling algorithm to replace one key in the sampling array with the key in the data to be processed; using the key in the sampling array as the sampling result. In the sampling mode, the second quantity threshold (ArraySize) needs to be preset in advance, and the second quantity threshold is greater than the first quantity threshold. After entering the sampling mode, if the cumulative number of processed items is less than the second quantity threshold, write the key extracted from the data to be processed into the sampling array. If the cumulative number of processed items is greater than or equal to the second quantity threshold, use the reservoir sampling algorithm to replace one key in the sampling array with the key extracted from the data to be processed, and finally use all the keys in the sampling array as the sampling result.
[0091] In an embodiment of the present invention, when entering the sampling mode, all keys in the counting table are written into the sampling array, so as to directly obtain the sampling results from the sampling array, and all keys in the sampling array are the sampling results. First, an attempt is made to perform accurate sampling. When the accumulated number of processed items is greater than or equal to the first quantity threshold, it will fallback to reservoir sampling.
[0092] Optionally, obtaining the sample data distribution of the data table according to the sampling results of the data to be processed includes: calculating the cumulative sampling amount of each sample respectively; wherein, the sampling results of the data to be processed include multiple samples; for each data table, dividing the data volume of the data to be processed by the cumulative sampling amount of each sample to obtain the weight of each sample in the data table, so as to obtain the sample data distribution of the data table. In an embodiment of the present invention, after the sampling of each executor is completed, the sampling results, the shard data volume C, and the sampling volume S are sent back to the driver. The driver assigns weights C / S to each sample and performs a summary and accumulation process on the sample weights corresponding to each executor (the weights of each sample are summarized and accumulated separately). Driver: In the Spark computing engine, the node responsible for scheduling, summarizing, and distributing tasks.
[0093] As Figure 1 shown, taking Table A as an example, the weights of each sample in each map task are summarized and accumulated respectively, so as to obtain the sample data distribution of Table A. Similarly, the sample data distribution of Table B can also be obtained.
[0094] Step 302, dividing the data to be processed into multiple shards, and respectively calculating the output data volume corresponding to the join operation of the multiple shards according to the sample data distribution of the data table.
[0095] The data to be processed can be divided into multiple shards by using a hash algorithm. As Figure 1 shown, Table A is divided into four shards A0, A1, A2, and A3 according to the hash algorithm, and Table B is divided into two shards B0 and B1; then, according to the sample data distribution of Table A and Table B calculated in step 301, the output data volume corresponding to the Join of each shard is calculated respectively.
[0096] Join connection: In SQL, it is used to combine rows from two or more tables to obtain the calculation logic of the matching relationship. According to semantics, it can be divided into: InnerJoin (inner join), LeftJoin (left join), RightJoin (right join), FullOuterJoin (full join), etc.
[0097] Optionally, calculate the output data volume corresponding to the join operation of the multiple shards respectively according to the sample data distribution of the data tables, including: for any one of the multiple shards, screen out the target samples assigned to the shard from the sampling results; wherein each of the data tables contains the target samples; calculate the output data volume corresponding to the join operation of the shard according to the weights of the respective target samples in the data tables. Through adaptive sampling, the sample data distributions of Table A and Table B are obtained. For a certain shard X, the target samples A_X and B_X that fall within shard X are screened out from the sampling results of Table A and Table B, and then the output data volume corresponding to the Join of shard X is calculated according to the weights of the respective target samples A_X and B_X. It should be noted that both Table A and Table B contain target samples, and there can be multiple target samples.
[0098] Optionally, calculate the output data volume corresponding to the join operation of the shard according to the weights of the respective target samples in the data tables, including: multiply the weights of the respective target samples in each of the data tables and accumulate the products to obtain the output data volume corresponding to the join operation of the shard. Specifically, the following formula can be used to calculate the output data volume corresponding to the Join of each shard:
[0099]
[0100] where w A (x) and w B (x) are the weights of sample x in the data of Table A and Table B respectively.
[0101] Optionally, multiply the weights of the respective target samples in each of the data tables and accumulate the products to obtain the output data volume corresponding to the join operation of the shard, including: if the type of the join operation is an inner join, multiply the weights of the respective target samples in each of the data tables and accumulate the products to obtain the output data volume corresponding to the join operation of the shard; if the type of the join operation is a left join, multiply the weights of the respective target samples in each of the data tables and accumulate the products, and then add the weights of the respective target samples in the left join data table to obtain the output data volume corresponding to the join operation of the shard; if the type of the join operation is a right join, multiply the weights of the respective target samples in each of the data tables and accumulate the products, and then add the weights of the respective target samples in the right join data table to obtain the output data volume corresponding to the join operation of the shard; if the type of the join operation is a full join, multiply the weights of the respective target samples in each of the data tables and accumulate the products, and then add the weights of the respective target samples in the left join data table and the right join data table to obtain the output data volume corresponding to the join operation of the shard. In order to accurately calculate the output data volume corresponding to the Join of the shard, different types of Join use different calculation formulas:
[0102] If JoinType is Inner Join:
[0103] Estimated Size(X) = ∑ x∈A_X∩B_X w A (x) * w B (x);
[0104] If JoinType is Left Join:
[0105] Estimated Size(X) = ∑ x∈A_X∩B_X w A (x) * w B x∈A_X w A x∈A_X∩B_X (x);
[0106] If JoinType is Right Join:
[0107] Estimated Size(X) = ∑ x∈A_X∩B_X w A (x) * w B x∈B_X w B x∈A_X∩B_X (x);
[0108] If JoinType is Full Outer Join:
[0109] Estimated Size(X) = ∑ x∈A_X∩B_X w A (x) * w B x∈A_X w A x∈B_X w B A B (x);
[0110] Among them, w A (x) and w B (x) are respectively the weights of the sample x in the data of table A and table B.
[0111] Step 303, according to the output data volume corresponding to the join operation of the multiple shards, screen out the target shard from the multiple shards, and perform data sharding on the target shard.
[0112] Since the output data volume corresponding to the Join of each shard is calculated, the shards can be sorted based on the output data volume, so as to screen out the target shard and perform data sharding on the target shard to control the size of the output data volume of the shard.
[0113] Optionally, step 303 may include: sorting the output data volumes corresponding to the connection operations of the multiple shards, and screening out the median shard from the multiple shards; screening out target shards from the multiple shards whose output data volume is greater than or equal to a preset multiple of the output data volume corresponding to the connection operation of the median shard; dividing each of the target shards into multiple shards. In an embodiment of the present invention, first, the individual shards are sorted based on the output data volume, then the median shard is screened out from these shards, and finally, target shards whose output data volume is greater than or equal to W times (a preset parameter) the output data volume corresponding to the connection operation of the median shard are screened out.
[0114] In an embodiment of the present invention, after screening out the target shards, a hash algorithm may be used to perform data sharding on the target shards to control the size of the output data volume of the shards.
[0115] According to the various embodiments described above, it can be seen that the embodiments of the present invention solve the technical problem of being unable to handle data expansion in the prior art by adopting an adaptive method to sample the data to be processed in the data table, obtaining the sample data distribution of the data table, calculating the output data volumes corresponding to the connection operations of multiple shards respectively according to the sample data distribution of the data table, thereby screening out the target shards and performing data sharding on them. The embodiments of the present invention obtain the sample data distribution of the data table through adaptive data sampling to obtain the overall distribution of the data, and then calculate the output data volumes of each allocation according to the sample data distribution, accurately screening out the target shards that need to be further data-sharded, thereby limitedly handling data expansion and avoiding data expansion.
[0116] Figure 4 is a schematic diagram of the main process of a data processing method according to a reference embodiment of the present invention. As another embodiment of the present invention, as Figure 4 shown, the steps of sampling the data to be processed in the data table in an adaptive manner may include:
[0117] Step 401, obtaining the data to be processed from the data table to obtain the key in the data to be processed.
[0118] Step 402, determining whether the cumulative number of processed items is greater than or equal to a first quantity threshold; if so, execute step 403; if not, execute step 406.
[0119] Step 403, determining whether the cumulative number of processed items is greater than or equal to a second quantity threshold; if so, execute step 404; if not, execute step 405.
[0120] Step 404: Use the reservoir sampling algorithm to replace one key in the sampling array with the key in the data to be processed, and use the key in the sampling array as the sampling result.
[0121] Step 405: Write the key in the data to be processed into the sampling array, and use the key in the sampling array as the sampling result.
[0122] Step 406: Write the key in the data to be processed into the counting table, and use the key in the counting table as the sampling result.
[0123] In addition, the specific implementation content of the data processing method in a referenceable embodiment of the present invention has been described in detail in the above-mentioned data processing method, so the repeated content will not be described here.
[0124] Figure 5 It is a schematic diagram of the main process of the data processing method according to another referenceable embodiment of the present invention. As another embodiment of the present invention, as Figure 5 shown, the steps of sampling the data to be processed in the data table in an adaptive manner may include:
[0125] Step 501: Sample the data to be processed in the data table in an adaptive manner.
[0126] Optionally, step 501 may include: obtaining the data to be processed from the data table to obtain the key in the data to be processed; determining whether the cumulative number of processed items is greater than or equal to the first quantity threshold; if not, writing the key in the data to be processed into the counting table, and using the key in the counting table as the sampling result; if so, sampling the data to be processed using the reservoir sampling algorithm; obtaining the sample data distribution of the data table according to the sampling result of the data to be processed. By presetting the first quantity threshold, the counting mode or the sampling mode can be adaptively adopted, and the sampling logic is executed in a distributed form at the execution end.
[0127] In the embodiment of the present invention, each time a piece of data to be processed is obtained from the data table, the key is extracted from this piece of data to be processed, and the cumulative number of processed items is recorded in a cumulative manner. If the cumulative number of processed items is less than the first quantity threshold, the counting mode is entered. In this mode, the key is written into the counting table (CountMap), and all keys in the counting table are used as the sampling result; if the cumulative number of processed items is greater than or equal to the first quantity threshold, that is, the size of the counting table reaches the preset first quantity threshold, the sampling mode is entered, that is, the reservoir sampling algorithm is used to sample the data to be processed, ensuring that each key in the entire stream is sampled and retained with equal probability.
[0128] Step 502: Calculate the cumulative sampling amount of each sample respectively.
[0129] Among them, the sampling results of the data to be processed include multiple samples, and the cumulative sampling amount of each sample is counted from the sampling results.
[0130] Step 503: For each data table, divide the data amount of the data to be processed by the cumulative sampling amount of each sample to obtain the weight of each sample in the data table, so as to obtain the sample data distribution of the data table.
[0131] After the sampling of each executor is completed, the sampling results, the shard data amount C, and the sampling amount S will be sent back to the driver. The driver assigns weights C / S to each sample and performs a summary and accumulation process on the sample weights corresponding to each executor (the weights of each sample are summarized and accumulated respectively), so as to obtain the sample data distribution of each data table.
[0132] Step 504: Divide the data to be processed into multiple shards.
[0133] The data to be processed can be divided into multiple shards by using a hash algorithm.
[0134] Step 505: For any one of the multiple shards, screen out the target samples assigned to the shard from the sampling results.
[0135] Each data table contains the target samples, and the target samples are assigned to the shard.
[0136] Step 506: Calculate the output data amount corresponding to the join operation of the shard according to the weights of the respective target samples in the data table.
[0137] Specifically, multiply the weights of each target sample in each data table and accumulate the products, so as to obtain the output data amount corresponding to the join operation of the shard. Specifically, the following formula can be used to calculate the output data amount corresponding to the Join of each shard:
[0138]
[0139] where w A (x) and w B (x) are the weights of sample x in the data of table A and table B respectively.
[0140] In order to accurately calculate the output data amount corresponding to the Join of the shard, different types of Join use different calculation formulas, which will not be elaborated here.
[0141] Step 507: Sort the output data volumes corresponding to the connection operations of the multiple shards, and select the median shard with the median value from the multiple shards.
[0142] Step 508: Select target shards from the multiple shards whose output data volumes are greater than or equal to a preset multiple of the output data volume corresponding to the connection operation of the median shard.
[0143] Step 509: Divide each of the target shards into multiple shards.
[0144] After selecting the target shards, a hash algorithm can be used to perform data sharding on the target shards to control the size of the output data volume of the shards.
[0145] In addition, the specific implementation content of the data processing method in another referenceable embodiment of the present invention has been described in detail in the above-mentioned data processing method, so the repeated content will not be described herein again.
[0146] In practice, the method provided by the embodiments of the present invention can effectively handle the problem of data expansion. The following is an example:
[0147] As Figure 6a shown, the input data (Shuffle Read Size) of shard 5 is 1.8 MiB, which is not much different from other shards; however, its output data volume (Output Size) is 75.5 MiB, far exceeding other shards (the median is 853.4 KiB). There are 10 shards in this example. The running time of the shard with data expansion is 17 seconds, far exceeding the running time of other shards, which is 2 seconds. Since the existing technical solutions only consider the size of the input data volume, the problem of data expansion cannot be solved.
[0148] As Figure 6b shown, after adopting the method provided by the embodiments of the present invention, the input data is segmented according to the estimated output data volume. From the original 10 shards, it is segmented into 72 shards; after segmentation, the output of all shards is more uniform, and the execution time of all shards does not exceed 2 seconds, so that the total running time is optimized from 17 seconds to 2 seconds, and the efficiency is greatly improved.
[0149] Figure 7 is a schematic diagram of the main modules of the data processing device according to the embodiments of the present invention. As Figure 7As shown, the data processing device 700 includes a sampling module 701, a calculation module 702, and a sharding module 703. Among them, the sampling module 701 is used to sample the data to be processed in the data table in an adaptive manner to obtain the sample data distribution of the data table. The calculation module 702 is used to divide the data to be processed into multiple shards and calculate the output data volume corresponding to the join operation of the multiple shards respectively according to the sample data distribution of the data table. The sharding module 703 is used to screen out the target shards from the multiple shards according to the output data volume corresponding to the join operation of the multiple shards and perform data sharding on the target shards.
[0150] Optionally, the sampling module 701 is further used for:
[0151] Obtain the data to be processed from the data table to obtain the key in the data to be processed;
[0152] Judge whether the cumulative number of processed items is greater than or equal to the first quantity threshold. If not, write the key in the data to be processed into the counting table and use the key in the counting table as the sampling result. If so, sample the data to be processed using the reservoir sampling algorithm;
[0153] Obtain the sample data distribution of the data table according to the sampling result of the data to be processed.
[0154] Optionally, the sampling module 701 is further used for:
[0155] Judge whether the cumulative number of processed items is greater than or equal to the second quantity threshold. If not, write the key in the data to be processed into the sampling array. If so, use the reservoir sampling algorithm to replace one key in the sampling array with the key in the data to be processed;
[0156] Use the key in the sampling array as the sampling result.
[0157] Optionally, the sampling module 701 is further used for:
[0158] Calculate the cumulative sampling volume of each sample respectively. Among them, the sampling result of the data to be processed includes multiple samples;
[0159] For each data table, divide the data volume of the data to be processed by the cumulative sampling volume of each sample to obtain the weight of each sample in the data table, so as to obtain the sample data distribution of the data table.
[0160] Optionally, the calculation module 702 is further used for:
[0161] For any one of the multiple slices, filter out the target samples assigned to the slice from the sampling results; wherein, each of the data tables contains the target samples.
[0162] Calculate the output data volume corresponding to the join operation of the slice according to the weights of the respective target samples in the data table.
[0163] Optionally, the calculation module 702 is further configured to:
[0164] Multiply the weights of the respective target samples in the respective data tables and accumulate the products to obtain the output data volume corresponding to the join operation of the slice.
[0165] Optionally, the calculation module 702 is further configured to:
[0166] If the type of the join operation is an inner join, multiply the weights of the respective target samples in the respective data tables and accumulate the products to obtain the output data volume corresponding to the join operation of the slice.
[0167] If the type of the join operation is a left join, multiply the weights of the respective target samples in the respective data tables and accumulate the products, and then add the weights of the respective target samples in the left join data table to obtain the output data volume corresponding to the join operation of the slice.
[0168] If the type of the join operation is a right join, multiply the weights of the respective target samples in the respective data tables and accumulate the products, and then add the weights of the respective target samples in the right join data table to obtain the output data volume corresponding to the join operation of the slice.
[0169] If the type of the join operation is a full join, multiply the weights of the respective target samples in the respective data tables and accumulate the products, and then add the weights of the respective target samples in the left join data table and the right join data table to obtain the output data volume corresponding to the join operation of the slice.
[0170] Optionally, the slicing module 703 is further configured to:
[0171] Sort the output data volumes corresponding to the join operations of the multiple slices, and filter out the median slice among the multiple slices.
[0172] Filter out the target slices from the multiple slices whose output data volume is greater than or equal to a preset multiple of the output data volume corresponding to the join operation of the median slice.
[0173] Divide each of the target slices into multiple slices.
[0174] According to the various embodiments described above, it can be seen that the embodiments of the present invention sample the data to be processed in the data table in an adaptive manner to obtain the sample data distribution of the data table, and calculate the output data volume corresponding to the join operations of multiple shards respectively according to the sample data distribution of the data table, so as to screen out the target shard and perform data sharding on it, thereby solving the technical problem that data expansion cannot be processed in the prior art. The embodiments of the present invention obtain the sample data distribution of the data table through adaptive data sampling to obtain the overall distribution of the data, and then calculate the output data volume of each allocated shard according to the sample data distribution, accurately screen out the target shard that needs to be sharded again, so as to limit the processing of data expansion and avoid data expansion.
[0175] It should be noted that the specific implementation content of the data processing device of the present invention has been described in detail in the above data processing method, so the repeated content will not be described here.
[0176] Figure 8 An exemplary system architecture 800 to which the data processing method or data processing device of the embodiments of the present invention can be applied is shown.
[0177] As Figure 8 shown, the system architecture 800 may include terminal devices 801, 802, 803, a network 804, and a server 805. The network 804 is used to provide a medium for communication links between the terminal devices 801, 802, 803 and the server 805. The network 804 may include various connection types, such as wired, wireless communication links, or fiber optic cables, etc.
[0178] Users can use the terminal devices 801, 802, 803 to interact with the server 805 through the network 804 to receive or send messages, etc. Various communication client applications may be installed on the terminal devices 801, 802, 803, such as shopping applications, web browser applications, search applications, instant messaging tools, email clients, social platform software, etc. (only as examples).
[0179] The terminal devices 801, 802, 803 may be various electronic devices with a display screen and supporting web browsing, including but not limited to smart phones, tablet computers, laptop portable computers, and desktop computers, etc.
[0180] The server 805 may be a server providing various services, such as a background management server that supports the shopping websites browsed by users using the terminal devices 801, 802, 803 (only as an example). The background management server may analyze and process data such as item information query requests received, and feedback the processing results to the terminal devices.
[0181] It should be noted that the data processing method provided by the embodiments of the present invention is generally executed by the server 805. Correspondingly, the data processing device is generally disposed in the server 805.
[0182] It should be understood that Figure 8 the numbers of the terminal devices, networks, and servers in
[0183] are merely illustrative. According to actual requirements, there can be any number of terminal devices, networks, and servers. Figure 9 The following refers to Figure 9 which shows a schematic structural diagram of a computer system 900 of a terminal device suitable for implementing the embodiments of the present invention.
[0184] As Figure 9 shown, the computer system 900 includes a central processing unit (CPU) 901, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 902 or a program loaded from a storage section 908 into a random access memory (RAM) 903. In the RAM 903, various programs and data required for the operation of the system 900 are also stored. The CPU 901, ROM 902, and RAM 903 are connected to each other via a bus 904. An input / output (I / O) interface 905 is also connected to the bus 904.
[0185] The following components are connected to the I / O interface 905: an input section 906 including a keyboard, a mouse, etc.; an output section 907 including a cathode ray tube (CRT), a liquid crystal display (LCD), etc. and a speaker, etc.; a storage section 908 including a hard disk, etc.; and a communication section 909 including a network interface card such as a LAN card, a modem, etc. The communication section 909 performs communication processing via a network such as the Internet. A drive 910 is also connected to the I / O interface 905 as required. A removable medium 911, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 910 as required so that a computer program read from it can be installed into the storage section 908 as required.
[0186] In particular, according to the embodiments disclosed in the present invention, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, the embodiments disclosed in the present invention include a computer program comprising a computer program carried on a computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such an embodiment, the computer program can be downloaded and installed from a network via the communication section 909, and / or installed from the removable medium 911. When the computer program is executed by the central processing unit (CPU) 901, the above-described functions defined in the system of the present invention are performed.
[0187] It should be noted that the computer-readable medium shown in the present invention can be a computer-readable signal medium or a computer-readable storage medium or any combination of the two. A computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of the computer-readable storage medium can include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present invention, the computer-readable storage medium can be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. In the present invention, a computer-readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, in which computer-readable program code is carried. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. A computer-readable signal medium can also be any computer-readable medium other than a computer-readable storage medium that can send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any appropriate medium, including but not limited to: wireless, wire, optical fiber, RF, etc., or any suitable combination of the above.
[0188] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functions, and operations of systems, methods, and computer programs that can be implemented according to various embodiments of the present invention. In this regard, each block in the flowchart or block diagram may represent a module, a segment of a program, or a part of code that contains one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than that marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram or flowchart, as well as combinations of blocks in the block diagram or flowchart, can be implemented by a dedicated hardware-based system that performs the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.
[0189] The modules described in the embodiments of the present invention can be implemented in software or in hardware. The described modules can also be provided in a processor. For example, it can be described as: a processor includes a sampling module, a calculation module, and a sharding module. In some cases, the names of these modules do not constitute a limitation on the module itself.
[0190] As another aspect, the present invention also provides a computer-readable medium, which can be included in the devices described in the above embodiments; or can exist separately without being assembled into the device. The above computer-readable medium carries one or more programs. When the one or more programs are executed by a device, the device implements the following method: sampling the data to be processed in a data table in an adaptive manner to obtain the sample data distribution of the data table; dividing the data to be processed into multiple shards, and respectively calculating the output data volume corresponding to the join operation of the multiple shards according to the sample data distribution of the data table; screening out target shards from the multiple shards according to the output data volume corresponding to the join operation of the multiple shards, and performing data sharding on the target shards.
[0191] According to the technical solution of the embodiment of the present invention, since the data to be processed in the data table is sampled in an adaptive manner to obtain the sample data distribution of the data table, and the output data volume corresponding to the join operations of multiple shards is calculated respectively according to the sample data distribution of the data table, so as to screen out the target shard and perform data sharding on it, the technical problem that the prior art cannot handle data expansion is overcome. The embodiment of the present invention obtains the sample data distribution of the data table through adaptive data sampling to obtain the overall distribution of the data, and then calculates the output data volume of each allocated shard according to the sample data distribution, accurately screens out the target shard that needs to be sharded again, so as to limit the processing of data expansion and avoid data expansion.
[0192] The above specific embodiments do not constitute a limitation on the protection scope of the present invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations and substitutions can occur depending on design requirements and other factors. Any modifications, equivalent replacements and improvements made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. A data processing method, characterized in that, Including: Sampling the data to be processed in the data table in an adaptive manner to obtain the sample data distribution of the data table; Dividing the data to be processed into multiple shards, and respectively calculating the output data volume corresponding to the join operation of the multiple shards according to the sample data distribution of the data table; According to the output data volume corresponding to the join operation of the multiple shards, screening out the target shards from the multiple shards, and performing data sharding on the target shards; Sampling the data to be processed in the data table in an adaptive manner to obtain the sample data distribution of the data table, including: Obtaining the data to be processed from the data table to obtain the key in the data to be processed; Judging whether the cumulative number of processed items is greater than or equal to the first quantity threshold; if not, writing the key in the data to be processed into the counting table, and using the key in the counting table as the sampling result; if so, sampling the data to be processed by using the reservoir sampling algorithm; Obtaining the sample data distribution of the data table according to the sampling result of the data to be processed; Sampling the data to be processed by using the reservoir sampling algorithm, including: Judging whether the cumulative number of processed items is greater than or equal to the second quantity threshold; if not, writing the key in the data to be processed into the sampling array; if so, using the reservoir sampling algorithm to replace a key in the sampling array with the key in the data to be processed; Using the key in the sampling array as the sampling result.
2. The method according to claim 1, wherein Obtaining the sample data distribution of the data table according to the sampling result of the data to be processed, including: Respectively calculating the cumulative sampling volume of each sample; wherein, the sampling result of the data to be processed includes multiple samples; For each data table, dividing the data volume of the data to be processed by the cumulative sampling volume of each sample to obtain the weight of each sample in the data table, so as to obtain the sample data distribution of the data table.
3. The method according to claim 2, wherein Respectively calculating the output data volume corresponding to the join operation of the multiple shards according to the sample data distribution of the data table, including: For any one of the multiple shards, screening out the target samples assigned to the shard from the sampling result; wherein each data table contains the target samples; Calculating the output data volume corresponding to the join operation of the shard according to the weights of the respective target samples in the data table.
4. The method according to claim 3, wherein Calculating the output data volume corresponding to the join operation of the shard according to the weights of the respective target samples in the data table, including: Multiplying the weights of the respective target samples in each data table and accumulating the products, so as to obtain the output data volume corresponding to the join operation of the shard.
5. The method according to claim 4, wherein Multiplying the weights of the respective target samples in each data table and accumulating the products, so as to obtain the output data volume corresponding to the join operation of the shard, including: If the type of the join operation is an inner join, multiplying the weights of the respective target samples in each data table and accumulating the products, so as to obtain the output data volume corresponding to the join operation of the shard; If the type of the join operation is a left join, multiply the weights of each target sample in each data table and accumulate the products, and then add the weights of each target sample in the left join data table, so as to obtain the output data volume corresponding to the join operation of the shard; If the type of the join operation is a right join, multiply the weights of each target sample in each data table and accumulate the products, and then add the weights of each target sample in the right join data table, so as to obtain the output data volume corresponding to the join operation of the shard; If the type of the join operation is a full join, multiply the weights of each target sample in each data table and accumulate the products, and then add the weights of each target sample in the left join data table and the right join data table, so as to obtain the output data volume corresponding to the join operation of the shard.
6. The method according to claim 1, characterized in that, According to the output data volumes corresponding to the join operations of the multiple shards, filter out target shards from the multiple shards, and perform data sharding on the target shards, including: Sort the output data volumes corresponding to the join operations of the multiple shards, and filter out the median shard of the median from the multiple shards; Filter out target shards from the multiple shards whose output data volume is greater than or equal to a preset multiple of the output data volume corresponding to the join operation of the median shard; Divide each of the target shards into multiple shards.
7. A data processing device, characterized in that, including: A sampling module, configured to sample the data to be processed in the data table in an adaptive manner to obtain the sample data distribution of the data table; A calculation module, configured to divide the data to be processed into multiple shards, and respectively calculate the output data volumes corresponding to the join operations of the multiple shards according to the sample data distribution of the data table; A sharding module, configured to filter out target shards from the multiple shards according to the output data volumes corresponding to the join operations of the multiple shards, and perform data sharding on the target shards; The sampling module is further configured to: Obtain the data to be processed from the data table to obtain the key in the data to be processed; Judge whether the cumulative number of processed items is greater than or equal to the first quantity threshold; if not, write the key in the data to be processed into the counting table, and use the key in the counting table as the sampling result; if so, sample the data to be processed by using the reservoir sampling algorithm; Obtain the sample data distribution of the data table according to the sampling result of the data to be processed; The sampling module is further configured to: Judge whether the cumulative number of processed items is greater than or equal to the second quantity threshold; if not, write the key in the data to be processed into the sampling array; if so, use the reservoir sampling algorithm to replace a key in the sampling array with the key in the data to be processed; Use the key in the sampling array as the sampling result.
8. An electronic device, characterized in that, including: One or more processors; A storage device, configured to store one or more programs, When the one or more programs are executed by the one or more processors, the one or more processors implement the method according to any one of claims 1-6.
9. A computer-readable medium having a computer program stored thereon, characterized in that, The program, when executed by the processor, implements the method according to any one of claims 1-6.
Citation Information
Patent Citations
Binary equivalent connection tilt optimization method based on distributed sensing
CN108804626A
Data inclination processing method and device, terminal equipment and storage medium
CN112000467A
Multi-table connection method and device
CN112711588A
Data processing method and device, computer equipment and storage medium
CN112905596A