Distributed data processing method, device, electronic device and storage medium

By determining the bucket field and data bucket serial number in distributed data processing, the problem of uneven data distribution is solved, and the data query efficiency and ease of use of the data detail layer are improved.

CN113886491BActive Publication Date: 2025-06-06GUANGZHOU HUYA TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111095167.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-09-17
Publication Date
2025-06-06
Estimated Expiration
2041-09-17

AI Technical Summary

Technical Problem

Existing distributed data processing methods are prone to the problems of duplicate data and uneven distribution of data in folders at different levels, resulting in excessive computing resources consumption during data query.

Method used

By determining the bucket field of a distributed data set, the data set is divided into multiple data subsets, and the bucket sequence number of the data bucket is determined based on the number of data records or data size of the data subset, ensuring that each piece of data is written to the data bucket corresponding to the appropriate target bucket sequence number.

Benefits of technology

It effectively reduces the data skew and uneven distribution, improves the rationality of data distribution, reduces the computing resource consumption during data query, and improves the ease of use of the data detail layer.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113886491B_ABST
    Figure CN113886491B_ABST
Patent Text Reader

Abstract

The present application provides a distributed data processing method, device, electronic device and storage medium, which belongs to the field of database. The method includes: determining the bucket field corresponding to the distributed data set, and determining the distributed data set as at least one data subset based on the bucket field; for each piece of data in the distributed data set, determining the bucket code of the data based on the partition storage information of the data; determining the bucket sequence number of the data bucket corresponding to each data subset based on the total number of data records or the total size of the distributed data set; for each piece of data in each data subset, determining the target bucket sequence number corresponding to the data from the bucket sequence number of the data bucket corresponding to the data subset based on the bucket code of the data, and writing the data into the data bucket corresponding to the target bucket sequence number. The implementation of the present application can store data with the same field value in the distributed data into one or more data buckets, reducing the situation of uneven data distribution.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of databases. Specifically, the present application relates to a distributed data processing method and device, an electronic device, and a computer-readable storage medium. Background Art

[0002] In the construction of data warehouses and data governance systems, in order to improve the usability of DWD (Data Warehouse Detail) and enhance data query efficiency, some dimensionality degradation techniques are usually used in the data detail layer to partition or bucket distributed data sets.

[0003] However, existing distributed data processing methods often result in duplicate data in folders at different levels, and the data is unevenly distributed, which results in more computing resources being consumed when using the data. Summary of the invention

[0004] The purpose of this application is to provide a distributed data processing method, device, electronic device and storage medium to solve at least one of the above technical problems. The solution provided in the embodiment of this application is as follows:

[0005] In a first aspect, the present application provides a distributed data processing method, comprising:

[0006] Determine a bucket field corresponding to the distributed data set, and determine the distributed data set into at least one data subset based on the bucket field, each of the data subsets corresponding to a field value of the bucket field;

[0007] For each piece of data in the above distributed data set, determine the bucket encoding of the data based on the partition storage information of the data;

[0008] Based on the total number of data records or the total size of the distributed data set, determine the bucket sequence number of the data bucket corresponding to each of the data subsets;

[0009] For each piece of data in each of the above data subsets, the target bucket number corresponding to the data is determined from the bucket number of the data bucket corresponding to the data subset based on the bucket code of the data, and the data is written into the data bucket corresponding to the above target bucket number.

[0010] In combination with the first aspect, in a first implementation of the first aspect, the determining the bucket sequence number of the data bucket corresponding to each of the data subsets based on the total number of data records of the distributed data set further includes:

[0011] Based on the total number of data records in the distributed data set and the preset number of buckets, determine the average number of data records in the bucket;

[0012] The number of data records of each of the data subsets is determined, and based on the number of data records of each of the data subsets and the average number of data records in the bucket, the bucket sequence number of the data bucket corresponding to each of the data subsets is determined.

[0013] In combination with the first implementation of the first aspect, in the second implementation of the first aspect, determining the bucket sequence number of the data bucket corresponding to each of the data subsets based on the number of data records in each of the data subsets and the average number of data records in the bucket includes:

[0014] Arrange the data subsets in descending order according to the number of data records, and determine the bucket number of the data bucket corresponding to each data subset based on the arrangement order;

[0015] For data subset i, if the number of data records corresponding to data subset i is K i If the number of data records in the bucket is equal to N, the bucket number of the data bucket corresponding to the data subset i is determined to be m+1, where when i is a positive integer greater than 1, m is the maximum bucket number in the data bucket corresponding to the previous data subset, and when i is equal to 1, m is equal to 0;

[0016] If the number of data records is K i If the number of data records in the bucket is greater than N, then the minimum bucket number of the data bucket corresponding to the data subset i is determined to be m+1, and the maximum bucket number is in, Indicates rounding up;

[0017] If the number of data records K i If the number of data records in the bucket is less than N, the bucket number of the data bucket corresponding to the data subset i to the data subset j is determined to be m+1, where j is the value such that The smallest positive integer greater than or equal to the average number of data records in the bucket N. r is the index of the data subset. K is the sum of the number of data records in all data subsets from data subset i to data subset j, r Represents the number of data records in data subset r, where r is a positive integer greater than i and less than or equal to j.

[0018] In combination with the first aspect, in a third implementation of the first aspect, determining the bucket sequence number of the data bucket corresponding to each of the data subsets based on the total data size of the distributed data set includes:

[0019] Determine the number of buckets based on the total data size of the distributed data set and the preset bucket data size;

[0020] The data size of each of the data subsets is determined, and based on the data size of each of the data subsets and the number of buckets, the bucket number of the data bucket corresponding to each of the data subsets is determined.

[0021] In combination with the third implementation of the first aspect, in a fourth implementation of the first aspect, determining the bucket sequence number of the data bucket corresponding to each of the data subsets based on the data size of each of the data subsets and the number of buckets includes:

[0022] Determine the proportion of the data size of each of the above data subsets to the total size of the above data;

[0023] Based on the data size ratio corresponding to each of the above data subsets and the above bucket number, the bucket sequence number of the data bucket corresponding to each of the above data subsets is determined, wherein the bucket sequence number of the data bucket corresponding to each of the above data subsets is continuous.

[0024] In combination with the first aspect, in a fifth implementation of the first aspect, for each piece of data in the distributed data set, determining the bucket encoding of the data based on the partition storage information of the data includes:

[0025] The bucket code of the data is determined based on the data row sequence number of the data, the partition index number of the partition where the data is located, and the total number of partitions.

[0026] In combination with the first aspect, in a sixth implementation of the first aspect, for each piece of data in each of the data subsets, determining the target bucket sequence number corresponding to the data from the bucket sequence number of the data bucket corresponding to the data subset based on the bucket code of the data includes:

[0027] Determine the number of buckets and the minimum bucket sequence number of the data buckets corresponding to the data subset;

[0028] The number of buckets corresponding to the data subset is used as the divisor, and the remainder of the bucket encoding of the data is obtained;

[0029] The sum of the remainder and the minimum bucket number is determined as the target bucket number of the data bucket corresponding to the data subset.

[0030] In combination with the first aspect, in a seventh implementation of the first aspect, further comprising:

[0031] Obtaining a structured query statement, and parsing the structured query statement to obtain a logical execution plan corresponding to the structured query statement;

[0032] Determine the query key information in the above logic execution plan, and determine the bucket storage information of the target data corresponding to the above query key information;

[0033] Based on the logical execution plan and the bucket storage information, a physical execution plan corresponding to the structured query statement is determined, and the physical execution plan is executed.

[0034] In a second aspect, the present application provides a distributed data processing device, including:

[0035] A data processing module, configured to determine a bucket field corresponding to a distributed data set, and to determine the distributed data set into at least one data subset based on the bucket field, each of the data subsets corresponding to a field value of the bucket field;

[0036] A bucket code determination module, used to determine the bucket code of each piece of data in the above distributed data set based on the partition storage information of the data;

[0037] A data bucket determination module, used to determine the bucket sequence number of the data bucket corresponding to each of the above data subsets based on the total number of data records or the total size of the above distributed data set;

[0038] The data writing module is used to determine the target bucket number corresponding to each data in each of the above data subsets from the bucket number of the data bucket corresponding to the data subset based on the bucket code of the data, and write the data into the data bucket corresponding to the above target bucket number.

[0039] In a third aspect, the present application provides an electronic device comprising a memory and a processor; the memory stores a computer program; and the processor is used to execute the method provided in the first aspect and any one of the aspects thereof when running the computer program.

[0040] In a fourth aspect, the present application provides a computer-readable storage medium, in which a computer program is stored. When the computer program is executed by a processor, the method provided in the first aspect and any one of the aspects thereof is executed.

[0041] Compared with the prior art, the technical solution provided by this application has the following beneficial effects:

[0042] The implementation of this application can store data with the same field values ​​in distributed data into one or more data buckets, reduce data skew and uneven distribution, make data distribution more reasonable, and reduce the consumption of computing resources when querying data, thereby improving the usability of the data detail layer. BRIEF DESCRIPTION OF THE DRAWINGS

[0043] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings required for use in describing the embodiments of the present application are briefly introduced below.

[0044] Figure 1 A flowchart of a distributed data processing method provided for an embodiment of the present application;

[0045] Figure 2a A schematic diagram of distributed data storage in related technologies;

[0046] Figure 2b A schematic diagram of storage of distributed data in an embodiment of the present application;

[0047] Figure 3 A schematic diagram of a process for determining a physical execution plan in an embodiment of the present application;

[0048] Figure 4 A schematic diagram of the structure of a distributed data processing device provided in one embodiment of the present application;

[0049] Figure 5 A schematic diagram of the structure of an electronic device provided in this application. DETAILED DESCRIPTION

[0050] The embodiments of the present application are described in detail below, and examples of the embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals throughout represent the same or similar elements or elements having the same or similar functions. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the present application, and cannot be interpreted as limiting the present invention.

[0051] It will be understood by those skilled in the art that, unless expressly stated, the singular forms "one", "said", and "the" used herein may also include plural forms. It should be further understood that the term "comprising" used in the specification of the present application refers to the presence of the features, integers, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or groups thereof. It should be understood that when we refer to an element as being "connected" or "coupled" to another element, it may be directly connected or coupled to the other element, or there may be an intermediate element. In addition, the "connection" or "coupling" used herein may include wireless connection or wireless coupling. The term "and / or" used herein includes all or any unit and all combinations of one or more associated listed items.

[0052] In order to make the objectives, technical solutions and advantages of the present application clearer, various optional implementation modes of the present application will be described in detail below in combination with specific embodiments and drawings.

[0053] Figure 1 A distributed data processing method provided by an embodiment of the present application is shown in the figure. The method can be performed by an electronic device provided by an embodiment of the present application. Specifically, the electronic device can be a server. The server can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server providing cloud computing services. The present application does not limit this. Specifically, the method includes the following steps S101-104:

[0054] Step S101: determine a bucket field corresponding to a distributed data set, and determine the distributed data set into at least one data subset based on the bucket field.

[0055] Specifically, the distributed data set is the data to be stored in buckets, such as the incremental data in the source data layer (Operational Data Store, ODS), the data in the Hive table, etc.

[0056] The bucket field is any common field corresponding to each data in the distributed data set, or a pre-specified field, such as the bucket field pre-defined in the Hive table. Therefore, the data corresponding to a field value of the bucket field in the distributed data set can be determined as a data subset, and then the distributed data set can be determined as at least one data subset based on the bucket field. A data subset corresponds to a field value of the bucket field, and the field values ​​corresponding to different data subsets are different.

[0057] Among them, a distributed dataset is an immutable, partitionable data set whose elements can be calculated in parallel. In addition, a distributed dataset can automatically tolerate faults and be location-aware, and is scalable, making it suitable for datasets calculated on large-scale or ultra-large-scale clusters.

[0058] Step S102: for each piece of data in the distributed data set, determine the bucket encoding of the data based on the partition storage information of the data.

[0059] Specifically, for each piece of data in a distributed data set, the partition storage information of the data may include the data row sequence number of the data in the iterator of the corresponding partition, the index number of the partition where it is located, and the total number of partitions corresponding to the distributed data set.

[0060] Furthermore, based on the data row number (starting from 0), the partition index number of the partition where the data is located, and the total number of partitions, the bucket code of the data can be determined. The calculation formula can be:

[0061] Bucket code = data row number * total number of partitions + current partition index number.

[0062] The bucket code of any data is used to determine the target bucket number of the data bucket into which the data is written, and any data has a unique long integer bucket code, that is, any data has a corresponding bucket code, and the bucket code is divisible.

[0063] Furthermore, each data in the distributed data set can be redefined based on the bucket code of each data, that is, the distributed key-value data corresponding to the distributed data set is determined, and the specified key value of any key-value pair is the bucket code and the field value corresponding to the corresponding data subset, and the value is the corresponding data. Based on this, each piece of data, bucket code and corresponding field value in the distributed data set can be uniformly defined and recorded.

[0064] Step S103: Based on the total number of data records or the total size of data in the distributed data set, determine the bucket sequence number of the data bucket corresponding to each data subset.

[0065] Specifically, based on the total number of data records in the distributed data set, the bucket sequence number of the data bucket corresponding to each data subset can be determined. Wherein, any data subset corresponds to at least one data bucket, and when each data subset corresponds to multiple data buckets, the bucket sequence numbers of the multiple data buckets corresponding to the data subset are continuous.

[0066] When determining the bucket sequence number of the data bucket corresponding to each data subset, the total number of data records of all data in the distributed data set and the number of data records of all data in each data subset may be determined first, and the preset number of buckets may be determined.

[0067] The preset number of buckets is a predefined number of data buckets when bucketing each data in a distributed data set.

[0068] Furthermore, based on the total number of data records of the distributed data set and the preset number of buckets, the average number of data records in a bucket can be determined. The average number of data records in a bucket indicates the number of data records in each data bucket when all data of the distributed data set are evenly written into data buckets of the preset number of buckets.

[0069] Furthermore, based on the number of data records in each data subset and the average number of data records in the bucket, the bucket number of the data bucket corresponding to each data subset can be determined. Specifically, for each data subset, the number of buckets of the data bucket corresponding to the data subset can be determined based on the number of data records in the data subset and the average number of data records in the bucket, and the bucket number of the data bucket corresponding to the data subset can be determined based on the number of buckets of the data bucket corresponding to the data subset.

[0070] After determining the number of data records of each data subset, the data subsets may be arranged in descending order of the number of data records, and the bucket number of the data bucket corresponding to each data subset may be determined in sequence based on the arrangement order.

[0071] For data subset i, if the number of data records K corresponding to data subset i iIf the number of data records in the bucket is equal to the average number of data records in the bucket N, the number of buckets corresponding to data subset i is determined to be 1, and the bucket number of the corresponding data bucket is m+1. When i is a positive integer greater than 1, m is the maximum bucket number in the data bucket corresponding to the previous data subset i-1 of data subset i. When i is equal to 1, that is, data subset i is the first data subset, the bucket number corresponding to data subset i is 1.

[0072] As an example, if the number of data records corresponding to data subset i is 10 and the average number of data records in a bucket is 10, when data subset i is the first data subset, the sequence number of the data bucket corresponding to data subset i can be determined to be 1. When data subset i is not the first data subset, the sequence number of the data bucket corresponding to data subset i can be determined to be the maximum bucket sequence number corresponding to data subset i-1 plus 1.

[0073] Furthermore, if the number of data records K corresponding to the data subset i is i If the number of data records in the bucket is greater than N, then the minimum bucket number of the data bucket corresponding to the data subset i is determined to be m+1, and the maximum bucket number is in, Indicates rounding up.

[0074] That is, determine the quotient of the number of data records Ki and the average number of records in the bucket N, and round up the quotient to get the number of data buckets corresponding to the data subset i: Similarly, when i is a positive integer greater than 1, m is the maximum bucket number in the data bucket corresponding to the previous data subset i-1 of data subset i, and the minimum bucket number of the data bucket corresponding to data subset i is m+1. Based on the minimum bucket number and the number of buckets corresponding to data subset i, The maximum bucket number corresponding to data subset i can be obtained. When i is equal to 1, that is, data subset i is the first data subset, the minimum bucket number of the data bucket corresponding to data subset i is 1, and the maximum bucket number is

[0075] As an example, if the number of data records corresponding to data subset i is 20 and the average number of records in a bucket is 8, the number of buckets corresponding to data subset i can be obtained by rounding up the quotient of the number of data records 20 and the average number of records in a bucket 8 to 3. In the case where data subset i is the first data subset, it can be determined that the minimum bucket number corresponding to data subset i is 1 and the maximum bucket number is 3. In the case where data subset i is not the first data subset, assuming that the maximum bucket number corresponding to data subset i-1 is 4, the minimum bucket number corresponding to data subset i is 5 and the maximum bucket number is 8.

[0076] Furthermore, if the number of data records K iIf the number of data records in the bucket is less than N, the bucket number of the data bucket corresponding to the data subset i to the data subset j is determined to be m+1, where j is the value such that The smallest positive integer greater than or equal to the average number of data records in the bucket N. r is the index of the data subset. K is the sum of the number of data records in all data subsets from data subset i to data subset j, r Represents the number of data records in data subset r, where r is a positive integer greater than i and less than or equal to j.

[0077] That is, the number of data records K in data subset i i When the number of data records is less than the average number of data records in the bucket N, it can be determined that data subset i needs to share the same data bucket with at least one other data subset. Therefore, the number of data records in data subset i and the number of data records in the subsequent data subsets can be added one by one in the order of the number of data records from large to small, until the sum of the number of data records is just equal to or greater than the average number of data records in the bucket N. At this time, the bucket numbers of the last data subset j, data subset i, and the data subsets between data subset i and data subset j whose data records are added are determined to be m+1. And when i is a positive integer greater than 1, m is the maximum bucket number in the data bucket corresponding to the previous data subset i-1 of data subset i, and the bucket numbers of data subset i to data subset j are all m+1. When i is equal to 1, that is, data subset i is the first data subset, the bucket numbers of the data buckets corresponding to data subset i to data subset j are all 1. Among them, the sum of the number of data records of the above-mentioned data subsets can be recorded based on a temporary accumulator, and the initial value of the temporary accumulator is 0. After determining the data bucket numbers corresponding to data subset i to data subset j, the initial value of the temporary accumulator needs to be reinitialized to 0.

[0078] As an example, if the number of data records corresponding to data subset i is 10 and the average number of records in the bucket is 26, since the number of data records corresponding to data subset i is less than the average number of records in the bucket, the number of data records corresponding to data subset i can be recorded based on the temporary accumulator, and the value of the temporary accumulator is 10 at this time. Furthermore, if the number of data records in data subset i+1 is 9, the sum of the number of data records of data subset i and data subset i+1, 19, is recorded based on the temporary accumulator, and it is determined whether the value of the temporary accumulator is greater than or equal to the average number of records in the bucket. At this time, the value of the temporary accumulator is less than the average number of records in the bucket, and the sum of the number of data records from data subset i to data subset i+2 is recorded based on the temporary accumulator. If the number of data records in data subset i+2 is 8, the value of the temporary accumulator is 27 at this time, which is just greater than the average number of records in the bucket, 26. In this case, it can be determined that the bucket numbers of data subset i, data subset i+1 and data subset i+2 are the same, and when data subset i is the first data subset, the bucket numbers corresponding to the above three data subsets are 1. When data subset i is not the first data subset, assuming that the maximum bucket number corresponding to data subset i-1 is 5, the bucket numbers corresponding to the above three data subsets are all 6.

[0079] Based on the above method, the number of data buckets of each data subset and the corresponding data bucket sequence number can be determined in descending order of the number of records, and for each data subset, the corresponding data bucket sequence numbers are continuous.

[0080] Optionally, based on the total data size of the distributed data set, the bucket sequence number of the data bucket corresponding to each data subset can be determined, wherein any data subset corresponds to at least one data bucket, and when each data subset corresponds to multiple data buckets, the bucket sequence numbers of the multiple data buckets corresponding to the data subset are continuous.

[0081] Among them, the data size can also be called file size or data volume. For the convenience of description, the total data size is uniformly used to describe the total data volume or total file size of the distributed data set, and the data size is used to describe the data volume or file size of each data subset.

[0082] When determining the bucket sequence number of the data bucket corresponding to each data subset, the total data size of all data in the distributed data set and the data size of all data in each data subset may be determined first, and a preset bucket data size may be determined.

[0083] The preset bucket data size is a predefined data size threshold of each data bucket when bucketing each data in the distributed data set.

[0084] Furthermore, based on the total data size of the distributed data set and the preset bucket data, the number of buckets can be determined. The number of buckets represents the number of data buckets required to write all data of the distributed data set into the data bucket if the data size threshold of each data bucket is the preset bucket data size.

[0085] Furthermore, based on the data size and the number of buckets of each data subset, the bucket number of the data bucket corresponding to each data subset can be determined. Specifically, for each data subset, the number of buckets of the data bucket corresponding to the data subset can be determined based on the data size and the number of buckets of the data subset, and then the bucket number of the data bucket corresponding to the data subset can be determined based on the number of buckets of the data bucket corresponding to the data subset.

[0086] Among them, each data subset can be arranged according to a certain basis, and the bucket sequence number of the data bucket corresponding to each data subset can be determined in sequence, so that the bucket sequence number of the data bucket corresponding to each data subset is continuous. For example, the data can be arranged in order from large to small in size, and the bucket sequence number of the data bucket corresponding to each data subset can be determined in this order.

[0087] Alternatively, for each data subset, the number of data buckets corresponding to the data subset may be determined based on the data size and the number of buckets of the data subset, and then the bucket numbers of the data buckets of each data subset pair may be determined and sorted to make them continuous.

[0088] When determining the bucket number of the data bucket corresponding to each data subset based on the data size and the number of buckets of each data subset, it is possible to determine the data proportion of the data size of each data subset to the total data size of the distributed data set, and then determine the bucket number of the data bucket corresponding to each data subset based on the data size proportion and the number of buckets corresponding to each data subset.

[0089] That is, the proportion of the data size corresponding to each data subset can be used as the proportion of the number of buckets of the data buckets corresponding to each data subset to the number of buckets, and then based on the proportion of the number of buckets of the data buckets corresponding to each data subset, the number of buckets of the data buckets corresponding to each data subset and the bucket sequence number of the data buckets corresponding to each data subset are determined. Among them, the bucket sequence numbers of the data buckets corresponding to each data subset are continuous, that is, the data buckets corresponding to each data subset are one or more continuous data buckets.

[0090] Among them, for each data subset, the number of data buckets corresponding to the data subset is an integer. Therefore, the product of the proportion of the number of data buckets corresponding to the data subset and the number of buckets can be rounded up to obtain the number of data buckets corresponding to the data subset.

[0091] As an example, suppose that each data subset is arranged in order of data size from large to small, and data subset 1, data subset 2 and data subset 3 are obtained, and the data sizes of data subset 1, data subset 2 and data subset 3 are 128M, 64M and 32M respectively. It can be seen that the total data size of the distributed data set is 224, the data size of data subset 1 accounts for 4 / 7 of the total data size, the data size of data subset 2 accounts for 2 / 7 of the total data size, and the data size of data subset 3 accounts for 1 / 7 of the total data size. If the number of buckets determined based on the preset bucket data size and the total data size is 5 for data subset 1, then the number of buckets corresponding to data subset 1 is The number of data buckets corresponding to data subset 2 is The number of data corresponding to data subset 3 is

[0092] The bucket numbers of the data buckets corresponding to data subset 1, data subset 2 and data subset 3 can be further determined continuously, that is, the bucket numbers corresponding to data subset 1 are 1-3, the bucket numbers corresponding to data subset 2 are 4-5, and the bucket number corresponding to data subset 3 is 6.

[0093] Optionally, the data of each data subset can be stored in ORC (Optimized Row Columnar) format, and a data bucket corresponding to each data subset can be regarded as an ORC data file. After determining the data buckets corresponding to each data subset based on any of the above methods, the data corresponding to each field value of the bucket field can be stored in one or more ORC data files, thereby avoiding problems such as data skew and a large number of small files. Based on this, after determining the data buckets and their serial numbers corresponding to each data subset based on any of the above methods, bucket dictionary data can be generated to uniformly record the bucketing of the data corresponding to each field value of the bucket field.

[0094] like Figure 2a As shown in the figure, assuming that the bucket field is EID, after processing the data in the distributed data set based on the existing technology, the data corresponding to the different field values ​​of the bucket field EID are stored in the same ORC file. Figure 2b The storage conditions shown are as follows: Figure 2bThe bucket dictionary data may be (EID1→[ORC-file1, ORC-file2], EID2→[ORC-file3], EID3→[ORC-file4], EID4→[ORC-file4]). That is, there are only data records with EID='EID1' in ORC-file1 and ORC-file2 data files. And EID='EID2' will only be stored in ORC-file3 file, and all data records with EID='EID3' and EID='EID4' will only exist in ORC-file4 data file and will not appear in other ORC data files.

[0095] Step S104: for each piece of data in each data subset, determine the target bucket number corresponding to the data from the bucket number of the data bucket corresponding to the data subset based on the bucket code of the data, and write the data into the data bucket corresponding to the target bucket number.

[0096] Specifically, for each piece of data in each data subset, the number of buckets and the minimum bucket number of the data bucket corresponding to the data subset can be determined first, and the number of buckets corresponding to the data subset is used as the divisor to obtain the corresponding remainder by taking the remainder of the bucket encoding of the data.

[0097] Among them, the remainder obtained by taking the remainder of the bucket encoding of the data is less than the number of data buckets corresponding to the data subset, and its minimum value is 0, and the maximum value is the number of data buckets corresponding to the data subset minus 1.

[0098] Furthermore, the remainder corresponding to the data subset and the minimum bucket sequence number are summed, and the sum of the minimum bucket sequence number and the remainder is determined as the target bucket sequence number corresponding to the data, and the data is written into the data bucket corresponding to the target bucket sequence number.

[0099] Among them, since the remainder is less than the number of buckets of the data bucket corresponding to the data subset, the bucket number corresponding to the sum of the minimum bucket number and the remainder is greater than or equal to the minimum bucket number corresponding to the data subset, and less than or equal to the maximum bucket number corresponding to the data subset. And in the case where each data in the data subset has a unique bucket code, the target bucket number corresponding to the data can be determined from the bucket numbers of each data bucket corresponding to the data subset.

[0100] As an example, the number of buckets corresponding to data subset i is 16, and the bucket numbers are 17 to 32. For any data in data subset i, the possible values ​​of the remainder obtained after taking the remainder of the bucket encoding of the data based on the number of buckets 16 corresponding to the data subset are 0 to 15, so the value of the sum of the minimum bucket number and the remainder is 17 to 32, so the sum of the minimum bucket number and the remainder can be determined as the target bucket number corresponding to data subset i.

[0101] Among them, the number of buckets and the minimum bucket sequence number corresponding to each data subset can be determined based on the bucket dictionary data, or can be obtained based on statistics of distributed key-value data, and there is no limitation here.

[0102] After step S104, the distributed data processing method provided in the embodiment of the present application may further include a data query process. Figure 3 The process is further described, wherein: Figure 3 A schematic diagram of the process of determining a physical execution plan is shown.

[0103] Specifically, an SQL statement (Structured Query Language) may be obtained, and the SQL statement may be parsed to obtain a logical execution plan corresponding to the SQL statement. Specifically, the SQL statement may be parsed by a Spark or Iceberg parser to convert the statement into a corresponding logical execution plan.

[0104] Furthermore, after obtaining the logical execution plan, the logical execution plan can be analyzed to determine the query key information in the logical execution plan. Specifically, the obtained logical execution plan can be parsed for partition information and constant information to obtain the query key information. For example, if the SQL expression is: select c2, c3 from table name where dt = '2020-01-01' and ety = 'custom' and f1 in ('A1', 'A2'), the query key information obtained after extracting the partition information and constant information is ([dt→[2020-01-01], ety→['custom'], f1→['A1', 'A2']]).

[0105] Furthermore, based on the obtained query key information, the bucket storage information of the target data to be queried by the SQL statement can be queried, that is, the HDFS (Hadoop Distributed File System) path where the target data is located. Specifically, the query can be performed based on multiple pieces of information in the query key information, such as: hdfs: / / cluster / hive / db / table_name / first-level bucket storage information / second-level bucket storage information / bucket_dict / dict_meta.parquet. Furthermore, based on the bucket storage information and the parsed logical execution plan obtained after parsing the logical execution plan, a physical execution plan corresponding to the SLQ statement is constructed and executed.

[0106] The implementation of this application can centrally store data with the same field value in distributed data into one or more data buckets, reducing data skew and uneven distribution, making data distribution more reasonable. And by optimizing SQL statements, it can reduce the consumption of computing resources when querying data and improve the usability of the data detail layer.

[0107] Corresponding to the distributed data processing method provided in the present application, the present application embodiment also provides a distributed data processing device 400, whose structural diagram is shown in FIG. Figure 4 As shown in , the distributed data processing device 400 includes: a data processing module 401, a bucket encoding determination module, a data bucket determination module and a data writing module.

[0108] The data processing module 401 is used to determine a bucket field corresponding to the distributed data set, and determine the distributed data set into at least one data subset based on the bucket field, each of which corresponds to a field value of the bucket field;

[0109] A bucket code determination module 402 is used to determine the bucket code of each piece of data in the above distributed data set based on the partition storage information of the data;

[0110] A data bucket determination module 403 is used to determine the bucket sequence number of the data bucket corresponding to each of the data subsets based on the total number of data records or the total size of the distributed data set;

[0111] The data writing module 404 is used to determine the target bucket number corresponding to each data in each of the above data subsets from the bucket number of the data bucket corresponding to the data subset based on the bucket code of the data, and write the data into the data bucket corresponding to the above target bucket number.

[0112] Optionally, the data bucket determination module 403 is used to:

[0113] Based on the total number of data records in the distributed data set and the preset number of buckets, determine the average number of data records in the bucket;

[0114] The number of data records of each of the data subsets is determined, and based on the number of data records of each of the data subsets and the average number of data records in the bucket, the bucket sequence number of the data bucket corresponding to each of the data subsets is determined.

[0115] Optionally, the data bucket determination module 403 is used to:

[0116] Arrange the data subsets in descending order according to the number of data records, and determine the bucket number of the data bucket corresponding to each data subset based on the arrangement order;

[0117] For data subset i, if the number of data records corresponding to data subset i is K i If the number of data records in the bucket is equal to N, the bucket number of the data bucket corresponding to the data subset i is determined to be m+1, where when i is a positive integer greater than 1, m is the maximum bucket number in the data bucket corresponding to the previous data subset, and when i is equal to 1, m is equal to 0;

[0118] If the number of data records K i If the number of data records in the bucket is greater than N, then the minimum bucket number of the data bucket corresponding to the data subset i is determined to be m+1, and the maximum bucket number is in, Indicates rounding up;

[0119] If the number of data records K i If the number of data records in the bucket is less than N, the bucket number of the data bucket corresponding to the data subset i to the data subset j is determined to be m+1, where j is the value such that The smallest positive integer greater than or equal to the average number of data records in the bucket N. r is the index of the data subset. K is the sum of the number of data records in all data subsets from data subset i to data subset j, r Represents the number of data records in data subset r, where r is a positive integer greater than i and less than or equal to j.

[0120] Optionally, the data bucket determination module 403 is used to:

[0121] Determine the number of buckets based on the total data size of the distributed data set and the preset bucket data size;

[0122] The data size of each of the data subsets is determined, and based on the data size of each of the data subsets and the number of buckets, the bucket number of the data bucket corresponding to each of the data subsets is determined.

[0123] Optionally, the data bucket determination module 403 is used to:

[0124] Determine the proportion of the data size of each of the above data subsets to the total size of the above data;

[0125] Based on the data size ratio corresponding to each of the above data subsets and the above bucket number, the bucket sequence number of the data bucket corresponding to each of the above data subsets is determined, wherein the bucket sequence number of the data bucket corresponding to each of the above data subsets is continuous.

[0126] Optionally, for each piece of data in the distributed data set, the bucket code determination module 402 is used to:

[0127] The bucket code of the data is determined based on the data row sequence number of the data, the partition index number of the partition where the data is located, and the total number of partitions.

[0128] Optionally, for each piece of data in each of the data subsets, the data writing module 404 is used to:

[0129] Determine the number of buckets and the minimum bucket sequence number of the data buckets corresponding to the data subset;

[0130] The number of buckets corresponding to the data subset is used as the divisor, and the remainder of the bucket encoding of the data is obtained;

[0131] The sum of the remainder and the minimum bucket number is determined as the target bucket number of the data bucket corresponding to the data subset.

[0132] Optionally, the apparatus 400 further includes an execution module, configured to:

[0133] Obtaining a structured query statement, and parsing the structured query statement to obtain a logical execution plan corresponding to the structured query statement;

[0134] Determine the query key information in the above logic execution plan, and determine the bucket storage information of the target data corresponding to the above query key information;

[0135] Based on the logical execution plan and the bucket storage information, a physical execution plan corresponding to the structured query statement is determined, and the physical execution plan is executed.

[0136] The device of the embodiment of the present application can execute the method provided by the embodiment of the present application, and the implementation principle is similar. The actions performed by each module in the device in each embodiment of the present application correspond to the steps in the method in each embodiment of the present application. For the detailed functional description of each module of the device, please refer to the description in the corresponding method shown in the previous text, which will not be repeated here.

[0137] The present application also provides an electronic device, which includes a memory and a processor; wherein a computer program is stored in the memory; and the processor is used to execute the method provided in any optional embodiment of the present application when running the computer program.

[0138] The present application also provides a computer-readable storage medium, in which a computer program is stored. When the computer program is executed by a processor, the method provided in any optional embodiment of the present application is executed.

[0139] As an alternative, Figure 5 A schematic diagram of the structure of an electronic device applicable to the embodiment of the present application is shown. Figure 5As shown, the electronic device 500 may include a processor 501 and a memory 503. The processor 501 and the memory 503 are connected, such as through a bus 502. Optionally, the electronic device 500 may also include a transceiver 504. It should be noted that in actual applications, the transceiver 504 is not limited to one, and the structure of the electronic device 500 does not constitute a limitation on the embodiments of the present application.

[0140] Processor 501 may be a CPU (Central Processing Unit), a general purpose processor, a DSP (Digital Signal Processor), an ASIC (Application Specific Integrated Circuit), an FPGA (Field Programmable Gate Array) or other programmable logic devices, transistor logic devices, hardware components or any combination thereof. It may implement or execute various exemplary logic blocks, modules and circuits described in conjunction with the disclosure of this application. Processor 501 may also be a combination that implements computing functions, such as a combination of one or more microprocessors, a combination of a DSP and a microprocessor, etc.

[0141] The bus 502 may include a path to transmit information between the above components. The bus 502 may be a PCI (Peripheral Component Interconnect) bus or an EISA (Extended Industry Standard Architecture) bus. The bus 502 may be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 5 Only one thick line is used in the diagram, but this does not mean that there is only one bus or only one type of bus.

[0142] The memory 503 can be a ROM (Read Only Memory) or other types of static storage devices that can store static information and instructions, a RAM (Random Access Memory) or other types of dynamic storage devices that can store information and instructions, or an EEPROM (Electrically Erasable Programmable Read Only Memory), a CD-ROM (Compact Disc Read Only Memory) or other optical disk storage, optical disk storage (including compressed optical disk, laser disk, optical disk, digital versatile disk, Blu-ray disk, etc.), a magnetic disk storage medium or other magnetic storage device, or any other medium that can be used to carry or store the desired program code in the form of instructions or data structures and can be accessed by a computer, but is not limited to these.

[0143] The memory 503 is used to store the application code for executing the solution of the present application, and the execution is controlled by the processor 501. The processor 501 is used to execute the application code (computer program) stored in the memory 503 to implement the content shown in any of the above method embodiments.

[0144] It should be understood that, although the steps in the flowchart of the accompanying drawings are displayed in sequence as indicated by the arrows, these steps are not necessarily executed in sequence in the order indicated by the arrows. Unless otherwise specified herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least a part of the steps in the flowchart of the accompanying drawings may include multiple sub-steps or multiple stages, and these sub-steps or stages are not necessarily executed at the same time, but can be executed at different times, and their execution order is not necessarily sequential, but can be executed in turn or alternately with other steps or at least a part of the sub-steps or stages of other steps.

[0145] The above descriptions are only some embodiments of the present invention. It should be pointed out that, for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present invention. These improvements and modifications should also be regarded as the scope of protection of the present invention.

Claims

1. A distributed data processing method, It is characterized in that The method comprises: Determine a bucket field corresponding to the distributed data set, and determine the distributed data set into at least one data subset based on the bucket field, each of the data subsets corresponding to a field value of the bucket field; For each piece of data in the distributed data set, determining a bucket code for the data based on partition storage information of the data; Determine the bucket sequence number of the data bucket corresponding to each of the data subsets based on the total number of data records or the total size of the distributed data set; For each piece of data in each of the data subsets, determine the target bucket number corresponding to the data from the bucket number of the data bucket corresponding to the data subset based on the bucket code of the data, and write the data into the data bucket corresponding to the target bucket number; The determining, based on the total number of data records of the distributed data set, the bucket sequence number of the data bucket corresponding to each of the data subsets comprises: Determine an average number of data records in a bucket based on the total number of data records in the distributed data set and a preset number of buckets; Determine the number of data records of each of the data subsets, arrange the data subsets in descending order of the number of data records, and determine the bucket sequence number of the data bucket corresponding to each of the data subsets based on the arrangement order; For data subset i, if the number of data records Ki corresponding to data subset i is equal to the average number of data records N in the bucket, then the bucket number of the data bucket corresponding to data subset i is determined to be m+1, where when i is a positive integer greater than 1, m is the maximum bucket number in the data bucket corresponding to the previous data subset, and when i is equal to 1, m is equal to 0; If the number of data records Ki is greater than the average number of data records in the bucket N, then the minimum bucket number of the data bucket corresponding to the data subset i is determined to be m+1 and the maximum bucket number is ,in, Indicates rounding up; If the number of data records Ki is less than the average number of data records in the bucket N, the bucket number of the data bucket corresponding to the data subset i to the data subset j is determined to be m+1, where j is such that The smallest positive integer greater than or equal to the average number of data records in the bucket N. r is the index of the data subset. is the sum of the number of data records of all data subsets from data subset i to data subset j, Kr represents the number of data records of data subset r, and r is a positive integer greater than i and less than or equal to j.

2. The method according to claim 1, It is characterized in that The determining, based on the total data size of the distributed data set, the bucket sequence number of the data bucket corresponding to each of the data subsets comprises: Determine the number of buckets based on the total data size of the distributed data set and a preset bucket data size; The data size of each of the data subsets is determined, and based on the data size of each of the data subsets and the number of buckets, the bucket sequence number of the data bucket corresponding to each of the data subsets is determined.

3. The method according to claim 2, It is characterized in that The determining, based on the data size of each of the data subsets and the number of buckets, the bucket sequence number of the data bucket corresponding to each of the data subsets comprises: Determine the data proportion of the data size of each of the data subsets to the total data size; Based on the data size ratio corresponding to each of the data subsets and the number of buckets, the bucket sequence number of the data bucket corresponding to each of the data subsets is determined, wherein the bucket sequence number of the data bucket corresponding to each of the data subsets is continuous.

4. The method according to claim 1, It is characterized in that For each piece of data in the distributed data set, determining the bucket encoding of the data based on the partition storage information of the data includes: The bucket code of the data is determined based on the data row sequence number of the data, the partition index number of the partition where the data is located, and the total number of partitions.

5. The method according to claim 1, It is characterized in that For each piece of data in each of the data subsets, determining the target bucket sequence number corresponding to the data from the bucket sequence number of the data bucket corresponding to the data subset based on the bucket code of the data includes: Determine the number of buckets and the minimum bucket sequence number of the data buckets corresponding to the data subset; The number of buckets corresponding to the data subset is used as the divisor, and the remainder of the bucket encoding of the data is obtained; The sum of the remainder and the minimum bucket sequence number is determined as the target bucket sequence number of the data bucket corresponding to the data subset.

6. The method according to claim 1, It is characterized in that The method further comprises: Obtaining a structured query statement, and parsing the structured query statement to obtain a logical execution plan corresponding to the structured query statement; Determine the query key information in the logical execution plan, and determine the bucket storage information of the target data corresponding to the query key information; Based on the logical execution plan and the bucket storage information, a physical execution plan corresponding to the structured query statement is determined, and the physical execution plan is executed.

7. A distributed data processing device, It is characterized in that The device comprises: A data processing module, configured to determine a bucket field corresponding to a distributed data set, and determine the distributed data set into at least one data subset based on the bucket field, each of the data subsets corresponding to a field value of the bucket field; A bucket code determination module, configured to determine, for each piece of data in the distributed data set, a bucket code of the data based on the partition storage information of the data; A data bucket determination module, used to determine the bucket sequence number of the data bucket corresponding to each of the data subsets based on the total number of data records or the total size of the distributed data set; A data writing module is used to determine, for each piece of data in each of the data subsets, a target bucket sequence number corresponding to the data from the bucket sequence number of the data bucket corresponding to the data subset based on the bucket code of the data, and write the data into the data bucket corresponding to the target bucket sequence number; The data bucket determination module is used to: Determine an average number of data records in a bucket based on the total number of data records in the distributed data set and a preset number of buckets; Determine the number of data records of each of the data subsets, arrange the data subsets in descending order of the number of data records, and determine the bucket sequence number of the data bucket corresponding to each of the data subsets based on the arrangement order; For data subset i, if the number of data records Ki corresponding to data subset i is equal to the average number of data records N in the bucket, then the bucket number of the data bucket corresponding to data subset i is determined to be m+1, where when i is a positive integer greater than 1, m is the maximum bucket number in the data bucket corresponding to the previous data subset, and when i is equal to 1, m is equal to 0; If the number of data records Ki is greater than the average number of data records in the bucket N, then the minimum bucket number of the data bucket corresponding to the data subset i is determined to be m+1 and the maximum bucket number is ,in, Indicates rounding up; If the number of data records Ki is less than the average number of data records in the bucket N, the bucket number of the data bucket corresponding to the data subset i to the data subset j is determined to be m+1, where j is such that The smallest positive integer greater than or equal to the average number of data records in the bucket N. r is the index of the data subset. is the sum of the number of data records of all data subsets from data subset i to data subset j, Kr represents the number of data records of data subset r, and r is a positive integer greater than i and less than or equal to j.

8. An electronic device, It is characterized in that including memory and processor; The memory stores a computer program; The processor is configured to execute the method according to any one of claims 1 to 6 when running the computer program.

9. A computer-readable storage medium, It is characterized in that The storage medium stores a computer program, and when the computer program is executed by a processor, the method according to any one of claims 1 to 6 is executed.

Citation Information

Patent Citations

  • Method, device and equipment for fusing data among multiple platforms

    CN110046638A

  • Data query method and device, computer equipment and storage medium

    CN112650759A