Data grouping statistics method and system for distributed database
By allowing data nodes to perform group statistics in a distributed database and pooling the results to the server, the problem of inefficiency in the server when processing large data sets is solved, achieving faster response time and higher efficiency.
Patent Information
- Application Number
- CN202010752066.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-07-30
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2040-07-30
AI Technical Summary
In distributed databases, when grouping large data sets, the server is less efficient and has a long response time.
The server obtains the data fragment type of the data set to be grouped, instructs N data nodes to perform group statistics according to this type, and obtains the target packet statistics result set. The server then collects these results to determine the final statistical result.
It reduces the amount of data that the server needs to process, reduces the server's response time, and improves the server's packet statistics efficiency.
Smart Images

Figure CN113254493B_ABST
Abstract
Description
Technical Field
[0001] The present application belongs to the technical field of data analysis, and in particular, relates to a data grouping statistics method and system for a distributed database. Background Art
[0002] Distributed databases store the data sets to be grouped in multiple data nodes (DNs). When grouping and counting data scattered on multiple DNs, the server needs to read the qualified data from all DNs and then perform grouping and counting in the server. However, when the data set to be grouped is large, the response time of grouping and counting through the server is long, resulting in low efficiency of server grouping and counting. Summary of the invention
[0003] The embodiment of the present application provides a data grouping statistics method and system for a distributed database, which can solve the problem of low grouping statistics efficiency of the server when performing grouping statistics on a large data set in the prior art.
[0004] In a first aspect, an embodiment of the present application provides a data grouping statistics method for a distributed database, the data grouping statistics method is applied to a data grouping statistics system for a distributed database, the data grouping statistics system includes a server and N data nodes, N is an integer greater than 1, and the data sets to be grouped are distributed on the N data nodes, the data grouping statistics method includes:
[0005] The server obtains the data shard type of the data set to be grouped;
[0006] The N data nodes perform one or two grouping statistics on their respective local data sets according to the data sharding type of the data set to be grouped, and obtain N target grouping statistical result sets, wherein the local data set of a data node is the data set of the data set to be grouped distributed on the data node, and one data node corresponds to one target grouping statistical result set;
[0007] The server performs statistics on the N target group statistical result sets to determine target statistical results.
[0008] In a second aspect, an embodiment of the present application provides a data grouping statistics system for a distributed database, the data grouping statistics system comprising a server and N data nodes, N being an integer greater than 1, and the data sets to be grouped are distributed on the N data nodes;
[0009] The server is used to obtain the data shard type of the data set to be grouped;
[0010] The N data nodes are used to perform one or two grouping statistics on their respective local data sets according to the data sharding type of the data set to be grouped, to obtain N target grouping statistics result sets, wherein the local data set of a data node is the data set of the data set to be grouped distributed on the data node, and one data node corresponds to one target grouping statistics result set;
[0011] The server is used to perform statistics on the N target group statistical result sets to determine target statistical results.
[0012] Compared with the prior art, the embodiments of the present application have the following beneficial effects: the present application performs data grouping statistics on the data set to be grouped distributed in the N data nodes according to the server and the N data nodes, obtains the data sharding type of the data set to be grouped through the server, and enables the N data nodes to perform one or two grouping statistics according to the data sharding type of the data set to be grouped, and the server then performs statistics on the grouping statistics results of the N data nodes to obtain the final statistical results. That is, the present application can perform grouping statistics on the data set to be grouped through the N data nodes, and when dealing with a large amount of data, there is no need to perform grouping through the server, which avoids the server from retrieving a large amount of data from the N data nodes to form a data set to be grouped, reduces the response time of the server, and improves the efficiency of server grouping statistics. BRIEF DESCRIPTION OF THE DRAWINGS
[0013] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative labor.
[0014] Figure 1 It is a flowchart of a data grouping and statistics method of a distributed database provided in Example 1 of the present application;
[0015] Figure 2 It is a flowchart of hash sharding data grouping statistics provided in Example 1 of the present application;
[0016] Figure 3 It is a flowchart of a data grouping and statistics method of a distributed database provided in Example 2 of the present application;
[0017] Figure 4 It is a flowchart of non-hash sharding big data grouping statistics provided in Example 2 of the present application;
[0018] Figure 5 It is a flowchart of non-hash fragmentation small data grouping statistics provided in Example 2 of the present application;
[0019] Figure 6 It is a flowchart of the class merging and grouping statistics method provided in Example 2 of the present application;
[0020] Figure 7 This is a schematic diagram of the network architecture of the data grouping statistics system of the distributed database provided in Example 3 of the present application. DETAILED DESCRIPTION
[0021] In the following description, specific details such as specific system structures, technologies, etc. are provided for the purpose of illustration rather than limitation, so as to provide a thorough understanding of the embodiments of the present application. However, it should be clear to those skilled in the art that the present application may also be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to prevent unnecessary details from obstructing the description of the present application.
[0022] It should be understood that when used in the present specification and the appended claims, the term "comprising" indicates the presence of described features, wholes, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, wholes, steps, operations, elements, components and / or combinations thereof.
[0023] It should also be understood that the term “and / or” used in the specification and appended claims refers to any and all possible combinations of one or more of the associated listed items, and includes these combinations.
[0024] As used in the specification and appended claims of this application, the term "if" can be interpreted as "when" or "uponce" or "in response to determining" or "in response to detecting", depending on the context. Similarly, the phrase "if it is determined" or "if [described condition or event] is detected" can be interpreted as meaning "uponce it is determined" or "in response to determining" or "uponce [described condition or event] is detected" or "in response to detecting [described condition or event]", depending on the context.
[0025] In addition, in the description of the present application specification and the appended claims, the terms "first", "second", "third", etc. are only used to distinguish the descriptions and cannot be understood as indicating or implying relative importance.
[0026] References to "one embodiment" or "some embodiments" etc. described in the specification of this application mean that one or more embodiments of the present application include specific features, structures or characteristics described in conjunction with the embodiment. Therefore, the statements "in one embodiment", "in some embodiments", "in some other embodiments", "in some other embodiments", etc. that appear in different places in this specification do not necessarily refer to the same embodiment, but mean "one or more but not all embodiments", unless otherwise specifically emphasized in other ways. The terms "including", "comprising", "having" and their variations all mean "including but not limited to", unless otherwise specifically emphasized in other ways.
[0027] The data grouping statistics method of a distributed database provided in the embodiment of the present application can be applied to terminal devices such as desktop computers, notebook computers, and ultra-mobile personal computers (UMPC). The embodiment of the present application does not impose any restrictions on the specific type of terminal devices.
[0028] It should be understood that the size of the serial numbers of the steps in the above embodiments does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present application.
[0029] In order to illustrate the technical solution of the present application, a specific embodiment is provided below for illustration.
[0030] See also Figure 1 , is a flow chart of a data grouping statistics method for a distributed database provided in Embodiment 1 of the present application. The data grouping statistics method can be used in a data grouping statistics system for a distributed database. The data grouping statistics system includes a server and N data nodes, where N is an integer greater than 1. Figure 1 As shown, the data grouping statistical method may include the following steps:
[0031] Step S101: The server obtains the data shard type of the data set to be grouped.
[0032] The data set to be grouped is distributed on N data nodes, and the data set to be grouped includes M field sets, where M is an integer greater than or equal to 1, such as Figure 2 As shown, the field set of the data set to be grouped in the data set to be grouped on DN1 (data node 1) (original data of DN1) includes name (Zhang San, Li Si, etc.), class (Class 1, Class 2), course (Chinese, Mathematics) and score (90, 96, etc.).
[0033] The data set to be grouped may be distributed on N data nodes in the form of a data table. The attributes of the data table include attribute information such as data sharding type. The server may obtain the data sharding type of the data set to be grouped by acquiring the attributes of the data table corresponding to the data set to be grouped. For example, the server sends an instruction for obtaining attribute information of the data set to be grouped to any N data nodes. The data node that receives the attribute information instruction feeds back the attribute information of the data set to be grouped to the server, so that the server obtains the data sharding type of the data set to be grouped.
[0034] Optionally, the server is a scheduling server, which is connected to N data nodes.
[0035] Step S102 : N data nodes perform one or two grouping statistics on their respective local data sets according to the data sharding type of the data set to be grouped, and obtain N target grouping statistics result sets.
[0036] Among them, the local data set of a data node is the data set to be grouped and distributed on the data node. N data nodes group and count their respective local data sets to obtain their respective target grouping statistical result sets. One data node corresponds to one target grouping statistical result set, and N data nodes correspond to N target grouping statistical result sets.
[0037] Optionally, the N data nodes perform one or two group statistics on their respective local data sets according to the data sharding type of the data set to be grouped, and obtain N target group statistics result sets including:
[0038] When the data sharding type of the data set to be grouped is hash sharding and the sharding field set of the hash sharding of the data set to be grouped is a subset of the target sharding field set, N data nodes perform group statistics on their respective local data sets to obtain N target grouping statistical result sets.
[0039] Among them, after obtaining the data sharding type of the data set to be grouped, the server detects whether the data sharding type of the data set to be grouped is hash sharding; if the data sharding type of the data set to be grouped is hash sharding, it detects whether the sharding field set of the hash sharding of the data set to be grouped belongs to a subset of the target sharding field set.
[0040] The target sharding field set is at least one of the M field sets. For example, the field set includes name, class, course and score. The target sharding field set can be class and course. The subset of the target sharding field set is class or course. If the data set to be grouped is hash sharding, the sharding field set of the hash sharding is class, then the sharding field set of the hash sharding of the data set to be grouped belongs to the subset of the target sharding field set. If the data set to be grouped is hash sharding, the sharding field set of the hash sharding is name, then the sharding field set of the hash sharding of the data set to be grouped does not belong to the subset of the target sharding field set.
[0041] When the data sharding type of the data set to be grouped is hash sharding, and the sharding field set of the hash sharding of the data set to be grouped is a subset of the target sharding field set, the server sends first result information to N data nodes. After receiving the first result information, the N data nodes perform group statistics on their respective local data sets to obtain N target grouping statistical result sets, wherein the first result information may refer to the data sharding type of the data set to be grouped is hash sharding, and the sharding field set of the hash sharding of the data set to be grouped is a subset of the target sharding field set.
[0042] like Figure 2 The figure shows the flow chart of data grouping statistics when the data set to be grouped is hash sharding and the sharding field set of the hash sharding is a subset of the target sharding field set. The sharding field set of the hash sharding is class. Each data node has its own local data set ( Figure 2 The original data in ( ) are grouped and counted according to "class" and "course". The statistical results are "class, course, total score and average score". d1, d2, and d3 represent the target group statistical result sets of DN1 (data node 1), DN2 (data node 2), and DN3 (data node 3) respectively. The target statistical result is obtained by gathering all the target group statistical result sets together.
[0043] Optionally, the data grouping statistical method further includes:
[0044] The server detects whether the data sharding type of the data set to be grouped is hash sharding;
[0045] If the data sharding type of the data set to be grouped is hash sharding, detecting whether the sharding field set of the hash sharding of the data set to be grouped belongs to a subset of the target sharding field set, obtaining a first detection result, and sending the first detection result to the N data nodes;
[0046] If the data sharding type of the data set to be grouped is not hash sharding, obtaining a second detection result, and sending the second detection result to the N data nodes;
[0047] Among them, the first detection result includes that the data sharding type of the data set to be grouped is hash sharding and the sharding field set of the hash sharding of the data set to be grouped is a subset of the target sharding field set, and the data sharding type of the data set to be grouped is hash sharding and the sharding field set of the hash sharding of the data set to be grouped does not belong to a subset of the target sharding field set. The second detection result refers to that the data sharding type of the data set to be grouped is not hash sharding.
[0048] The server sends the first detection result or the second detection result to the data node, and the data node executes corresponding steps according to information in the first detection result or the second detection result.
[0049] Step S103: the server collects N target group statistical result sets to determine the target statistical result.
[0050] After the N data nodes obtain their respective target group statistical result sets, the target group statistical result sets are sent to the server, and the server aggregates the N target group statistical result sets, for example, integrates the N target group statistical result sets to obtain the target statistical result. In one embodiment, the server sends a target group statistical result set acquisition instruction to the data node, waits for the data node to complete the group statistics to obtain the target group statistical result set, and the data node responds to the target group statistical result set acquisition instruction and feeds the target group statistical result set back to the server.
[0051] The embodiment of the present application can perform group statistics on the data set to be grouped through N data nodes. When dealing with a large amount of data, there is no need to group through a server, thus avoiding the server from retrieving a large amount of data from N data nodes to form the data set to be grouped, reducing the response time of the server and improving the efficiency of server group statistics.
[0052] See also Figure 3 , is a flow chart of a data grouping and statistics method of a distributed database provided in Embodiment 2 of the present application, such as Figure 3 As shown, the data grouping statistical method may include the following steps:
[0053] Step S301: The server obtains the data shard type of the data set to be grouped.
[0054] Step S302, when the data sharding type of the data set to be grouped is not hash sharding, or the sharding field set of the hash sharding of the data set to be grouped does not belong to a subset of the target sharding field set, the N data nodes perform the first grouping statistics on their respective local data sets to obtain N initial grouping statistics result sets.
[0055] Among them, after obtaining the data sharding type of the data set to be grouped, the server detects whether the data sharding type of the data set to be grouped is hash sharding; if the data sharding type of the data set to be grouped is hash sharding, it detects whether the sharding field set of the hash sharding of the data set to be grouped belongs to a subset of the target sharding field set.
[0056] When the data sharding type of the data set to be grouped is hash sharding, and the sharding field set of the hash sharding of the data set to be grouped does not belong to a subset of the target sharding field set, or the data sharding type of the data set to be grouped is not hash sharding, the server sends second result information to N data nodes. After receiving the second result information, the N data nodes perform group statistics on their respective local data sets to obtain N initial group statistical result sets, wherein the second result information may refer to the data sharding type of the data set to be grouped is hash sharding, and the sharding field set of the hash sharding of the data set to be grouped does not belong to a subset of the target sharding field set, or the data sharding type of the data set to be grouped is not information corresponding to hash sharding.
[0057] like Figure 4 The figure shows a flow chart of a data set to be grouped that is not a hash shard, where each data node performs a local data set ( Figure 4 The original data in the data set are grouped and counted by "class" and "course", and the statistical results are "class, course, total score and number". d1 represents the initial group statistical result set of DN1 (data node 1), d2 represents the initial group statistical result set of DN2 (data node 2), and d3 represents the initial group statistical result set of DN3 (data node 3).
[0058] Step S303, when the total number of rows in the N initial grouping statistical result sets is greater than the row number threshold, the N data nodes perform hash sharding on their respective initial grouping statistical result sets according to the target sharding field set to obtain their respective hash sharding result sets.
[0059] Among them, after the N data nodes obtain their respective initial group statistical result sets, they determine whether it is necessary to perform hash sharding on their respective initial group statistical result sets according to the total number of rows in the N initial group statistical result sets. Since the total number of rows is large and the data volume of the initial group statistical result sets is large, in order to ensure the statistical efficiency, when the total number of rows is greater than the row number threshold, the respective initial group statistical result sets are hash sharded to obtain a hash shard result set, and a data node performs group statistics on its own hash shard result set to obtain the target group statistical result set of the data node.
[0060] Optionally, the N data nodes perform hash sharding on their respective initial grouping statistical result sets according to the target sharding field set, and obtain their respective hash sharding result sets including:
[0061] N data nodes use hash sharding to distribute each row of data in their initial grouping statistical result sets to corresponding data nodes according to the target sharding field set, and obtain their own hash sharding result sets.
[0062] Each row of data in the initial group statistics result set may be a set of initial group statistics results, such as Figure 4 As shown, in the initial grouping statistical result set d1 of DN1, "Class 1 Chinese 94 1" is a row of data. The initial grouping statistical result set d1 includes 6 rows of data. Each row of data is distributed to a corresponding data node according to the target sharding field set.
[0063] like Figure 4 As shown in the figure, if the total number of rows of the initial group statistics result set of d1, d2, and d3 is greater than the row number threshold, the data node performs hash sharding on its own initial group statistics result set. Figure 4 DN1 distributes "Class 4 Chinese 90 1" and "Class 4 Mathematics 76 1" to DN2. Finally, the hash sharding result set in the three data nodes is a hash sharding grouped by "class".
[0064] Optionally, the N data nodes distribute each row of data in their respective initial grouping statistical result sets to corresponding data nodes using hash sharding according to the target sharding field set, and obtain their respective hash sharding result sets including:
[0065] N data nodes obtain the hash value of the target data in each row of data of their respective initial grouping statistical result sets, and calculate the remainder of the hash value of the target data in each row of data of their respective initial grouping statistical result sets, and distribute each row of data to the data node corresponding to the remainder among the N data nodes to obtain their respective hash sharding result sets, wherein the target data refers to the data whose field set in each row of data is the target sharding field set.
[0066] The data node reads a row of data, calculates the hash value of the target data in the row of data, finds the remainder of the hash value, and decides which data node to distribute the row of data corresponding to the hash value to based on the remainder. The target data with similar hash values can be distributed to one data node as required. For example, the data of "Class 1" and "Class 2" can be distributed to one data node. The target data can refer to the data in the row of data whose field set is the target shard field set. For example, if the row of data is "Class 1 Chinese 94 1" and the target shard field set is class, then the target data of the row of data is "Class 1". The hash value of "Class 1" is calculated. If the hash value corresponds to data node 1, the data node will distribute "Class 1 Chinese 94 1" to data node 1. Figure 4As shown in the figure, the hash values of "Class 1" and "Class 2" both correspond to DN1. DN2 distributes the three data of "Class 1 Mathematics 98 1", "Class 2 Chinese 80 1" and "Class 2 Mathematics 86 1" to DN1, and combines the three data of "Class 1 Chinese 94 1", "Class 2 Chinese 90 1" and "Class 2 Mathematics 96 1" retained on DN1 to form the hash shard result set of DN1.
[0067] Step S304, the N data nodes perform a second grouping statistics on their respective hash sharding result sets to obtain N target grouping statistics result sets.
[0068] Among them, a data node performs group statistics on its own hash sharding result set. The group statistics can use the "class merging group statistics" algorithm to obtain the target group statistics result set of the data node, such as Figure 4 As shown, DN1 performs group statistics on its own hash sharding result set to obtain ds1, ds1 is the target group statistics result set of DN1, ds2 is the target group statistics result set of DN2, and ds3 is the target group statistics result set of DN3. The server aggregates ds1, ds2, and ds3 together to obtain the target statistics result.
[0069] Step S305: the server collects N target group statistical result sets to determine the target statistical result.
[0070] Among them, the contents of step S301 and step S305 are the same as those of step S101 and step S103 in the above-mentioned embodiment 1, and the description of step S101 and step S103 may be referred to, and will not be repeated here.
[0071] Optionally, the data grouping statistical method further includes:
[0072] When the total number of rows is less than or equal to the row number threshold, the server performs group statistics on the N initial group statistical result sets to obtain the target statistical results.
[0073] Among them, after the N data nodes obtain their respective initial group statistics result sets, they determine whether to perform hash sharding on their respective initial group statistics result sets according to the total number of rows in the N initial group statistics result sets. Since the total number of rows is small and the data volume of the initial group statistics result sets is not large, the server can be directly used to perform group statistics on the initial group statistics result sets, which can not only ensure the group statistics performance, but also save the communication time between the server and the data nodes.
[0074] like Figure 5 The figure shows a flow chart of a data set to be grouped that is not a hash shard, where each data node performs a local data set ( Figure 5The original data in the data set is grouped and counted by "class" and "course", and the statistical results are "class, course, total score and number". d1 represents the initial group statistical result set of DN1 (data node 1), d2 represents the initial group statistical result set of DN2 (data node 2), and d3 represents the initial group statistical result set of DN3 (data node 3). If the total number of rows in the initial group statistical result set of d1, d2, and d3 is less than or equal to the row number threshold, the server directly performs group statistics on d1, d2, and d3 to obtain the target statistical result.
[0075] Optionally, group statistics are performed on the N initial group statistics result sets to obtain target statistics results including:
[0076] Sort each row of data in the N initial grouping statistical result sets in ascending order according to the sharding field set;
[0077] Get the N first rows of data from N initial group statistics result sets;
[0078] The data with the smallest fragment field set in the N first rows of data are grouped into the same group and counted to obtain a group statistics result;
[0079] Remove the data with the smallest sharding field set from the corresponding initial grouping statistical result set, and take the data after the data with the smallest sharding field set in the corresponding initial grouping statistical result set as the first row of data;
[0080] Traverse each row of data in the N initial group statistical result sets, and obtain all the group statistical results, which are the target statistical results.
[0081] The server performs group statistics on d1, d2, and d3. The group statistics can be processed using the "class merging group statistics" algorithm, such as Figure 6 As shown, the data in the initial grouping statistical result set is sorted in ascending order by "class" and "course", and the data is processed using a similar "merge" idea:
[0082] (1) Step 1: Read the first row of data from data sets d1, d2, and d3 respectively, and find the smallest data obtained when comparing data based on "class" and "course".
[0083] If the comparison is based on "class", "Class 1" can be used as the smallest data. If the comparison is based on "course", "Chinese" can be used as the smallest data. If the comparison is based on "class" and "course", such as Figure 6 As shown, the smallest data is "Class 1 Chinese 94 1" in d1, which is considered as a group for statistics: the total score is 94, the number of people is 1, so the average score = 94 / 1 = 94, and the statistical result of the group "Class 1 Chinese 9494" is output to the result set, and then the first row of data is removed from d1.
[0084] (2) Step 2: Read the first row of data from data sets d1, d2, and d3 respectively, and find the smallest data obtained when comparing data based on "class" and "course".
[0085] If we compare based on "class" and "course", Figure 6 As shown, the smallest data is "Class 1 Mathematics 98 1" in d2. It is considered as a group for statistics: the total score is 98, the number of people is 1, so the average score = 98 / 1 = 98. The statistical result of the group "Class 1 Mathematics 98 98" is output to the result set, and then the first row of data is removed from d2.
[0086] (3) Step 3: Read the first row of data from data sets d1, d2, and d3 respectively, and find the smallest data obtained when comparing data based on "class" and "course".
[0087] If we compare based on "class" and "course", Figure 6 As shown, the smallest data is "Class 2 Chinese 90 1" from d1 and "Class 2 Chinese 80 1" from d2. These two data are regarded as a group for statistics: total score 90+80=170, number of people 1+1=2, so average score=170 / 2=85. The group statistical result "Class 2 Chinese 170 85" is output to the result set, and then the first row of data is removed from d1 and d2 respectively.
[0088] And so on, the first row of data of d1, d2, and d3 is processed in a loop until all the data of d1, d2, and d3 are processed, and all the removed first rows of data are counted together to obtain the target grouping result.
[0089] Optionally, the data grouping statistical method further includes:
[0090] The server obtains the total number of rows of N initial group statistical result sets, and determines whether the total number of rows is greater than a row number threshold.
[0091] Among them, after the N data nodes obtain their respective initial group statistics result sets, each data node can send the number of rows of its own initial group statistics result set to the server, and the server summarizes and obtains the total number of rows of the N initial group statistics result sets. In one embodiment, after the data node obtains the initial group statistics result set, the data node sends a group statistics completion instruction to the server, the server receives the group statistics completion instruction, and sends a row number acquisition instruction to the data node, the data node calculates the number of rows of its own initial group statistics result set according to the row number acquisition instruction, and feeds the row number back to the server.
[0092] The embodiment of the present application can implement data grouping statistics of the data set to be grouped that is not hashed and distributed on N data nodes through N data nodes. When dealing with a large amount of data, there is no need to group them through the server, thus avoiding the server from retrieving the data set to be grouped consisting of a large amount of data from N data nodes, reducing the response time of the server and improving the efficiency of server grouping statistics.
[0093] Corresponding to the data grouping statistical method in the above embodiment, Figure 7 A schematic diagram of the network architecture of a data grouping statistics system of a distributed database provided in Example 3 of the present application is shown. For ease of explanation, only the parts related to the embodiment of the present application are shown.
[0094] See also Figure 7 , the data grouping statistics system includes: a server and N data nodes, N is an integer greater than 1, and the data sets to be grouped are distributed in the N data nodes;
[0095] The server is used to obtain the data shard type of the data set to be grouped;
[0096] N data nodes, used to perform one or two grouping statistics on their respective local data sets according to the data sharding type of the data set to be grouped, to obtain N target grouping statistics result sets, wherein the local data set of a data node is the data set of the data set to be grouped distributed on the data node, and one data node corresponds to one target grouping statistics result set;
[0097] The server is used to collect N target group statistical result sets and determine the target statistical results.
[0098] Optionally, the N data nodes are specifically used for:
[0099] When the data sharding type of the data set to be grouped is hash sharding, and the sharding field set of the hash sharding of the data set to be grouped is a subset of the target sharding field set, group statistics are performed on the respective local data sets to obtain N target grouping statistical result sets.
[0100] Optionally, the N data nodes are specifically used for:
[0101] When the data sharding type of the data set to be grouped is not hash sharding, or the sharding field set of the hash sharding of the data set to be grouped does not belong to the subset of the target sharding field set, group statistics are performed on the respective local data sets to obtain N initial grouping statistical result sets;
[0102] When the total number of rows in N initial group statistical result sets is greater than the row number threshold, hash sharding is performed on each initial group statistical result set according to the target sharding field set to obtain each hash sharding result set;
[0103] Perform group statistics on the respective hash sharding result sets to obtain N target group statistics result sets.
[0104] Optionally, the server is also used to:
[0105] When the total number of rows is less than or equal to the row number threshold, group statistics are performed on the N initial group statistical result sets to obtain the target statistical result.
[0106] Optionally, when the total number of rows is less than or equal to the row number threshold, the server is specifically configured to:
[0107] Sort each row of data in the N initial grouping statistical result sets in ascending order according to the sharding field set;
[0108] Get the N first rows of data from N initial group statistics result sets;
[0109] The data with the smallest fragment field set in the N first rows of data are grouped into the same group and counted to obtain a group statistics result;
[0110] Remove the data with the smallest sharding field set from the corresponding initial grouping statistical result set, and take the data after the data with the smallest sharding field set in the corresponding initial grouping statistical result set as the first row of data;
[0111] Traverse each row of data in the N initial group statistical result sets, and obtain all the group statistical results, which are the target statistical results.
[0112] Optionally, the N data nodes are specifically used for:
[0113] According to the target sharding field set, hash sharding is used to distribute each row of data in the respective initial grouping statistical result set to the corresponding data node to obtain the respective hash sharding result set.
[0114] Optionally, the N data nodes are specifically used for:
[0115] Obtain the hash value of the target data in each row of data of the respective initial grouping statistical result set, and calculate the remainder of the hash value of the target data in each row of data of the respective initial grouping statistical result set, distribute each row of data to the data node corresponding to the remainder among N data nodes, and obtain the respective hash sharding result sets, wherein the target data refers to the data whose field set in each row of data is the target sharding field set.
[0116] Optionally, the server is also used to:
[0117] The total number of rows of N initial grouping statistical result sets is obtained, and it is determined whether the total number of rows is greater than a row number threshold.
[0118] Optionally, the server is also used to:
[0119] Check whether the data sharding type of the data set to be grouped is hash sharding;
[0120] If the data sharding type of the data set to be grouped is hash sharding, detecting whether the sharding field set of the hash sharding of the data set to be grouped belongs to a subset of the target sharding field set, obtaining a first detection result, and sending the first detection result to the N data nodes;
[0121] If the data sharding type of the data set to be grouped is not hash sharding, obtaining a second detection result, and sending the second detection result to the N data nodes;
[0122] Among them, the first detection result includes that the data sharding type of the data set to be grouped is hash sharding and the sharding field set of the hash sharding of the data set to be grouped is a subset of the target sharding field set, and the data sharding type of the data set to be grouped is hash sharding and the sharding field set of the hash sharding of the data set to be grouped does not belong to a subset of the target sharding field set. The second detection result refers to that the data sharding type of the data set to be grouped is not hash sharding.
[0123] It should be noted that the information interaction, execution process, etc. between the above-mentioned server and N data nodes are based on the same concept as the method embodiment of the present application. Their specific functions and technical effects can be found in the method embodiment part and will not be repeated here.
[0124] The above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. These modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present application, and should all be included in the protection scope of the present application.
Claims
1. A data grouping statistics method for a distributed database, the data grouping statistics method is applied to a data grouping statistics system for a distributed database, the data grouping statistics system comprises a server and N data nodes, N is an integer greater than 1, and is characterized in that: The data sets to be grouped are distributed on the N data nodes, and the data grouping statistical method includes: The server obtains the data shard type of the data set to be grouped; The N data nodes perform one or two grouping statistics on their respective local data sets according to the data sharding type of the data set to be grouped, and obtain N target grouping statistical result sets, wherein the local data set of a data node is the data set of the data set to be grouped distributed on the data node, and one data node corresponds to one target grouping statistical result set; The server performs statistics on the N target group statistical result sets to determine target statistical results; The N data nodes perform one or two grouping statistics on their respective local data sets according to the data sharding type of the data set to be grouped, and obtain N target grouping statistics result sets including: When the data sharding type of the data set to be grouped is hash sharding and the sharding field set of the hash sharding of the data set to be grouped belongs to a subset of the target sharding field set, the N data nodes perform grouping statistics on their respective local data sets to obtain N target grouping statistical result sets; When the data sharding type of the data set to be grouped is not hash sharding, or the sharding field set of the hash sharding of the data set to be grouped does not belong to the subset of the target sharding field set, the N data nodes perform the first grouping statistics on their respective local data sets to obtain N initial grouping statistics result sets; When the total number of rows in the N initial grouping statistical result sets is greater than the row number threshold, the N data nodes perform hash sharding on their respective initial grouping statistical result sets according to the target sharding field set to obtain their respective hash sharding result sets; The N data nodes perform a second grouping statistics on their respective hash sharding result sets to obtain the N target grouping statistical result sets; The data grouping statistical method also includes: When the total number of rows is less than or equal to the row number threshold, the server performs group statistics on the N initial group statistical result sets to obtain a target statistical result; The performing group statistics on the N initial group statistics result sets to obtain target statistics results comprises: Sort each row of data in the N initial grouping statistical result sets in ascending order of the sharding field set; Obtain the N first rows of data of the N initial grouping statistical result sets; The data with the smallest fragment field set in the N first rows of data are grouped into the same group and counted to obtain a grouping statistical result; The data with the smallest sharding field set is removed from the corresponding initial grouping statistical result set, and the data after the data with the smallest sharding field set in the corresponding initial grouping statistical result set is used as the first row of data; Traversing each row of data of the N initial group statistical result sets, all group statistical results are obtained as target statistical results.
2. The data grouping statistical method according to claim 1, characterized in that: The N data nodes perform hash sharding on their respective initial grouping statistical result sets according to the target sharding field set, and obtain their respective hash sharding result sets including: The N data nodes distribute each row of data in their respective initial grouping statistical result sets to corresponding data nodes using hash sharding according to the target sharding field set, and obtain their respective hash sharding result sets.
3. The data grouping statistical method according to claim 2, characterized in that: The N data nodes distribute each row of data in their respective initial grouping statistical result sets to corresponding data nodes using hash sharding according to the target sharding field set, and obtain their respective hash sharding result sets including: The N data nodes obtain the hash value of the target data in each row of data of their respective initial grouping statistical result sets, and calculate the remainder of the hash value of the target data in each row of data of their respective initial grouping statistical result sets, and distribute each row of data to the data nodes corresponding to the remainder among the N data nodes to obtain their respective hash sharding result sets, wherein the target data refers to the data whose field set in each row of data is the target sharding field set.
4. The data grouping statistical method according to claim 1, further comprising: The server obtains the total number of rows of the N initial grouping statistical result sets, and determines whether the total number of rows is greater than a row number threshold.
5. The data grouping statistical method according to claim 1, characterized in that: The data grouping statistical method also includes: The server detects whether the data sharding type of the data set to be grouped is hash sharding; If the data sharding type of the data set to be grouped is hash sharding, detecting whether the sharding field set of the hash sharding of the data set to be grouped belongs to a subset of the target sharding field set, obtaining a first detection result, and sending the first detection result to the N data nodes; If the data sharding type of the data set to be grouped is not hash sharding, obtaining a second detection result, and sending the second detection result to the N data nodes; Among them, the first detection result includes that the data sharding type of the data set to be grouped is hash sharding and the sharding field set of the hash sharding of the data set to be grouped is a subset of the target sharding field set, and the data sharding type of the data set to be grouped is hash sharding and the sharding field set of the hash sharding of the data set to be grouped does not belong to a subset of the target sharding field set. The second detection result means that the data sharding type of the data set to be grouped is not hash sharding.
6. A data grouping statistics system for a distributed database, the data grouping statistics system comprising a server and N data nodes, N being an integer greater than 1, characterized in that: The data sets to be grouped are distributed on the N data nodes, and the data grouping statistics system is applied to the method described in any one of claims 1 to 5, and the data grouping statistics system includes: The server is used to obtain the data shard type of the data set to be grouped; The N data nodes are used to perform one or two grouping statistics on their respective local data sets according to the data sharding type of the data set to be grouped, to obtain N target grouping statistics result sets, wherein the local data set of a data node is the data set of the data set to be grouped distributed on the data node, and one data node corresponds to one target grouping statistics result set; The server is used to perform statistics on the N target group statistical result sets to determine target statistical results.
Citation Information
Patent Citations
Distributed query method and system for complex task of querying massive structured data
CN102521406A
Aggregation framework system architecture and method
US20170262517A1