Method and apparatus for processing data
By generating a high-frequency value hash table to identify and pre-aggregate high-frequency data, the problems of low data aggregation rate and data skew in group push-down technology are solved, and data processing performance and query efficiency are optimized.
Patent Information
- Application Number
- PCT/CN2024/127793
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-12-28
- Filing Date
- 2024-10-28
- Publication Date
- 2025-07-03
AI Technical Summary
In the group push-down technology, the data processing performance caused by low data aggregation rate is poor, and the adaptive GBY technology leads to data tilt problems, slowing down the data query process.
The high-frequency value hash table is generated by sampling data, high-frequency data is identified and pre-aggregated. The low-frequency data is not pre-aggregated, which balances the aggregation overhead and data volume, and eliminates the data tilt problem.
Effectively improve data query performance, eliminate or alleviate data skew, optimize data distribution, and reduce the overhead of pre-aggregation operations.
Smart Images

Figure CN2024127793_03072025_PF_FP_ABST
Abstract
Description
Data processing method and device Technical Field
[0001] One or more embodiments of this specification relate to the field of terminal technology, and in particular, to a data processing method and device. Background Art
[0002] GPD (Group By Pushdown) is an optimization method for database calculation aggregation in parallel scenarios. It refers to pushing the GBY (Group By) operator down to the local computer. Before data transmission, the GBY operator is used to pre-aggregate the data locally. The pre-aggregated data is then distributed to different worker threads to complete the final data aggregation.
[0003] GPD has excellent scalability and can reduce data distribution costs. However, when the GBY push operator has a low aggregation rate, data processing performance is poor. Therefore, related technologies have introduced adaptive GBY technology. This technology directly sends data to upper-layer nodes without further aggregation when the GBY push operator detects poor aggregation results. However, this can lead to data skew, slowing down the entire data query process.
[0004] Summary of the Invention
[0005] In order to eliminate or alleviate the data skew problem that occurs when the aggregation effect is poor during the group push-down process, the embodiments of this specification provide a data processing method, device, electronic device and storage medium.
[0006] In a first aspect, one or more embodiments of the present specification provide a data processing method, including: determining a data aggregation rate based on sampled data in a group push-down process; generating a high-frequency value hash table based on the sampled data when the data aggregation rate does not meet the aggregation effect, the high-frequency value hash table including the hash value of the high-frequency data in the sampled data; performing pre-aggregation processing on the high-frequency data in the data to be processed based on the high-frequency value hash table, and distributing the pre-aggregation processing results to the target thread of the aggregation node.
[0007] In one or more embodiments of the present specification, determining the data aggregation rate based on the sampled data in the group push-down process includes: determining the data aggregation rate based on a first data volume of the sampled data in the group push-down process, and a second data volume after pre-aggregation processing of the sampled data.
[0008] In one or more embodiments of the present specification, the data aggregation rate is determined based on the sampled data in the group push-down process, including: periodically acquiring the sampled data with a first preset data volume, and executing the process of determining the data aggregation rate based on the sampled data until the data aggregation rate of the current period does not meet the aggregation effect.
[0009] In one or more embodiments of the present specification, the generation of a high-frequency value hash table based on the sampled data includes: in the sampled data with a total data volume of T, calculating the corresponding hash value based on the grouping key of each data, and counting the first data volume that is the same as the hash value of the i-th data, where 1≤i≤T; based on the total data volume, the first data volume, and the number of threads for grouped push-down, determining the second data volume of other data different from the hash value of the i-th data and evenly distributing it to each thread; based on the ratio of the first data volume to the second data volume, determining whether the i-th data is high-frequency data; and generating the high-frequency value hash table based on all high-frequency data in the sampled data and their hash values.
[0010] In one or more embodiments of the present specification, determining whether the i-th data is high-frequency data based on the ratio of the first data amount to the second data amount includes: in response to the ratio being greater than or equal to a preset threshold, determining that the i-th data is high-frequency data; in response to the ratio being less than the preset threshold, determining that the i-th data is low-frequency data.
[0011] In one or more embodiments of the present specification, the high-frequency data in the data to be processed is pre-aggregated based on the high-frequency value hash table, and the pre-aggregation processing result is distributed to the target thread of the aggregation node, including: for each data in the data to be processed, a corresponding hash value is calculated based on the grouping key of the data; the hash value of the data is matched with the hash value in the high-frequency value hash table, and in response to a consistent match, the data is determined to be high-frequency data; the high-frequency data is pre-aggregated to obtain the pre-aggregation processing result, and the target thread is determined based on the grouping key of the high-frequency data; and the pre-aggregation processing result is distributed to the target thread.
[0012] In one or more embodiments of the present specification, the method further includes: periodically acquiring the data to be processed with a second preset data volume, and determining the frequency value parameters of each high-frequency data in the high-frequency value hash table based on the data to be processed, until the frequency value parameters of all high-frequency data in the high-frequency value hash table do not meet the set requirements, and the frequency value parameters represent the frequency of occurrence of the high-frequency data in the data to be processed; in response to the existence of at least one high-frequency data in the high-frequency value hash table with a frequency value parameter that meets the set requirements, executing the process of pre-aggregating the high-frequency data in the data to be processed based on the high-frequency value hash table, and distributing the pre-aggregation processing results to the target thread of the aggregation node; in response to the frequency value parameters of all high-frequency data in the high-frequency value hash table not meeting the set requirements, distributing the data to be processed directly to the aggregation node.
[0013] In a second aspect, one or more embodiments of the present specification provide a data processing device, including: an aggregation rate determination module, configured to determine the data aggregation rate based on the sampled data in the group push-down process; a hash table generation module, configured to generate a high-frequency value hash table based on the sampled data when the data aggregation rate does not meet the aggregation effect, the high-frequency value hash table including the hash value of the high-frequency data in the sampled data; a pre-aggregation module, configured to perform pre-aggregation processing on the high-frequency data in the data to be processed based on the high-frequency value hash table, and distribute the pre-aggregation processing results to the target thread of the aggregation node.
[0014] In one or more embodiments of the present specification, the aggregation rate determination module is configured to determine the data aggregation rate based on a first data volume of the sampled data during group pushdown and a second data volume after pre-aggregation processing of the sampled data.
[0015] In one or more embodiments of the present specification, the aggregation rate determination module is configured to: periodically obtain the sampled data with a first preset data volume, and execute the process of determining the data aggregation rate based on the sampled data until the data aggregation rate of the current period does not meet the aggregation effect.
[0016] In one or more embodiments of the present specification, the hash table generation module is configured to: in the sampled data with a total data volume of T, calculate the corresponding hash value based on the grouping key of each data, and count the first data volume that is the same as the hash value of the i-th data, where 1≤i≤T; based on the total data volume, the first data volume and the number of threads for group push-down, determine the second data volume of other data different from the hash value of the i-th data and distribute it evenly to each thread; based on the ratio of the first data volume to the second data volume, determine whether the i-th data is high-frequency data; based on all high-frequency data in the sampled data and their hash values, generate the high-frequency value hash table.
[0017] In one or more embodiments of the present specification, the hash table generation module is configured to: in response to the ratio being greater than or equal to a preset threshold, determine that the i-th data is high-frequency data; in response to the ratio being less than the preset threshold, determine that the i-th data is low-frequency data.
[0018] In one or more embodiments of the present specification, the pre-aggregation module is configured to: for each data in the data to be processed, calculate a corresponding hash value based on the grouping key of the data; match the hash value of the data with the hash value in the high-frequency value hash table, and in response to a consistent match, determine that the data is high-frequency data; perform pre-aggregation processing on the high-frequency data to obtain the pre-aggregation processing result, and determine the target thread based on the grouping key of the high-frequency data; and distribute the pre-aggregation processing result to the target thread.
[0019] In one or more embodiments of the present specification, the pre-aggregation module is configured to: periodically obtain the data to be processed with a second preset data volume, and determine the frequency value parameters of each high-frequency data in the high-frequency value hash table based on the data to be processed, until the frequency value parameters of all high-frequency data in the high-frequency value hash table do not meet the set requirements, and the frequency value parameters represent the frequency of occurrence of the high-frequency data in the data to be processed; in response to the existence of at least one high-frequency data in the high-frequency value hash table with a frequency value parameter that meets the set requirements, execute the process of pre-aggregating the high-frequency data in the data to be processed based on the high-frequency value hash table, and distribute the pre-aggregation processing results to the target thread of the aggregation node; in response to the frequency value parameters of all high-frequency data in the high-frequency value hash table not meeting the set requirements, distribute the data to be processed directly to the aggregation node.
[0020] In a third aspect, one or more embodiments of this specification provide an electronic device comprising: a processor; and a memory storing computer instructions, wherein the computer instructions are used to enable the processor to execute the method described in any embodiment of the first aspect.
[0021] In a fourth aspect, one or more embodiments of this specification provide a storage medium storing computer instructions, wherein the computer instructions are used to enable a computer to execute the method according to any embodiment of the first aspect.
[0022] The data processing method of one or more embodiments of the present specification includes determining a data aggregation rate based on sampled data during a group push-down process, generating a high-frequency value hash table based on the sampled data if the data aggregation rate does not meet the aggregation effect, performing pre-aggregation processing on the high-frequency data in the to-be-processed data based on the high-frequency value hash table, and distributing the pre-aggregation processing results to the target thread of the aggregation node. In the embodiments of the present specification, if it is detected that the data aggregation effect is poor during the sampling phase, a high-frequency value hash table is obtained based on the hash value of the high-frequency data from the data statistics during the sampling phase, so that only the high-frequency data in the high-frequency value hash table is pre-aggregated during the processing phase, eliminating or alleviating the data skew problem of the upper-level aggregation node, effectively balancing the aggregation overhead and the amount of data distributed, and improving data query performance. BRIEF DESCRIPTION OF THE DRAWINGS
[0023] FIG1 is a schematic diagram showing a principle of a data processing method according to an exemplary embodiment of the present specification.
[0024] FIG2 is a schematic diagram showing the principle of a data processing method according to an exemplary embodiment of the present specification.
[0025] FIG3 is a flowchart of a data processing method according to an exemplary embodiment of the present specification.
[0026] FIG4 is a schematic diagram showing a principle of a data processing method according to an exemplary embodiment of the present specification.
[0027] FIG5 is a flowchart of a data processing method according to an exemplary embodiment of the present specification.
[0028] FIG6 is a flowchart of a data processing method according to an exemplary embodiment of this specification.
[0029] FIG. 7 is a flowchart of a data processing method according to an exemplary embodiment of the present specification.
[0030] FIG8 is a structural block diagram of a data processing device according to an exemplary embodiment of this specification.
[0031] FIG9 is a structural block diagram of an electronic device according to an exemplary embodiment of this specification. DETAILED DESCRIPTION
[0032] Exemplary embodiments will be described in detail herein, examples of which are illustrated in the accompanying drawings. When the following description refers to the drawings, identical numerals in different figures represent identical or similar elements unless otherwise indicated. The embodiments described in the following exemplary embodiments are not intended to represent all embodiments consistent with one or more embodiments of this specification. Rather, they are merely examples of apparatuses and methods consistent with certain aspects of one or more embodiments of this specification, as detailed in the appended claims.
[0033] It should be noted that in other embodiments, the steps of the corresponding method are not necessarily performed in the order shown and described in this specification. In some other embodiments, the method may include more or fewer steps than those described in this specification. In addition, a single step described in this specification may be broken down into multiple steps for description in other embodiments; and multiple steps described in this specification may be combined into a single step for description in other embodiments.
[0034] In addition, the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this manual are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with the relevant laws, regulations and standards of relevant countries and regions, and provide corresponding operation entrances for users to choose to authorize or refuse.
[0035] GPD (Group By Pushdown) is an optimization method for distributed databases to calculate aggregation in parallel scenarios. It refers to pushing the GBY (Group By) operator down to the local computer. Before data transmission, the GBY operator is used to pre-aggregate the data locally. The pre-aggregated data is then distributed to different worker threads to complete the final data aggregation.
[0036] For example, as shown in Figure 1, node 1 is a lower-level node, and nodes 2 and 3 are upper-level nodes. In GPD technology, the GBY operator can be pushed down to the lower-level node 1, so that before the data is transmitted from node 1 to the upper-level nodes 2 and 3, the data can be pre-aggregated in node 1 using the downward-pushing GBY operator (one-stage GBY). Pre-aggregation refers to grouping the data according to the grouping key (GBY key, Group By Keys), performing aggregation operations on each group of data, and then sending the pre-aggregation results to each thread of the upper-level node according to the grouping key. Each thread of the upper-level node receives the data sent by the lower-level node according to the GBY key, that is, the data processed by each thread of the upper-level node has the same GBY key value. Each thread of the upper-level node performs the final aggregation processing on the data therein to obtain the corresponding data query results.
[0037] GPD has excellent scalability and can effectively reduce data distribution costs, but it does incur the additional cost of pre-aggregating data for the GBY pushdown operator. This can lead to poor data processing performance, especially when the data aggregation rate at lower-level nodes is low. For example, when pre-aggregating data, lower-level nodes aggregate data based on the same grouping key. If there is little data with the same grouping key, not only does this increase the overhead of pre-aggregation, but the data aggregation rate is also very low, resulting in poor data processing performance.
[0038] In related technologies, in order to solve this problem, some databases introduce adaptive GBY technology. That is, during the group push process, if the lower-level node detects that the data aggregation effect is not good, it will directly distribute the data to the upper-level node and no longer perform the pre-aggregation operation, thereby avoiding the additional overhead of performing the pre-aggregation operation.
[0039] However, as shown in Figure 1, if the lower-level nodes distribute data directly to the upper-level nodes, if there is high-frequency data in the data of the lower-level nodes, that is, there is a lot of data with the same grouping key, since the pre-aggregation operation is no longer performed, the high-frequency data will be directly distributed to a thread of the upper-level node, resulting in a large amount of data in the thread and a small amount of data in the other threads, forming a data skew phenomenon, which in turn slows down the entire data query process.
[0040] Based on the defects of the above-mentioned related technologies, the embodiments of this specification provide a data processing method, device, electronic device and storage medium, which aims to detect that when the data aggregation effect is not good during the group push-down process, based on the high-frequency value hash table statistically calculated in the sampling stage, aggregate the high-frequency data in the subsequent data, and no longer aggregate the low-frequency data, so as to balance the aggregation overhead and data volume, and eliminate or alleviate the data skew problem caused by high-frequency data.
[0041] The system architecture of the embodiments of this specification can be seen in Figure 1. The data processing method of the embodiments of this specification is mainly applied to the lower-layer nodes, namely, the first-stage GBY nodes. For ease of explanation, the lower-layer first-stage GBY nodes are defined as "pre-aggregation nodes", and the GBY operators pushed down to the pre-aggregation nodes are defined as lower-layer GBY operators. At the same time, the upper-layer second-stage GBY nodes are defined as "aggregation nodes", and the GBY operators located at the aggregation nodes are defined as upper-layer GBY operators. This specification will be explained in accordance with this definition and will not be repeated here.
[0042] In the pre-aggregation node, all data in the group push-down process can be regarded as two stages, namely the sampling stage and the processing stage.
[0043] As shown in Figure 2, in the sampling phase, the lower-level pre-aggregation node determines whether the aggregation effect is good or bad based on the sampling phase data. If the aggregation effect is good, it can be maintained in the sampling phase. Conversely, if the aggregation effect is not good, it enters the processing phase.
[0044] The work in the sampling stage mainly includes two parts. The first is to judge the data aggregation effect. When the data aggregation effect is good, the data processing process can be maintained in the sampling stage. When the data aggregation effect is not good, it can enter the processing stage. The second is to calculate the high-frequency value hash table based on the data statistics in the sampling stage when the data aggregation effect is not good. The high-frequency value hash table is used to record the data that appears more frequently in the data in the sampling stage, that is, the hash value of the high-frequency data.
[0045] The processing stage requires the use of the high-frequency value hash table of the sampling stage to screen the data and determine whether the data is high-frequency data. If a certain data is high-frequency data, pre-aggregation processing is performed on this data. If a certain data is low-frequency data, no pre-aggregation processing is required.
[0046] Based on the above, it's clear that if data aggregation isn't effective, continuing to pre-aggregate the data won't improve performance. In fact, the additional overhead of pre-aggregation will degrade query performance. If data is distributed directly without pre-aggregation, subsequent high-frequency data can easily cause data skew at upper-level aggregation nodes.
[0047] Therefore, in the embodiments of this specification, if the sampling phase detects that the data aggregation effect is poor, a high-frequency value hash table is obtained based on the hash value of the high-frequency data from the data statistics of the sampling phase. Therefore, in the processing phase, only the high-frequency data in the high-frequency value hash table is pre-aggregated, and the low-frequency data is no longer pre-aggregated. Because the high-frequency data is pre-aggregated, the amount of aggregation result data distributed to the upper-level aggregation node thread is small, eliminating or alleviating the data skew problem of the upper-level aggregation node. In addition, because only the high-frequency data is pre-aggregated, the low-frequency data is no longer pre-aggregated, which reduces the overhead of the pre-aggregation operation, effectively balancing the aggregation overhead and data volume, and improving data query performance.
[0048] As shown in FIG3 , in some embodiments, the data processing method exemplified in this specification includes steps S310 to S330 .
[0049] S310: Determine a data aggregation rate based on sampled data in a group push-down process.
[0050] As shown in Figure 1, during the GPD group push process, the data of the lower-level pre-aggregation node needs to be pre-aggregated locally based on the lower-level GBY operator, so the amount of data before and after pre-aggregation will change. The aggregation rate can reflect the data compression ratio before and after data pre-aggregation, and then reflect the data pre-aggregation effect.
[0051] In the embodiments of this specification, during the sampling phase, during the data pre-aggregation process, a certain amount of sampled data may be obtained, and the amount of the sampled data before and after the pre-aggregation process may be counted. Based on the amount of the sampled data before and after the pre-aggregation process, a corresponding data aggregation ratio may be calculated. The specific process for calculating the data aggregation ratio is described in the embodiments below.
[0052] In some embodiments, during the sampling phase, sampled data may be periodically acquired at a fixed data volume, and then a data aggregation ratio for the current period may be calculated based on the sampled data. If the data aggregation ratio satisfies the aggregation effect, it indicates that the current data pre-aggregation effect is high, and the sampling phase is continued. The data aggregation ratio is then recalculated based on the sampled data of the next period, and the processing phase is entered until the data aggregation ratio no longer satisfies the aggregation effect, indicating that the current data aggregation effect is poor. This is described in the embodiments below in this specification and will not be elaborated on here.
[0053] At the same time, in the implementation of this specification, during the sampling phase, the hash value of each data can be counted. For example, a hash operation can be performed based on the grouping key (Group By Keys) of each data to obtain the corresponding hash value. The purpose of counting the statistical hash value is to calculate a high-frequency value hash table including high-frequency data based on the sampled data.
[0054] S320: When the data aggregation rate does not meet the aggregation effect, generate a high-frequency value hash table based on the sampled data.
[0055] As mentioned above, during the sampling phase, if the data aggregation rate is determined to be insufficient for the desired aggregation effect, indicating poor data aggregation, the processing phase is necessary. Simultaneously, a high-frequency value hash table is generated based on the sampled data from the sampling phase. The high-frequency value hash table records the hash values of the high-frequency data that appears most frequently in the sampled data.
[0056] For example, in an example, during the sampling stage, the hash value of each data can be counted. When it is determined that the data aggregation rate does not meet the aggregation effect, the hash value of each data in the statistical sampling data can be used to calculate which data is high-frequency data that may cause data skew, and the hash values of these high-frequency data can be recorded to obtain a high-frequency value hash table.
[0057] It is worth noting that the high-frequency data recorded in the high-frequency value hash table refers to data that may cause data skew in the upper-level aggregation node. Therefore, when calculating whether a certain data is high-frequency data, it is necessary not only to consider the frequency of the hash value of the data in all the sampled data, but also to consider the difference in data volume after the data is distributed to different threads with other data. The greater the difference in data volume, the more serious the data skew.
[0058] Therefore, in the implementation manner of this specification, when calculating whether certain data is high-frequency data based on sampled data, a preset threshold can be set in advance, and then the frequency of occurrence of the hash value of the data in the sampled data is counted, and then the ratio of the frequency of occurrence of the data to the average amount of data distributed to each thread by other data is calculated, and the ratio is compared with the preset threshold. If it exceeds the preset threshold, it means that the data is high-frequency data, otherwise it is low-frequency data. This specification explains the implementation manner below.
[0059] After all high-frequency data included in the sampled data is calculated through the above process, a high-frequency value hash table can be generated based on the hash values of all high-frequency data.
[0060] S330 , performing pre-aggregation processing on the high-frequency data in the data to be processed based on the high-frequency value hash table, and distributing the pre-aggregation processing result to the target thread of the aggregation node.
[0061] Combined with the above, it can be seen that when it is determined that the data aggregation rate does not meet the aggregation effect, the processing stage is entered. The goal of the processing stage is to screen the data in the data to be processed in the processing stage based on the high-frequency data recorded in the high-frequency value hash table. If the hash value of a certain data in the data to be processed matches the same hash value in the high-frequency value hash table, it means that the data is high-frequency data, and thus the data can be pre-aggregated. Conversely, if the hash value of a certain data in the data to be processed does not match the same hash value in the high-frequency value hash table, it means that the data is low-frequency data, and thus there is no need for pre-aggregation processing, and it can be directly distributed to the target thread of the upper-level aggregation node.
[0062] In the embodiments of this specification, after pre-aggregating high-frequency data, a pre-aggregation result is obtained, which is then distributed to the target thread of the upper-level aggregation node based on the data's group by key. It is worth noting that the aggregation and pre-aggregation described in the embodiments of this specification include, but are not limited to, summation, average, count, maximum, minimum, median, variance, standard deviation, etc., and this specification does not impose any restrictions on this.
[0063] In some implementations, considering that the sampled data in the sampling phase cannot accurately represent the actual data distribution during the entire GPD group push-down period, the high-frequency value hash table generated in the sampling phase can be further tested in combination with the data in the processing phase to determine whether the high-frequency data recorded in the high-frequency value hash table meets the high-frequency index. If all the high-frequency data recorded in the high-frequency value hash table does not meet the high-frequency index, then there is no need to continue the above-mentioned data screening process, and the data can be directly distributed to the upper-level aggregation node. This specification will explain the implementation below and will not be described in detail here.
[0064] It can be understood that the data processing method of the embodiment of the present disclosure, compared with the adaptive GBY scheme in the related art, does not send data directly to the upper-level node when the data aggregation effect is not good. Instead, it uses the high-frequency value hash table in the sampling stage to screen the high-frequency data in the processing stage, and pre-aggregates the high-frequency data that may cause data skew in the upper-level node, thereby eliminating or alleviating the data skew problem, and sends the low-frequency data directly to the upper-level node.
[0065] From the above, it can be seen that in the implementation mode of this specification, when it is detected that the data aggregation effect is not good in the sampling stage, the hash value of the high-frequency data is obtained based on the data statistics of the sampling stage to obtain a high-frequency value hash table, so that in the processing stage, only the high-frequency data in the high-frequency value hash table is pre-aggregated, eliminating or alleviating the data skew problem of the upper-level aggregation node, effectively balancing the aggregation overhead and the amount of data distributed, and improving data query performance.
[0066] FIG4 shows a group push-down data processing process in some implementations of this specification. The data processing method of the implementations of this specification will be described below with reference to the example in FIG4 .
[0067] In some embodiments, during the sampling phase, sampled data may be periodically acquired with a first preset data volume T, and then the sampled data of the data volume T may be combined to calculate whether the data aggregation rate of the current period satisfies the aggregation effect.
[0068] As shown in FIG4 , the amount of sampled data obtained in the sampling phase 1 is a first preset data amount T, and then the data aggregation rate Ragg is calculated based on the sampled data of the period, which is expressed as:
[0069] In formula (1), R agg Represents the data aggregation rate, T represents the first data volume before the sampled data is pre-aggregated, and M represents the second data volume after the sampled data is pre-aggregated. Combining formula (1), we can see that the data aggregation rate R agg The smaller the data pre-aggregation effect, the better. On the contrary, the data aggregation rate R agg The larger the value, the worse the data aggregation effect.
[0070] In addition, the data aggregation rate R can be pre-set agg Set the threshold R thre , the threshold R thre Indicates that the data aggregation rate meets the critical value of aggregation effect.
[0071] In sampling phase 1, the data aggregation rate R corresponding to sampling phase 1 can be calculated by combining formula (1): agg Then, the data aggregation rate R agg With threshold R thre For comparison, assume that the data aggregation rate R in sampling stage 1 is agg Less than the threshold R thre , indicating that the data aggregation effect of sampling phase 1 is better, thus continuing to the sampling phase of the next cycle, that is, sampling phase 2.
[0072] For sampling stage 2, similarly combined with the above formula (1), based on the sampling data with a data volume of T collected in sampling stage 2, the data aggregation rate R corresponding to sampling stage 2 can be calculated: agg Then, the data aggregation rate R agg With threshold R thre For comparison, assume that the data aggregation rate R in sampling stage 2 is agg Greater than or equal to the threshold R thre , indicating that the data aggregation effect of sampling stage 2 is poor.
[0073] In the implementation manner of this specification, when it is determined that the data aggregation rate in the sampling stage does not meet the aggregation effect, the processing stage is entered. At the same time, a high-frequency value hash table needs to be generated based on the sampled data in the sampling stage.
[0074] In some embodiments, as shown in FIG4 , the sampling data used to generate the high-frequency value hash table may include only sampling data with a data volume of T in sampling stage 2, or may include sampling data with a data volume of 2T in both sampling stage 1 and sampling stage 2. This specification does not impose any restrictions on this.
[0075] Hereinafter, the process of generating a high-frequency value hash table is described by taking the sampled data with a data volume of T in the sampling stage 2 as an example in conjunction with the embodiment of FIG5 .
[0076] As shown in FIG. 5 , in some embodiments, the data processing method exemplified in this specification, the process of generating a high-frequency value hash table based on the sampled data T, includes steps S510 to S540 .
[0077] S510 . In the sampled data with a total data volume of T, a corresponding hash value is calculated based on the grouping key of each data, and a first data volume that is the same as the hash value of the i-th data is counted, where 1≤i≤T.
[0078] S520 : Based on the total data volume, the first data volume, and the number of threads for group pushdown, determine a second data volume of other data having a different hash value from the i-th data and evenly distributing it to each thread.
[0079] S530 : Determine whether the i-th data is high-frequency data based on the ratio of the first data amount to the second data amount.
[0080] S540: Generate a high-frequency value hash table based on all high-frequency data in the sampled data and their hash values.
[0081] In the implementation manner of this specification, the total amount of sampled data is T. For each data in the sampled data, a hash operation can be performed on the grouping key (Group By Keys) of each data to obtain a corresponding hash value, thereby obtaining the hash values of T data.
[0082] In some implementations, the hash value of each data may be counted during the sampling phase, so that when it is determined to enter the processing phase, the hash value of each data in the sampled data may be obtained, which will not be described in detail in this specification.
[0083] Based on the principle of hash operation, for data with the same grouping key, the hash value obtained after hash operation should also be the same. Therefore, for T data, some of them may contain data with the same grouping key, and thus the obtained hash value will also contain many identical hash values.
[0084] For any data i (1≤i≤T) in the sampled data, the frequency ki of occurrence of the hash value of data i in the hash values of all data can be counted. The frequency ki of occurrence can be understood as: the amount of data in the hash values of all data that is the same as the hash value of data i. This data amount is the first data amount ki described in this specification.
[0085] Combined with the above, it can be seen that data skew of the upper-level aggregation node refers to the uneven distribution of data volume of each thread. Therefore, when calculating whether the i-th data is high-frequency data, it is necessary to compare the difference between the first data volume ki and the average data volume of other data evenly distributed to each thread (that is, the second data volume Li). The greater the difference, the greater the risk of data skew.
[0086] In the implementation manner of this specification, the second data amount Li is expressed as:
[0087] In formula (2), T represents the total amount of sampled data, k i represents the first data amount of data with the same hash value as the i-th data, so (Tk i ) represents the amount of data in the sampled data that has a different hash value from the i-th data. dop represents the number of threads in the upper aggregation node. For example, in the example in Figure 1, the number of threads in the upper aggregation node is 4.
[0088] From this we can understand that L i The average amount of data distributed evenly to each thread, which is different from the hash value of the i-th data, is also the second data amount L in this specification. i .
[0089] After obtaining the first data amount ki and the second data amount L i Then, using the first data volume ki and the second data volume L i The difference between the two is expressed as the ratio:
[0090] In formula (3), ratio represents the first data volume ki and the second data volume L i ratio.
[0091] In some embodiments, the first data amount ki and the second data amount L may be pre-set. iThe ratio ratio sets a corresponding preset threshold ratio_thre, and the preset threshold ratio_thre represents a critical value of high-frequency data.
[0092] For the i-th data, if the corresponding first data volume ki and the second data volume L i The ratio ratio is greater than or equal to the preset threshold ratio_thre, indicating that the first data volume ki and the second data volume L i The difference between them is large, so after data distribution, the data volume of the thread where the i-th data is located will far exceed the average data volume of other threads, so the risk of data skew is very high. Therefore, it can be determined that the i-th data is high-frequency data.
[0093] On the contrary, if the first data volume ki corresponding to the i-th data is equal to the second data volume L i The ratio ratio is less than the preset threshold ratio_thre, indicating that the first data volume ki and the second data volume L i The difference between them is small, so after data distribution, the data volume of the thread where the i-th data is located is close to the average data volume of other threads, and the risk of data skew is low. Therefore, it can be determined that the i-th data is low-frequency data.
[0094] By performing high-frequency detection on a sample of data of size T through the above method, all high-frequency data included in the sampled data can be obtained. Then, based on these high-frequency data and their hash values, a high-frequency value hash table is generated. That is, the high-frequency value hash table records each high-frequency data in the sampled data and its corresponding hash value.
[0095] From the above, it can be seen that in the implementation mode of this specification, high-frequency data is determined based on the difference between the frequency of occurrence of the hash value of data i and the average amount of data evenly distributed to each thread by other data, and the data distribution characteristics of data skew are fully considered. The high-frequency value hash table generated can provide an efficient data basis for subsequent data skew processing.
[0096] Continuing with FIG4 , in sampling phase 2 , it is determined that the data aggregation rate does not meet the aggregation effect, i.e., the processing phase is entered. At the same time, a high-frequency value hash table is generated through the above-mentioned method process. The data processing process of the processing phase is described below in conjunction with FIG6 .
[0097] As shown in FIG6 , in some embodiments, the data processing method exemplified in this specification, the data processing process in the processing stage includes steps S610 to S630 .
[0098] S610: For each data in the data to be processed, obtain a corresponding hash value based on the grouping key of the data.
[0099] In the implementation manner of this specification, the data to be processed is the data in the processing stage. For each data in the data to be processed, a hash operation can be performed on the grouping key (Group By Keys) of the data to obtain a corresponding hash value.
[0100] S620: Match the hash value of the data with the hash value in the high-frequency hash table, and in response to a consistent match, determine that the data is high-frequency data.
[0101] In the implementation of this specification, taking a certain data as an example, after calculating the hash value of the data, the hash value can be matched with the hash value in the high-frequency value hash table. If the same hash value is matched in the high-frequency value hash table, it means that the data is high-frequency data. Conversely, if the same hash value is not matched in the high-frequency value hash table, it means that the data is not high-frequency data, but low-frequency data.
[0102] S630: Perform pre-aggregation processing on the high-frequency data to obtain a pre-aggregation processing result, and determine a target thread based on the grouping key of the high-frequency data.
[0103] In the implementation manner of this specification, after determining the high-frequency data and low-frequency data based on the aforementioned method process, only the high-frequency data can be pre-aggregated. The basis of the pre-aggregation processing is the grouping key (Group By Keys) of the data, that is, the data with the same grouping key is pre-aggregated. Those skilled in the art can understand this and will not be elaborated in this specification.
[0104] For high-frequency data, the corresponding pre-aggregation processing results are obtained after pre-aggregation processing, and then sent to the upper-level aggregation node. When distributing data, it is necessary to determine the thread to which the pre-aggregation processing results need to be distributed based on the data's group by key (Group By Keys), that is, to determine the target thread of the upper-level aggregation node, and then distribute the pre-aggregation data results to the target thread.
[0105] For low-frequency data, there is no need for pre-aggregation processing. The corresponding target thread is directly determined based on the grouping key (Group By Keys) of the data, and the low-frequency data is distributed to the target thread.
[0106] As shown in FIG1 , after receiving the data, each thread of the upper aggregation node can complete the final aggregation processing of the data through the upper GBY operator and feed back the data query result, which will not be described in detail in this specification.
[0107] In some implementations, the sampled data during the sampling phase may not accurately represent the actual data distribution during the entire GPD group pushdown period. For example, if a certain data item is high-frequency data during the sampling phase but appears very infrequently during the processing phase, and if all high-frequency data in the high-frequency value hash table generated based on the sampled data no longer meets the high-frequency indicator, then the high-frequency data screening process based on the high-frequency value hash table during the processing phase will result in unnecessary performance loss, leading to reduced data query performance.
[0108] Therefore, in the implementation manner of this specification, during the processing stage, the data to be processed can be periodically obtained with a second preset data volume, and each high-frequency data in the high-frequency value hash table can be detected through the data to be processed to determine whether these high-frequency data still meet the high-frequency indicators. If all high-frequency data do not meet the high-frequency indicators, the above-mentioned processing stage process is stopped, and the data is sent directly to the upper-level aggregation node without performing the pre-aggregation operation. This is explained below in conjunction with Figure 7.
[0109] As shown in FIG. 7 , in some embodiments, the data processing method exemplified in this specification includes steps S710 to S730 .
[0110] S710: Periodically obtain data to be processed with a second preset data volume, and determine frequency value parameters of each high-frequency data in the high-frequency value hash table based on the data to be processed, until the frequency value parameters of all high-frequency data in the high-frequency hash table do not meet the set requirements.
[0111] S720: In response to a frequency value parameter of at least one high-frequency data stored in the high-frequency value hash table meeting a set requirement, pre-aggregate the high-frequency data in the to-be-processed data based on the high-frequency value hash table, and distribute the pre-aggregation result to a target thread of the aggregation node.
[0112] S730 : In response to the frequency value parameters of all high-frequency data in the high-frequency value hash table not meeting the set requirements, the data to be processed is directly distributed to the aggregation node.
[0113] In the embodiment of this specification, during the processing phase, the data to be processed in the processing phase may be acquired periodically with a fixed data volume, where the fixed data volume is the second preset data volume. For example, in the example of FIG. 4 , the second preset data volume may be 10 TB, meaning that 10 TB of data to be processed is acquired per cycle.
[0114] After obtaining 10T of data to be processed, the frequency value parameter can be calculated for the high-frequency data in the high-frequency value hash table based on the data to be processed. The frequency value parameter can be the ratio in the aforementioned formula (3). In other words, the process of calculating the frequency value parameter of a certain high-frequency data is similar to the aforementioned formulas (2) to (3), except that the data used to calculate the frequency value parameter ratio is the 10T of data to be processed. The rest of the method process is exactly the same and will not be repeated in this specification.
[0115] In the implementation manner of this specification, a corresponding threshold value may be pre-set for the frequency value parameter, and the threshold value may be the aforementioned preset threshold value ratio_thre.
[0116] It can be understood that for each high-frequency data in the high-frequency value hash table, after calculating its corresponding frequency value parameter ratio based on 10T of data to be processed, the frequency value parameter ratio can be compared with the preset threshold ratio_thre. If the frequency value parameter ratio is greater than or equal to the preset threshold ratio_thre, it means that the high-frequency data meets the set requirements and is still high-frequency data. Conversely, if the calculated frequency value parameter ratio of a certain high-frequency data in the high-frequency value hash table is less than the preset threshold ratio_thre, it means that the high-frequency data no longer meets the set requirements and has become low-frequency data.
[0117] After determining whether each high-frequency data in the high-frequency value hash table meets the set requirements through the above method process, if there is at least one high-frequency data in the high-frequency value hash table with a frequency value parameter that meets the set requirements, it means that there is at least one high-frequency data in the high-frequency value hash table, so that the current processing stage can be maintained and the method process shown in the above Figure 6 can continue to be executed.
[0118] On the contrary, if the frequency value parameters of all high-frequency data in the high-frequency hash table do not meet the set requirements, it means that all high-frequency data in the high-frequency value hash table have become low-frequency data. Then, the process of screening high-frequency data based on the high-frequency value hash table in the processing stage will bring unnecessary performance loss. Therefore, combined with Figure 4, the method process shown in Figure 6 above can be no longer executed, and the data can be directly distributed to each thread of the upper-level aggregation node.
[0119] Of course, those skilled in the art will understand that in other embodiments, when more than a certain proportion of high-frequency data in the high-frequency hash table does not meet the set requirements, the method process shown in Figure 6 may be stopped. For example, when more than half of the high-frequency data in the high-frequency hash table does not meet the set requirements, the lower-level pre-aggregation node can directly distribute the data to the various threads of the upper-level aggregation node without performing pre-aggregation processing.
[0120] From the above, it can be seen that in the implementation mode of this specification, the high-frequency data in the high-frequency value hash table is periodically checked during the processing stage. When the high-frequency data in the hash table no longer meets the set requirements, the data is directly distributed to the upper-level aggregation node and the pre-aggregation operation is no longer performed, thereby avoiding the performance loss caused by the process of high-frequency data screening based on the high-frequency value hash table and improving data query performance.
[0121] In some embodiments, this specification provides a data processing device, as shown in Figure 8, the data processing device includes: an aggregation rate determination module 10, configured to determine the data aggregation rate based on the sampled data in the group push-down process; a hash table generation module 20, configured to generate a high-frequency value hash table based on the sampled data when the data aggregation rate does not meet the aggregation effect, the high-frequency value hash table including the hash value of the high-frequency data in the sampled data; a pre-aggregation module 30, configured to perform pre-aggregation processing on the high-frequency data in the data to be processed based on the high-frequency value hash table, and distribute the pre-aggregation processing results to the target thread of the aggregation node.
[0122] In one or more embodiments of the present specification, the aggregation rate determination module 10 is configured to determine the data aggregation rate based on a first data volume of the sampled data during group push-down and a second data volume after pre-aggregation processing of the sampled data.
[0123] In one or more embodiments of the present specification, the aggregation rate determination module 10 is configured to: periodically obtain the sampled data with a first preset data volume, and execute the process of determining the data aggregation rate based on the sampled data until the data aggregation rate of the current period does not meet the aggregation effect.
[0124] In one or more embodiments of the present specification, the hash table generation module 20 is configured to: in the sampled data with a total data volume of T, calculate the corresponding hash value based on the grouping key of each data, and count the first data volume that is the same as the hash value of the i-th data, where 1≤i≤T; based on the total data volume, the first data volume and the number of threads for group push-down, determine the second data volume of other data different from the hash value of the i-th data and distribute it evenly to each thread; based on the ratio of the first data volume to the second data volume, determine whether the i-th data is high-frequency data; based on all high-frequency data in the sampled data and their hash values, generate the high-frequency value hash table.
[0125] In one or more embodiments of the present specification, the hash table generation module 20 is configured to: in response to the ratio being greater than or equal to a preset threshold, determine that the i-th data is high-frequency data; in response to the ratio being less than the preset threshold, determine that the i-th data is low-frequency data.
[0126] In one or more embodiments of the present specification, the pre-aggregation module 30 is configured to: for each data in the data to be processed, calculate a corresponding hash value based on the grouping key of the data; match the hash value of the data with the hash value in the high-frequency value hash table, and in response to a consistent match, determine that the data is high-frequency data; perform pre-aggregation processing on the high-frequency data to obtain the pre-aggregation processing result, and determine the target thread based on the grouping key of the high-frequency data; and distribute the pre-aggregation processing result to the target thread.
[0127] In one or more embodiments of the present specification, the pre-aggregation module 30 is configured to: periodically obtain the data to be processed with a second preset data volume, and determine the frequency value parameters of each high-frequency data in the high-frequency value hash table based on the data to be processed, until the frequency value parameters of all high-frequency data in the high-frequency value hash table do not meet the set requirements, and the frequency value parameters represent the frequency of occurrence of the high-frequency data in the data to be processed; in response to the existence of at least one high-frequency data in the high-frequency value hash table with a frequency value parameter that meets the set requirements, execute the process of pre-aggregating the high-frequency data in the data to be processed based on the high-frequency value hash table, and distribute the pre-aggregation processing results to the target thread of the aggregation node; in response to the frequency value parameters of all high-frequency data in the high-frequency value hash table not meeting the set requirements, distribute the data to be processed directly to the aggregation node.
[0128] In some embodiments, this specification provides an electronic device, including: a processor; and a memory storing computer instructions, wherein the computer instructions are used to enable the processor to execute the method described in any of the above embodiments.
[0129] In some embodiments, this specification provides a storage medium storing computer instructions, wherein the computer instructions are used to enable a computer to execute the method described in any of the above embodiments.
[0130] FIG9 is a schematic structural diagram of an electronic device provided by an exemplary embodiment. Referring to FIG9 , at the hardware level, the device includes a processor 702, an internal bus 704, a network interface 706, a memory 708, and a non-volatile memory 710, and may also include hardware required for other scenarios. One or more embodiments of this specification may be implemented based on software, such as the processor 702 reading the corresponding computer program from the non-volatile memory 710 into the memory 708 and then running it. Of course, in addition to software implementation, one or more embodiments of this specification do not exclude other implementation methods, such as logic devices or a combination of software and hardware, etc., that is, the execution subject of the following processing flow is not limited to each logic unit, but may also be hardware or logic devices.
[0131] The systems, devices, modules, or units described in the above embodiments may be implemented by computer chips or entities, or by products having certain functions. A typical implementation device is a computer, which may be in the form of a personal computer, laptop computer, cellular phone, camera phone, smartphone, personal digital assistant, media player, navigation device, email transceiver, game console, tablet computer, wearable device, or any combination of these devices.
[0132] In a typical configuration, a computer includes one or more processors (CPU), input / output interfaces, network interfaces, and memory.
[0133] Memory may include non-permanent storage in a computer-readable medium, random access memory (RAM) and / or non-volatile memory in the form of read-only memory (ROM) or flash RAM. Memory is an example of a computer-readable medium.
[0134] Computer-readable media include permanent and non-permanent, removable and non-removable media that can be used to store information using any method or technology. Information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassettes, disk storage, quantum memory, graphene-based storage media or other magnetic storage devices, or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer-readable media does not include transitory media such as modulated data signals and carrier waves.
[0135] It should also be noted that the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, commodity, or apparatus that includes a series of elements includes not only those elements but also other elements not explicitly listed, or includes elements inherent to such process, method, commodity, or apparatus. In the absence of further limitations, an element defined by the phrase "comprises a ..." does not exclude the presence of other identical elements in the process, method, commodity, or apparatus that includes the element.
[0136] The foregoing description of specific embodiments of this specification describes the process. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims may be performed in an order different from that described in the embodiments and still achieve the desired results. Furthermore, the processes depicted in the accompanying drawings do not necessarily require the specific order shown or the sequential order to achieve the desired results. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0137] The terms used in one or more embodiments of this specification are for the purpose of describing specific embodiments only and are not intended to limit one or more embodiments of this specification. The singular forms "a," "an," "the," and "the" used in one or more embodiments of this specification and the appended claims are also intended to include plural forms unless the context clearly indicates otherwise. It should also be understood that the term "and / or" used herein refers to and includes any or all possible combinations of one or more associated listed items.
[0138] It should be understood that although the terms first, second, third, etc. may be used in one or more embodiments of this specification to describe various information, such information should not be limited to these terms. These terms are only used to distinguish information of the same type from each other. For example, without departing from the scope of one or more embodiments of this specification, first information may also be referred to as second information, and similarly, second information may also be referred to as first information. Depending on the context, the word "if" as used herein may be interpreted as "at the time of" or "when" or "in response to determining."
[0139] The above description is merely a preferred embodiment of one or more embodiments of this specification and is not intended to limit one or more embodiments of this specification. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of one or more embodiments of this specification shall be included in the scope of protection of one or more embodiments of this specification.
Claims
1. A data processing method, comprising: Determining a data aggregation rate based on sampled data during a grouping pushdown process; Generating a high-frequency value hash table based on the sampled data when the data aggregation rate does not meet the aggregation effect, the high-frequency value hash table including hash values of high-frequency data in the sampled data; Based on the high-frequency value hash table, performing pre-aggregation processing on high-frequency data in the data to be processed and distributing the pre-aggregation processing result to a target thread of an aggregation node.
2. The data processing method according to claim 1, wherein the determining the data aggregation rate based on the sampled data during the grouping pushdown process comprises: Determining the data aggregation rate based on a first data volume of the sampled data during the grouping pushdown process and a second data volume after performing pre-aggregation processing on the sampled data.
3. The data processing method according to claim 1, wherein the determining the data aggregation rate based on the sampled data during the grouping pushdown process comprises: Periodically obtaining the sampled data in a first preset data volume and performing the process of determining the data aggregation rate based on the sampled data until the data aggregation rate in the current period does not meet the aggregation effect.
4. The data processing method according to any one of claims 1 to 3, wherein the generating the high-frequency value hash table based on the sampled data comprises: In the sampled data with a total data volume of T, calculating a corresponding hash value based on the grouping key of each data and counting a first data volume that is the same as the hash value of the i-th data, where 1 ≤ i ≤ T; Based on the total data volume, the first data volume, and the number of threads of the grouping pushdown, determining a second data volume for evenly distributing other data different from the hash value of the i-th data to each thread; Determining whether the i-th data is high-frequency data based on a ratio of the first data volume to the second data volume; Generating the high-frequency value hash table based on all high-frequency data and their hash values in the sampled data.
5. The data processing method according to claim 4, wherein the determining whether the i-th data is high-frequency data based on the ratio of the first data volume to the second data volume comprises: Responding to the ratio being greater than or equal to a preset threshold, determining that the i-th data is high-frequency data; Responding to the ratio being less than the preset threshold, determining that the i-th data is low-frequency data.
6. The data processing method according to claim 1, wherein the performing pre-aggregation processing on high-frequency data in the data to be processed based on the high-frequency value hash table and distributing the pre-aggregation processing result to a target thread of an aggregation node comprises: For each data in the data to be processed, calculating a corresponding hash value based on the grouping key of the data; Matching the hash value of the data with the hash values in the high-frequency value hash table, and responding to a match, determining that the data is high-frequency data; Performing pre-aggregation processing on the high-frequency data to obtain the pre-aggregation processing result and determining the target thread based on the grouping key of the high-frequency data; Distributing the pre-aggregation processing result to the target thread.
7. The data processing method according to claim 1, further comprising: Periodically obtain the data to be processed with a second preset data volume, and determine the frequency value parameters of each high-frequency data in the high-frequency value hash table based on the data to be processed until the frequency value parameters of all high-frequency data in the high-frequency value hash table do not meet the set requirements, where the frequency value parameter represents the occurrence frequency of the high-frequency data in the data to be processed; In response to that there is at least one high-frequency data in the high-frequency value hash table whose frequency value parameter meets the set requirements, execute the process of pre-aggregating the high-frequency data in the data to be processed based on the high-frequency value hash table and distributing the pre-aggregation processing result to the target thread of the aggregation node; In response to that the frequency value parameters of all high-frequency data in the high-frequency value hash table do not meet the set requirements, directly distribute the data to be processed to the aggregation node.
8. A data processing device, comprising: An aggregation rate determination module configured to determine a data aggregation rate based on sampled data during a grouping pushdown process; A hash table generation module configured to generate a high-frequency value hash table based on the sampled data when the data aggregation rate does not meet the aggregation effect, where the high-frequency value hash table includes hash values of high-frequency data in the sampled data; A pre-aggregation module configured to pre-aggregate the high-frequency data in the data to be processed based on the high-frequency value hash table and distribute the pre-aggregation processing result to the target thread of the aggregation node.
9. An electronic device, comprising: A processor; And A memory storing computer instructions for causing the processor to execute the method according to any one of claims 1 to 7.
10. A storage medium storing computer instructions for causing a computer to execute the method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Parallel data partitioning method and device, electronic equipment and storage medium
CN111770025A
Distributed oblique flow processing method and system based on high-frequency key value counting
CN112783644A
Aggregate query method and device
CN116501756A
Data processing method and device
CN117807086A
Reduction of Volume of Reporting Data Using Multiple Datasets
US20180285373A1