Data processing method, device, system, electronic device and storage medium

By identifying and merging similar deduplication operators and generating new execution plans, the problems of low computational efficiency and large memory usage of traditional multiple deduplication are solved, efficient batch deduplication marking is achieved, and the performance and stability of analytical databases are improved.

CN116431660BActive Publication Date: 2025-08-26ALIBABA CLOUD COMPUTING CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310379498.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-03-31
Publication Date
2025-08-26
Estimated Expiration
2043-03-31

AI Technical Summary

Technical Problem

Traditional multi-deduplication calculations are inefficient in analytical databases, with many calculations, large memory occupies, affect machine stability, and require manual rewriting of query plans, resulting in poor performance.

Method used

By identifying and merging similar deduplication operators, generating similar operator sets, generating new execution plans, implementing batch deduplication mark calculations, reducing the number of calculations and memory usage.

Benefits of technology

On the basis of not increasing data volume and memory usage, improve computing performance, reduce query time, and improve user experience and overall database processing performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116431660B_ABST
    Figure CN116431660B_ABST
Patent Text Reader

Abstract

The present application provides a data processing method, device, system, electronic device and storage medium. The data processing method includes: determining similar deduplication operators in the deduplication operators of the target workload; merging similar deduplication operators into a similar operator set; generating a corresponding execution plan based on the similar operator set; and sending the execution plan to the execution device. The execution plan includes information about the similar operator set and the multiple deduplication marking operators corresponding to the similar operator set. The multiple deduplication marking operators are used to perform batch deduplication marking calculations on the data to be processed in the database, and the data to be processed includes data corresponding to the similar operator set. According to the embodiments of the present application, data processing performance can be improved and memory usage can be reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of data processing technology, and in particular to a data processing method, device, system, electronic device and storage medium. Background Art

[0002] The Distinct operator (deduplication operator) is a native SQL (Structured Query Language) statement commonly used in analytical databases. Its function is to deduplicate the specified data. In actual application scenarios, multiple deduplication operators are usually used in a query for calculation (i.e., multiple deduplication calculations). In the traditional multiple deduplication calculation process, each deduplication operator is usually used one by one to calculate the deduplication mark, or the data is expanded into multiple copies based on the Grouping Set method, and then the deduplication mark is calculated uniformly. The one-by-one calculation method has more calculation times and poor computing performance. The method of expanding the data and then calculating uniformly requires manual rewriting of the query plan, which is less efficient. In addition, the larger amount of data after expansion leads to a larger memory usage, which can easily cause memory shortage and even machine instability. Summary of the Invention

[0003] The embodiments of the present application provide a data processing method, device, system, electronic device and storage medium to solve the problems existing in the prior art.

[0004] In a first aspect, an embodiment of the present application provides a data processing method, applied to an optimization device, the method comprising:

[0005] determining similar deduplication operators among deduplication operators of the target workload;

[0006] Merge similar deduplication operators into a similar operator set;

[0007] Generate a corresponding execution plan based on the similarity operator set; the execution plan includes information about the similarity operator set and a multiple deduplication operator corresponding to the similarity operator set, the multiple deduplication operator being used to perform batch deduplication calculations on data to be processed in the database, the data to be processed including data corresponding to the similarity operator set;

[0008] Send the execution plan to the execution device.

[0009] In a second aspect, an embodiment of the present application provides a data processing method, applied to an execution device, the method comprising:

[0010] Based on the execution plan provided by the optimization device, batch deduplication marking is performed on the data to be processed in the database to obtain deduplication marking information; the execution plan is obtained by the data processing method provided by the first aspect of the embodiment of the present application, and the data to be processed includes data corresponding to each similar operator set in the target workload;

[0011] Batch deduplication processing is performed on the data to be processed based on the deduplication tag information.

[0012] In a third aspect, an embodiment of the present application provides a data processing apparatus for use in optimizing a device, the apparatus comprising:

[0013] a similarity determination module, configured to determine similar deduplication operators among the deduplication operators of the target workload;

[0014] Operator merging module, used to merge similar deduplication operators into a similar operator set;

[0015] A plan generation module is used to generate a corresponding execution plan based on the similarity operator set; the execution plan includes information about the similarity operator set and the multiple deduplication marking operator corresponding to the similarity operator set, and the multiple deduplication marking operator is used to perform batch deduplication marking calculations on the data to be processed in the database, and the data to be processed includes the data corresponding to the similarity operator set;

[0016] The plan sending module is used to send the execution plan to the execution device.

[0017] In a fourth aspect, an embodiment of the present application provides a data processing apparatus, applied to an execution device, the apparatus comprising:

[0018] a deduplication marking module for batch deduplication marking of the data to be processed in the database based on an execution plan to obtain deduplication marking information; the execution plan is obtained by the data processing device provided by the third aspect of the embodiment of the present application, and the data to be processed includes data corresponding to each similarity operator set;

[0019] The deduplication processing module is used to perform batch deduplication processing on the data to be processed based on the deduplication mark information.

[0020] In a fifth aspect, an embodiment of the present application provides an electronic device, comprising: a memory, a processor, and a computer program stored in the memory, wherein the processor implements any one of the methods provided in the embodiment of the present application when executing the computer program.

[0021] In a sixth aspect, an embodiment of the present application provides a data processing system, comprising: a communication-connected optimization device and an execution device;

[0022] At least one of the optimization device and the execution device is the electronic device provided in the fifth aspect of the embodiment of the present application.

[0023] In a seventh aspect, an embodiment of the present application provides a computer-readable storage medium storing a computer program, which, when executed by a processor, implements any one of the methods provided in the embodiment of the present application.

[0024] Compared with the prior art, this application has the following advantages:

[0025] According to the technical solution of the embodiment of the present application, the similarities of each deduplication operator in the target workload can be discovered, and optimization opportunities for deduplication marking calculations can be automatically identified based on the discovered similarities. When optimization opportunities are identified (i.e., similar deduplication operators are determined), multiple similar deduplication operators can be merged to generate multiple deduplication marking operators corresponding to the merged similar operator set, so as to rewrite the execution plan of the target workload and form a new execution plan to realize batch deduplication marking calculations for the data specified by the similar operator set. This can effectively reduce the number of deduplication marking calculations and query time without increasing the amount of data or memory usage, and thus effectively improve computing performance without affecting machine stability, thereby improving the overall processing performance of the database and enhancing the user experience of analytical workloads.

[0026] The above description is only an overview of the technical solution of this application. In order to more clearly understand the technical means of this application, it can be implemented in accordance with the contents of the specification. In order to make the above and other purposes, features and advantages of this application more obvious and easy to understand, the specific implementation methods of this application are listed below. BRIEF DESCRIPTION OF THE DRAWINGS

[0027] In the accompanying drawings, unless otherwise specified, the same reference numerals throughout the multiple drawings represent the same or similar components or elements. These drawings are not necessarily drawn to scale. It should be understood that these drawings only depict some embodiments according to the present application and should not be regarded as limiting the scope of the present application.

[0028] Figure 1 A schematic diagram of the data processing solution provided for this application;

[0029] Figure 2 A flowchart of a data processing method provided in an embodiment of the present application;

[0030] Figure 3 A flowchart of another data processing method provided in an embodiment of the present application;

[0031] Figure 4 This is an example diagram of deduplication in an embodiment of the present application;

[0032] Figure 5A schematic diagram of the structural framework of a data processing device provided in an embodiment of the present application;

[0033] Figure 6 A schematic diagram of the structural framework of another data processing device provided in an embodiment of the present application; and

[0034] Figure 7 A schematic diagram of the structural framework of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0035] Hereinafter, only certain exemplary embodiments are briefly described. As will be appreciated by those skilled in the art, the described embodiments may be modified in various ways without departing from the spirit or scope of the present application. Therefore, the drawings and description are to be regarded as illustrative in nature and not restrictive.

[0036] First, some technical terms involved in the embodiments of this application are introduced as follows:

[0037] Analytical database: A database used for analytical workloads that can process large-scale data and perform analytical calculations.

[0038] A deduplication operator is a native SQL statement that deduplicates specified data, retaining only one row of duplicate data and removing all other rows. The structure of a deduplication operator is as follows: distinct SELECTED_COLUMN, where SELECTED_COLUMN is the column specified by the deduplication operator and the statement requests deduplication of that column.

[0039] Multiple deduplication calculations: In the same workload, such as the same query, multiple deduplication operators are used for calculations.

[0040] Similarity of deduplication operators: If multiple deduplication operators specify (i.e., calculate) the same data column, then these operators are considered similar and can be used as a group of similar deduplication operators. The data columns specified by each deduplication operator can be filtered according to their respective filter conditions, and data columns that do not meet the filter conditions can be filled with NULL. For example, if there are two deduplication operators, the first of which specifies the data column C1 selected by the first filter condition FILTER1, and the second specifies the data column C1 selected by the second filter condition FILTER2, then the two deduplication operators specify the same data column and can be used as a group of similar deduplication operators.

[0041] Data redistribution: In a multi-concurrency query process, data is redistributed to multiple concurrencies using specific data columns as key values, so that the same data is distributed to the same concurrency. This process is called data redistribution.

[0042] In the era of big data, all industries generate vast amounts of data. Analyzing this data, such as analyzing past annual sales figures or user behavior, can further guide production optimization and application direction. This type of data processing is called analytical workloads and is handled by specialized analytical databases. Analytical databases often need to process massive amounts of data all at once, such as summing all transactions in a year to calculate the annual transaction volume. This means that a single query often consumes significant resources and time. Therefore, performance optimization in analytical databases has a significant impact on user experience.

[0043] In analytical databases, the deduplication operator is a commonly used SQL native statement. The calculation process of the deduplication operator usually involves complex operations such as data redistribution and deduplication marking, which has a high computational cost. Traditional multiple deduplication calculations are often inefficient and require a large amount of query time, which can easily become a query bottleneck, resulting in longer response times and a decreased user experience. It may also increase memory usage and affect machine stability.

[0044] In response to the defects of the traditional multiple deduplication calculation method, research has found that in actual application scenarios, when a target workload, such as a query, involves the use of multiple deduplication operators, there are often similar deduplication operators among these deduplication operators. These deduplication operators with high similarity can be used in batches to perform deduplication mark calculations, and then batch deduplication processing can be performed on the data to be processed in the database. Based on the similarity between deduplication operators, this application designs a new data processing solution that can avoid one-by-one calculations and data expansion. On the basis of effectively reducing the number of calculations to improve performance, it can reduce the amount of data, reduce memory usage, and improve machine stability. In addition, it can automatically rewrite query plans to improve query efficiency. The technical solution of this application can be applied to Figure 1 The data processing system shown can achieve efficient and memory-saving deduplication processing by improving the processing methods at both ends of the optimization device and the execution device in the data processing system.

[0045] To facilitate understanding of the technical solutions of the embodiments of the present application, the following describes the related technologies of the embodiments of the present application. The following related technologies can be combined with the technical solutions of the embodiments of the present application as optional solutions, and all of them fall within the scope of protection of the embodiments of the present application.

[0046] The present application embodiment provides a data processing method that can be applied to optimize equipment, such as Figure 2 As shown, the method may include the following steps: S201, determining similar deduplication operators among the deduplication operators of the target workload (e.g., query); S202, merging the similar deduplication operators into a similar operator set; S203, generating a corresponding execution plan based on the similar operator set; S204, sending the execution plan to the execution device. The execution plan includes information about the similar operator set and a multi-mark distinguishing operator (MultiMarkDistinct operator) corresponding to the similar operator set. The multi-mark distinguishing operator can be used to perform batch deduplication calculations on the data to be processed in the database. The data to be processed may include data corresponding to the similar operator set.

[0047] The above-mentioned database may be an analytical database, and the above-mentioned batch deduplication marking may mean that at least part of the data to be processed is deduplicated synchronously without the need to deduplicate one by one.

[0048] The data processing method for optimization equipment provided in the embodiment of the present application can discover the similarities of each deduplication operator in the target workload, and automatically identify optimization opportunities for deduplication marking calculation based on the discovered similarities. When the optimization opportunity is identified (that is, similar deduplication operators are determined), multiple similar deduplication operators can be merged to generate multiple deduplication marking operators corresponding to the merged similar operator set, so as to rewrite the execution plan of the target workload and form a new execution plan to realize batch deduplication marking calculation of the data specified by the similar operator set. It can effectively reduce the number of deduplication marking calculations without increasing the amount of data or the memory usage, reduce the query time and the response time to the user, and effectively improve the computing performance without affecting the stability of the machine, thereby improving the overall processing performance of the database and enhancing the user experience of analytical workloads.

[0049] In one embodiment, in the above-mentioned step S201, similar deduplication operators among the deduplication operators are determined based on the similarities and differences of the data specified by each deduplication operator in the target workload, which may include: determining whether the data corresponding to each deduplication operator in the target workload has common data, that is, data that can be used by multiple deduplication operators; in the case that the data corresponding to multiple deduplication operators are common data, determining that the multiple deduplication operators are a group of similar deduplication operators, otherwise the multiple deduplication operators are not similar deduplication operators.

[0050] In one example, it can be determined whether the data columns filtered out by the filtering conditions corresponding to each deduplication operator of the target workload (i.e., the data columns that meet the filtering conditions) are the same; when the data columns filtered out by the filtering conditions corresponding to multiple deduplication operators are the same, the data in the same data columns are common data, and the same data columns can be called common data columns, thereby determining that the multiple deduplication operators are a group of similar deduplication operators, otherwise the multiple deduplication operators are not similar deduplication operators.

[0051] A group of similar deduplication operators can be combined to form a similar operator set, and multiple groups of similar deduplication operators can be combined to form multiple similar operator sets. The data corresponding to a similar operator set can include shared data corresponding to each similar deduplication operator in the similar operator set. The filtering conditions corresponding to each similar deduplication operator in a similar operator set can filter out a shared data column, and the filtering conditions corresponding to each similar deduplication operator in multiple similar operator sets can filter out multiple shared data columns.

[0052] In this way, similar deduplication operators can be determined based on the common data corresponding to each deduplication operator. The similar deduplication operators determined in this way provide a data basis for batch deduplication marking because the corresponding data are the same.

[0053] In one embodiment, in the above-mentioned step S203, generating a corresponding execution according to a similar operator set may include: taking the common data corresponding to the similar operator set and the validity marking information corresponding to the common data as input information, and taking the deduplication marking information corresponding to the common data as output information, to construct a multiple deduplication marking operator; the multiple deduplication marking operator is used to synchronously deduplication the common data corresponding to each similar operator set according to the validity marking information corresponding to each similar operator set, and obtain the deduplication marking information corresponding to each similar operator set.

[0054] The shared data (such as shared data columns) corresponding to a similar operator set can reflect the similarity of each deduplication operator in the similar operator set. The multiple deduplication marking operators constructed based on the shared data and its corresponding validity marking information can synchronously deduplication the shared data corresponding to each similar operator set according to the validity marking information corresponding to the shared data. There is no need to perform deduplication calculations one by one, which can effectively reduce the amount of calculation on the basis of realizing multiple deduplication calculations. For a similar operator set, the multiple deduplication marking operators deduplication the shared data corresponding to the similar operator set, which is equivalent to deduplication the data corresponding to all deduplication operators in the similar operator set, which can further reduce the amount of calculation.

[0055] In one embodiment, in step S203, generating a corresponding execution plan based on a similar operator set may further include: for each similar operator set, determining whether each data in the shared data corresponding to the similar operator set is valid based on the filtering conditions corresponding to each similar deduplication operator in the similar operator set, and marking each data in the shared data for validity, thereby obtaining a set of validity marking information corresponding to the shared data. The execution plan in the embodiment of the present application may also include validity marking information.

[0056] Based on this implementation, the embodiment of the present application can mark the shared data based on the filtering conditions, that is, validity marking, and send the results of the validity marking to the execution device as part of the execution plan. The execution device can perform secondary marking based on the results of the validity marking, that is, deduplication marking.

[0057] In one example, each similar deduplication operator in the similarity operator set corresponds to a filtering condition. Different deduplication operators may correspond to different filtering conditions. Based on the filtering conditions corresponding to any similar deduplication operator, a corresponding set of validity mark information may be obtained. N similar deduplication operators corresponding to N filtering conditions may obtain corresponding N sets of validity mark information. The validity mark information may be recorded in the form of a data column, that is, the validity mark information may be a validity mark column, and the N sets of validity mark information may be N sets of validity mark columns.

[0058] The validity of a row of data can be determined based on whether the row of data meets the corresponding filtering conditions. If a row of data meets the corresponding filtering conditions, for example, the filtering conditions are to filter data greater than 1, and the row of data is greater than 1, then the row of data can be considered as valid row data. If a row of data does not meet the corresponding filtering conditions, for example, the filtering conditions are to filter data greater than 1, and the row of data is less than 1, then the row of data can be considered as invalid row data.

[0059] In the above method, according to the filtering conditions corresponding to each similar deduplication operator in the similar operator set, it is determined whether each row of data in the same data column corresponding to the similar operator set is valid, and the validity of each row of data in the same data column is marked. This can be achieved through the corresponding project operator (projection operator). For example, a project operator can be generated for the filtering conditions corresponding to each deduplication operator in the similar operator set. The calculation rule of the project operator can be to record the judgment result of whether the data column meets the filtering conditions, for example, it is recorded as True when it meets the conditions and False when it does not meet the conditions. In this way, the validity marking of the same data column can be achieved through the project operator.

[0060] In one embodiment, in step S203, generating corresponding executions based on similar operator sets may further include: synchronously redistributing the shared data corresponding to each similar operator set. The execution plan in the embodiment of the present application may further include the shared data after data redistribution.

[0061] The calculation process of the deduplication operator usually involves data redistribution. The traditional calculation method is to use each deduplication operator one by one for calculation. When there are N (a positive integer greater than 1) deduplication operators in a target workload, the traditional calculation method requires N data redistributions. The embodiment of the present application can be based on the similarity between the deduplication operators. The data columns corresponding to each similar operator set are synchronously redistributed, which can effectively reduce the number of data redistributions and thus reduce the amount of calculation. For a similar operator set, the same data columns corresponding to the similar operator set can reflect the similarity of each deduplication operator in the similar operator set. Data redistribution of the same data columns corresponding to the similar operator set is equivalent to completing the data redistribution of all deduplication operators in the similar operator set through a unified data redistribution, which can further reduce the number of data distributions and further reduce the amount of calculation.

[0062] Based on the same technical concept, the embodiment of the present application also provides a data processing method that can be applied to an execution device, such as Figure 3 As shown, the method may include the following steps: S301, batch deduplication marking of the data to be processed in the database based on the execution plan to obtain deduplication marking information; S302, batch deduplication processing of the data to be processed based on the deduplication marking information. The execution plan may be obtained by the data processing method for the optimization device provided in the embodiment of the present application and provided by the optimization device, and the data to be processed may include data corresponding to each similar operator set in the target workload.

[0063] Based on the execution plan provided by the optimization device, the data processing method applied to the execution device provided in the embodiment of the present application can perform batch deduplication on the data to be processed in the database based on the multiple deduplication operators, that is, deduplication can be performed synchronously on at least part of the data to be processed. This can effectively reduce the number of deduplication calculations and query time without increasing the amount of data or memory usage, and can effectively improve computing performance without affecting machine stability, thereby improving the overall processing performance of the database.

[0064] In one embodiment, the data corresponding to the similar operator set may include data in the same data column corresponding to each similar deduplication operator in the similar operator set.

[0065] Correspondingly, in step S301, batch deduplication of the to-be-processed data in the database based on the execution plan to obtain deduplication information may include: inputting the shared data and validity marking information corresponding to each similar operator set in the execution plan into a multiple deduplication operator in the execution plan, and using the multiple deduplication operator to simultaneously deduplication the same data columns corresponding to each similar operator set to obtain the deduplication information corresponding to each similar operator set. This allows deduplication to be performed on the validity marking results, thereby improving the accuracy and efficiency of deduplication.

[0066] In one embodiment, the same data columns corresponding to each similar operator set are de-duplicated synchronously by multiple de-duplication operators, which may include: for the common data corresponding to each similar operator set, the multiple de-duplication operators determine whether each data in the common data is valid based on the validity marking information corresponding to the common data; for invalid data, the invalid data is marked as non-duplicate data by the multiple de-duplication operators, which serves as the de-duplication marking information for the invalid data; for valid data, the valid data is determined by the multiple de-duplication operators whether it is duplicate data in the current common data, and the valid data is marked as duplicate data or non-duplicate data based on the determination result, which serves as the de-duplication marking information for the valid data. If a certain data is marked as non-duplicate data, it means that the data appears for the first time in the valid data in the common data and is data that needs to be retained. If a certain data is marked as duplicate data, it means that the data does not appear for the first time (i.e., appears repeatedly) in the valid data in the common data or is invalid data and is data that needs to be removed. The common data may be the data of a data column, which may be called a common data column, and the data in the common data may be the data in each row of the common data column.

[0067] For a set of validity mark information, whether the marked object is valid, invalid, duplicate or non-duplicate can be determined based on a data set (Set) combined with data traversal. For example, a corresponding data set can be set for a common data column to store valid non-duplicate data, and each validity mark information in the set of validity mark information can be traversed. If the validity mark information shows that the corresponding row data is invalid data, the row data can be marked as duplicate data. If the validity mark information shows that the corresponding row data is valid data, it can be determined whether the row data already exists in the corresponding data set. If it exists, it means that the row data is not the first time it appears in the valid data and is marked as duplicate data. If it does not exist, it means that the row data appears for the first time in the valid data and can be marked as non-duplicate data and inserted into the data set. After the traversal is completed, the data in the data set is the data that needs to be retained. A row of data in a common data column usually corresponds to multiple validity mark columns. For a row of data, its corresponding validity mark columns can be traversed in turn to determine the corresponding deduplication mark information. Different groups of validity marking information can be used to construct data sets respectively, so as to better control the size of the data set, avoid the data set being too large, effectively reduce memory usage, and improve computing performance.

[0068] Both the validity mark information and the deduplication mark information can be recorded in the form of data columns. A set of validity mark information can be a validity mark column. Based on a validity mark column, a corresponding deduplication mark column can be obtained. Figure 4 An example of deduplication based on common data columns and validity mark columns in an embodiment of the present application is shown. Figure 4 Common_column in the table represents the common data column corresponding to a similarity operator set. Predicate_column_1 represents the validity mark column obtained based on the filter condition corresponding to the first similarity deduplication operator in the similarity operator set (hereinafter referred to as the first filter condition). Predicate_column_2 represents the validity mark column obtained based on the filter condition corresponding to the second similarity deduplication operator in the similarity operator set (hereinafter referred to as the second filter condition). Mark_column_1 represents the deduplication mark column corresponding to Predicate_column_1, and Mark_column_2 represents the deduplication mark column corresponding to Predicate_column_2. In the validity mark column, valid row data can be marked with True or T, and invalid row data can be marked with False or F. In the deduplication mark column, non-duplicate data can be marked with True or T, and duplicate data can be marked with False or F.

[0069] Reference Figure 4For example, there are five rows of data in the common data column Common_column, namely a, b, a, b, b. You can start from the first row and traverse each row of data and the corresponding validity mark column. Taking the first row of data a as an example, first traverse its corresponding first validity mark column Predicate_column_1, and obtain the mark information F of the first row of data a in Predicate_column_1, indicating that the first row of data a does not meet the first filtering condition and is invalid. In the first deduplication mark column Mark_column_1, mark the first row of data a as F, and then traverse the second validity mark column Predicate_column_2 corresponding to the first row of data a, and obtain the mark information T of the first row of data a in Predicate_column_2, indicating that the first row of data a is valid row data that meets the second filtering condition. At this time, it is necessary to determine whether the first row of data a appears for the first time in the same data column Common_column. Figure 4 As can be seen, the first row of data a appears for the first time in the same data column Common_column, so it is non-duplicate data. The first row of data a can be marked as T in the second deduplication mark column Mark_column_2. The traversal method for other rows of data is similar to the traversal method for the first row of data a, so it will not be repeated here.

[0070] In the Figure 4 After all the row data in are traversed, we can get Figure 4 The first deduplication mark column Mark_column_1 and the second deduplication mark column Mark_column_2 shown in the figure can remove duplicate data based on the first deduplication mark column Mark_column_1 and the second deduplication mark column Mark_column_2 to complete the deduplication process. Figure 5 In the figure, the mark information of the third and fourth rows of the first deduplication mark column Mark_columN_1 are both T, which means the data needs to be retained, and the other row data marked as F are all data that need to be removed. The mark information of the first and second rows of the second deduplication mark column Mark_columN_2 are both T, which means the data needs to be retained, and the other row data marked as F are all data that need to be removed.

[0071] In one embodiment, the shared data input to the multiple deduplication marking operators is the shared data after data redistribution. After data redistribution to ensure that duplicate data is assigned to the same concurrent thread, the shared data is then input to the multiple deduplication marking operators. This allows deduplication processing to be performed within the same concurrent thread, which improves deduplication efficiency. Data redistribution can be performed by an optimization device, and the shared data after data redistribution can be provided by the optimization device to the current execution device.

[0072] Based on the same technical concept, the embodiment of the present application also provides a data processing device that can be used to optimize equipment, such as Figure 5 As shown, the device may include: a similarity determination module 501, an operator merging module 502, a plan generating module 503 and a plan sending module 504.

[0073] The similarity determination module 501 is used to determine similar deduplication operators among the deduplication operators of the target workload; the operator merging module 502 is used to merge similar deduplication operators into a similar operator set; the plan generation module 503 is used to generate a corresponding execution plan based on the similar operator set; and the plan sending module 504 is used to send the execution plan to the execution device. The execution plan can be used to perform batch deduplication marking calculations on unprocessed data in the database, where the unprocessed data includes data corresponding to the similar operator set.

[0074] In one embodiment, the similarity determination module 501 can be used to: determine whether there is common data corresponding to each deduplication operator in the target workload; when the data corresponding to multiple deduplication operators are common data, determine that the multiple deduplication operators are a group of similar deduplication operators; the data corresponding to the similar operator set includes the common data corresponding to each similar deduplication operator in the similar operator set.

[0075] In one embodiment, the plan generation module 503 can be used to: take the common data corresponding to the similar operator set and the validity mark information corresponding to the common data as input information, and take the deduplication mark information corresponding to the common data as output information to construct a multiple deduplication mark operator; the multiple deduplication mark operator is used to synchronously deduplication mark the common data corresponding to each similar operator set according to the validity mark information corresponding to each similar operator set, and obtain the deduplication mark information corresponding to each similar operator set.

[0076] In one embodiment, the plan generation module 503 may also be configured to: for each similarity operator set, determine whether each data item in the shared data corresponding to the similarity operator set is valid based on the filtering conditions corresponding to each similar deduplication operator in the similarity operator set, and perform validity marking on each data item in the shared data to obtain a set of validity marking information corresponding to the shared data. The execution plan may also include the validity marking information.

[0077] In one embodiment, the plan generation module 503 may also be used to synchronously redistribute the same data corresponding to each similar operator set. The execution plan may also include the shared data after the data redistribution.

[0078] Based on the same technical concept, the embodiment of the present application also provides a data processing device, which can be applied to an execution device, such as Figure 6As shown, the apparatus may include: a deduplication marking module 601 and a deduplication processing module 602 .

[0079] The deduplication marking module 601 is configured to batch deduplication the data to be processed in the database based on the execution plan to obtain deduplication marking information. The deduplication processing module 602 is configured to batch deduplication the data to be processed based on the deduplication marking information. The execution plan may be obtained by the data processing apparatus for the optimization device provided in an embodiment of the present application, and the data to be processed includes data corresponding to each similarity operator set.

[0080] In one embodiment, the data corresponding to the similar operator set includes the shared data corresponding to each similar deduplication operator in the similar operator set.

[0081] Correspondingly, the deduplication marking module 601 can be used to: input the common data and validity marking information corresponding to each similar operator set in the execution plan into the multiple deduplication marking operators in the execution plan, and synchronously deduplication the common data corresponding to each similar operator set through the multiple deduplication marking operators to obtain the deduplication marking information corresponding to each similar operator set.

[0082] In one embodiment, when deduplication is performed on the shared data corresponding to each similar operator set synchronously through multiple deduplication operators, the deduplication module 601 can be used to: for the shared data corresponding to each similar operator set, the multiple deduplication operators determine whether each data in the shared data is valid based on the validity marking information corresponding to the shared data; for invalid data, the multiple deduplication operators are used to mark the invalid data as non-duplicate data as the deduplication information of the invalid data; for valid data, the multiple deduplication operators are used to determine whether the valid data is duplicate data in the current shared data, and the valid data is marked as duplicate data or non-duplicate data according to the determination result as the deduplication information of the valid data.

[0083] In one embodiment, the common data input into the multiple de-duplication marking operators is the common data after data redistribution.

[0084] The functions of each module in each device in the embodiments of the present application can be found in the corresponding description in the above method, and have corresponding beneficial effects, which will not be repeated here.

[0085] Based on the same technical concept, the embodiment of the present application also provides an electronic device that can be used as an optimization device or an execution device, such as Figure 7As shown, the electronic device includes: a memory 701 and a processor 702. The memory 701 stores a computer program that can be run on the processor 702. When the processor 702 executes the computer program, any one of the methods in the above embodiments is implemented. The number of the memory 701 and the processor 702 can be one or more.

[0086] The electronic device also includes:

[0087] The communication interface 703 is used to communicate with external devices and perform data exchange transmission.

[0088] If the memory 701, processor 702, and communication interface 703 are implemented independently, the memory 701, processor 702, and communication interface 703 can be connected to each other via a bus and communicate with each other. The bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, or an Extended Industry Standard Architecture (EISA) bus. The bus can be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 7 Only one thick line is used in the diagram, but this does not mean that there is only one bus or one type of bus.

[0089] Optionally, in a specific implementation, if the memory 701, the processor 702 and the communication interface 703 are integrated on a chip, the memory 701, the processor 702 and the communication interface 703 can communicate with each other through an internal interface.

[0090] Based on the same technical concept, Figure 1 The embodiment of the present application further provides a data processing system, which may include an optimization device and an execution device connected in communication, and at least one of the optimization device and the execution device may be as follows: Figure 7 The electronic device shown. The optimization device can be configured to perform the following Figure 2 The optimizer of the data processing method shown in FIG. 1 can be set on the execution device to execute the data processing method as shown in FIG. Figure 3 The execution engine of the data processing method shown.

[0091] There can be only one optimization device and multiple execution devices. Each execution device can be connected to the optimization device for communication to receive an execution plan provided by the optimization device.

[0092] Based on the same technical concept, an embodiment of the present application also provides a computer-readable storage medium, which stores a computer program, and when the program is executed by a processor, the method provided in the embodiment of the present application is implemented.

[0093] Based on the same technical concept, an embodiment of the present application also provides a chip, which includes a processor for calling and running instructions stored in the memory from the memory, so that a communication device equipped with the chip executes the method provided in the embodiment of the present application.

[0094] Based on the same technical concept, an embodiment of the present application also provides a chip, including: an input interface, an output interface, a processor and a memory. The input interface, the output interface, the processor and the memory are connected through an internal connection path. The processor is used to execute the code in the memory. When the code is executed, the processor is used to execute the method provided in the embodiment of the application.

[0095] It should be understood that the processor may be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor may be a microprocessor or any conventional processor, etc. It is worth noting that the processor may be a processor that supports the Advanced RISC Machines (ARM) architecture.

[0096] Furthermore, optionally, the above-mentioned memory may include a read-only memory and a random access memory. The memory may be a volatile memory or a non-volatile memory, or may include both volatile and non-volatile memories. Among them, the non-volatile memory may include a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), or a flash memory. The volatile memory may include a random access memory (RAM), which is used as an external cache. By way of example but not limitation, many forms of RAM are available. For example, static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDR SDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous link DRAM (SLDRAM) and direct memory bus random access memory (DR RAM).

[0097] In the above embodiments, all or part of the embodiments may be implemented using software, hardware, firmware, or any combination thereof. When implemented using software, all or part of the embodiments may be implemented in the form of a computer program product. A computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions according to the present application are generated in whole or in part. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions may be stored in a computer-readable storage medium or transferred from one computer-readable storage medium to another.

[0098] In the description of this specification, the reference terms "one embodiment," "some embodiments," "example," "specific example," or "some examples" mean that the specific features, structures, materials, or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present application. Moreover, the specific features, structures, materials, or characteristics described may be combined in any appropriate manner in any one or more embodiments or examples. In addition, those skilled in the art may combine and combine different embodiments or examples described in this specification, as well as features of different embodiments or examples, unless they are mutually inconsistent.

[0099] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features being referred to. Thus, a feature defined as "first" or "second" may explicitly or implicitly include at least one such feature. Throughout the description of this application, "plurality" means two or more, unless otherwise specifically defined.

[0100] Any process or method described in the flowchart or otherwise described herein can be understood to represent a module, segment or portion of code comprising one or more executable instructions for implementing the steps of a specific logical function or process. The scope of the preferred embodiments of the present application includes other implementations in which the functions may be performed in a different order than shown or discussed, including performing the functions substantially simultaneously or in reverse order depending on the functions involved.

[0101] The logic and / or steps described in the flowchart or otherwise described herein, for example, can be considered as an ordered list of executable instructions for implementing logical functions, and can be embodied in any computer-readable medium for use by an instruction execution system, apparatus or device (such as a computer-based system, a system including a processor or other system that can fetch instructions from an instruction execution system, apparatus or device and execute instructions), or used in combination with such instruction execution systems, apparatuses or devices.

[0102] It should be understood that various parts of the present application can be implemented using hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented using software or firmware stored in a memory and executed by a suitable instruction execution system. All or part of the steps of the above embodiment method can be completed by instructing the relevant hardware through a program, which can be stored in a computer-readable storage medium. When the program is executed, it includes one or a combination of the steps of the method embodiment.

[0103] In addition, the functional units in the various embodiments of the present application may be integrated into a single processing module, or each unit may exist physically separately, or two or more units may be integrated into a single module. The aforementioned integrated modules may be implemented in the form of hardware or in the form of software functional modules. If the aforementioned integrated modules are implemented in the form of software functional modules and sold or used as independent products, they may also be stored in a computer-readable storage medium. The storage medium may be a read-only memory, a magnetic disk, or an optical disk, etc.

[0104] The above is merely an exemplary embodiment of the present application, but the scope of protection of the present application is not limited thereto. Any person skilled in the art can easily conceive of various modifications or substitutions within the technical scope described in this application, and such modifications or substitutions should be included in the scope of protection of the present application. Therefore, the scope of protection of the present application should be based on the scope of protection of the claims.

Claims

1. A data processing method, characterized in that: Applied to optimize equipment, including: determining similar deduplication operators among deduplication operators of the target workload; Merging the similar deduplication operators into a similar operator set; Generate a corresponding execution plan according to the similarity operator set; the execution plan includes information about the similarity operator set and a multiple deduplication marking operator corresponding to the similarity operator set, the multiple deduplication marking operator is used to perform batch deduplication marking calculations on the data to be processed in the database, the data to be processed includes data corresponding to the similarity operator set; Sending the execution plan to an execution device; Generating a corresponding execution plan according to the similarity operator set includes: The multiple deduplication marking operator is constructed by taking the common data corresponding to the similarity operator set and the validity marking information corresponding to the common data as input information and taking the deduplication marking information corresponding to the common data as output information; the multiple deduplication marking operator is used to synchronously deduplication the common data corresponding to each similarity operator set according to the validity marking information corresponding to each similarity operator set, and obtain the deduplication marking information corresponding to each similarity operator set; For each similar operator set, based on the filtering conditions corresponding to each similar deduplication operator in the similar operator set, determine whether each data in the shared data corresponding to the similar operator set is valid, and mark the validity of each data in the shared data to obtain a set of validity marking information corresponding to the shared data; the execution plan also includes the validity marking information.

2. The data processing method according to claim 1, wherein: Determining similar deduplication operators among the deduplication operators of the target workload includes: Determining whether there is common data among the data corresponding to each deduplication operator in the target workload; In a case where the data corresponding to the multiple deduplication operators are common data, determining the multiple deduplication operators as a group of similar deduplication operators; The data corresponding to the similarity operator set includes the shared data corresponding to each similar deduplication operator in the similarity operator set.

3. The data processing method according to claim 1 or 2, characterized in that: The similarity operator set generates a corresponding execution plan, further comprising: Synchronously redistribute the shared data corresponding to each similar operator set; the execution plan also includes the shared data after data redistribution.

4. A data processing method, characterized in that: Applicable to execution equipment, including: Deduplication marking is performed on the data to be processed in the database in batches based on an execution plan provided by the optimization device, thereby obtaining deduplication marking information; the execution plan is obtained by the data processing method according to any one of claims 1 to 3, and the data to be processed includes data corresponding to each similar operator set in the target workload; Batch deduplication processing is performed on the data to be processed based on the deduplication mark information.

5. The data processing method according to claim 4, wherein: The data corresponding to each similar operator set includes the shared data corresponding to each similar deduplication operator in each similar operator set; The step of batch deduplication marking of the data to be processed in the database based on the execution plan to obtain deduplication marking information includes: The common data and validity marking information corresponding to each similar operator set in the execution plan are input into the multiple deduplication marking operator in the execution plan, and the common data corresponding to each similar operator set are synchronously deduplicated through the multiple deduplication marking operator to obtain the deduplication marking information corresponding to each similar operator set.

6. The data processing method according to claim 5, wherein: The step of synchronously deduplicating the shared data corresponding to each similar operator set by using the multiple deduplication operators includes: For the shared data corresponding to each similarity operator set, the multiple deduplication marking operator determines whether each data in the shared data is valid according to the validity marking information corresponding to the shared data; For invalid data, marking the invalid data as non-duplicate data by the multiple deduplication marking operators as deduplication marking information of the invalid data; For valid data, the multiple deduplication marking operators are used to determine whether the valid data is duplicate data in the current shared data, and the valid data is marked as duplicate data or non-duplicate data according to the determination result as the deduplication marking information of the valid data.

7. The data processing method according to claim 5, characterized in that: The common data input to the multiple deduplication marking operators is the common data after the data in the execution plan is redistributed.

8. A data processing device, characterized in that: Applied to optimize equipment, including: a similarity determination module, configured to determine similar deduplication operators among the deduplication operators of the target workload; An operator merging module, configured to merge the similar deduplication operators into a similar operator set; a plan generation module, configured to generate a corresponding execution plan based on the similarity operator set; the execution plan includes information about the similarity operator set and a multiple deduplication marking operator corresponding to the similarity operator set, the multiple deduplication marking operator being used to perform batch deduplication marking calculations on data to be processed in a database, the data to be processed including data corresponding to the similarity operator set; A plan sending module is used to send the execution plan to the execution device; The plan generation module is further configured to take the common data corresponding to the similar operator sets and the validity marking information corresponding to the common data as input information, and take the deduplication marking information corresponding to the common data as output information, to construct the multiple deduplication marking operator; the multiple deduplication marking operator is configured to synchronously deduplication the common data corresponding to each similar operator set based on the validity marking information corresponding to each similar operator set, thereby obtaining the deduplication marking information corresponding to each similar operator set; For each similar operator set, based on the filtering conditions corresponding to each similar deduplication operator in the similar operator set, determine whether each data in the shared data corresponding to the similar operator set is valid, and mark the validity of each data in the shared data to obtain a set of validity marking information corresponding to the shared data; the execution plan also includes the validity marking information.

9. A data processing device, characterized in that: Applicable to execution equipment, including: a deduplication marking module, configured to perform batch deduplication marking on the data to be processed in the database based on an execution plan to obtain deduplication marking information; the execution plan is obtained by the data processing device according to claim 8, and the data to be processed includes data corresponding to each similarity operator set; A deduplication processing module is used to perform batch deduplication processing on the data to be processed based on the deduplication mark information.

10. An electronic device, characterized in that: include: A memory, a processor, and a computer program stored in the memory, wherein the processor implements the data processing method according to any one of claims 1 to 7 when executing the computer program.

11. A data processing system, characterized in that: include: Optimization equipment and execution equipment for communication connections; At least one of the optimization device and the execution device is the electronic device according to claim 10.

12. A computer-readable storage medium, characterized in that A computer program is stored, and when the computer program is executed by a processor, the data processing method according to any one of claims 1 to 7 is implemented.

Citation Information

Patent Citations

  • Data processing method and device

    CN103870308A

  • Data processing method and device, medium and electronic equipment

    CN114442940A