Data Processing Method, Apparatus, Device, and Storage Medium

By using a column-by-column sorting method in the column-by-column database, the row number and data value of the data table are adjusted, and the data sorting performance of the database is improved.

CN114925067BActive Publication Date: 2025-06-20ALIBABA CLOUD COMPUTING CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202210552150.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-05-20
Publication Date
2025-06-20
Estimated Expiration
2042-05-20

AI Technical Summary

Technical Problem

The prior art consumes a large amount of computing resources and memory resources in the process of data sorting of multiple data tables, reducing the data sorting performance of the database.

Method used

By obtaining the multiple columns of data to be sorted in the column database, adjusting the row number and data value according to the data value of the i-th column data in the multiple column data, obtaining the first sorting result, and adjusting the row number according to the i+1 column data value under the same data value, obtaining the sorting result of the row number, and finally determining the ordered data of the multiple column data based on the sorting result of the row number.

Benefits of technology

The column-by-column sorting method is adopted to reduce the consumption of computing resources and memory resources, improve the data sorting performance of the database, and reduce memory access overhead.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114925067B_ABST
    Figure CN114925067B_ABST
Patent Text Reader

Abstract

The present application discloses a data processing method, apparatus, device, and storage medium. The data processing method adjusts the row numbers of the rows where the i-th column data is located and the data corresponding to the row numbers in the i-th column data by comparing the data values of the i-th column data in multiple columns of data to obtain a first sorting result. Then, in the case where the first sorting result includes the same data values, the row numbers of at least two rows where the same data values are located are adjusted according to the data values of the target data in the (i + 1)-th column corresponding to the row numbers where the same data values are located to obtain a sorting result of the row numbers, and based on the sorting result of the row numbers, the ordered data of the multiple columns of data is determined. According to the data processing method provided by the embodiments of the present application, the traditional row-by-row sorting method is abandoned, and column-by-column sorting is adopted, which can reduce the consumption of computing resources and memory resources, and at the same time can also reduce the memory access overhead and improve the data sorting performance of the database.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer technology, and in particular, to a data processing method, apparatus, device, and storage medium. Background Art

[0002] A database is a warehouse that organizes, stores, and manages data according to a data structure. It usually includes multiple data tables. In practical applications, the data in the data tables is sorted according to actual needs.

[0003] In related technologies, the data in multiple data tables can be sorted according to the row numbers in the data tables. However, due to the large amount of data in multiple data tables, the data sorting process not only consumes a large amount of computing resources and memory resources, but also reduces the sorting performance of accessing the database. Summary of the Invention

[0004] Embodiments of this application provide a data processing method, apparatus, device, and storage medium, which can reduce the consumption of computing resources and memory resources in data sorting and improve the data sorting performance of the database.

[0005] According to the first aspect of the embodiments of this application, a data processing method is provided, including:

[0006] Obtain multiple columns of data to be sorted in a columnar database;

[0007] Adjust the row number of the row where the i-th column data is located and the data corresponding to the row number in the i-th column data according to the data value of the i-th column data in the multiple columns of data, to obtain a first sorting result, where i is an integer greater than or equal to 1;

[0008] In the case where the first sorting result includes the same data values, adjust the row numbers of at least two rows where the same data values are located according to the data value of the target data in the (i + 1)-th column corresponding to the row numbers of the same data values, to obtain a sorting result of the row numbers;

[0009] Based on the sorting result of the row numbers, determine the ordered data of the multiple columns of data.

[0010] According to the second aspect of the embodiments of this application, a data processing apparatus is provided, including:

[0011] An obtaining module, configured to obtain multiple columns of data to be sorted in a columnar database;

[0012] An adjusting module, configured to adjust the row number of the row where the i-th column data is located and the data corresponding to the row number in the i-th column data according to the data value of the i-th column data in the multiple columns of data, to obtain a first sorting result, where i is an integer greater than or equal to 1;

[0013] The adjustment module is further configured to, when the first sorting result includes the same data values, adjust the row numbers where at least two same data values are located according to the data values of the target data in the (i + 1)-th column corresponding to the row numbers where the same data values are located, so as to obtain the sorting result of the row numbers;

[0014] The determination module is configured to determine the ordered data of the multi-column data based on the sorting result of the row numbers.

[0015] According to the third aspect of the embodiments of the present application, there is provided a computer device, including: a memory and a processor;

[0016] The memory is configured to store a computer program;

[0017] The processor is configured to execute the computer program stored in the memory, and when the computer program runs, the processor is caused to execute the steps of the data processing method as shown in the first aspect.

[0018] According to the fourth aspect of the embodiments of the present application, there is provided a computer-readable storage medium, on which a program or an instruction is stored. When the program or the instruction is executed by a computer device, the computer device is caused to execute the steps of the data processing method as shown in the first aspect.

[0019] According to the fifth aspect of the embodiments of the present application, there is provided a computer program product, including a computer program. When the computer program is executed by a computer device, the computer device is caused to execute the steps of the data processing method as shown in the first aspect.

[0020] According to the data processing method, device, equipment, and storage medium in the embodiments of the present application, by comparing the data values of the data in the i-th column in the multi-column data, adjusting the row numbers of the rows where the data in the i-th column is located and the data corresponding to the row numbers in the i-th column data, the first sorting result is obtained. Then, when the first sorting result includes the same data values, according to the data values of the target data in the (i + 1)-th column corresponding to the row numbers where the same data values are located, the row numbers where at least two same data values are located are adjusted to obtain the sorting result of the row numbers, and based on the sorting result of the row numbers, the ordered data of the multi-column data is determined. Thus, the traditional row-by-row sorting method is abandoned, and column-by-column sorting is adopted, which improves the memory access speed in the column database with columnar storage. And, since the memory distance of the data in the same column in the column database is shorter, so adopting column-by-column sorting is beneficial to reducing the consumption of computing resources and memory resources for rearrangement and misarrangement, and at the same time can also reduce the memory access overhead and improve the data sorting performance of the database. Description of the Drawings

[0021] The present application can be better understood from the following description of the specific embodiments in conjunction with the accompanying drawings. Among them, the same or similar reference numerals represent the same or similar features.

[0022] Figure 1 It is a schematic diagram showing a data sorting process in the related art;

[0023] Figure 2 It is a schematic diagram showing another data sorting process in the related art;

[0024] Figure 3 It is a schematic diagram showing a data processing architecture according to an embodiment;

[0025] Figure 4 It is a schematic diagram showing a data processing flow according to an embodiment;

[0026] Figure 5 It is a flowchart showing a data processing method according to an embodiment;

[0027] Figure 6 It is a schematic diagram showing the structure of a data processing device according to an embodiment;

[0028] Figure 7 It is a schematic diagram showing the hardware structure of a computer device according to an embodiment. Detailed implementation manners

[0029] The features and exemplary embodiments of various aspects of the present application will be described in detail below. For the purpose of making the objectives, technical solutions and advantages of the present application clearer and more understandable, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only configured to explain the present application and are not configured to limit the present application. For those skilled in the art, the present application can be implemented without some of these specific details. The following description of the embodiments is only to provide a better understanding of the present application by showing examples of the present application.

[0030] It should be noted that in this article, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or device. Without further limitation, the elements defined by the statement "including..." do not exclude the existence of additional identical elements in the process, method, article or device including the elements.

[0031] Online analytical processing (OLAP) is a data processing method in computer technology that quickly solves multi-dimensional analysis problems. As a result, OLAP databases related to OLAP often need to sort massive amounts of data to improve memory access efficiency. However, the data sorting process is a very computationally resource-intensive process. Although there has been relatively in-depth and comprehensive research on the data sorting algorithm itself, there is currently no specific optimization method for optimizing the sorting algorithm in OLAP databases.

[0032] In related technologies, the data in the database is represented in segments. Usually, a data segment is called a Page. As Figure 1 shown, taking two columns of data as an example, such as the data in column C1 and column C2, the data in column C1 and column C2 are divided into 3 Pages according to RowId. For example, Page1 includes the data in column C1 and column C2 corresponding to RowId 0 and 1, Page2 includes the data in column C1 and column C2 corresponding to RowId 2 and 3, and Page3 includes the data in column C1 and column C2 corresponding to RowId 4. Among them, RowId is used to represent the row number of the overall data in the database. Based on this, the current sorting method is to first sort RowId without swapping the row data corresponding to RowId. Finally, the ordered RowId is materialized into real ordered data. However, during the process of sorting RowId, it is necessary to find the row data corresponding to RowId from multiple data tables in the database and compare these real row data respectively. Due to the excessive amount of data in multiple data tables, each time a comparison is made, each element between the row data needs to be compared one by one, making the RowId sorting process not only consume a large amount of computational resources and memory resources, but also be extremely time-consuming.

[0033] In addition, in order to shorten the sorting time, prefix sorting can be used to optimize the above sorting method. As Figure 2 shown, compared with the previous method, there is an additional process of generating prefixes. Among them, the prefix is a set of row data corresponding to RowId, such as the set of row data "1A" corresponding to RowId 0 and the set of row data "2B" corresponding to RowId 1. Then, the real set of row data in the prefix can be directly compared. Although the prefix will be compressed to different degrees, making the comparison speed of the prefix lower than the speed of comparing each row data one by one, a large amount of computational resources and memory resources will still be consumed during the process of creating and maintaining the prefix, affecting the data sorting performance of the database.

[0034] Based on this, the embodiments of the present application provide a high-performance data processing method applicable to a database using columnar storage. By comparing the data values of the i-th column data in multiple columns of data, the row numbers of the rows where the i-th column data is located and the data corresponding to the row numbers in the i-th column data are adjusted to obtain a first sorting result. When the first sorting result includes the same data values, according to the data values of the target data in the (i + 1)-th column corresponding to the row numbers where the same data values are located, the row numbers where at least two same data values are located are adjusted to obtain a sorting result of the row numbers. Then, based on the sorting result of the row numbers, the ordered data of the multiple columns of data is determined. Thus, the data processing method provided by the embodiments of the present application abandons the traditional row-by-row sorting method and uses column-by-column sorting instead, greatly reducing the memory access overhead and improving the memory access speed in the columnar database. Moreover, since the memory distance of the data in the same column in the columnar database is shorter, using column-by-column sorting is beneficial to reducing the consumption of computing resources and memory resources for rearrangement and misarrangement, and at the same time can also reduce the memory access overhead and the overhead caused by branch prediction failure, improving the data sorting performance of the database.

[0035] Therefore, the data processing architecture provided by the embodiments of the present application will be described in detail below in conjunction with Figure 3 .

[0036] In one or more possible embodiments, as Figure 3 shown, the data processing architecture 10 proposed by the embodiments of the present application may be a columnar database 101, a materialization module 102, and a memory 103. The data processing architecture will be described in detail below respectively.

[0037] The columnar database 101 is a database using columnar storage corresponding to OLAP. The data in the columnar database is stored based on columns as basic logical storage units. The data in one column (simply referred to as column data in the embodiments of the present application) exists in a continuous storage form in the storage medium. The columnar database 101 is used to store multiple columns of data, and row numbers are set for each column of data in the multiple columns of data corresponding to rows. In one example, the data processing architecture 10 may, when obtaining a sorting instruction, obtain the multiple columns of data to be sorted corresponding to the sorting instruction from the columnar database 101 according to the sorting instruction. In another example, the columnar database 101 stores column data of at least two data types, such as column data with null values and non-null values. At this time, the data processing architecture 10 may, when obtaining a sorting instruction, obtain the multiple columns of data corresponding to each data type from the columnar database 101 according to the sorting instruction, so as to process the multiple columns of data of each data type separately.

[0038] The materialization module 102 is used to materialize the column data in the columnar database 101, that is, store the data of the corresponding rows of each column in the columnar database 101 into the continuous memory 103. In one example, the materialization module 102 is used to store the data of the corresponding rows of the i-th column in the columnar database 101 into a continuous memory in the order of the row numbers arranged in the columnar database 101. In another example, the materialization module 102 can also be used to store the column data other than the i-th column data and the target data of the (i + 1)-th column in the multiple column data into the memory according to the sorting result of the row numbers to obtain ordered data.

[0039] The memory 103 is used to store the data output by the materialization module 102. In one example, the memory 103 is used to store the data of the corresponding rows of the i-th column. In another example, the memory 103 is used to store the ordered data after sorting the multiple column data.

[0040] Based on the data processing architecture as Figure 3 shown, the following combines the attached Figure 4 Taking the columnar database 101 as the database in the OLAP scenario and the multiple column data as two column data as an example, the data processing method provided by the embodiments of the present application will be described in detail.

[0041] As Figure 4 shown, when the data processing architecture 10 obtains a sorting instruction, it can obtain two columns of data to be sorted corresponding to the sorting instruction from the columnar database 101. Here, the C1 column data and the C2 column data are divided into 3 Pages according to the RowId, such as the C1 column data (1 and 2, 1 and 2, 3) and the C2 column data (A and B, E and E, F). Each column data corresponding row is set with a row number RowId (0, 1, 2, 3, and 4). Exemplarily, when the RowId is 0, indicating the 0th row, it corresponds to the C1 column data (1) and the C2 column data (A). Similarly, when the RowId is 1, indicating the 1st row, it corresponds to the C1 column data (2) and the C2 column data (B).

[0042] Next, the data processing architecture 10 materializes the first column data in the order of the RowId arrangement to obtain the data (1, 2, 1, 2, and 3) corresponding to the row numbers in the first column data stored in the continuous memory 103. Furthermore, according to the ascending order of the data values of the first column data, such as the C1 column data, the row numbers of the rows where the C1 column data is located and the data corresponding to the row numbers in the first column data are adjusted to obtain the first sorting result. Here, in combination with Figure 4, if the data value of the data in column C1(2) is greater than the data in column C1(1) and less than the data in column C1(3), then swap row number 1 and row number 2, and swap the data in column C1(2) corresponding to row number 1 and the data in column C1(1) corresponding to row number 2, to obtain the adjusted row numbers (0, 2, 1, 3, and 4), and the data in column C1(1, 1, 2, 2, and 3) corresponding to the row numbers.

[0043] Furthermore, the data processing architecture 10 detects whether the first sorting result includes the same data values. In the case where the first sorting result includes the same data values, such as the data values of the data in column C1 corresponding to row number 0 and row number 2 are both 1, and the data values of the data in column C1 corresponding to row number 1 and row number 3 are both 2, from the data in column 2, obtain the target data (A), (E), (B), and (E) of the rows where the same data values are located, materialize the target data (A), (E), (B), and (E), to obtain the target data (A E), (B E) of column 2 stored in the continuous memory 103. Compare the data values of A and E respectively, and compare the data values of B and E. At this time, according to the order of the data values of the target data (AE) of column 2, adjust the order of row number 0 and row number 2. If the data value of A is greater than the data value of E, then swap row number 0 and row number 2. Similarly, if the data value of B is less than the data value of E, then obtain the final row number sorting result (2, 0, 1, 3, and 4).

[0044] Then, based on the row number sorting result 401(2, 0, 1, 3, and 4), the data processing architecture 10 adjusts the data in column 2 to obtain the sorted result data, and materialize the column data (E, A, B, and E) other than the target data of column 2 with RowId being 4, that is, (F), in the sorted result data, to obtain the true ordered data 402.

[0045] In summary, the data processing method provided by the embodiment of the present application abandons the traditional row-by-row sorting method and adopts column-by-column sorting. Since the memory distance of the data in the same column in the column database is shorter, column-by-column sorting is beneficial to improving the memory access speed in the column database with columnar storage. And, the embodiment of the present application adopts a column-by-column materialization process on the basis of column-by-column comparison. This process will make the calculation and memory overhead of the materialization of the subsequent columns lower and lower as the sorting of the column data progresses. Thus, while ensuring the advantages of materialization, this materialization process reduces the overhead of parsing the ordered data according to RowId, and at the same time, it is more memory-intensive and more friendly to memory access, which can reduce the consumption of computing resources and memory resources in data sorting and improve the data sorting performance of the database.

[0046] It should be noted that the data processing method provided in the embodiments of the present application can be applied not only to the data sorting scenario of the columnar database corresponding to OLAP as shown above, but also to any data sorting scenario for columnar databases. The specific scenario to which this data processing method is applied is not limited herein.

[0047] According to the above architecture and application scenarios, the data processing method provided in the embodiments of the present application will be described in detail below in conjunction with Figure 5 respectively.

[0048] Figure 5 FIG. is a flowchart showing a data processing method according to an embodiment.

[0049] As Figure 5 shown, the data processing method can be applied to the data processing architecture as Figure 3 shown, and specifically may include:

[0050] Step 510, obtaining multi-column data to be sorted in a columnar database; Step 520, adjusting the row numbers of the rows where the i-th column data is located and the data corresponding to the row numbers in the i-th column data according to the data values of the i-th column data in the multi-column data to obtain a first sorting result, where i is an integer greater than or equal to 1; Step 530, in the case that the first sorting result includes the same data values, adjusting the row numbers of at least two rows where the same data values are located according to the data values of the target data in the (i + 1)-th column corresponding to the row numbers where the same data values are located to obtain a sorting result of the row numbers; Step 540, determining the ordered data of the multi-column data based on the sorting result of the row numbers.

[0051] The above steps will be described in detail below, as specifically shown below.

[0052] Regarding Step 510, in one or more possible examples, the multi-column data includes column data of at least two data types. Based on this, this Step 510 may specifically include:

[0053] According to at least two data types, respectively obtain the column data corresponding to each data type to process the column data of each data type separately.

[0054] Exemplarily, the multi-column data may include column data of the Null data type and column data of the non-Null data type. Further, the i-th column data may include column data of the Null data type and column data of the non-Null data type. In this way, by respectively obtaining the column data corresponding to each data type, when comparing the data values column by column, the Null values in the i-th column data can be compared first and then the non-Null values in the i-th column data can be processed. Thus, the complex Null judgment logic for comparing row by row is avoided, and the overhead of branch prediction failure is reduced.

[0055] Regarding step 520, in one or more possible examples, this step 520 may specifically include steps 5201 to 5203.

[0056] Step 5201: Materialize the data in the i-th column in the order of row number arrangement to obtain the data corresponding to the row number in the data of the i-th column stored in memory.

[0057] Furthermore, in the case where multiple columns of data are divided into at least two data segments according to row numbers, the data in the i-th column of each of the at least two data segments may be materialized in the ascending order of row numbers to obtain the data corresponding to the row number in the data of the i-th column stored in continuous memory.

[0058] Exemplarily, as Figure 3 shown, since the data in each column of data exists in segments, it is possible to materialize the data in the i-th column of each data segment, and the materialization process may be to store the data corresponding to the row number in the i-th column of data into a continuous block of memory.

[0059] Step 5202: Arrange the data values of the data corresponding to the row number in the i-th column in ascending order to obtain a first order.

[0060] Step 5203: Sort the row numbers of the rows where the i-th column of data is located and the data corresponding to the row number in the i-th column according to the first order to obtain a first sorting result.

[0061] Exemplarily, sort the data values of the data corresponding to the row number in the materialized i-th column of data. During the sorting process, both the RowId and the specific data corresponding to the row number in the i-th column are exchanged.

[0062] Since the data processing method provided in the embodiments of the present application exchanges both the RowId and the data values of the data corresponding to the row number in the i-th column, as the subsequent column-by-column exchange progresses, the data will become more and more concentrated, and the memory access will be more and more friendly. Regarding step 530, in one or more possible examples, after obtaining the first sorting result, there may be a possibility that some data values are the same, resulting in the inability to obtain the order of the row numbers where the same data values are located. Therefore, it is necessary to continue to compare the data values of the target data in the (i + 1)-th column of the row numbers where the same data values are located.

[0063] Based on this, when the first sorting result includes the same data values, it is necessary to obtain the data in the (i + 1)-th column corresponding to the row numbers where the same data values are located in order to adjust the row numbers where at least two same data values are located and obtain the sorted result of the row numbers. Based on this, before the step 530 of adjusting the row numbers where at least two same data values are located according to the data values of the target data in the (i + 1)-th column corresponding to the row numbers where the same data values are located to obtain the sorted result of the row numbers, the data processing method may further include step 550 and step 560.

[0064] Step 550: Obtain the target data corresponding to the row numbers where the same data values are located from the data in the (i + 1)-th column of the multi-column data;

[0065] Step 560: Materialize the target data to obtain the target data in the (i + 1)-th column stored in continuous memory.

[0066] Here, the materialization process provided in the embodiments of the present application adopts a column-by-column materialization process based on column-by-column comparison. As the sorting of column data progresses, the calculation and memory overhead of the materialization of the subsequent columns will become lower and lower. Thus, while ensuring the advantages of materialization, this materialization process reduces the overhead of parsing ordered data according to RowId, and at the same time, the memory is more aggregated, the memory access is more friendly, which can reduce the consumption of computing resources and memory resources in data sorting and improve the data sorting performance of the database.

[0067] Based on the above-provided materialization process, the embodiments of the present application further provide a way to trigger the materialization process based on the materialization granularity. Based on this, in another possible embodiment, before step 550, the data processing method may further include:

[0068] Determine whether the number of the same data values meets the materialization condition;

[0069] When it is determined that the number of the same data values meets the materialization condition, obtain the target data corresponding to the row numbers where the same data values are located from the data in the (i + 1)-th column of the multi-column data.

[0070] Exemplarily, the materialization granularity for triggering the materialization of the data in the (i + 1)-th column can be set, that is, determine whether the number of the same data values is equal to a preset materialization threshold such as 2. When the number of the same data values is 2, trigger the materialization of the target data corresponding to the row numbers where the same data values are located in the (i + 1)-th column of data.

[0071] On the contrary, when it is determined that the number of the same data values does not meet the materialization condition, obtain the data in the (i + 1)-th column;

[0072] Based on the data values of the (i + 1)-th column data, adjust the row numbers of the rows where the (i + 1)-th column data is located and the data in the target column data corresponding to the row numbers, to obtain a second sorting result, where the target column data includes the i-th column data and the (i + 1)-th column data;

[0073] In the case where the second sorting result includes the same data values and the number of the same data values included in the second sorting result satisfies the materialization condition, based on the data values of the (i = i + 1)-th column data of the row numbers where the same data values in the second sorting result are located, adjust the row numbers of at least two rows where the same data values are located, to obtain a sorting result of the row numbers.

[0074] Exemplarily, still referring to the above example, in the case where it is determined that the number of the same data values is greater than a preset materialization threshold such as 2, still use the adjustment method for the i-th column data, that is, based on the data values of the (i + 1)-th column data, adjust the row numbers of the rows where the (i + 1)-th column data is located and the data in the i-th column data corresponding to the row numbers, and let i = i + 1, and adjust multiple columns of data in sequence according to the above method until it is determined that the number of the same data values is equal to the preset materialization threshold such as 2, and trigger the materialization of the target data of the row numbers where the same data values in the (i = i + 1)-th column data are located.

[0075] In addition, the materialization granularity involved in the materialization condition, that is, the preset materialization threshold, can also be adjusted according to actual needs. In this way, the materialization condition in the embodiments of the present application is obtained through at least one of the following methods: determined by the data types in the columnar database, preset by the user.

[0076] Therefore, the embodiments of the present application can provide a method for dynamically adjusting the materialization granularity, which can reduce the overhead of parsing ordered data according to the RowId while ensuring the materialization advantages, and at the same time is more aggregated in memory, more friendly to memory access, can reduce the consumption of computing resources and memory resources in data sorting, and improve the data sorting performance efficiency of the database.

[0077] Based on this, step 530 may specifically include:

[0078] Arrange the data values of the (i + 1)-th column target data in ascending order to obtain a second order;

[0079] According to the second order, sort the row numbers of at least two rows where the same data values are located to obtain a sorting result of the row numbers.

[0080] Regarding step 540, in one or more possible examples, step 540 may specifically include:

[0081] Based on the sorting result of the row numbers, adjust the column data other than the i-th column data in multiple columns of data to obtain sorted result data;

[0082] Materialize the target objects in the sorted result data to obtain the sorted data stored in memory;

[0083] Among them, the target objects include the column data other than the i-th column data and the target data of the (i + 1)-th column in the multi-column data.

[0084] Exemplarily, as Figure 4 shown, since the data in the first column is 1 when RowId is 0 and 2, after adjusting the order of the corresponding data A and E in the same row number in the second column data, since the data in the first column is 1, there is no need to adjust the first column data, and only the order of data A and E in the second column needs to be adjusted. Of course, if the data values of data A and E in the second column still cannot determine their existing order, continue to compare the data corresponding to RowId 0 and 2 in the third column data. In addition, the principle of materialized objects is the same, that is, since the data of the row numbers with the same data value has been materialized in step 560, after obtaining the sorting result of the row numbers, only the column data other than the i-th column data and the target data of the (i + 1)-th column needs to be materialized. In this way, it is beneficial to reduce the consumption of computing resources and memory resources for rearrangement and misarrangement, and at the same time, it can also reduce the memory access overhead and improve the data sorting performance of the database.

[0085] In summary, by comparing the data values of the i-th column data in the multi-column data, adjusting the row numbers of the rows where the i-th column data is located and the data corresponding to the row numbers in the i-th column data, a first sorting result is obtained. Then, in the case where the first sorting result includes the same data values, according to the data values of the target data of the (i + 1)-th column in the row numbers where the same data values are located, adjust the row numbers of at least two rows where the same data values are located to obtain the sorting result of the row numbers, and based on the sorting result of the row numbers, determine the sorted data of the multi-column data. Thus, the traditional row-by-row sorting method is abandoned, and column-by-column sorting is adopted to improve the memory access speed in the column database with columnar storage. And, since the memory distance of the same column data in the column database is shorter, so column-by-column sorting

[0086] is beneficial to reduce the consumption of computing resources and memory resources for rearrangement and misarrangement, and at the same time, it can also reduce the memory access overhead and improve the data sorting performance of the database.

[0087] In addition, the embodiment of the present application adopts a column-by-column materialization process on the basis of column-by-column comparison. As the sorting of column data progresses, the calculation and memory overhead of the materialization of subsequent columns will become lower and lower. Therefore, while ensuring the advantages of materialization, this materialization process reduces the overhead of parsing ordered data according to RowId. At the same time, the memory is more aggregated, the memory access is more friendly, and it can further reduce the consumption of computing resources and memory resources in data sorting, improving the data sorting performance of the database. It should be clear that the present application is not limited to the specific configurations and processes described in the above embodiments and shown in the figures. For the convenience and conciseness of description, the detailed description of known methods is omitted here, and the specific working processes of the systems, modules, and units described above can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated here.

[0088] Based on the same inventive concept, the embodiment of the present application provides a data processing device corresponding to the above-mentioned data processing method. Specifically, it will be described in detail in combination with Figure 6 this.

[0089] Figure 6 FIG. is a schematic structural diagram of a data processing device according to an embodiment.

[0090] As Figure 6 shown, the data processing device 60 is applied to the data processing architecture as shown in Figure 3 shown. The data processing device 60 may specifically include:

[0091] An acquisition module 601, configured to acquire multiple columns of data to be sorted in a columnar database;

[0092] An adjustment module 602, configured to adjust the row number of the row where the i-th column data is located and the data corresponding to the row number in the i-th column data according to the data value of the i-th column data in the multiple columns of data, to obtain a first sorting result, where i is an integer greater than or equal to 1;

[0093] The adjustment module 602 is further configured to, in the case where the first sorting result includes the same data values, adjust the row numbers of at least two rows where the same data values are located according to the data value of the target data in the (i + 1)-th column corresponding to the row numbers of the same data values, to obtain a sorting result of the row numbers;

[0094] A determination module 603, configured to determine the ordered data of the multiple columns of data based on the sorting result of the row numbers.

[0095] Based on this, the data processing device 60 provided by the embodiment of the present application will be described in detail below.

[0096] In one or more possible embodiments, the data processing device 60 provided by the embodiment of the present application may further include a first materialization module and a first sorting module; wherein,

[0097] The first materialization module is used to materialize the data in the i-th column according to the arrangement order of the line numbers, so as to obtain the data corresponding to the line numbers in the data of the i-th column stored in the memory;

[0098] The adjustment module 602 can also be used to arrange the data values of the data corresponding to the line numbers in the i-th column in ascending order to obtain the first order;

[0099] The first sorting module is used to sort the line numbers of the rows where the i-th column data is located and the data corresponding to the line numbers in the i-th column according to the first order to obtain the first sorting result.

[0100] In another or multiple possible embodiments, the first materialization module can specifically be used to, when the multi-column data is divided into at least two data segments according to the line numbers, materialize the data in the i-th column of each of the at least two data segments in ascending order of the line numbers, so as to obtain the data corresponding to the line numbers in the data of the i-th column stored in the continuous memory.

[0101] In yet another or multiple possible embodiments, the data processing device 60 provided in the embodiments of the present application may further include a second materialization module; wherein,

[0102] The acquisition module 601 can also be used to acquire the target data of the line numbers where the same data values are located from the data in the (i + 1)-th column of the multi-column data;

[0103] The second materialization module is used to materialize the target data to obtain the target data in the (i + 1)-th column stored in the continuous memory.

[0104] In still another or multiple possible embodiments, the data processing device 60 provided in the embodiments of the present application may further include a third materialization module; wherein,

[0105] The determination module 603 is further used to determine whether the number of the same data values meets the materialization condition;

[0106] The third materialization module is used to, when it is determined that the number of the same data values meets the materialization condition, acquire the target data of the line numbers where the same data values are located from the data in the (i + 1)-th column of the multi-column data.

[0107] In still another or multiple possible embodiments, the acquisition module 601 can also be used to acquire the data in the (i + 1)-th column when it is determined that the number of the same data values does not meet the materialization condition;

[0108] The adjustment module 602 is further configured to adjust the row number of the row where the (i + 1)-th column data is located and the data corresponding to the row number in the target column data based on the data value of the (i + 1)-th column data, so as to obtain a second sorting result, where the target column data includes the i-th column data and the (i + 1)-th column data;

[0109] The adjustment module 602 is further configured to, when the second sorting result includes the same data values and the number of the same data values included in the second sorting result meets the materialization condition, adjust the row numbers of at least two same data values according to the data value of the (i = i + 1)-th column data of the row numbers where the same data values included in the second sorting result are located, so as to obtain a sorting result of the row numbers.

[0110] It should be noted that the materialization condition in the embodiments of the present application is obtained through at least one of the following methods: determined by the data types in the columnar database, preset by the user.

[0111] In yet another or multiple possible embodiments, the data processing device 60 provided in the embodiments of the present application may further include a second sorting module; wherein,

[0112] The second sorting module is configured to arrange the data values of the (i + 1)-th column target data in ascending order to obtain a second order;

[0113] The second sorting module is further configured to sort the row numbers of at least two same data values according to the second order to obtain a sorting result of the row numbers.

[0114] In yet another or multiple possible embodiments, the data processing device 60 provided in the embodiments of the present application may further include a fourth materialization module; wherein,

[0115] The adjustment module 602 may further be configured to adjust the column data other than the i-th column data in the multiple column data based on the sorting result of the row numbers to obtain sorted result data;

[0116] The fourth materialization module is configured to materialize the target objects in the sorted result data to obtain ordered data stored in the memory;

[0117] Wherein, the target objects include the column data other than the i-th column data and the (i + 1)-th column target data in the multiple column data.

[0118] In yet another or multiple possible embodiments, the obtaining module 601 may specifically be configured to, when the multiple column data includes column data of at least two data types, respectively obtain the column data corresponding to each data type according to the at least two data types, so as to process the column data of each data type separately.

[0119] The data processing method provided by the embodiments of the present application abandons the traditional row-by-row sorting method and adopts column-by-column sorting. Since the memory distance of the same column data in the column database is shorter, column-by-column sorting is beneficial to improving the memory access speed in the column database with columnar storage. In addition, the embodiments of the present application adopt a column-by-column materialization process on the basis of column-by-column comparison. As the column data sorting progresses, the calculation and memory overhead of the materialization of the subsequent columns will become lower and lower. Therefore, while ensuring the materialization advantages, this materialization process reduces the overhead of parsing ordered data according to the RowId, and at the same time, it is more aggregated in memory, more friendly to memory access, can reduce the consumption of computing resources and memory resources in data sorting, and improve the data sorting performance of the database.

[0120] Figure 7 It is a schematic diagram of the hardware structure of a computer device according to an embodiment.

[0121] As Figure 7 shown, the computer device 700 includes an input device 701, an input interface 702, a processor 703, a memory 704, an output interface 705, and an output device 706.

[0122] The input interface 702, the processor 703, the memory 704, and the output interface 705 are connected to each other through a bus 710. The input device 701 and the output device 706 are respectively connected to the bus 710 through the input interface 702 and the output interface 705, and then connected to other components of the computer device 700. Specifically, the input device 701 receives external input information and transmits the input information to the processor 703 through the input interface 702; the processor 703 processes the input information based on the computer-executable instructions stored in the memory 704 to generate output information, temporarily or permanently stores the output information in the memory 704, and then transmits the output information to the output device 706 through the output interface 705; the output device 706 outputs the output information to the outside of the computer device 700 for user use.

[0123] In one embodiment, Figure 7 the computer device 700 shown can be implemented as a data processing device, which may include: a memory configured to store a program; a processor configured to run the program stored in the memory to execute the data processing method described in the above embodiments.

[0124] In one embodiment, the memory can also be used to store multiple columns of data, a first sorting result, a second sorting result, and the calculation results of each step in the data processing process described above Figures 3 to 5 combined.

[0125] According to an embodiment of the present application, the process described above with reference to the flowchart can be implemented as a computer-readable storage medium. For example, an embodiment of the present application includes a computer-readable storage medium that stores a program or instructions thereon. When the program or instructions are executed by a computer device, the computer device is caused to execute the steps of the above method.

[0126] According to an embodiment of the present application, the process described above with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of the present application includes a computer program product that tangibly includes a computer program on a machine-readable medium. The computer program includes program code for performing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from a network, and / or installed from a removable storage medium.

[0127] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When it runs on a computer, the computer is caused to execute the methods described in the above various embodiments. When loading and executing the computer program instructions on a computer, the processes or functions according to the embodiments of the present application are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center in a wired manner (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or wirelessly (such as infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that can be accessed by a computer, or a data storage device such as a server or data center that includes one or more integrated available media. The available medium can be a magnetic medium (such as a floppy disk, hard disk, magnetic tape), an optical medium (such as a DVD), or a semiconductor medium (such as a solid-state drive), etc.

[0128] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place, or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. A person of ordinary skill in the art can understand and implement it without creative effort.

[0129] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present application.

Claims

1. A data processing method, comprising: Obtain multi-column data to be sorted in a columnar database; Materialize the data in the i-th column among the multi-column data according to the row number arrangement order, to obtain the data corresponding to the row number in the data of the i-th column stored in continuous memory, where i is an integer greater than or equal to 1; Arrange the data values of the data corresponding to the row number in the i-th column data in ascending order to obtain a first order; Sort the row numbers of the rows where the i-th column data is located and the data corresponding to the row number in the i-th column data according to the first order to obtain a first sorting result In the case where the first sorting result includes the same data values and the number of the same data values satisfies the materialization condition, Obtain the target data of the row numbers where the same data values are located from the (i + 1)-th column data of the multi-column data; Materialize the target data to obtain the (i + 1)-th column target data stored in the continuous memory; Arrange the data values of the (i + 1)-th column target data in ascending order to obtain a second order; Sort the row numbers where the at least two same data values are located according to the second order to obtain a sorting result of the row numbers; Based on the sorting result of the row numbers, adjust the column data other than the i-th column data in the multi-column data to obtain sorted result data; Materialize the target object in the sorted result data to obtain ordered data stored in memory; Wherein, the target object includes the column data other than the i-th column data and the (i + 1)-th column target data in the multi-column data.

2. The method according to claim 1, wherein, The multi-column data is divided into at least two data segments according to the row numbers; the step of materializing the data in the i-th column among the multi-column data according to the row number arrangement order to obtain the data corresponding to the row number in the data of the i-th column stored in continuous memory includes: Materialize the data in the i-th column of each of the at least two data segments in ascending order of the row numbers to obtain the data corresponding to the row number in the data of the i-th column stored in continuous memory.

3. The method according to claim 1, the method further comprising: Determine whether the number of the same data values satisfies the materialization condition.

4. The method according to claim 1 or 3, wherein, The method further includes: In the case where it is determined that the number of the same data values does not satisfy the materialization condition, obtain the (i + 1)-th column data; Based on the data values of the (i + 1)-th column data, adjust the row numbers of the rows where the (i + 1)-th column data is located and the data corresponding to the row number in the target column data, to obtain a second sorting result, where the target column data includes the i-th column data and the (i + 1)-th column data; In the case where the second sorting result includes the same data values and the number of the same data values in the second sorting result satisfies the materialization condition, according to the data values of the (i = i + 1)-th column data of the row numbers where the same data values included in the second sorting result are located, adjust the row numbers where the at least two same data values are located to obtain a sorting result of the row numbers.

5. The method according to claim 1, wherein, The materialization condition is obtained by at least one of the following methods: Determined by the data types in the columnar database, preset by the user.

6. The method according to claim 1, wherein, The multi-column data includes column data of at least two data types; The obtaining of multi-column data to be sorted in the columnar database includes: According to the at least two data types, respectively obtain the column data corresponding to each data type, so as to process the column data of each data type respectively.

7. A data processing apparatus, comprising: An obtaining module, configured to obtain multi-column data to be sorted in the columnar database; An adjustment module, configured to materialize the data in the i-th column of the multi-column data according to the row number arrangement order, so as to obtain the data corresponding to the row number in the data of the i-th column stored in continuous memory, where i is an integer greater than or equal to 1; Arrange the data values of the data corresponding to the row number in the i-th column data in ascending order to obtain a first order; According to the first order, sort the row numbers of the rows where the i-th column data is located and the data corresponding to the row number in the i-th column data to obtain a first sorting result; The adjustment module is further configured to, when the first sorting result includes the same data values and the number of the same data values meets the materialization condition, Obtain the target data of the row numbers where the same data values are located from the data in the (i + 1)-th column of the multi-column data; Materialize the target data to obtain the target data in the (i + 1)-th column stored in the continuous memory; Arrange the data values of the target data in the (i + 1)-th column in ascending order to obtain a second order; According to the second order, sort the row numbers where the at least two same data values are located to obtain a sorting result of the row numbers; A determination module, configured to adjust the column data other than the data in the i-th column in the multi-column data based on the sorting result of the row numbers to obtain sorted result data; Materialize the target object in the sorted result data to obtain ordered data stored in memory; where the target object includes the column data other than the data in the i-th column and the target data in the (i + 1)-th column in the multi-column data.

8. A computer device, comprising: A memory and a processor, The memory is used to store a computer program; The processor is configured to execute the computer program stored in the memory, and when the computer program runs, the processor executes the steps of the data processing method according to any one of claims 1 to 6.

9. A computer-readable storage medium, having a program or instructions stored thereon, and when the program or instructions are executed by a computer device, causing the computer device to perform the steps of the data processing method according to any one of claims 1 to 6.

10. A computer program product, comprising a computer program, and when the computer program is executed by a computer device, causing the computer device to perform the steps of the data processing method according to any one of claims 1 to 6.

Citation Information

Patent Citations

  • Physic-chemical method and device of column storage database

    CN106354829A