A data sorting method, storage medium and device

Through the chain-heap sorting method, the heap structure and the data structure of the linked list node are used to judge the element addition conditions based on the first sorting key, avoiding invalid operations, and solving the problem of wasted CPU execution time in the existing top-N heap sorting algorithm, improving the sorting efficiency and accuracy.

CN119782289BActive Publication Date: 2025-07-18CETC JINCANG (BEIJING) TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510292383.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-12
Publication Date
2025-07-18
Estimated Expiration
2045-03-12

AI Technical Summary

Technical Problem

The existing top-N heap sorting algorithm has the problem of wasting CPU execution time due to the large amount of data, which reduces the overall sorting efficiency, especially in meaningless comparison and repeated replacement operations for data that are not the first sorting key.

Method used

Using the chain-heap sorting method, the data structure including the heap structure and linked list nodes is pre-constructed, and whether the conditions for joining the linked list are met only based on the first sorting key of the element to be sorted are determined, and the elements are temporarily stored in the linked list when the conditions are met, avoiding invalid node replacement operations.

Benefits of technology

It improves the overall sorting execution efficiency, reduces the replacement and comparison operations in the heap sorting process, and improves the performance and sorting accuracy of the database system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119782289B_ABST
    Figure CN119782289B_ABST
Patent Text Reader

Abstract

The present invention relates to database technology, and particularly to a data sorting method, a storage medium and a device. The data sorting method includes: obtaining elements to be sorted and a pre-constructed data structure, wherein the data structure includes a heap structure and a linked list with heap nodes in the heap structure as linked list nodes; based on the first sorting key of the elements to be sorted, determining whether the elements to be sorted meet the condition for joining the linked list; if so, adding the elements to be sorted to the linked list. The data sorting method of the present invention can, based on the first sorting key of the elements to be sorted, temporarily store the elements to be sorted that meet the conditions in a linked list with heap nodes in the heap structure as linked list nodes, avoiding the problem of low execution efficiency caused by repeatedly replacing the top node of the heap during the heap sorting process, reducing the node replacement operation during the heap sorting process, thereby improving the overall sorting execution efficiency, and further enhancing the performance of the database system.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to database technology, and in particular, to a data sorting method, a storage medium, and a device. Background Art

[0002] When the database executes the order by + limit statement, the top-N heap sorting algorithm in the sort node will be triggered when the set memory size is sufficient, and finally the first N qualified tuples are obtained from the dataset to be processed. The general operation of the sorting node in the database executor is as follows: when it is determined that top-N sorting is required, in order to ensure the correct comparison order of the tuples for building a max heap (or min heap), the overall sorting direction is first reversed; each time a tuple data is read from the lower node (if it is a column store database, a row of data will be read from a batch of data); initially, the new tuples are sequentially placed at the end of the heap and then rotated to build a heap of size N; once the number of tuples in the heap reaches N, subsequently, the sorting keys of the newly introduced tuple and the top tuple of the heap are compared in sequence to determine whether the newly introduced tuple can replace the top tuple of the heap; when it cannot be replaced, the newly introduced tuple is directly released; when replacement is required, the top tuple of the heap is released, the new tuple is placed at the top of the heap, and then the heap is reorganized; after all subsequent data is read, the top tuple of the final heap memory is sequentially exchanged with the last tuple, the heap length is reduced, and the heap is reorganized until the heap length is zero to complete the sorting; the sorted tuples are sequentially returned to the upper node to finally obtain the execution result of the order by + limit statement. In the general algorithm, once a comparison is required for a certain sorting key, the program will read the actual value from the memory for comparison (if it is a complex sorting key, the actual value will be searched and copied from the memory, and this process is more time-consuming).

[0003] However, in the process of sequentially comparing the sorting keys of the newly introduced tuple and the top tuple of the heap, in some cases, the comparison operation is meaningless. For example, the data that needs to be placed in the heap is judged by comparing non-first sorting keys in the early stage and is released by judging the first sorting key in the later stage. Therefore, the existing top-N heap sorting algorithm has unnecessary extraction and comparison operations for non-first sorting key data, as well as repeated replacement operations for the data in the heap and other invalid operations, and there is a problem of waste of CPU execution time due to excessive execution of invalid operations in the case of a large amount of data, and further there is a problem of reducing the execution efficiency of the overall sorting process. Summary of the Invention

[0004] In view of the above problems, a data sorting method, a storage medium, and a device are proposed to overcome the above problems or at least partially solve the above problems.

[0005] An object of the present invention is to provide a data sorting method to reduce the node replacement operation during the heap sorting process, thereby improving the overall sorting execution efficiency.

[0006] A further object of the present invention is to optimize the data structure and improve the sorting accuracy.

[0007] Another further object of the present invention is to reduce the comparison operation during the heap sorting process to further improve the overall sorting execution efficiency.

[0008] In particular, the present invention provides a data sorting method, including:

[0009] Obtaining the elements to be sorted and a pre-constructed data structure, where the data structure includes a heap structure and a linked list with the heap nodes in the heap structure as the linked list nodes;

[0010] Based on the first sorting key of the elements to be sorted, determining whether the elements to be sorted meet the condition for joining the linked list;

[0011] If so, adding the elements to be sorted to the linked list.

[0012] Furthermore, different heap nodes in the heap structure are used to store elements with different first sorting keys, and different linked list nodes in each linked list are used to store elements with the same first sorting key; and

[0013] The step of determining whether the elements to be sorted meet the condition for joining the linked list based on the first sorting key of the elements to be sorted includes:

[0014] Determining whether the first sorting key of the elements to be sorted is the same as the first sorting key of the elements stored in any heap node in the heap structure;

[0015] If the same, confirming that the elements to be sorted meet the condition for joining the linked list;

[0016] If not the same, confirming that the elements to be sorted do not meet the condition for joining the linked list.

[0017] Furthermore, each node in the data structure is used to store one element; and

[0018] After the step of adding the elements to be sorted to the linked list, the data sorting method further includes:

[0019] Taking the number of elements stored in the nodes other than the top heap linked list in the data structure as the number of non-top heap elements, where the top heap linked list is the linked list with the top heap node of the heap structure as the linked list node;

[0020] Determining whether the number of non-top heap elements reaches a set boundary value, where the boundary value is the target acquisition quantity for obtaining target elements from the dataset to be processed;

[0021] If so, release all elements of the top-linked list of the heap.

[0022] Furthermore, the step of obtaining the elements to be sorted includes:

[0023] Determine whether there are elements to be processed, where the elements to be processed are the elements in the dataset to be processed that have not been added to the data structure and have not been released;

[0024] If there are, obtain the elements to be processed;

[0025] Determine whether the number of elements stored in the data structure reaches the boundary value;

[0026] If not, use the obtained elements to be processed as the elements to be sorted;

[0027] If it reaches, based on the first sorting key of the element to be processed, determine whether the element to be processed meets the condition for adding to the data structure;

[0028] If it meets, use the obtained elements to be processed as the elements to be sorted.

[0029] Furthermore, after the step of determining whether there are elements to be processed, the data sorting method further includes:

[0030] In the case where there are no elements to be processed, sort the elements stored in the data structure based on the non-first sorting keys of the elements stored in the data structure until the number of elements stored in the data structure is the same as the boundary value, so as to obtain all target elements.

[0031] Furthermore, the step of determining whether the element to be processed meets the condition for adding to the data structure based on the first sorting key of the element to be processed includes:

[0032] In the case where the type of the heap structure is a min heap, determine whether the first sorting key of the element to be processed is not less than the first sorting key of the element stored in the top node of the heap. If not less than, confirm that the element to be processed meets the condition for adding to the data structure. If less than, confirm that the element to be processed does not meet the condition for adding to the data structure; or

[0033] In the case where the type of the heap structure is a max heap, determine whether the first sorting key of the element to be processed is not greater than the first sorting key of the element stored in the top node of the heap. If not greater than, confirm that the element to be processed meets the condition for adding to the data structure. If greater than, confirm that the element to be processed does not meet the condition for adding to the data structure.

[0034] Furthermore, after the step of determining whether the element to be processed meets the condition for adding to the data structure based on the first sorting key of the element to be processed, the data sorting method further includes:

[0035] Release the element to be processed in case the element to be processed does not meet the condition for adding to the data structure.

[0036] Furthermore, the data structure is constructed through the following steps:

[0037] Construct an index array structure based on the first sorting key of all elements to be processed in the dataset to be processed. The index array structure includes multiple arrays, and the multiple arrays correspond one-to-one with all elements to be processed. Each array includes the first sorting key of the corresponding element to be processed, a row index, and a next-hop index. The row index is used to represent the memory location of the element to be processed, and the next-hop index is used to point to the position of another element to be processed with the same first sorting key as the first sorting key of the element to be processed in the index array structure to simulate a linked list;

[0038] Set a head pointer and a tail pointer respectively pointing to the heads of the multiple arrays in the index array structure. The head pointer is used to represent that the element to be processed corresponding to the array it points to is the top of the heap structure of the heap structure, and the tail pointer is used to represent that the array it points to is the tail of the heap structure to simulate the heap structure.

[0039] According to another aspect of the present invention, there is also provided a machine-readable storage medium, on which a machine-executable program is stored. When the machine-executable program is executed by a processor, the steps of any one of the above data sorting methods are implemented.

[0040] According to still another aspect of the present invention, there is also provided a computer device, including a memory, a processor, and a machine-executable program stored on the memory and running on the processor. When the processor executes the machine-executable program, the steps of any one of the above data sorting methods are implemented.

[0041] The data sorting method of the present invention optimizes the data storage structure for heap sorting by pre-constructing a data structure including a heap structure and a linked list with the heap nodes in the heap structure as linked list nodes. After obtaining the elements to be sorted and the pre-constructed data structure, based on the first sorting key of the elements to be sorted, it is determined whether the elements to be sorted meet the condition for adding to the linked list. In case the elements to be sorted meet the condition for adding to the linked list, the elements to be sorted are added to the linked list, avoiding directly replacing the nodes in the heap structure when the elements to be sorted meet the condition for adding to the data structure, and reducing the node replacement operation in the heap sorting process. Thus, the data sorting method of the present invention can temporarily store the elements to be sorted that meet the conditions in a linked list with the heap nodes in the heap structure as linked list nodes based on the first sorting key of the elements to be sorted, avoiding the problem of low execution efficiency caused by repeatedly replacing the top node of the heap during the heap sorting process, reducing the replacement operation in the heap sorting process, thereby improving the overall sorting execution efficiency and further enhancing the performance of the database system.

[0042] Furthermore, in the data sorting method of the present invention, by using different heap nodes in the heap structure to store elements with different first sorting keys and different linked list nodes in each linked list to store elements with the same first sorting key, the data structure is optimized. In addition, the data sorting method of the present invention also determines whether the first sorting key of the element to be sorted is the same as the first sorting key of the element stored in any heap node in the heap structure, and only when the first sorting key is the same as the first sorting key of the element stored in any heap node in the heap structure, it is confirmed that the element to be sorted meets the condition for joining the linked list, realizing the judgment of whether the element to be sorted meets the condition for joining the linked list based on the first sorting key of the element to be sorted, ensuring that the first sorting keys of the elements stored in each linked list are the same, thereby improving the sorting accuracy.

[0043] Even further, in the data sorting method of the present invention, when the number of non-heap top elements reaches the set boundary value, all elements of the heap top linked list are released, avoiding the comparison operation of the non-first sorting keys of all elements of the heap top linked list, realizing the reduction of waste caused by invalid operations. At the same time, the storage amount of elements in the heap structure is effectively limited, thereby further improving the overall sorting execution efficiency and then enhancing the performance of the database system.

[0044] From the following detailed description of the specific embodiments of the present invention in conjunction with the accompanying drawings, those skilled in the art will become more clear about the above and other objects, advantages and features of the present invention. Description of the Drawings

[0045] Some specific embodiments of the present invention will be described in detail hereinafter with reference to the accompanying drawings in an exemplary but not restrictive manner. The same reference numerals in the drawings denote the same or similar components or parts. Those skilled in the art should understand that these drawings are not necessarily drawn to scale. In the drawings:

[0046] Figure 1 is a schematic diagram of an execution scheme of a top-N heap sorting algorithm in the prior art;

[0047] Figure 2 is a schematic diagram of the data sorting method according to an embodiment of the present invention;

[0048] Figure 3 is a schematic diagram of an application scenario of the data sorting method according to an embodiment of the present invention;

[0049] Figure 4 is a schematic diagram of the data structure in the initial state in the data sorting method according to an embodiment of the present invention;

[0050] Figure 5Schematic diagram of the data structure in the data sorting method according to an embodiment of the present invention in a state where the N = bound condition is not satisfied;

[0051] Figure 6 Schematic diagram of the data structure in the data sorting method according to an embodiment of the present invention in a state where the N = bound condition is satisfied;

[0052] Figure 7 Schematic diagram of the data structure in the data sorting method according to an embodiment of the present invention after releasing the top - heap element;

[0053] Figure 8 Schematic diagram of the data structure in the data sorting method according to an embodiment of the present invention after adding the element to be sorted to the top - heap linked list;

[0054] Figure 9 Schematic diagram of the data structure in the data sorting method according to an embodiment of the present invention after releasing the element to be processed;

[0055] Figure 10 Schematic diagram of the data structure in the data sorting method according to an embodiment of the present invention after releasing the top - heap element;

[0056] Figure 11 Schematic diagram of the data structure in the data sorting method according to an embodiment of the present invention in a completed construction state;

[0057] Figure 12 Schematic diagram of the arrow operation in the data sorting method according to an embodiment of the present invention;

[0058] Figure 13 Flowchart of the data sorting method according to an embodiment of the present invention;

[0059] Figure 14 Control flowchart of the data sorting method according to an embodiment of the present invention;

[0060] Figure 15 Schematic diagram of the structure of a machine - readable storage medium according to an embodiment of the present invention; and

[0061] Figure 16 Schematic diagram of the structure of a computer device according to an embodiment of the present invention. Detailed implementation manners

[0062] Exemplary embodiments of the present invention will be described in more detail below with reference to the accompanying drawings. Although the exemplary embodiments of the present invention are shown in the drawings, it should be understood that the present invention can be implemented in various forms and should not be limited by the embodiments set forth herein. On the contrary, these embodiments are provided so that the present disclosure can be more thoroughly understood and the scope of the present invention can be fully conveyed to those skilled in the art.

[0063] Figure 1 is a schematic diagram of an execution scheme of a top-N heap sorting algorithm in the prior art. First of all, it should be noted that top-N heap sorting is an algorithm for quickly finding the top N largest or smallest elements in a set of data. The basic idea is to maintain a min heap of size N (for finding the smallest N elements) or a max heap (for finding the largest N elements). In addition, in the existing execution scheme of the top-N heap sorting algorithm, once a comparison needs to be made for a certain sorting key, the program will read the actual value from memory for comparison (if it is a complex sorting key, the actual value will be searched from memory for copying, which is a more time-consuming process). However, in some cases, the comparison operation is meaningless. For example, as Figure 1 shown, in the top-N algorithm, there may be data that is determined to be put into the heap by comparing non-first sorting keys in the early stage and whose first sorting key is released in the later stage. The comparison operation of the non-first sorting key of this data is a meaningless and invalid operation. At the same time, the replacement operation of this data with the top data in the heap is also a meaningless and invalid operation.

[0064] Using the above method, multiple sorting keys of the data are extracted during the early comparison. Therefore, the execution process of the top-N heap sorting algorithm has unnecessary extraction and comparison operations on non-first sorting key data, as well as repeated replacement operations on the data in the heap and other invalid operations. In the case of a large amount of data, there is a problem of waste of CPU execution time due to too many invalid operations, and thus there is a problem of reducing the execution efficiency of the overall sorting process.

[0065] To solve the above problems, an embodiment of the present invention proposes a data sorting method to improve and optimize the overall algorithm. Figure 2 is a schematic diagram of a data sorting method according to an embodiment of the present invention. As Figure 2 shown, the data sorting method of this embodiment generally may include:

[0066] Step S202, obtaining the elements to be sorted and a pre-constructed data structure, where the data structure includes a heap structure and a linked list with the heap nodes in the heap structure as linked list nodes. It should be noted that the heap structure is a complete binary tree structure, and the internal container of the heap structure can be an array. In addition, in this embodiment, the linked list can be formed with the heap nodes in the heap structure as linked list nodes.

[0067] Step S204: Based on the first sorting key of the element to be sorted, determine whether the element to be sorted meets the condition for adding to the linked list. If so, execute Step S206.

[0068] Step S206: Add the element to be sorted to the linked list. That is, when the element to be sorted meets the condition for adding to the linked list, instead of extracting non-first sorting keys, the element to be sorted is temporarily stored in the data structure.

[0069] It should be noted that the data sorting method of the present invention can be used in the top-N heap sorting algorithm to obtain the first N target elements from the dataset to be processed. The target acquisition quantity for obtaining the target elements from the dataset to be processed can be referred to as the boundary value (which can be denoted as bound). Correspondingly, the upper limit value of the number of heap nodes in the heap structure in the data structure can be set as the boundary value, so that after the entire sorting process ends, the elements in the heap structure can be read as the first N target elements obtained from the dataset to be processed.

[0070] In addition, in this embodiment, the data sorting method of the present invention adopts a data structure combining a binary tree and a linked list to construct a chained heap. The construction logic of the standard heap sorting algorithm is used for the data in the heap structure, and the data that meets the condition for adding to the linked list is saved in the linked list, which increases the upper limit of the number of elements stored in the data structure.

[0071] The data sorting method of the present invention pre-constructs a data structure including a heap structure and a linked list with the heap nodes in the heap structure as linked list nodes, optimizes the data storage structure for heap sorting, and after obtaining the element to be sorted and the pre-constructed data structure, based on the first sorting key of the element to be sorted, determines whether the element to be sorted meets the condition for adding to the linked list, and when the element to be sorted meets the condition for adding to the linked list, adds the element to be sorted to the linked list, avoiding directly replacing the nodes in the heap structure when the element to be sorted meets the condition for adding to the data structure, and reducing the node replacement operation in the heap sorting process.

[0072] Therefore, the data sorting method of the present invention can, based on the first sorting key of the element to be sorted, temporarily store the element to be sorted that meets the condition in a linked list with the heap nodes in the heap structure as linked list nodes, so as to improve and optimize the overall algorithm, avoid the problem of low execution efficiency caused by repeatedly replacing the top node of the heap during the heap sorting process, reduce the replacement operation in the heap sorting process, achieve reducing the waste caused by invalid operations, thereby further improving the overall sorting execution efficiency and then enhancing the performance of the database system.

[0073] In some embodiments, in the above step S202, each node in the data structure is used to store an element. Additionally, different heap nodes in the heap structure of the data structure are used to store elements with different first sorting keys, and different linked list nodes in each linked list are used to store elements with the same first sorting key. That is, elements with different first sorting keys are stored in different heap nodes in the heap structure, and elements with the same first sorting key are stored in the linked list. Thus, the data sorting method of the present invention optimizes the data structure and improves the sorting accuracy.

[0074] In some embodiments, the elements to be sorted are the elements to be added to the data structure obtained from the dataset to be processed. Correspondingly, the step of obtaining the elements to be sorted in the above step S202 may include the following steps: determining whether there are elements to be processed; if there are, obtaining the elements to be processed; determining whether the number of elements stored in the data structure reaches the boundary value; if it does not reach, using the obtained elements to be processed as the elements to be sorted; if it reaches, based on the first sorting key of the elements to be processed, determining whether the elements to be processed meet the conditions for adding to the data structure; if they meet, using the obtained elements to be processed as the elements to be sorted; if they do not meet, sorting the elements stored in the data structure based on the non-first sorting keys of the elements stored in the data structure until the number of elements stored in the data structure is the same as the boundary value to obtain all target elements. It should be noted that the elements to be processed are the elements in the dataset to be processed that have not been added to the data structure and have not been released.

[0075] For example, for the execution statement 'SELECT * FROM table ORDER BY a DESC, b DESC LIMIT 3;', the dataset to be processed is the dataset composed of column a and column b in table table, the elements to be processed are the data in column a and column b that have not been added to the data structure for sorting and have not been directly released, and the elements to be sorted are the data in column a and column b that meet the conditions for adding to the data structure.

[0076] Specifically, the step of determining whether the elements to be processed meet the conditions for adding to the data structure based on the first sorting key of the elements to be processed may include the following steps: in the case where the type of the heap structure is a min heap, determining whether the first sorting key of the elements to be processed is not less than the first sorting key of the element stored in the heap top node; if it is not less, confirming that the elements to be processed meet the conditions for adding to the data structure; if it is less, confirming that the elements to be processed do not meet the conditions for adding to the data structure; or in the case where the type of the heap structure is a max heap, determining whether the first sorting key of the elements to be processed is not greater than the first sorting key of the element stored in the heap top node; if it is not greater, confirming that the elements to be processed meet the conditions for adding to the data structure; if it is greater, confirming that the elements to be processed do not meet the conditions for adding to the data structure.

[0077] In this embodiment, the types of heap structures may include min - heaps and max - heaps. When the type of the heap structure is a min - heap, the element stored in the heap top node (which can be called the heap top element) is the smallest element in the heap structure, and can be used to obtain the top N largest target elements from the dataset to be processed using the top - N heap sorting algorithm. At this time, the heap top element is the minimum value among the top N largest target elements. When the type of the heap structure is a max - heap, the heap top element is the largest element in the heap structure, and can be used to obtain the top N smallest target elements from the dataset to be processed using the top - N heap sorting algorithm. At this time, the heap top element is the maximum value among the top N smallest target elements.

[0078] Therefore, when the type of the heap structure is a min - heap, if the first sorting key of the element to be processed is not less than the first sorting key of the heap top element, it indicates that the element to be processed belongs to the top N largest target elements among the currently obtained elements, and thus it is confirmed that the element to be processed meets the condition for adding to the data structure. When the type of the heap structure is a max - heap, if the first sorting key of the element to be processed is not greater than the first sorting key of the heap top element, it indicates that the element to be processed belongs to the top N smallest target elements among the currently obtained elements, and thus it is confirmed that the element to be processed meets the condition for adding to the data structure.

[0079] In addition, the step of sorting the elements stored in the data structure based on the non - first sorting keys of the elements stored in the data structure can be called the head - cutting operation (which can be denoted as Head - cut Operation). When performing the head - cutting operation, it is necessary to sort the next sorting key. Specifically, different sorting methods can be selected for sorting the next sorting key according to the size of the data volume to be sorted and the data type. For example, reconstruct the data structure based on the non - first sorting keys of the elements stored in the data structure, and obtain the first several target elements. Thus, when extracting the next sorting key in the linked list sequence during the final sorting, quick sorting may achieve higher efficiency in the best and average cases.

[0080] In some embodiments, the above step S204 may include the following steps: determining whether the first sorting key of the element to be sorted is the same as the first sorting key of the element stored in any heap node in the heap structure; if the same, confirming that the element to be sorted meets the condition for joining the linked list; if not the same, confirming that the element to be sorted does not meet the condition for joining the linked list. It should be noted that the element stored in the heap node can be referred to as the element in the heap. Correspondingly, the above step of determining whether the first sorting key of the element to be sorted is the same as the first sorting key of the element stored in any heap node in the heap structure can be executed as the following steps: determining whether there is an element in the heap whose first sorting key is the same as the first sorting key of the element to be sorted in the heap structure; if so, confirming that the first sorting key of the element to be sorted is the same as the first sorting key of the element stored in any heap node in the heap structure; if not, confirming that the first sorting key of the element to be sorted is not the same as the first sorting key of the element stored in any heap node in the heap structure.

[0081] Thus, the data sorting method of the present invention also determines whether the element to be sorted meets the condition for joining the linked list based on the first sorting key of the element to be sorted, and only when the first sorting key is the same as the first sorting key of the element stored in any heap node in the heap structure, confirms that the element to be sorted meets the condition for joining the linked list, realizing the judgment of whether the element to be sorted meets the condition for joining the linked list based on the first sorting key of the element to be sorted, ensuring that the first sorting keys of the elements stored in each linked list are the same, thereby improving the sorting accuracy.

[0082] In some embodiments, after the above step of determining whether the element to be processed meets the condition for joining the data structure based on the first sorting key of the element to be processed, the data sorting method of the present invention may further include the following steps: releasing the element to be processed when the element to be processed does not meet the condition for joining the data structure. Thus, the data structure always stores the first N target elements among the currently obtained elements, thereby improving the sorting accuracy.

[0083] In addition, in some embodiments, the above step S206 may include the following steps: adding the element to be sorted to the linked list where the element in the heap with the same first sorting key as the element to be sorted is located. Thus, the element to be sorted and the elements in the heap with the same first sorting key are temporarily stored in the data structure, and do not directly continue with the comparison and replacement operations of non-first sorting keys, avoiding the problem of low execution efficiency caused by repeatedly replacing the top node of the heap during the heap sorting process, reducing the replacement operations during the heap sorting process, and improving the sorting efficiency.

[0084] In some embodiments, after the above step S206, the data sorting method of the present invention may further include the following steps: taking the number of elements stored in the nodes other than the top heap linked list in the data structure as the non-top heap element number, where the top heap linked list is a linked list with the top node of the heap structure as the linked list node; determining whether the non-top heap element number reaches a set boundary value, where the boundary value is the target acquisition number for obtaining the target element from the data set to be processed; if so, releasing all elements of the top heap linked list.

[0085] Thus, after the non-top heap element number reaches the set boundary value, the data sorting method of the present invention releases all elements of the top heap linked list, avoiding the comparison operation of non-first sorting keys for all elements of the top heap linked list, realizing the reduction of waste caused by invalid operations. At the same time, it effectively limits the element storage amount in the heap structure, thereby further improving the overall sorting execution efficiency and further enhancing the performance of the database system.

[0086] Figure 3 It is a schematic diagram of the application scenario of the data sorting method according to an embodiment of the present invention. The following combines Figure 3 to describe the application scenario of the data sorting method of the present invention.

[0087] As Figure 3 shown, the data sorting method of the present invention adopts a data structure combining a binary tree and a linked list to construct a heap, adopts the construction logic of the standard heap sorting algorithm for data with different key values, and stores data with the same key value in the linked list. Therefore, the data sorting method of the present invention can be abbreviated as chain-heap sorting (which can be denoted as Linked-heapSort), and the data structure of the present invention can be abbreviated as chain heap (which can be denoted as linked-heap).

[0088] In addition, as Figure 3 shown, the linked list can use the heap node in the heap structure as the head node. That is to say, the head node of each linked list is the heap node in the heap structure. Specifically, the linked list with the top node of the heap structure as the linked list node can be called the top heap linked list, the number of elements stored in all nodes other than the top heap linked list in the data structure can be called the non-top heap element number (which can be denoted as subN), and the number of elements stored in the data structure can be called the total number of elements (which can be denoted as N).

[0089] In a specific embodiment, after inserting a new node into the chained heap, if the number of non-top elements in the chained heap is equal to the boundary value, all elements (including the top element) in the top linked list of the heap are released, so that the total number of elements in the chained heap is reduced to within the boundary value. After constructing the chained heap, if the total number of elements is more than the boundary value, the first top linked list needs to be sorted based on the non-first sorting key and then truncated, and the redundant elements are released until the sorting is completed. Thus, after the entire sorting process ends, the data sorting method of the present invention can obtain the first N target elements from the dataset to be processed.

[0090] Using the above method, as Figure 3 shown, each time the determination process of the new element only makes one comparison for the first sorting key, and the values with the same first sorting key are placed in the same linked list. When the linked list needs to be released, it is directly released in batches, and there is no need to compare the data in the linked list again to obtain the correct result. After the heap is constructed, only the data in the linked list where it is determined that there is data to be left is sorted for the subsequent sorting keys, thus avoiding the unnecessary comparison process of the non-first sorting keys.

[0091] Figure 4 is a schematic diagram of the data structure in the initial state in the data sorting method according to an embodiment of the present invention. Figure 5 is a schematic diagram of the data structure in the state where the N = bound condition is not satisfied in the data sorting method according to an embodiment of the present invention. Figure 6 is a schematic diagram of the data structure in the state where the N = bound condition is satisfied in the data sorting method according to an embodiment of the present invention. Figure 7 is a schematic diagram of the data structure after releasing the top element of the heap in the data sorting method according to an embodiment of the present invention. Figure 8 is a schematic diagram of the data structure after adding the element to be sorted to the top linked list in the data sorting method according to an embodiment of the present invention. Figure 9 is a schematic diagram of the data structure after releasing the element to be processed in the data sorting method according to an embodiment of the present invention. Figure 10 is a schematic diagram of the data structure after releasing the top element of the heap in the data sorting method according to an embodiment of the present invention. Figure 11 is a schematic diagram of the data structure in the completed construction state in the data sorting method according to an embodiment of the present invention. Figure 12 is a schematic diagram of the arrow operation in the data sorting method according to an embodiment of the present invention. The following combines Figures 4 to 12 , and details of the multi-key sorting execution operation of the data sorting method of the present invention will be described.

[0092] In some embodiments, as Figure 4As shown in the figure, the data structure in the data sorting method of the present invention can be constructed through the following steps: construct an index array structure based on the first sorting key of all elements to be processed in the dataset to be processed, where the index array structure includes multiple arrays, the multiple arrays correspond one-to-one with all elements to be processed, and each array includes the first sorting key of the corresponding element to be processed, a row index, and a next-hop index. The row index is used to represent the memory location of the element to be processed, and the next-hop index is used to point to the position of another element to be processed with the same first sorting key as the element to be processed in the index array structure to simulate a linked list; set a head pointer and a tail pointer respectively pointing to the heads of the multiple arrays in the index array structure. The head pointer is used to represent that the element to be processed corresponding to the pointed array is the top of the heap structure, and the tail pointer is used to represent that the pointed array is the tail of the heap structure to simulate the heap structure.

[0093] Specifically, as Figure 4 shown, assume that there are 4 columns and 8 rows of data stored in the memory, and there are two sorting keys. Initially, the required index array structure has been constructed for the actual top-N sorting operation. Each array in the index array structure consists of three pieces of information: a sorting key, a row index, and a next-hop index. The row index represents the corresponding position relationship of the element in the actual memory. The sorting key can be the actual data of a certain key extracted from the memory, initially the actual value of the first sorting key. The next-hop index points to the position of the element with the same value in this data structure to simulate a linked list in sequence. In a specific embodiment, setting the next-hop index to 0 indicates that this element is the end of the linked list, and setting the next-hop index to -1 indicates that this element is not in the linked heap and needs to be released.

[0094] Furthermore, set a head pointer (denoted as head) to point to the top of the heap structure, set a tail pointer (denoted as tail) to represent the tail of the heap structure (also representing the length of the heap structure), and set an insertion pointer (denoted as temp) to point to the next element to be inserted into the heap operation. In addition, set bound to represent the boundary size, N to represent the total number of elements in the current linked heap (including all elements in the linked list), and subN to represent the total number of elements excluding the elements in the linked list at the top of the heap. For example, when executing the statement 'SELECT* FROM table ORDER BY a DESC, b DESC LIMIT 3;', bound is equal to 3. It should be noted that when subN is equal to bound, it means that the entire linked list at the top of the heap can be released. The subsequent "swap pointer" operation means swapping all the contents of the arrays pointed to by the two pointers (including the row index, sorting key, and next-hop index).

[0095] In some embodiments, as Figures 4 to 12As described above, when executing the statement 'SELECT * FROM table ORDER BY a DESC, b DESC LIMIT 3;', the multi-key sorting execution process of the data sorting method of the present invention may include the following steps:

[0096] 1. First, construct a heap with N = 3. As Figure 4 shown, when inserting an element into the chained heap, it is necessary to first determine whether it can be added to the linked list. If so, directly add it to the linked list and move temp; otherwise, swap temp and tail, organize the heap, and move tail backward. Specifically, the step of determining whether it can be added to the linked list can be specifically executed as determining whether the temp key value is the same as the head key value. If not, swap temp and tail, organize the heap, and then move tail backward. The chained heap becomes Figure 5 the structure shown.

[0097] 2. At this time, N < bound, and temp cannot be added to the chain. Therefore, swap temp and tail, organize the heap, and move tail backward. The chained heap becomes Figure 6 the structure shown.

[0098] 3. At this time, the condition N = bound is satisfied. Therefore, the subsequent operation needs to determine whether temp can be inserted into the chained heap. As Figure 6 shown, the temp key value is greater than the head value and cannot be inserted into the linked list. Therefore, the heap top replacement operation needs to be executed. In actual operation, first mark head as to be released (set the next-hop index to -1), swap temp and head, organize the heap, and move temp forward. The chained heap becomes Figure 7 the structure shown.

[0099] 4. At this time, the temp key value is the same as the head key value. Therefore, temp can be added to the linked list of head, and set the next-hop index of head to temp (the index number in the array, not the row index), indicating that the temp position is added to the linked list of head. Then move temp forward. The chained heap becomes Figure 8 the structure shown.

[0100] 5. At this time, when comparing temp with head, it does not meet the condition for inserting into the chained heap. Directly mark temp as to be released and move temp forward. The chained heap becomes Figure 9 the structure shown.

[0101] 6. At this time, temp is compared with head. It meets the conditions for insertion into the linked list and is not inserted into the top heap linked list. Therefore, the following operations are required: ① Add temp to the corresponding chain and then move temp forward; ② Since subN = bound at this time, the top heap needs to be released and all elements of the top heap linked list are marked as to be released; ③ Since the release of the top heap causes the heap structure to become shorter, move tail forward, swap head and tail, organize the heap, then swap tail and temp, and move temp forward. The chained heap becomes Figure 10 the structure shown. Thus, the data sorting method of the present invention can optimize the algorithm. Compared with the standard algorithm, all elements in the released linked list do not compare other sorting keys, reducing the number of comparisons and improving the sorting execution efficiency.

[0102] 7. Now the last element is scanned. It is found that it can be inserted into the chained heap and cannot be inserted into the linked list, and after insertion, subN < bound, so the top heap element does not need to be released. Therefore, swap temp and tail, and move tail backward after organizing the heap. At this time, temp < tail, and the construction of the chained heap is completed. The memory format after completion is Figure 11 as shown.

[0103] 8. After the chained heap is constructed, it is found that N > bound. Therefore, the head cutting operation needs to be performed to cut off N - bound elements from the top heap linked list. That is, sort the top heap linked list in turn through the subsequent sorting keys and place the whole at the end of the result array until the sorting is completed. Specifically, as Figure 12 shown, transfer the top heap linked list to a temporary array, read the next sorting key value of the corresponding row, and perform the chained heap sorting of bound = bound - subN again. After sorting, store the data in the temporary array back to the original position of the linked list in turn to complete the head cutting operation.

[0104] So far, the construction of the entire chained heap is completed, and the elements stored in the chained heap are the target elements. When the memory needs to be released, release the memory from the memory in turn according to the marked row numbers (the row numbers corresponding to the elements with the next-hop index set to -1).

[0105] Therefore, in actual operation, the chained heap sort only uses an array to simulate the structure of a binary tree plus a linked list, which can effectively utilize the memory space. Through the chained heap sort, a large number of ineffective read and comparison operations on the actual information of non-first sorting keys are effectively avoided, thereby improving the execution efficiency of the overall top-N sort. Therefore, the data sorting method of the present invention can improve and optimize the overall algorithm, avoid the problem of low execution efficiency caused by repeatedly replacing and comparing operations on the top node of the heap during the heap sort process, reduce ineffective operations such as replacement and comparison operations during the heap sort process, realize the reduction of waste caused by ineffective operations, thereby further improving the overall sorting execution efficiency, and further enhancing the performance of the database system.

[0106] Figure 13 It is a flowchart of the data sorting method according to an embodiment of the present invention. The following will specifically describe Figure 13 the process steps of this embodiment.

[0107] Step S1302, initialize the top element of the heap and the boundary value. It should be noted that the step of initializing the top element of the heap can be specifically executed as initializing the top node of the heap structure in the pre-constructed data structure. Specifically, a heap structure of the large top heap type is used for the top-N heap sort to find the first N smallest elements in the dataset, and the top element is the element with the largest sorted key value stored in the heap structure; a heap structure of the small top heap type is used for the top-N heap sort to find the first N largest elements in the dataset, and the top element is the element with the smallest sorted key value stored in the heap structure. In addition, the boundary value is the target acquisition quantity for obtaining the first N target elements from the dataset to be processed when performing the top-N heap sort task, which can be denoted as bound.

[0108] Step S1304, determine whether there is a new element. If so, execute step S1306; if not, execute step S1324.

[0109] Step S1306, obtain the new element.

[0110] Step S1308, determine whether the number of elements stored in the data structure reaches the boundary value. If so, execute step S1320; if not, execute step S1310. It should be noted that the number of elements stored in the data structure can be denoted as N.

[0111] Step S1310, determine whether the new element is the same as the element in the heap. If so, execute step S1312; if not, execute step S1314. It should be noted that the element in the heap refers to the element stored in the heap structure. In addition, this step can be specifically executed as determining whether the first sorting key of the new element is the same as the first sorting key of any element in the heap structure.

[0112] Step S1312: Add the new element to the corresponding linked list and execute step S1316. It should be noted that the first sorting key of each element in the corresponding linked list is the same as that of the new element.

[0113] Step S1314: Add the new element to the heap structure and execute step S1316. It should be noted that the first sorting key of each element in the heap structure is different.

[0114] Step S1316: Determine whether the number of non-heap-top elements reaches the boundary value. If so, execute step S1318; if not, return to step S1304. It should be noted that the number of non-heap-top elements refers to the number of elements stored in the nodes other than the heap-top linked list in the data structure, and the heap-top linked list is a linked list with the heap-top node of the heap structure as the linked list node.

[0115] Step S1318: Release all elements of the heap-top linked list to reconstruct the heap structure and return to step S1304.

[0116] Step S1320: Determine whether the new element meets the condition for adding to the data structure. If so, execute step S1310; if not, execute step S1322.

[0117] Step S1322: Release the new element and return to step S1304.

[0118] Step S1324: Sort the elements in the heap-top linked list based on the non-first sorting keys of the elements stored in the data structure.

[0119] Step S1326: Determine whether the number of elements stored in the data structure is greater than the boundary value. If so, execute step S1328; if not, execute step S1330.

[0120] Step S1328: Truncate and release the redundant elements in the heap-top linked list so that the number of elements stored in the data structure is the same as the boundary value, and execute step S1330.

[0121] Step S1330: Obtain the elements stored in the data structure as all target elements. Thus, the execution process of the top-N heap sorting task is completed, and this process ends.

[0122] Therefore, the data sorting method of the present invention realizes a heap sorting optimization scheme applied to multi-key top-N sorting, reduces the number of element comparisons in the execution logic compared with the standard heap sorting algorithm, and only performs data comparison when it is necessary to compare, which can improve the sorting efficiency as a whole.

[0123] Figure 14It is a control flow chart of a data sorting method according to an embodiment of the present invention. Taking the execution of a top-N sorting task in a database with a columnar storage structure as an example, and the execution statement is 'SELECT * FROM table ORDER BY a DESC, b DESC LIMIT 3;', the control flow of this embodiment will be specifically described in combination with Figure 14 to specifically describe the control flow of this embodiment.

[0124] Step S1402, create an index array for the in-memory data. It should be noted that the in-memory data refers to all the data in columns a and b of table table.

[0125] Step S1404, initialize pointers head, tail, and temp, as well as variables N and subN. It should be noted that head points to the top of the heap structure, tail represents the tail of the heap structure (also representing the length of the heap structure), temp points to the next element to be inserted into the data structure or released, N represents the total number of elements in the current data structure, and subN represents the total number of elements excluding the elements in the top-linked list. When subN is equal to bound, it means that the entire top-linked list can be released. The subsequent "swap pointer" operation means swapping all the contents of the arrays pointed to by two pointers (including row index, sorting key value, and next-hop index).

[0126] Step S1406, determine whether temp < tail. If so, execute step S1430. If not, execute step S1408. It should be noted that temp < tail means there is a new element to be inserted into the data structure or released.

[0127] Step S1408, determine whether N < bound. If so, execute step S1410. If not, execute step S1416. It should be noted that bound represents the boundary value. In this embodiment, bound is equal to 3.

[0128] Step S1410, determine whether temp has the same value as the heap node. If so, execute step S1412. If not, execute step S1414. It should be noted that temp having the same value as the heap node means that the first sorting key of the element corresponding to the array pointed to by temp is the same as the first sorting key of any element in the heap structure.

[0129] Step S1412, add temp to the corresponding linked list, update N and subN, move temp forward, and return to step S1406. It should be noted that if temp is added to the top-linked list, perform the operation of N + 1, that is, N = N + 1; if temp is added to a non-top-linked list, perform the operation of subN++, that is, subN = subN + 1.

[0130] Step S1414, swap temp and tail, reconstruct the heap, update N and subN, move tail backward, and return to Step S1406. It should be noted that the first sorting keys of each element in the heap structure are different from each other.

[0131] Step S1416, compare temp with head, and determine whether temp can be put into the data structure. If so, execute Step S1418; if not, execute Step S1428. It should be noted that for a max heap, when the first sorting key of the element corresponding to the array pointed to by temp is less than the first sorting key of the heap top element, it is confirmed that temp can be put into the data structure; for a min heap, when the first sorting key of the element corresponding to the array pointed to by temp is greater than the first sorting key of the heap top element, it is confirmed that temp can be put into the data structure.

[0132] Step S1418, determine whether temp has the same value as the heap node. If so, execute Step S1420; if not, execute Step S1426.

[0133] Step S1420, add temp to the corresponding linked list, update N and subN, move temp forward, and execute Step S1422.

[0134] Step S1422, determine whether subN is equal to bound. If so, execute Step S1424; if not, return to Step S1406.

[0135] Step S1424, release head and all elements in the heap top linked list, move tail forward, swap head and tail, reconstruct the heap, update N and subN, swap tail and temp, move temp forward, and return to Step S1406. It should be noted that updating N and subN includes recording N = bound and recalculating subN. Specifically, the step of releasing head and all elements in the heap top linked list can be specifically executed as setting the next-hop index in the array corresponding to head and all elements in the heap top linked list to -1 to indicate that they need to be released finally.

[0136] Step S1426, release head and all elements in the heap top linked list, swap head and temp, reconstruct the heap, update subN, move temp forward, and return to Step S1406.

[0137] Step S1428, mark temp as to be released, move temp forward, and return to Step S1406. Specifically, this step can be specifically executed as setting the next-hop index in the array pointed to by temp to -1 to indicate that it needs to be released finally.

[0138] Step S1430: Sort the top - list and truncate and release the redundant elements in the top - list, so that the number of elements stored in the data structure is the same as the boundary value, and then execute Step S1432.

[0139] Step S1432: Obtain the elements stored in the data structure as all target elements. Thus, the execution process of the top - N heap sorting task is completed, and this process ends.

[0140] Therefore, the data sorting method of the present invention can be used for the heap sorting of multi - key top - N sorting. Logically, compared with the standard heap sorting algorithm, it reduces the number of element comparisons, and only performs data comparisons when necessary, thereby improving the overall sorting efficiency.

[0141] Furthermore, the data sorting method of the present invention can be applied to the sorting nodes of column - store databases. Due to the reduction in the number of comparisons, it effectively reduces the number of executions of extracting actual values from memory, especially reducing the operation consumption of locating and copying actual data from memory for complex sorting keys, thereby improving the overall sorting execution efficiency of the nodes.

[0142] This embodiment also provides a machine - readable storage medium and a computer device. Figure 15 FIG. is a schematic structural diagram of a machine - readable storage medium 10 according to an embodiment of the present invention. Figure 16 FIG. is a schematic structural diagram of a computer device 20 according to an embodiment of the present invention.

[0143] The machine - readable storage medium 10 stores a machine - executable program 11 thereon. When the machine - executable program 11 is executed by a processor, it implements the data sorting method of any of the above - mentioned embodiments.

[0144] The computer device 20 may include a memory 220, a processor 210, and a machine - executable program 11 stored on the memory 220 and running on the processor 210. When the processor 210 executes the machine - executable program 11, it implements the data sorting method of any of the above - mentioned embodiments.

[0145] It should be noted that the logic and / or steps represented in the flowchart or described in other ways herein, for example, can be considered as a definite sequence list of executable instructions for implementing logical functions, and can be specifically implemented in any machine - readable storage medium for use by an instruction execution system, apparatus, or device (such as a computer - based system, a system including a processor, or other systems that can fetch and execute instructions from the instruction execution system, apparatus, or device), or in combination with these instruction execution systems, apparatus, or devices.

[0146] For the description of this embodiment, the machine-readable storage medium 10 can be any device that can contain, store, communicate, propagate, or transport a program for use by or in connection with an instruction execution system, apparatus, or device. More specific examples (non-exhaustive list) of the machine-readable storage medium 10 include the following: an electrical connection portion (electronic device) having one or more wirings, a portable computer disk cartridge (magnetic device), a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber device, and a portable compact disc read-only memory (CDROM). Additionally, the machine-readable storage medium 10 can even be paper or other suitable media on which the program can be printed, since the program can be obtained electronically, for example, by optically scanning the paper or other media, followed by editing, interpretation, or otherwise processing as appropriate, and then storing it in a computer memory.

[0147] It should be understood that various parts of the present invention can be implemented by hardware, software, firmware, or a combination thereof. In the above-described embodiments, multiple steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system.

[0148] The computer device 20 can be, for example, a server, a desktop computer, a laptop computer, a tablet computer, or a smartphone. In some examples, the computer device 20 can be a cloud computing node. The computer device 20 can be described in the general context of computer system-executable instructions, such as program modules, executed by a computer system. Generally, program modules can include routines, programs, object programs, components, logic, data structures, etc. that perform specific tasks or implement specific abstract data types. The computer device 20 can be implemented in a distributed cloud computing environment where tasks are executed by remote processing devices linked through a communication network. In a distributed cloud computing environment, program modules can be located on local or remote computing system storage media including storage devices.

[0149] The computer device 20 can include a processor 210 adapted to execute stored instructions and a memory 220 that provides temporary storage space for the operation of the instructions during operation. The processor 210 can be a single-core processor, a multi-core processor, a computing cluster, or any other number of other configurations. The memory 220 can include a random access memory (RAM), a read-only memory, a flash memory, or any other suitable storage system.

[0150] The flowcharts provided in this embodiment are not intended to indicate that the operations of the method will be performed in any specific order, or that all operations of the method are included in every case. In addition, the method may include additional operations. Within the scope of the technical concept provided by the method of this embodiment, additional changes may be made to the above method.

[0151] At this point, those skilled in the art should recognize that although many exemplary embodiments of the present invention have been shown and described in detail herein, many other variations or modifications that conform to the principles of the present invention can still be directly determined or derived from the content disclosed in the present invention without departing from the spirit and scope of the present invention. Therefore, the scope of the present invention should be understood and determined to cover all such other variations or modifications.

Claims

1. A data sorting method, comprising: Obtaining elements to be sorted and a pre-constructed data structure, wherein the data structure includes a heap structure and a linked list with the heap nodes in the heap structure as linked list nodes, and different heap nodes in the heap structure are used to store elements with different first sorting keys, and different linked list nodes in each linked list are used to store elements with the same first sorting key; Judging whether the first sorting key of the element to be sorted is the same as the first sorting key of any element stored in any heap node in the heap structure; If so, adding the element to be sorted to the linked list.

2. The data sorting method according to claim 1, wherein, Each node in the data structure is used to store one element; And After the step of adding the element to be sorted to the linked list, the data sorting method further includes: Taking the number of elements stored in the nodes other than the top heap linked list in the data structure as the number of non-top heap elements, wherein the top heap linked list is a linked list with the top heap node of the heap structure as a linked list node; Judging whether the number of non-top heap elements reaches a set boundary value, wherein the boundary value is the target acquisition quantity for obtaining target elements from the data set to be processed; If so, releasing all elements of the top heap linked list.

3. The data sorting method according to claim 2, wherein The step of obtaining the element to be sorted includes: Judging whether there is an element to be processed, wherein the element to be processed is an element in the data set to be processed that has not been added to the data structure and has not been released; If there is, obtaining the element to be processed; Judging whether the number of elements stored in the data structure reaches the boundary value; If not, taking the obtained element to be processed as the element to be sorted; If it reaches, judging whether the element to be processed meets the condition of being added to the data structure based on the first sorting key of the element to be processed; If it meets, taking the obtained element to be processed as the element to be sorted.

4. The data sorting method according to claim 3, wherein After the step of judging whether there is an element to be processed, the data sorting method further includes: In the case that there is no element to be processed, sorting the elements stored in the data structure based on the non-first sorting keys of the elements stored in the data structure until the number of elements stored in the data structure is the same as the boundary value to obtain all the target elements.

5. The data sorting method according to claim 3, wherein The step of judging whether the element to be processed meets the condition of being added to the data structure based on the first sorting key of the element to be processed includes: In the case that the type of the heap structure is a min heap, judging whether the first sorting key of the element to be processed is not less than the first sorting key of the element stored in the top heap node. If not less, it is confirmed that the element to be processed meets the condition of being added to the data structure. If less, it is confirmed that the element to be processed does not meet the condition of being added to the data structure; or When the type of the heap structure is a max heap, determine whether the first sorting key of the element to be processed is not greater than the first sorting key of the element stored in the heap top node. If it is not greater, confirm that the element to be processed meets the condition for adding to the data structure. If it is greater, confirm that the element to be processed does not meet the condition for adding to the data structure.

6. The data sorting method according to claim 3, wherein After the step of determining whether the element to be processed meets the condition for adding to the data structure based on the first sorting key of the element to be processed, the data sorting method further includes: When the element to be processed does not meet the condition for adding to the data structure, release the element to be processed.

7. The data sorting method according to claim 1, wherein, The data structure is constructed through the following steps: Construct an index array structure based on the first sorting keys of all elements to be processed in the data set to be processed. The index array structure includes multiple arrays, and the multiple arrays correspond one by one to all the elements to be processed. Each array includes the first sorting key of the corresponding element to be processed, a row index, and a next-hop index. The row index is used to represent the memory location of the element to be processed, and the next-hop index is used to point to the position of another element to be processed with the same first sorting key as the first sorting key of the element to be processed in the index array structure to simulate the linked list. Set a head pointer and a tail pointer respectively pointing to the heads of the multiple arrays in the index array structure. The head pointer is used to indicate that the element to be processed corresponding to the pointed array is the heap top of the heap structure, and the tail pointer is used to indicate that the pointed array is the tail of the heap structure to simulate the heap structure.

8. A machine-readable storage medium, on which a machine-executable program is stored. When the machine-executable program is executed by a processor, the data sorting method according to any one of claims 1 to 7 is implemented.

9. A computer device, including a memory, a processor, and a machine-executable program stored on the memory and running on the processor. When the processor executes the machine-executable program, the data sorting method according to any one of claims 1 to 7 is implemented.

Citation Information

Patent Citations

  • Information screening method and device and computer equipment

    CN111737263A

  • Data structure construction method and device, equipment and storage medium

    CN117632950A