Data sorting method and device based on heterogeneous computing architecture and storage medium

Through heterogeneous computing architecture, the data to be sorted is distributed to different hardware for hierarchical sorting, which solves the performance bottleneck problem in large-scale data processing and achieves more efficient data sorting.

CN120596056APending Publication Date: 2025-09-05ZHEJIANG DAHUA TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510538902.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-25
Publication Date
2025-09-05

AI Technical Summary

Technical Problem

When processing large-scale data, existing technologies are limited by hardware, and data sorting and processing are prone to performance bottlenecks, affecting data processing efficiency.

Method used

A method based on heterogeneous computing architecture is used to transfer the data to be sorted to the first hardware with strong parallel processing capabilities for preliminary sorting to obtain candidate data, which is then transferred to the second hardware with weaker parallel processing capabilities for further sorting to alleviate the performance bottleneck of the second hardware.

Benefits of technology

By distributing the sorting task to multiple hardware devices, the load on a single hardware is reduced, and the efficiency and performance of data sorting are improved. It is suitable for application scenarios such as feature retrieval and data mining.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120596056A_ABST
    Figure CN120596056A_ABST
Patent Text Reader

Abstract

The invention discloses a data sorting method and device based on a heterogeneous computing architecture and a storage medium, the method is applied to a data processing system, the data processing system comprises first hardware and second hardware, the parallel processing capacity of the first hardware is larger than that of the second hardware, and the parallel processing capacity of the second hardware is larger than that of the first hardware. The method comprises the following steps: responding to the situation that the to-be-sorted data volume of received to-be-sorted data is not matched with the parallel processing capability of second hardware, and transmitting the to-be-sorted data to first hardware; sorting the to-be-sorted data according to the first hardware to obtain candidate data; and transmitting the candidate data to second hardware, and sorting the candidate data according to the second hardware to obtain target data. According to the scheme, the data processing pressure of the second hardware can be relieved, and the data sorting efficiency is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of data sorting technology, and in particular to a data sorting method, device, and storage medium based on a heterogeneous computing architecture. Background Art

[0002] Data sorting method refers to the method of arranging the data to be sorted in a specified order.

[0003] Data sorting methods are often used in applications such as feature retrieval and data mining. For example, the Top-K sorting algorithm finds the top K largest or smallest elements in a set of data, avoiding the performance overhead of a complete sort.

[0004] However, when processing large-scale data, due to the influence of hardware level, the data sorting and processing process is prone to performance bottlenecks, affecting data processing efficiency. Summary of the Invention

[0005] This application at least provides a data sorting method, apparatus, device, and computer-readable storage medium based on a heterogeneous computing architecture.

[0006] In a first aspect, the present application provides a data sorting method based on a heterogeneous computing architecture, the method being applied to a data processing system, the data processing system comprising first hardware and second hardware, the parallel processing capability of the first hardware being greater than the parallel processing capability of the second hardware, the method comprising: in response to a mismatch between the amount of received data to be sorted and the parallel processing capability of the second hardware, transferring the data to be sorted to the first hardware; sorting the data to be sorted according to the first hardware to obtain candidate data; transmitting the candidate data to the second hardware, and sorting the candidate data according to the second hardware to obtain target data.

[0007] In one embodiment, the sorting of the data to be sorted according to the first hardware to obtain candidate data includes: grouping the data to be sorted to obtain multiple data groups to be sorted; sorting the multiple data groups to be sorted according to a preset sorting algorithm to obtain a sorted data group; and determining the candidate data group from the sorted data groups, wherein the candidate data group includes the candidate data.

[0008] In one embodiment, the grouping processing of the data to be sorted to obtain multiple data groups to be sorted includes: determining a data sharing ratio based on the parallel processing capabilities of the first hardware and the parallel processing capabilities of the second hardware; and grouping processing of the data to be sorted based on the data sharing ratio to obtain multiple sorted data groups.

[0009] In one embodiment, the sorting process of the plurality of data groups to be sorted according to a preset sorting algorithm to obtain the sorted data groups includes: performing pooling process on each data group to be sorted to obtain the characteristic value of each data group to be sorted; and sorting the respective data groups to be sorted according to the preset sorting algorithm and the characteristic value of each data group to be sorted to obtain the sorted data groups.

[0010] In one embodiment, the performing pooling processing on each data group to be sorted to obtain the characteristic value of each data group to be sorted includes: performing maximum pooling processing on each data group to be sorted to obtain the characteristic value of each data group to be sorted; or performing minimum pooling processing on each data group to be sorted to obtain the characteristic value of each data group to be sorted.

[0011] In one embodiment, the sorting of each data group to be sorted according to the preset sorting algorithm and the characteristic value of each data group to be sorted to obtain the sorted data group includes: obtaining a preset number of initial data groups from each data group to be sorted; constructing an initial data group stack according to the characteristic value of each initial data group; the initial data group stack includes a top data group; traversing the remaining data groups except the initial data groups in each data group to be sorted, and comparing the characteristic value of the remaining data group currently traversed with the characteristic value of the top data group to obtain a first comparison result; in response to the first comparison result meeting the first heap entry condition, replacing the top data group according to the remaining data group currently traversed to obtain the sorted data group.

[0012] In one embodiment, the candidate data is sorted according to the second hardware to obtain target data, including: obtaining a preset number of initial data from each candidate data; constructing an initial data pile based on each initial data; the initial data pile includes top data; traversing the remaining data except the initial data in each candidate data, and comparing the currently traversed remaining data with the top data to obtain a second comparison result; in response to the second comparison result meeting the second heap entry condition, replacing the top data according to the currently traversed remaining data to obtain the target data.

[0013] In one embodiment, in response to the fact that the amount of the received data to be sorted does not match the parallel processing capability of the second hardware, before transmitting the data to be sorted to the first hardware, the method further includes: obtaining the amount of the data to be sorted of the data to be sorted; and performing a matching judgment process on the amount of the data to be sorted and the parallel processing capability of the second hardware.

[0014] The second aspect of the present application provides a data sorting device based on a heterogeneous computing architecture, including: a data transmission module, used to transmit the to-be-sorted data to the first hardware in response to the amount of to-be-sorted data received not matching the parallel processing capability of the second hardware; a first sorting module, used to sort the to-be-sorted data according to the first hardware to obtain candidate data; and a second sorting module, used to transmit the candidate data to the second hardware and sort the candidate data according to the second hardware to obtain target data.

[0015] A third aspect of the present application provides an electronic device, including a memory and a processor, wherein the processor is used to execute program instructions stored in the memory to implement the above-mentioned data sorting method based on heterogeneous computing architecture.

[0016] In a fourth aspect, the present application provides a computer-readable storage medium having program instructions stored thereon. When the program instructions are executed by a processor, the above-mentioned data sorting method based on a heterogeneous computing architecture is implemented.

[0017] The above scheme matches the amount of received data to be sorted with the parallel processing capability of the second hardware. When the amount of data to be sorted does not match the parallel processing capability of the second hardware, the data to be sorted is transferred to the first hardware, and the data to be sorted is pre-processed by the first hardware to reduce the data processing pressure of the second hardware; wherein, the data to be sorted is sorted according to the first hardware to obtain candidate data; and then the candidate data is transferred to the second hardware, and then the candidate data can be sorted according to the second hardware to obtain target data. In this way, when the parallel processing capability of the second hardware is difficult to support the processing of the received data to be sorted, the data to be sorted can be processed by the first hardware first to reduce the load of the second hardware, alleviate the performance bottleneck of the second hardware, and improve the data sorting efficiency.

[0018] It is to be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the present application. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] The drawings herein are incorporated into and constitute a part of the specification. These drawings illustrate embodiments consistent with the present application and, together with the specification, are used to illustrate the technical solutions of the present application.

[0020] Figure 1 It is a flowchart of an exemplary embodiment of the data sorting method based on heterogeneous computing architecture of the present application;

[0021] Figure 2 This is a schematic diagram of an exemplary data processing system in the data sorting method based on heterogeneous computing architecture of the present application;

[0022] Figure 3 This is a schematic diagram of an exemplary sorting process in the data sorting method based on heterogeneous computing architecture of the present application;

[0023] Figure 4 This is a schematic diagram of improving the sorting effect in the data sorting method based on heterogeneous computing architecture of the present application;

[0024] Figure 5 is a block diagram of a data sorting device based on a heterogeneous computing architecture, shown in an exemplary embodiment of the present application;

[0025] Figure 6 This is a structural diagram of an embodiment of an electronic device of the present application;

[0026] Figure 7 It is a structural diagram of an embodiment of a computer-readable storage medium of the present application. DETAILED DESCRIPTION

[0027] The following describes the embodiments of the present application in detail with reference to the accompanying drawings.

[0028] In the following description, for the purpose of explanation rather than limitation, specific details such as specific system structures, interfaces, and technologies are provided to facilitate a thorough understanding of the present application.

[0029] The term "and / or" in this article is simply a description of the association relationship of associated objects, indicating that three relationships can exist. For example, A and / or B can represent three situations: A exists alone, A and B exist at the same time, and B exists alone. In addition, the character " / " in this article generally indicates that the associated objects are in an "or" relationship. In addition, "many" in this article means two or more than two. In addition, the term "at least one" in this article means any combination of at least two of any one or more of a plurality of. For example, including at least one of A, B, and C can mean including any one or more elements selected from the set consisting of A, B, and C.

[0030] To facilitate understanding, an exemplary description of one of the applicable scenarios of this application is now given.

[0031] In the existing data sorting process, the selection, deployment and optimization of the full sorting algorithm or the Top-K sorting algorithm are completed on a certain hardware platform such as CPU or GPU. Some common optimization methods are to improve the utilization rate of hardware resources, and some optimization methods are to reduce the amount of data to be sorted based on the method of filtering and pruning, thereby improving the performance of the sorting algorithm. The heterogeneous computing architecture of this application refers to the use of multiple different types of processors or hardware components (such as CPU, GPU, FPGA, NPU, etc.) in a computing system to improve computing performance and efficiency through collaborative work. This application is mainly explained by taking CPU and GPU as examples.

[0032] Intelligent algorithms deployed in heterogeneous computing architectures typically employ a multi-task pipeline design involving computation and data transfer. For example, in image retrieval scenarios, data comparison and sorting is one such task. Because sorting depends on previous and subsequent data, concurrent processing is difficult, and performance bottlenecks can easily occur when dealing with large amounts of data. Optimizing the sorting process is particularly crucial for optimizing the overall performance of intelligent algorithms.

[0033] It is understood that the sorting algorithms used in the data sorting process can include a variety of algorithms, which are not limited here. The Top-K sorting algorithm aims to find the top K largest or smallest elements from a set of data, especially when processing large-scale data, which reduces the performance overhead compared to complete sorting. It is also commonly used in application scenarios such as feature retrieval and data mining. This application mainly uses the Top-K algorithm as an example for illustration.

[0034] See also Figure 1 , Figure 1 This is a flow chart of an exemplary embodiment of the data sorting method based on heterogeneous computing architecture of the present application. The method of the present application is applied to the data processing system, and can refer to Figure 2 As shown, Figure 2 This is a schematic diagram of an exemplary data processing system in the data sorting method based on a heterogeneous computing architecture of the present application. The data processing system may include at least first hardware and second hardware, and the parallel processing capability of the first hardware is greater than the parallel processing capability of the second hardware. For example, the first hardware may be a GPU and the second hardware may be a CPU; or the first hardware may be an NPU and the second hardware may be a CPU, etc., which are not limited here. Specifically, the following steps may be included:

[0035] Step S110 : In response to the fact that the amount of the received data to be sorted does not match the parallel processing capability of the second hardware, the data to be sorted is transmitted to the first hardware.

[0036] The parallel processing capability of the hardware refers to the computing power of the hardware when sorting the data to be sorted. It is understood that the parallel processing capability (computing power) of the present application can be characterized by various indicators, such as the number of floating-point operations per second, the amount of information data that can be processed per second, etc., which are not limited here.

[0037] For example, when receiving data to be sorted, if the data to be sorted is to be sent to the second hardware for processing, a matching determination can be made between the amount of data to be sorted and the parallel processing capability of the second hardware. If the amount of data to be sorted does not match the parallel processing capability of the second hardware, the data to be sorted is transferred to the first hardware. If the amount of data to be sorted matches the parallel processing capability of the second hardware, the sorting process can be directly performed by the second hardware.

[0038] As another example, the data processing system of the present application may include one or more first hardware components, which is not limited herein. Upon receiving the data to be sorted, if the data to be sorted is sent to a first hardware component for processing, the sorting process may be performed directly by any of the first hardware components. Alternatively, after matching the amount of data to be sorted with the parallel processing capabilities of each piece of first hardware component, a matching first hardware component may be selected from the various pieces of first hardware components to perform the sorting process.

[0039] In another exemplary embodiment, when receiving the data to be sorted, if the data to be sorted is sent to the second hardware for processing, the amount of the data to be sorted can be matched with the parallel processing capability of the second hardware. If the amount of the data to be sorted does not match the parallel processing capability of the second hardware, the amount of the data to be sorted is matched with the parallel processing capability of each piece of first hardware, and the first hardware with matching capability is selected from the various pieces of first hardware for sorting processing.

[0040] The process of determining whether the hardware's parallel processing capability matches the amount of data to be sorted can be implemented using preset judgment rules. For example, if the parallel processing capability of a certain hardware is xGB and the amount of data to be sorted is yGB, then when xGB is greater than yGB, it can be determined that the capability matches, and when xGB is less than or equal to yGB, it can be determined that the capability does not match. This is not limited here. Optionally, a fault tolerance space can be introduced as a reference. That is, in order to prevent the hardware from devoting all of its currently available computing power to data sorting, an idle parameter can be set to ensure that the hardware still has a certain amount of idle computing power left while being able to sort the data. For example, if the parallel processing capability of a certain hardware is xGB, the amount of data to be sorted is yGB, and the idle parameter is zGB, then it can be determined that the capability matches only when xGB is greater than (y+z)GB. Furthermore, if there are multiple hardware devices whose parallel processing capabilities match the amount of data to be sorted, the hardware with the smallest parallel processing capability can be selected from the multiple matching hardware devices to perform data sorting. This can reduce resource fragmentation in the hardware and avoid resource waste.

[0041] Step S120 , sorting the data to be sorted according to the first hardware to obtain candidate data.

[0042] In conjunction with the above steps, after the first hardware sorts the data to be sorted, candidate data can be obtained. The candidate data can be all or part of the data to be sorted (equivalent to the number of candidate data being less than or equal to the amount of data to be sorted), which is not limited here.

[0043] Step S130 : transmitting the candidate data to the second hardware, and sorting the candidate data according to the second hardware to obtain the target data.

[0044] In conjunction with the above steps, after obtaining candidate data sorted by the first hardware, the candidate data can be transferred to the second hardware, and the second hardware is used to further sort the candidate data to obtain the target data. The target data can be all or part of the candidate data (equivalent to the target data being less than or equal to the candidate data), which is not limited here.

[0045] It can be seen that the present application matches the amount of received data to be sorted with the parallel processing capability of the second hardware. When the amount of data to be sorted does not match the parallel processing capability of the second hardware, the data to be sorted is transferred to the first hardware, and the data to be sorted is pre-processed by the first hardware to reduce the data processing pressure of the second hardware. The data to be sorted is sorted according to the first hardware to obtain candidate data. The candidate data is then transferred to the second hardware, and then the candidate data can be sorted according to the second hardware to obtain target data. In this way, when the parallel processing capability of the second hardware is unable to support the processing of the received data to be sorted, the data to be sorted can be processed by the first hardware first to reduce the load of the second hardware, alleviate the performance bottleneck of the second hardware, and improve the efficiency of data sorting.

[0046] Based on the above embodiment, the embodiment of the present application describes the steps of sorting the data to be sorted according to the first hardware to obtain candidate data. Specifically, the method of this embodiment includes the following steps:

[0047] The data to be sorted are grouped to obtain a plurality of data groups to be sorted; the plurality of data groups to be sorted are sorted according to a preset sorting algorithm to obtain sorted data groups; and a candidate data group is determined from the sorted data groups, wherein the candidate data group includes the candidate data.

[0048] In conjunction with the aforementioned embodiment, in the process of sorting the data to be sorted according to the first hardware in the present application, the data to be sorted may be sorted directly, or the data to be sorted may be grouped to obtain multiple data groups to be sorted, and then the data groups to be sorted are sorted. Each data group to be sorted may include multiple data to be sorted.

[0049] The method of sorting the data group to be sorted by the first hardware can realize rough sorting of the data; then the sorted data group can be sent to the second hardware, and the method of sorting the data in the sorted data group by the second hardware can realize fine sorting of the data.

[0050] For example, the preset sorting algorithm may be a Top-K algorithm, which can obtain K sorted data groups after sorting multiple data groups to be sorted. All or part of the sorted data groups can be determined as candidate data groups.

[0051] For example, by grouping the data to be sorted, multiple data groups to be sorted and a data group number for each data group to be sorted can be obtained. After the rough sorting process by the first hardware, all or part of the data groups can be extracted from the sorted data groups according to the data group numbers as new data to be sorted (i.e., candidate data to be sorted by the second hardware).

[0052] For another example, if the amount of data to be sorted is M, it can be divided into M / S data groups according to the preset data allocation ratio S. After the first hardware performs Top-K sorting on the data groups to be sorted, K candidate data groups can be obtained. The current candidate data amount is K*S, which reduces the amount of data that needs to be sorted by the second hardware.

[0053] Based on the above embodiment, the embodiment of the present application exemplifies the steps of grouping the data to be sorted to obtain multiple data groups to be sorted. Specifically, the method of this embodiment includes the following steps:

[0054] The data sharing ratio is determined according to the parallel processing capabilities of the first hardware and the second hardware; the data to be sorted is grouped and processed according to the data sharing ratio to obtain a plurality of sorted data groups.

[0055] With reference to the foregoing embodiments, in the group processing process of the present application, the data to be sorted may be grouped and processed according to a preset data sharing ratio, or the data sharing ratio may be determined based on the parallel processing capability of the first hardware and the parallel processing capability of the second hardware (or the preset data sharing ratio may be adjusted).

[0056] Specifically, the larger the data sharing ratio S is, the smaller the number of data groups to be sorted is, which is equivalent to a smaller computing load on the first hardware and a larger computing load on the second hardware. Therefore, an initial data sharing ratio can be determined in advance based on the initial parallel processing capability of the first hardware and / or the initial parallel processing capability of the second hardware. Subsequently, the data sharing ratio can be adjusted based on the real-time parallel processing capability of the first hardware and / or the real-time parallel processing capability of the second hardware each time a data sorting task is received, and this is not limited here. For example, the data sharing ratio can be negatively correlated with the currently available parallel processing capability of the first hardware, or the data sharing ratio can be positively correlated with the currently available parallel processing capability of the second hardware, or the data sharing ratio can be negatively correlated with the capability gap between the currently available parallel processing capabilities of the first hardware and the second hardware.

[0057] For example, the data allocation ratio S of the present application can be determined according to a fixed ratio mode or a dynamic ratio mode, which is not limited here. Assume that the amount of data to be sorted is M, and it is necessary to obtain the top K data with the highest numerical ranking. Then, every S data can be divided into a group at a fixed interval, resulting in M / S data groups. By adjusting the size of S, the sorting performance can be adjusted.

[0058] If the first hardware sorting in the sorting process takes time t1, the second hardware sorting takes time t2.

[0059] (1) Fixed ratio mode:

[0060] Because each data group needs to select its own characteristic value on the first hardware, multi-threaded concurrent processing can be used, so the size of S can be configured as a power of 2, such as 4, 8, 16, 32, ..., 64, etc.

[0061] Generally speaking, the idle resources available for sorting tasks on the second hardware are greater than the idle resources available for sorting tasks on the first hardware. The size of S can be adjusted according to the adjustment instructions temporarily received, and the appropriate ratio can be determined based on one or more test results as a fixed ratio (data sharing ratio), thereby improving the overall performance.

[0062] (2) Dynamic ratio mode:

[0063] Record the three parameters M, K, and S, as well as the actual time t1 and t2 consumed by the first and second hardware, respectively, to account for differences in computing power in the actual operating environment. A parameter α representing computing power can be introduced to adjust the model. Assume that α1 and α2 represent the computing power of the first and second hardware, respectively. Fine-tune the S ratio decision based on the mathematical model.

[0064] 1. Determine model parameters

[0065] There are three parameters M, K and S. Task 1 is to select the top-K from M / S data sets, and task 2 is to select the top-K from K×S data sets.

[0066] 2. Build a mathematical model

[0067] Assuming that the time complexity of heap sort is O(NlogK), and introducing the computing power parameter α to represent the computing power difference in the actual operating environment. The mathematical expression of the model can be:

[0068] t1=(c1*M / S*logK) / α1

[0069] t2=(c1*K*S*logK) / α2

[0070] Where c1 and c2 are constants, representing the computing power overhead in the actual operating environment.

[0071] During the application process, the goal is to make t1 and t2 as close as possible, that is, t1 ≈ t2, which means (c1*M / S*logK) / α1 ≈ (c1*K*S*logK) / α2. Assuming c1 ≈ c2, we ignore the difference introduced by the additional computing power overhead of the two hardware platforms. Substituting the data into the model, we can further organize it into M / (S*α1) ≈ K*S / α2. This can be approximated as:

[0072] 3. Verification and Adjustment

[0073] Use the approximate parameter S to verify that the actual time consumption of t1 and t2 is similar. Optionally, the parameter S can be further adjusted based on the measured data.

[0074] Based on the above embodiment, the embodiment of the present application describes the steps of sorting multiple data groups to be sorted according to a preset sorting algorithm to obtain a sorted data group. Specifically, the method of this embodiment includes the following steps:

[0075] Perform pooling processing on each data group to be sorted to obtain the characteristic value of each data group to be sorted; and perform sorting processing on each data group to be sorted according to a preset sorting algorithm and the characteristic value of each data group to be sorted to obtain a sorted data group.

[0076] With reference to the foregoing embodiments, in the process of sorting a plurality of data groups to be sorted in the present application, each data group to be sorted can be pooled separately to obtain the characteristic value of each data group to be sorted, so as to sort each data group to be sorted according to the characteristic value of each data group to be sorted.

[0077] Based on the above embodiment, the embodiment of the present application exemplarily illustrates the steps of performing pooling processing on each data group to be sorted to obtain the feature value of each data group to be sorted. Specifically, the method of this embodiment includes the following steps:

[0078] Performing maximum pooling processing on each data group to be sorted to obtain the eigenvalue of each data group to be sorted; or performing minimum pooling processing on each data group to be sorted to obtain the eigenvalue of each data group to be sorted.

[0079] In conjunction with the above embodiments, the pooling process of the present application may include maximum pooling or minimum pooling during a certain data sorting process. Among them, by performing maximum pooling on each data group to be sorted (selecting the numerical value corresponding to the largest data to be sorted in the data group to be sorted), the maximum value (eigenvalue) of each data group to be sorted can be obtained; by performing minimum pooling on each data group to be sorted (selecting the numerical value corresponding to the smallest data to be sorted in the data group to be sorted), the minimum value (eigenvalue) of each data group to be sorted can be obtained.

[0080] The specific use of maximum pooling or minimum pooling needs to be selected according to the actual application scenario. For example, if the received representation of the request to be sorted requires the data to be sorted to be sorted in descending order, then the maximum pooling process can be selected to sort according to the maximum value of each data group to be sorted. If the received representation of the request to be sorted requires the data to be sorted to be sorted in descending order, then the minimum pooling process can be selected to sort according to the minimum value of each data group to be sorted.

[0081] In summary, after obtaining the data group to be sorted, concurrent operations can be performed on the first hardware. The first hardware performs a maximum pooling (or minimum pooling) operation on each data group to be sorted to obtain its maximum value (or minimum value). Then, sorting processing can be performed according to the maximum value data of each data group to be sorted.

[0082] Based on the above embodiment, the embodiment of the present application describes the steps of sorting each data group to be sorted according to a preset sorting algorithm and the characteristic value of each data group to be sorted to obtain a sorted data group. Specifically, the method of this embodiment includes the following steps:

[0083] Obtain a preset number of initial data groups from each data group to be sorted; construct an initial data group stack based on the characteristic values ​​of each initial data group; the initial data group stack includes a top data group; traverse the remaining data groups except the initial data groups in each data group to be sorted, and compare the characteristic values ​​of the remaining data groups currently traversed with the characteristic values ​​of the top data group to obtain a first comparison result; in response to the first comparison result meeting the first heap entry condition, replace the top data group according to the remaining data groups currently traversed to obtain a sorted data group.

[0084] In conjunction with the above embodiments, the preset sorting algorithm of the present application can be implemented by constructing heap data. The heap data can include a max heap (maximum heap) and a min heap (minimum heap), which corresponds to the data sorting method (for example, a min heap can be used to find larger data, and a max heap can be used to find smaller data).

[0085] For example, reference may be made to Figure 3 , Figure 3 This is an exemplary sorting process diagram in the data sorting method based on heterogeneous computing architecture of the present application. A preset number of initial data groups are obtained from the data group to be sorted (for example, the first K data groups are obtained). An initial data group heap is constructed according to the characteristic values ​​of each initial data group (a large top heap or a small top heap can be constructed according to the sorting requirements of the actual application scenario). The top data group is also the top element in the initial data group heap. Then the remaining data groups can be traversed in sequence, and the characteristic values ​​of the remaining data groups currently traversed are compared with the characteristic values ​​of the top data group to obtain a first comparison result. In response to the first comparison result meeting the first heap entry condition, the top data group can be replaced according to the remaining data groups currently traversed; in response to the first comparison result not meeting the first heap entry condition, the processing of the remaining data groups is skipped; until the traversal and comparison of the remaining data groups are completed, a sorted data group can be obtained.

[0086] Among them, the first condition for entering the heap can be to determine whether the characteristic value of the remaining data group currently traversed is greater than the characteristic value of the top data group; or to determine whether the characteristic value of the remaining data group currently traversed is less than the characteristic value of the top data group. The specific selection needs to be made according to the sorting requirements of the actual application scenario.

[0087] For example, in the scenario of building a small top heap, if the eigenvalue of the remaining data group currently traversed is greater than the eigenvalue of the top data group of the heap, the processing of this data group can be skipped; if the eigenvalue of the remaining data group currently traversed is less than or equal to the eigenvalue of the top data group of the heap, the top data group and the remaining data group currently traversed can be replaced to adjust the small top heap.

[0088] Similarly, in the scenario of building a max heap, if the eigenvalue of the currently traversed remaining data group is smaller than the eigenvalue of the top data group, processing of the data group can be skipped; if the eigenvalue of the currently traversed remaining data group is greater than or equal to the eigenvalue of the top data group, the top data group and the currently traversed remaining data group can be replaced to adjust the max heap.

[0089] Based on the above embodiment, the embodiment of the present application describes the steps of sorting candidate data according to the second hardware to obtain target data. Specifically, the method of this embodiment includes the following steps:

[0090] Obtain a preset number of initial data from each candidate data; construct an initial data pile based on each initial data; the initial data pile includes top data; traverse the remaining data except the initial data in each candidate data, and compare the currently traversed remaining data with the top data of the pile to obtain a second comparison result; in response to the second comparison result meeting the second heap entry condition, replace the top data of the pile according to the currently traversed remaining data to obtain the target data.

[0091] With reference to the foregoing embodiments, the algorithm used in the present application for sorting in the second hardware may be the same as or different from the algorithm used in the first hardware for sorting, and this is not limited here. For application scenarios such as feature retrieval, the Top-K algorithm is preferably used, so this application is mainly explained using the example of the first hardware and the second hardware both using the Top-K algorithm. It should be noted that the K value in the Top-K algorithm used by the first hardware and the second hardware may be the same or different, and this is not limited here. In some cases, since the computing power consumption is relatively large in the coarse sorting stage of the first hardware, the K value of the coarse sorting stage can be set to be greater than the K value of the fine sorting stage.

[0092] Exemplarily, a preset number of initial data are obtained from the candidate data (for example, the first K data are obtained). An initial data heap is constructed according to the numerical values ​​corresponding to each initial data (a max-top heap or a min-top heap can be constructed according to the sorting requirements of the actual application scenario). The top data of the heap is also the top element in the initial data heap. The remaining data can then be traversed in sequence, and the numerical values ​​corresponding to the remaining data currently traversed are compared with the numerical values ​​corresponding to the top data of the heap to obtain a second comparison result. In response to the second comparison result meeting the second heap entry condition, the top data of the heap can be replaced according to the remaining data currently traversed, and then a heap adjustment can be performed; in response to the second comparison result not meeting the second heap entry condition, the processing of the remaining data is skipped; until the traversal and comparison of the remaining data are completed, the K elements in the final heap can be used as target data.

[0093] Based on the above embodiment, the present embodiment of the application describes the steps before transferring the data to be sorted to the first hardware in response to the amount of the received data to be sorted not matching the parallel processing capability of the second hardware. Specifically, the method of this embodiment includes the following steps:

[0094] Acquire the amount of data to be sorted; and perform matching judgment processing on the amount of data to be sorted and the parallel processing capability of the second hardware.

[0095] In conjunction with the above embodiment, before transmitting the data to be sorted to the first hardware, it is also possible to determine whether the amount of the data to be sorted matches the parallel processing capability of the second hardware, so as to indicate whether the second hardware can directly process the data to be sorted.

[0096] Specifically, referring to the description of the aforementioned embodiment, the judgment process of the present application may include, but is not limited to, obtaining the amount of data to be sorted; performing a matching judgment process on the amount of data to be sorted and the parallel processing capability of the second hardware to obtain a matching judgment result; and then selecting, based on the matching judgment result, whether the first hardware or the second hardware should prioritize processing the data.

[0097] In summary, the data sorting method based on heterogeneous computing architecture of this application breaks down the sorting task on a hardware into two levels of sorting pipelines, which are executed by two different hardware respectively. By distributing part of the computing power load to another relatively idle hardware, the overall sorting performance is systematically improved. You can refer to the example Figure 4 , Figure 4 This is a schematic diagram of the improvement of the sorting effect in the data sorting method based on the heterogeneous computing architecture of this application. Among them, sorting 1 refers to the first hardware sorting, and sorting 2 refers to the second hardware sorting. The data sorting method of this application can be for one sorting task or multiple sorting tasks, which is not limited here. Figure 4It can be seen that when facing multiple sorting tasks, since the first hardware and the second hardware can execute their respective sorting tasks concurrently, the effect of optimizing the time consumption is more obvious.

[0098] Therefore, compared with the prior art, the data sorting method of the present application takes the systematization of the solution into consideration more. It is not limited to the deployment and single-point optimization on a single hardware. It considers more how to make full use of multiple hardware units in a heterogeneous computing architecture. Through the design of a two-stage sorting pipeline solution, the sorting computing load is proportionally unloaded from a single hardware platform to other idle hardware in the heterogeneous computing architecture, making full use of the idle computing power of the system and achieving the purpose of optimizing the performance of the overall solution. In the re-sorting process of the fine sorting stage, the data to be sorted that has been filtered out of the data group is retrieved, so the output result will not produce a loss of accuracy, which is suitable for application scenarios that require high data accuracy, such as feature retrieval. Optionally, the use of the Top-K heap sorting algorithm can take up less memory space when processing large-scale data to be sorted.

[0099] It should be further explained that the execution subject of the data sorting method based on the heterogeneous computing architecture may be a data sorting device based on the heterogeneous computing architecture. For example, the data sorting method based on the heterogeneous computing architecture may be executed by a terminal device or a server or other processing device, wherein the terminal device may be a user equipment (UE), a computer, a mobile device, a user terminal, a terminal, a cellular phone, a cordless phone, a personal digital assistant (PDA), a handheld device, a computing device, a vehicle-mounted device, a wearable device, etc. In some possible implementations, the data sorting method based on the heterogeneous computing architecture may be implemented by a processor calling computer-readable instructions stored in a memory.

[0100] Figure 5 FIG is a block diagram of a data sorting device based on a heterogeneous computing architecture, shown as an exemplary embodiment of the present application. Figure 5 As shown, the exemplary data sorting device 500 based on heterogeneous computing architecture includes: a data transmission module 510, a first sorting module 520 and a second sorting module 530. Specifically:

[0101] The data transmission module 510 is configured to transmit the to-be-sorted data to the first hardware in response to a mismatch between the amount of the received to-be-sorted data and the parallel processing capability of the second hardware.

[0102] The first sorting module 520 is configured to sort the data to be sorted according to the first hardware to obtain candidate data.

[0103] The second sorting module 530 is configured to transmit the candidate data to the second hardware, and sort the candidate data according to the second hardware to obtain target data.

[0104] In this exemplary data sorting device based on a heterogeneous computing architecture, the amount of received data to be sorted is matched with the parallel processing capability of the second hardware. When the amount of data to be sorted does not match the parallel processing capability of the second hardware, the data to be sorted is transmitted to the first hardware, and the first hardware pre-processes the data to be sorted, thereby reducing the data processing pressure of the second hardware. The data to be sorted is sorted according to the first hardware to obtain candidate data. The candidate data is then transmitted to the second hardware, and the candidate data can be sorted according to the second hardware to obtain target data. In this way, when the parallel processing capability of the second hardware is unable to support the processing of the received data to be sorted, the first hardware can prioritize the processing of the data to be sorted to reduce the load on the second hardware, alleviate the performance bottleneck of the second hardware, and improve data sorting efficiency.

[0105] It should be noted that the apparatus provided in the above embodiments and the methods provided in the above embodiments are based on the same concept. The specific manner in which the various modules and units perform their operations has been described in detail in the method embodiments and will not be repeated here. In actual applications, the apparatus provided in the above embodiments can, as needed, allocate the above functions to different functional modules, i.e., divide the internal structure of the apparatus into different functional modules to perform all or part of the functions described above. This is not a limitation herein.

[0106] Among them, the functions of each module can be found in the embodiment of the data sorting method based on heterogeneous computing architecture, which will not be repeated here.

[0107] See also Figure 6 , Figure 6 1 is a schematic diagram of the structure of an embodiment of an electronic device of the present application. Electronic device 100 includes memory 101 and processor 102. Processor 102 is configured to execute program instructions stored in memory 101 to implement the steps of any of the aforementioned embodiments of the data sorting method based on a heterogeneous computing architecture. In a specific implementation scenario, electronic device 100 may include, but is not limited to, a microcomputer and a server. Furthermore, electronic device 100 may also include mobile devices such as laptops and tablet computers, which are not limited herein.

[0108] Specifically, the processor 102 is used to control itself and the memory 101 to implement the steps in any of the above-mentioned data sorting method embodiments based on a heterogeneous computing architecture. The processor 102 can also be referred to as a CPU (Central Processing Unit). The processor 102 may be an integrated circuit chip having signal processing capabilities. The processor 102 can also be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA) or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor, etc. In addition, the processor 102 can be implemented by an integrated circuit chip.

[0109] In this exemplary electronic device, by matching the amount of received data to be sorted with the parallel processing capability of the second hardware, when the amount of data to be sorted does not match the parallel processing capability of the second hardware, the data to be sorted is transmitted to the first hardware, and the first hardware first pre-processes the data to be sorted, thereby reducing the data processing pressure of the second hardware; wherein, the data to be sorted is sorted according to the first hardware to obtain candidate data; then the candidate data is transmitted to the second hardware, and then the candidate data can be sorted according to the second hardware to obtain target data. In this way, when the parallel processing capability of the second hardware is unable to support the processing of the received data to be sorted, the first hardware can prioritize the processing of the data to be sorted to reduce the load of the second hardware, alleviate the performance bottleneck of the second hardware, and improve data sorting efficiency.

[0110] See also Figure 7 , Figure 7 The computer-readable storage medium 110 stores program instructions 111 that can be executed by a processor, and the program instructions 111 are used to implement the steps of any of the above-mentioned data sorting method embodiments based on heterogeneous computing architecture.

[0111] In this exemplary storage medium, by running the program instructions in the storage medium, the amount of the received data to be sorted is matched with the parallel processing capability of the second hardware. When the amount of the data to be sorted does not match the parallel processing capability of the second hardware, the data to be sorted is transmitted to the first hardware, and the first hardware first pre-processes the data to be sorted to reduce the data processing pressure of the second hardware; wherein, the data to be sorted is sorted according to the first hardware to obtain candidate data; then the candidate data is transmitted to the second hardware, and then the candidate data can be sorted according to the second hardware to obtain target data. In this way, when the parallel processing capability of the second hardware is difficult to support the processing of the received data to be sorted, the first hardware can prioritize the processing of the data to be sorted to reduce the load of the second hardware, alleviate the performance bottleneck of the second hardware, and improve the efficiency of data sorting.

[0112] In some embodiments, the functions or modules included in the device provided by the embodiments of the present disclosure can be used to execute the method described in the above method embodiments. The specific implementation can refer to the description of the above method embodiments. For the sake of brevity, it will not be repeated here.

[0113] The above description of the various embodiments tends to emphasize the differences between the various embodiments. The same or similar aspects can be referenced with each other and will not be repeated herein for the sake of brevity.

[0114] In the several embodiments provided in this application, it should be understood that the disclosed methods and devices can be implemented in other ways. For example, the device implementation methods described above are only schematic. For example, the division of modules or units is only a logical function division. There may be other division methods in actual implementation. For example, units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, and the indirect coupling or communication connection of devices or units can be electrical, mechanical or other forms.

[0115] In addition, the functional units in the various embodiments of the present application can be integrated into a processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of a software functional unit. If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application is essentially or the part that contributes to the prior art or all or part of the technical solution can be embodied in the form of a software product, which is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, server, or network device, etc.) or a processor to execute all or part of the steps of the various embodiments of the present application. The aforementioned storage medium includes various media that can store program code, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk.

Claims

1. A data sorting method based on a heterogeneous computing architecture, characterized in that: The method is applied to a data processing system, the data processing system including first hardware and second hardware, the parallel processing capability of the first hardware being greater than the parallel processing capability of the second hardware, the method including: In response to an amount of the received data to be sorted not matching the parallel processing capability of the second hardware, transmitting the data to be sorted to the first hardware; sorting the data to be sorted according to the first hardware to obtain candidate data; The candidate data is transmitted to the second hardware, and the candidate data is sorted according to the second hardware to obtain target data.

2. The method according to claim 1, characterized in that The step of sorting the data to be sorted according to the first hardware to obtain candidate data includes: performing grouping processing on the data to be sorted to obtain a plurality of data groups to be sorted; Sorting multiple data groups to be sorted according to a preset sorting algorithm to obtain a sorted data group; A candidate data group is determined from the sorted data group, the candidate data group including the candidate data.

3. The method according to claim 2, characterized in that The grouping process of the data to be sorted to obtain a plurality of data groups to be sorted includes: Determining a data sharing ratio according to the parallel processing capability of the first hardware and the parallel processing capability of the second hardware; The data to be sorted is grouped according to the data allocation ratio to obtain multiple sorted data groups.

4. The method according to claim 2, characterized in that The step of sorting the plurality of data groups to be sorted according to a preset sorting algorithm to obtain a sorted data group includes: Perform pooling processing on each data group to be sorted to obtain the characteristic value of each data group to be sorted; The data groups to be sorted are sorted according to the preset sorting algorithm and the characteristic value of each data group to be sorted to obtain the sorted data groups.

5. The method according to claim 4, characterized in that The pooling process is performed on each data group to be sorted to obtain the characteristic value of each data group to be sorted, including: Perform maximum pooling on each data group to be sorted to obtain the eigenvalues ​​of each data group to be sorted; Alternatively, a minimum pooling process is performed on each data group to be sorted to obtain a feature value of each data group to be sorted.

6. The method according to claim 4, characterized in that The step of sorting each data group to be sorted according to the preset sorting algorithm and the characteristic value of each data group to be sorted to obtain the sorted data group includes: Obtaining a preset number of initial data groups from each data group to be sorted; Constructing an initial data group stack according to the characteristic values ​​of each initial data group; the initial data group stack includes a top data group; Traversing the remaining data groups except the initial data group in each data group to be sorted, and comparing the characteristic value of the currently traversed remaining data group with the characteristic value of the top data group to obtain a first comparison result; In response to the first comparison result meeting the first heap-entry condition, the top data group of the heap is replaced according to the remaining data groups currently traversed to obtain the sorted data group.

7. The method according to claim 1, characterized in that The sorting of the candidate data according to the second hardware to obtain target data includes: Obtain a preset number of initial data from each candidate data; Constructing an initial data pile according to each initial data; the initial data pile includes the top data of the pile; Traversing the remaining data except the initial data in each candidate data, and comparing the currently traversed remaining data with the top data of the heap to obtain a second comparison result; In response to the second comparison result meeting the second heap-entry condition, the top data of the heap is replaced according to the remaining data currently traversed to obtain the target data.

8. The method according to claim 1, characterized in that Before transmitting the received data to be sorted to the first hardware in response to the data amount of the received data to be sorted not matching the parallel processing capability of the second hardware, the method further includes: Obtaining the amount of data to be sorted of the data to be sorted; A matching judgment process is performed on the amount of data to be sorted and the parallel processing capability of the second hardware.

9. An electronic device, characterized in that: The method comprises a memory and a processor, wherein the processor is configured to execute program instructions stored in the memory to implement the method according to any one of claims 1 to 8.

10. A computer-readable storage medium having program instructions stored thereon, characterized in that: When the program instructions are executed by a processor, the method according to any one of claims 1 to 8 is implemented.