Data sorting method for tensor processor, tensor processor system and medium
By assigning core ids to elements on the TPU and using vector instructions to sort them in parallel, the problem of unstable traditional sorting algorithms on the TPU is solved, efficient and stable data sorting is achieved, and the stability needs of databases and other application scenarios are met.
Patent Information
- Application Number
- CN202511106133.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-08
- Publication Date
- 2025-09-05
- Estimated Expiration
- 2045-08-08
AI Technical Summary
The existing traditional sorting algorithm cannot be directly applied on tensor processors (TPUs), and the traditional double-tone sorting algorithm lacks stability on the TPU, resulting in the relative order of equal elements that may change after sorting, affecting the accuracy of query results.
By assigning core id to each element and passing the core id during the sorting process, combining TPU's vector instructions for parallel comparison and exchange, recursive division and merge are adopted to ensure that the order of equal elements remains unchanged.
It realizes efficient and stable sorting on the TPU, ensuring that equal elements maintain their original order after sorting, and meets the stability needs of database and other application scenarios.
Smart Images

Figure CN120596057A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of data processing technology, and in particular to a data sorting method for a tensor processor, a tensor processor system, and a medium. Background Art
[0002] The demand for efficient and stable parallel sorting algorithms is increasingly urgent in the field of big data processing, especially in scenarios such as deep learning, which require the acceleration of Tensor Processing Units (TPUs). TPU chips, with their highly parallel architecture and instruction set designed for vector computations, offer significant advantages in processing matrix operations. However, their instruction set lacks native support for the single-element comparison and swap operations found in traditional sorting algorithms, making it impossible to directly port classic algorithms such as quick sort and merge sort.
[0003] Bitonic sort is an efficient parallel sorting algorithm suitable for sorting networks. By constructing multiple comparators, it can easily implement parallel computations. Theoretically, it is suitable for TPU platforms. However, in practice, traditional bitonic sorting algorithms do not guarantee sort stability. That is, the relative order of equal elements may change before and after sorting. For example, in database sorting, if two records have the same sort key but different other attributes (such as timestamps or IDs), a sorting algorithm that cannot guarantee stability may cause the order of these records to change, affecting the accuracy of query results. Therefore, a solution that ensures sort stability is needed. Summary of the Invention
[0004] Embodiments of the present invention provide a data sorting method for a tensor processor, a tensor processor system, and a medium to solve the problem of improving sorting stability.
[0005] In a first aspect, an embodiment of the present invention provides a data sorting method for a tensor processor, comprising: Obtain input data and the storage capacity of a single TPU register, and assign a core ID to each element in the input data to mark its initial position; When the amount of the input data is less than or equal to the storage capacity of a single TPU register, performing a single register level sorting operation to obtain a bitonic sequence of the single register; When the amount of the input data is greater than the storage capacity of a single TPU register, recursively dividing the input data into multiple subsequences by bitonic sorting, performing the single-register-level sorting operation on each subsequence, and merging the sorted subsequences to obtain a multi-register bitonic sequence; wherein, when performing the single-register-level sorting operation on each subsequence and merging the sorted subsequences, the core ID of the element is passed; After obtaining the bitune sequence, the order of equal elements is adjusted according to the core id.
[0006] In one possible implementation, the single register level sorting operation includes: Use mask instructions to separate data into two registers by odd and even bits; Use the Max floating point instruction and Min floating point instruction in the TPU vector instruction set to compare the elements of corresponding positions in the two registers, and store the larger and smaller values in the max register and min register respectively; The select instruction is combined with the mask operation to make the max register retain only the odd-numbered elements and the min register retain only the even-numbered elements, and the elements of these two registers are merged to form a bitune sequence.
[0007] In a possible implementation, merging the sorted subsequences includes: The order of the elements of the two subsequences is swapped according to the bitune sorting logic; Use the Max floating point instruction and Min floating point instruction in the TPU vector instruction set to compare the elements at the same position in the two subsequences, and store the larger and smaller values in the max register and min register respectively; The single register level sorting operation is performed on the elements in the max register and the min register respectively to obtain the bitonal sequence of the max register and the bitonal sequence of the min register.
[0008] In a possible implementation, swapping the order of elements of the two subsequences according to the bitonal sorting logic includes: Swap the elements of the two subsequences in order to obtain an increasing subsequence and a decreasing subsequence.
[0009] In a possible implementation, recursively dividing the input data into a plurality of subsequences by bitonic sorting includes: The input data is bitonally sorted and recursively divided into subsequences whose length is less than or equal to a set length; wherein the set length is equal to the storage capacity of a single TPU register.
[0010] In a possible implementation, before performing bitonic sorting on the input data and recursively dividing the data into a plurality of subsequences, the method further includes: Dividing the input data into blocks of powers of 2; When there is a block whose data length is less than the set length, padding the input data to the set length; Assign a core id marking the initial position to each element in the supplemented input data.
[0011] In a possible implementation, after adjusting the order of equal elements according to the core ID, the method further includes: Remove padding values from the input data according to core id.
[0012] In a possible implementation, adjusting the order of equal elements according to core IDs includes: Determine whether the core id order of equal elements is consistent with the original order; When the judgment result is inconsistent, the positions of equal elements are swapped according to the core id.
[0013] In a second aspect, an embodiment of the present invention provides a tensor processor system, comprising: a plurality of TPU registers, a vector instruction execution module, and a core ID management module; Among them, the storage capacity of each TPU register is the same; The core id management module is used to assign a core id marking an initial position to each element in the input data; The vector instruction execution module is used to execute the method in the first aspect or any possible implementation of the first aspect.
[0014] In a third aspect, an embodiment of the present invention provides a computer-readable storage medium storing a computer program, which, when executed by a processor, implements the method in the first aspect or any possible implementation of the first aspect.
[0015] In the embodiment of the present invention, single register processing or recursive partitioning and merging are respectively adopted according to the size of the data, thereby fully utilizing the parallel processing capability of the TPU and improving the sorting efficiency. By recursively partitioning, a large data set is decomposed into small blocks suitable for TPU register processing, and the vector instructions of the TPU are used for parallel comparison and exchange, so that the bitonic sorting algorithm can be efficiently implemented on the TPU. By assigning a core id to each element and passing the core id during the sorting process, the instability problem of the traditional bitonic sorting algorithm is solved, so that equal elements can maintain their original order after sorting. Finally, the order of equal elements is adjusted according to the core id, which further ensures the stability of the sorting, thereby meeting the demand for stable sorting in application scenarios such as databases. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] Figure 1 is a flowchart of an implementation method for a data sorting method for a tensor processor provided by an embodiment of the present invention; Figure 2 Schematic diagram of the structure of the tensor processor system provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0017] Sorting is a fundamental problem in computer science, widely used in various data processing and analysis tasks. With the advent of the big data era, higher requirements are placed on the efficiency, stability, and parallelism of sorting algorithms. The Tensor Processing Unit (TPU), a processor designed specifically for tasks like deep learning, possesses powerful parallel computing capabilities. However, because its instruction set is primarily oriented towards vector operations, the direct application of traditional sorting algorithms on the TPU is limited.
[0018] Bitonic sort is an efficient parallel sorting algorithm suitable for sorting networks. It can easily implement parallel computing by constructing multiple comparators. The key to bitonic sorting is to construct a bitonic sequence, that is, the elements in a sequence are first arranged in ascending order and then start descending after a specific index. By recursively dividing and merging bitonic sequences, an ordered sequence can eventually be obtained. However, traditional bitonic sorting algorithms do not guarantee the stability of the sort, that is, the relative order of equal elements may change before and after sorting. Applying the bitonic sorting algorithm on the TPU provides some research solutions, which are implemented in the following ways: 1. Adaptation of vector instructions: Because the TPU's instruction set is primarily oriented toward vector operations, these solutions attempt to convert the compare-and-swap operations used in bitonic sorting into vector instructions. However, this conversion typically focuses solely on parallelism without considering sorting stability. Traditional sorting algorithms (such as quick sort and merge sort) rely more heavily on comparing and swapping individual elements. This instruction set mismatch prevents traditional sorting algorithms from being directly applied on the TPU.
[0019] 2. Parallel processing: Leveraging the multi-core parallel computing capabilities of the TPU, these solutions assign each stage of bitonic sorting (such as splitting and merging) to different cores for processing. While this parallel processing can significantly increase sorting speed, without additional measures to ensure sort stability, the relative order of equal elements may change during the sorting process.
[0020] However, these solutions have significant limitations. Because they don't address sort stability, they can produce erroneous results in certain application scenarios. For example, in database sorting, if two records have the same sort key but different other attributes (such as timestamps or IDs), maintaining their relative order is crucial. If the sorting algorithm doesn't guarantee stability, the order of these records may shift, affecting the accuracy of query results.
[0021] This paper proposes a method for achieving stable sorting on a TPU based on the bitonic sorting algorithm. This method not only utilizes the parallel computing capability of the TPU to accelerate the sorting process, but also ensures the stability of the sorting by introducing mechanisms such as core ID.
[0022] The embodiments of the present invention will be described in detail below with reference to the accompanying drawings.
[0023] Figure 1 Flowchart for implementing the data sorting method for a tensor processor provided by an embodiment of the present invention. Figure 1 As shown, the following steps are included: S101: Obtain input data and the storage capacity of a single TPU register, and assign a core ID marking the initial position to each element in the input data.
[0024] In practice, TPU register storage capacity can be varied: the current mainstream configuration is 1024 elements (determined by an 8x128-bit register structure), with some high-performance models supporting 2048 elements (using a 16x128-bit extended architecture). This capacity flexibility requires hardware-adaptive sorting algorithms. This application focuses on a TPU register with a storage capacity of 1024 elements.
[0025] During the data preprocessing phase, the system obtains the input data set and detects the capacity of a single TPU register. Each element is assigned a unique core ID to mark its original position (e.g., index 0 to 2047). The core purpose is to address the stability flaws of traditional bitonic sorting, accurately tracking the source and status of each element during the sorting process. This ensures that when comparing and exchanging elements, not only the size of the element value is considered, but also their relative order and source information. This way, even if two element values are equal, the original relative order can be maintained after sorting. For example, when two element values are equal (e.g., elements A
[10] =3.5 and A
[20] =3.5), the core ID can record their initial order (id:10 vs. id:20), providing a traceability basis for the final order adjustment.
[0026] S102 , when the amount of input data is less than or equal to the storage capacity of a single TPU register, performing a single register level sorting operation to obtain a bitonic sequence of the single register.
[0027] There is no instruction in the TPU register to directly process a register sort. Therefore, when the amount of input data is less than or equal to the storage capacity of a single TPU register, the input data is loaded into the register and a single register level sort operation is directly performed.
[0028] S103: When the amount of input data exceeds the storage capacity of a single TPU register, recursively divide the input data into multiple subsequences by bitonic sorting, perform a single-register-level sorting operation on each subsequence, and merge the sorted subsequences to obtain a multi-register bitonic sequence; wherein, when performing the single-register-level sorting operation on each subsequence and merging the sorted subsequences, the core ID of the element is transferred.
[0029] When the amount of input data exceeds the storage capacity of a single TPU register, a single register cannot load the excess data at one time. Therefore, the input data is recursively divided into multiple subsequences by bitonic sorting to avoid data overflow. For example, on a 1024-capacity TPU, 5000 data are divided into 5 sub-blocks (4 blocks of 1024 elements + 1 block of 904 elements), and the tail block is padded with 120 minimum values to 1024. After recursively reaching the length of all sub-blocks ≤ 1024, a single-register-level sorting operation is triggered. This hierarchical processing fully utilizes the TPU's parallel pipeline to avoid the risk of memory overflow.
[0030] During the subsequence merging and sorting process, the core ID is always transmitted synchronously with the elements to ensure that the tracking chain is not broken.
[0031] S104: After obtaining the bitune sequence, the order of equal elements is adjusted according to the core ID.
[0032] After obtaining a bitonic sequence, if the system detects equal elements in the output bitonic sequence (for example, an element with value 3.5 appears at positions 50 and 51), the order of the equal elements is adjusted based on their core IDs to ensure stability. For example, if the core IDs of two elements are compared (assuming id:20 and id:10), and if id:20 > id:10 (i.e., the original order had id:10 first), the positions of the two elements are swapped. This operation ensures that the final output satisfies global order (3.5 ≤ 3.5) while maintaining the original relative order of equal elements (id:10 always comes before id:20), resolving the issue of distorted results caused by ordering errors in scenarios such as database sorting.
[0033] In this embodiment, single-register processing or recursive partitioning and merging are used, depending on the data size, to fully utilize the parallel processing capabilities of the TPU and improve sorting efficiency. Recursive partitioning breaks down large data sets into small blocks suitable for TPU register processing, and the TPU's vector instructions are used for parallel comparison and exchange, enabling efficient implementation of the bitonic sorting algorithm on the TPU. By assigning a core ID to each element and passing the core ID during the sorting process, the instability problem of the traditional bitonic sorting algorithm is resolved, allowing equal elements to maintain their original order after sorting. Finally, the order of equal elements is adjusted based on the core ID, further ensuring the stability of the sorting, thereby meeting the requirements for stable sorting in application scenarios such as databases.
[0034] In one possible implementation, a single register-level sort operation includes: Use mask instructions to separate data into two registers by odd and even bits; Use the Max floating point instruction and Min floating point instruction in the TPU vector instruction set to compare the elements of corresponding positions in the two registers, and store the larger and smaller values in the max register and min register respectively; The select instruction is combined with the mask operation to make the max register retain only the odd-numbered elements and the min register retain only the even-numbered elements, and the elements of these two registers are merged to form a bitune sequence.
[0035] Taking a 1024-bit TPU register as an example (8×128-bit structure), we will explain the single-register-level sorting operation in detail. Assume the input data consists of 800 floating-point numbers (less than the register capacity), and each element has been assigned a core ID to record its original position (e.g., IDs 0 to 799). In this case, the single-register operation is directly triggered, and the specific process is as follows: Step 1: Mask control for odd-even separation Generate two 128-bit vector masks: Odd mask: binary sequence [1,0,1,0,...,1,0] (512 1s and 512 0s alternating) Even mask: binary sequence [0,1,0,1,...,0,1] (512 0s and 512 1s alternating) Use the TPU's permute instruction to separate the 800 elements by index position: Odd-indexed elements (positions 1, 3, 5, ...) are loaded into register A (actually 400 valid data + 112 filled with 0); even-indexed elements (positions 0, 2, 4, ...) are loaded into register B (400 valid data + 112 filled with 0); The key significance of parity separation is to convert serial data into a parallel processable vector structure, laying the foundation for subsequent batch comparison.
[0036] Step 2: Parallel comparison of vector instructions Call TPU hardware instructions: Max floating point (RegA, RegB) → Output the larger value to the max register (e.g. A[1]=5.2>B[1]=4.7, then max[1]=5.2) Min floating point (RegA, RegB) → Output the smaller value to the min register (min[1]=4.7) This operation completes the parallel comparison of 512 pairs of elements (i.e., 1024 / 2) at one time, which is significantly faster than the traditional serial comparison (which requires 512 cycles).
[0037] Step 3: Data reorganization of Select command Apply a fixup mask to the max and min registers: Use an odd-numbered mask ([1,0,1,0,...,1,0]) for the max register to retain only its odd-numbered elements (i.e., positions 1,3,5,... of the original input sequence); use an even-numbered mask ([0,1,0,1,...,0,1]) for the min register to retain only its even-numbered elements (positions 0,2,4,... of the original input sequence). The concat instruction is used to merge the data of the two registers to form a bitonal sequence (for example, the first half of the merged sequence is an increasing odd-numbered large value, and the second half is a decreasing even-numbered small value).
[0038] In this embodiment, a mask instruction is used to separate the parity bits of the data, and the TPU-specific Max floating point instruction and Min floating point instruction are combined to complete parallel vector comparison, converting traditional scalar comparison-based operations into single-cycle batch processing, significantly reducing the number of comparison instruction calls; the coordinated operation of the select instruction and the mask ensures the accurate reorganization of the parity bit data, forming a data layout that strictly conforms to the characteristics of the bitonal sequence.
[0039] In a possible implementation, merging the sorted subsequences includes: The order of the elements of the two subsequences is swapped according to the bitune sorting logic; Use the Max floating point instruction and Min floating point instruction in the TPU vector instruction set to compare the elements at the same position in the two subsequences, and store the larger and smaller values in the max register and min register respectively; Perform single register level sorting operations on the elements in the max register and the min register respectively to obtain the bitonal sequence of the max register and the bitonal sequence of the min register.
[0040] In one possible implementation, the order of the elements of the two subsequences is swapped according to the bitonal sorting logic, including: Swap the elements of the two subsequences in order to obtain an increasing subsequence and a decreasing subsequence.
[0041] Taking two subsequences of length 1024 (corresponding to the data of two registers) as an example, the specific operations are as follows: ① Order adjustment: Swap the order of the elements of the two subsequences to ensure that the merged subsequences conform to the bitonic sorting logic. This will make the merged sequence appear to be increasing first and then decreasing (or vice versa), preparing for subsequent comparison and sorting.
[0042] ② Comparison and screening: Max and Min floating point instructions are used to compare data at corresponding positions in the two subsequences. These instructions compare elements at the same position in the two subsequences in parallel. The Max floating point instruction stores the larger value in the Max register, while the Min floating point instruction stores the smaller value in the Min register. This process fully utilizes the vector computing capabilities of the TPU, efficiently completing comparison operations on large amounts of data.
[0043] ③ Intra-block sorting (same as single-register-level sorting): After obtaining the max and min registers, the 1024 data in these two registers are sorted at the bottom level. The specific method is similar to that of processing a single register (1024 data). That is, for each register, the data is first separated by odd and even bits using instructions with the help of a mask, and stored in new temporary registers. Then, the Max floating point and Min floating point instructions are used to compare the data in the corresponding positions of these temporary registers, and the larger and smaller values are stored in the new max and min temporary registers, respectively. Then, the select instruction combined with the mask operation is used on the data of the new max and min temporary registers, so that the max temporary register only retains odd-bit data and the min temporary register only retains even-bit data. Finally, the data of these two temporary registers are merged to complete the further sorting of the 1024 data in each register.
[0044] Repeat the above steps to bitonalize the subsequences, merging two adjacent 1024-bit data blocks into a bitonal sequence. As the recursion continues, these bitonal sequences are gradually merged. Each merge involves sequence adjustment, comparison and filtering, and intra-block sorting, ultimately completing the bitonal sequence construction for the entire multi-register data.
[0045] In this embodiment, a dual-tuning structure is constructed by reversing the order of subsequences. This allows subsequent parallel comparisons of Max and Min floating point instructions to directly generate max and min sequences that satisfy the recursive merge condition. Independent single-register sorting is performed on the max and min sequences, fully leveraging the multi-core parallel capabilities of the TPU. This ensures sequence order while avoiding the serial bottlenecks found in traditional merge algorithms.
[0046] In one possible implementation, the input data is bitonally sorted and recursively divided into multiple subsequences, including: The input data is bitonally sorted and recursively divided into subsequences whose length is less than or equal to a set length; where the set length is equal to the storage capacity of a single TPU register.
[0047] The shorter bitonic sequences obtained through recursive partitioning retain the fundamental characteristics of a bitonic sequence: the elements are initially arranged in increasing order, then begin descending after a certain index (or vice versa). Because the partitioning is performed evenly by length, the shorter bitonic sequences have a similar distribution of elements to the original sequence, albeit on a smaller scale. These shorter bitonic sequences continue to be sorted through parallel comparison and swap operations during the subsequent sorting process, ultimately achieving the sorting of the entire dataset.
[0048] In this embodiment, the termination condition of the recursive partitioning is set to that the subsequence length is equal to the TPU register capacity, ensuring that each recursive node can directly call the hardware-optimized single-register sorting, avoiding the additional adaptation cost caused by unconventional data blocks.
[0049] In a possible implementation, before performing bitonic sorting on the input data and recursively dividing the data into multiple subsequences, the method further includes: Divide the input data into blocks of powers of 2; When there is a block whose data length is less than the set length, the input data is padded and filled to the set length; Assign a core id marking the initial position to each element in the supplemented input data.
[0050] Optionally, filling the input data includes: using a minimum value in the input data as a filling value, and filling the input data based on the filling value.
[0051] Optionally, assign a core ID to each element in the supplemented input data to mark its initial position, including: Assign positive values to the initial element core id in the input data; assign negative values to the core id of the fill value.
[0052] In other possible implementations, a core ID marking an initial position is assigned to each element in the supplemented input data, including: Assign values to all core ID elements in the input data in sequence, and then add a unified prefix or suffix to the core ID of the fill value.
[0053] For example, let's consider a TPU with a register capacity of 1024 (8×128-bit structure) and process an input data set (e.g., 5000 floating-point numbers). The algorithm requires the data size to be a power of 2 (e.g., 1024, 2048). However, the original data of 5000 does not meet this requirement, so the preprocessing process is initiated: Step 1: Calculate the power of 2 in blocks Calculate the minimum coverage of 5000 to the power of 2: 2 12 =4096<5000,2 13=8192>5000 Block splitting decision: prioritize allocating complete blocks → 8192 is too large and wasteful, so a combination strategy is chosen: 4 standard blocks: 4×1024=4096 elements 1 tail block: 5000-4096=904 elements This blocking strategy balances efficiency and resources: the complete block is directly aligned to the register, and the tail block is processed independently to avoid space waste.
[0054] Step 2: Tail block filling and core ID allocation Detect tail block 904 < register capacity 1024 Filling operation: add 1024-904=120 minima (-3.4×10 38 , TPU single floating point minimum) Core ID allocation design: The original element is assigned a positive integer id (0 to 4999) Filler elements are assigned negative ids (-1 to -120) Among them, the negative ID mechanism achieves double protection, clearly identifies invalid data, avoids polluting the sorting results, and maintains the element position traceability chain (for example, the fill value always maps the negative interval of the ID).
[0055] In this embodiment, a power-of-two block partitioning strategy is adopted to meet the recursive constraints of bitonic sorting. By filling insufficient blocks with data values and assigning negative core IDs, the integrity of the algorithm is maintained and invalid data is clearly identified. The negative core ID mechanism provides a precise anchor point for subsequent data cleaning, ensuring that the original data is not contaminated.
[0056] In a possible implementation, after adjusting the order of equal elements according to core ID, the following is further included: Remove padding values from input data based on core id.
[0057] As in the previous example, the original data of 5000 elements is divided into four full blocks (each containing 1024 elements) and one padded tail block (904 real data elements + 120 padded minimum values). All padded elements are assigned negative core IDs (-1 to -120). After the bitonic sort recursive operation is completed, the padded value removal phase begins.
[0058] Scan the data stream of the final sorted result (for example, the output sequence length is 6144) and detect the core id attribute of each element: a positive integer id (0-4999) identifies the original data, and a negative integer id (-1 to -120) identifies the padding value. Delete the element corresponding to the negative integer id directly, that is, remove the padding value in the input data.
[0059] In this example, we quickly identify and remove padding values based on negative core IDs, eliminating the impact of preprocessing on the final result and ensuring the purity of the output data. This operation forms a closed-loop management with the core ID system, validating the robustness of the stability mechanism in data addition and deletion scenarios.
[0060] In one possible implementation, the order of equal elements is adjusted according to the core ID, including: Determine whether the core id order of equal elements is consistent with the original order; When the judgment result is inconsistent, the positions of equal elements are swapped according to the core id.
[0061] In this embodiment, the stability deviation that may occur in the sorting process is corrected in a targeted manner through local verification and exchange of the core ID sequence of adjacent elements. This linear adjustment strategy only acts on a subset of equal elements, avoiding the waste of resources caused by global reordering, and ultimately achieving a strictly stable sorting.
[0062] It should be understood that the size of the serial numbers of the steps in the above embodiments does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.
[0063] The following are device embodiments of the present invention. For details not fully described therein, reference may be made to the corresponding method embodiments described above.
[0064] Figure 2 A schematic diagram of the structure of a tensor processor system provided by an embodiment of the present invention is shown. For ease of explanation, only the parts related to the embodiment of the present invention are shown, which are described in detail as follows: like Figure 2 As shown, the tensor processor system includes: multiple TPU registers, a vector instruction execution module and a core ID management module.
[0065] Among them, the storage capacity of each TPU register is the same.
[0066] The core id management module is used to assign a core id to each element in the input data to mark its initial position.
[0067] A vector instruction execution module is used to execute the data sorting method for a tensor processor provided by any of the aforementioned method embodiments.
[0068] In this embodiment, single-register processing or recursive partitioning and merging are used, depending on the data size, to fully utilize the parallel processing capabilities of the TPU and improve sorting efficiency. Recursive partitioning breaks down large data sets into small blocks suitable for TPU register processing, and the TPU's vector instructions are used for parallel comparison and exchange, enabling efficient implementation of the bitonic sorting algorithm on the TPU. By assigning a core ID to each element and passing the core ID during the sorting process, the instability problem of the traditional bitonic sorting algorithm is resolved, allowing equal elements to maintain their original order after sorting. Finally, the order of equal elements is adjusted based on the core ID, further ensuring the stability of the sorting, thereby meeting the requirements for stable sorting in application scenarios such as databases.
[0069] For the sake of convenience and brevity, the division of the above functional modules / units is only used as an example. In actual applications, the above functions can be assigned to different functional modules / units as needed. The above modules / units can be implemented in the form of hardware, software, or a combination of hardware and software.
[0070] An embodiment of the present invention further provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the methods in the above-mentioned method embodiments.
[0071] In the above embodiments, the descriptions of each embodiment have their own focus. For parts not described or recorded in detail in one embodiment, please refer to the relevant descriptions of other embodiments. Unless otherwise specified or there is a logical conflict, the terms and / or descriptions between different embodiments are consistent and can be referenced to each other. The technical features of different embodiments can be combined to form new embodiments based on their inherent logical relationships.
[0072] The embodiments described above are only used to illustrate the technical solutions of the present invention, rather than to limit the same. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. These modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present invention, and should all be included in the scope of protection of the present invention.
Claims
1. A data sorting method for a tensor processor, characterized in that: include: Obtain input data and the storage capacity of a single tensor processor (TPU) register, and assign a core ID marking an initial position to each element in the input data; When the amount of the input data is less than or equal to the storage capacity of a single TPU register, performing a single register level sorting operation to obtain a bitonic sequence of the single register; When the amount of the input data is greater than the storage capacity of a single TPU register, recursively dividing the input data into multiple subsequences by bitonic sorting, performing the single-register-level sorting operation on each subsequence, and merging the sorted subsequences to obtain a multi-register bitonic sequence; wherein, when performing the single-register-level sorting operation on each subsequence and merging the sorted subsequences, the core ID of the element is passed; After obtaining the bitune sequence, the order of equal elements is adjusted according to the core id.
2. The data sorting method for a tensor processor according to claim 1, wherein: The single register level sorting operation includes: Use mask instructions to separate data into two registers by odd and even bits; Use the Max floating point instruction and Min floating point instruction in the TPU vector instruction set to compare the elements of corresponding positions in the two registers, and store the larger and smaller values in the max register and min register respectively; The select instruction is combined with the mask operation to make the max register retain only the odd-numbered elements and the min register retain only the even-numbered elements, and the elements of these two registers are merged to form a bitune sequence.
3. The data sorting method for a tensor processor according to claim 1, wherein: The merging of the sorted subsequences includes: The order of the elements of the two subsequences is swapped according to the bitune sorting logic; Use the Max floating point instruction and Min floating point instruction in the TPU vector instruction set to compare the elements at the same position in the two subsequences, and store the larger and smaller values in the max register and min register respectively; The single register level sorting operation is performed on the elements in the max register and the min register respectively to obtain the bitonal sequence of the max register and the bitonal sequence of the min register.
4. The data sorting method for a tensor processor according to claim 3, wherein: The step of swapping the order of the elements of the two subsequences according to the bitonal sorting logic includes: Swap the elements of the two subsequences in order to obtain an increasing subsequence and a decreasing subsequence.
5. The data sorting method for a tensor processor according to claim 1, wherein: The bitonic sorting and recursively dividing the input data into a plurality of subsequences comprises: The input data is bitonally sorted and recursively divided into subsequences whose length is less than or equal to a set length; wherein the set length is equal to the storage capacity of a single TPU register.
6. The data sorting method for a tensor processor according to claim 5, wherein: Before performing bitonic sorting on the input data and recursively dividing the input data into a plurality of subsequences, the method further includes: Dividing the input data into blocks of powers of 2; When there is a block whose data length is less than the set length, padding the input data to the set length; Assign a core id marking the initial position to each element in the supplemented input data.
7. The data sorting method for a tensor processor according to claim 6, wherein: After adjusting the order of equal elements according to core ID, the following steps are further included: Remove padding values from the input data according to core id.
8. The data sorting method for a tensor processor according to claim 1, wherein: The adjustment of the order of equal elements according to coreid includes: Determine whether the core id order of equal elements is consistent with the original order; When the judgment result is inconsistent, the positions of equal elements are swapped according to the core id.
9. A tensor processor system, characterized in that: include: Multiple TPU registers, vector instruction execution module and core ID management module; Among them, the storage capacity of each TPU register is the same; The core id management module is used to assign a core id marking an initial position to each element in the input data; The vector instruction execution module is used to execute the method according to any one of claims 1 to 8.
10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the method according to any one of claims 1 to 8 is implemented.
Citation Information
Patent Citations
Data sorting processing device and method and storage medium
CN111913955A
Data processing method and device and computer readable storage medium
CN114356512A
Comparator network-based topK data parallel acquisition method and device
CN117215780A
FPGA (Field Programmable Gate Array)-based dual-tone sorting method and FPGA
CN118192928A
K-selection using parallel processing
EP3355207A1
Cited By
Storage bank sorting method, electronic equipment and readable storage medium
CN121722444A
Memory bank ordering method, electronic device, and readable storage medium
CN121722444B