Data processing circuit, data processing method and related product

By performing merge sorting on multiple data streams through data processing circuitry, the efficiency problem of sparse processing on devices with limited hardware resources is solved, and efficient data fusion and processing are achieved.

CN114691559BActive Publication Date: 2026-04-28ANHUI CAMBRICON INFORMATION TECH CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
ANHUI CAMBRICON INFORMATION TECH CO LTD
Filing Date
2020-12-25
Publication Date
2026-04-28

AI Technical Summary

Technical Problem

Existing hardware and/or instruction sets cannot effectively support sparsification and related processing, making it difficult to apply deep learning technology on devices with limited hardware resources.

Method used

A data processing circuit is provided, including a control circuit, a storage circuit, and a processing circuit, for performing merge sorting processing on multiple data streams to be merged, supporting data fusion operations, and simplifying and accelerating the processing.

Benefits of technology

By implementing merge sort in hardware, processing efficiency is improved, and the application of sparsed deep learning models on devices with limited hardware resources is supported.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114691559B_ABST
    Figure CN114691559B_ABST
Patent Text Reader

Abstract

The present disclosure discloses a data processing circuit, a data processing method and related products. The data processing circuit can be implemented as a computing device included in a combined processing device, which can further include an interface device and other processing devices. The computing device interacts with the other processing devices to jointly complete a user-specified computing operation. The combined processing device can further include a storage device connected to the computing device and the other processing devices, respectively, for storing data of the computing device and the other processing devices. The scheme of the present disclosure provides a hardware implementation of data fusion related operations, which can simplify processing and improve the processing efficiency of the machine.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure generally relates to the field of data processing. More specifically, this disclosure relates to data processing circuits, data processing methods, chips, and circuit boards. Background Technology

[0002] In recent years, the rapid development of deep learning has led to leaps in the performance of algorithms in fields such as computer vision and natural language processing. However, deep learning algorithms are computationally and storage-intensive tools. As information processing tasks become increasingly complex and the requirements for real-time performance and accuracy of algorithms continue to rise, neural networks are often designed to be deeper and deeper, resulting in ever-increasing computational and storage demands. This makes it difficult to directly apply existing deep learning-based artificial intelligence technologies to devices with limited hardware resources, such as mobile phones, satellites, or embedded devices.

[0003] Therefore, the compression, acceleration, and optimization of deep neural network models have become extremely important. Numerous studies have attempted to reduce the computational and storage requirements of neural networks without compromising model accuracy, which is of great significance for the engineering application of deep learning technology in embedded and mobile devices. Sparsity is one such method for lightweighting models.

[0004] Network parameter sparsification reduces redundant components in large networks through appropriate methods, thereby lowering the network's computational and storage requirements. Existing hardware and / or instruction sets cannot effectively support sparsification processing and / or related processing. Summary of the Invention

[0005] In order to at least partially solve one or more of the technical problems mentioned in the background art, the present disclosure provides a data processing circuit, a data processing method, a chip, and a board.

[0006] In a first aspect, this disclosure discloses a data processing circuit, including a control circuit, a storage circuit, and a processing circuit, wherein: the control circuit is configured to control the storage circuit and the processing circuit to perform merge sorting processing on multiple streams of data to be merged; the storage circuit is configured to store information, the information including at least information before and / or after the merge sorting processing; and the processing circuit is configured to, under the control of the control circuit, merge the multiple streams of data to be merged into one ordered merged data stream.

[0007] In a second aspect, this disclosure provides a chip that includes the data processing circuitry of any of the embodiments of the first aspect.

[0008] In a third aspect, this disclosure provides a board including the chip of any of the embodiments of the second aspect above.

[0009] In a fourth aspect, this disclosure provides a method for processing data using a data processing circuit, the data processing circuit including a control circuit, a storage circuit, and a computation circuit, the method comprising: the control circuit reading multiple streams of data to be merged from the storage circuit; the computation circuit merging the multiple streams of data to be merged into a single ordered merged data stream; and outputting the merged data stream to the storage circuit.

[0010] Through the data processing circuit, method for processing data using the data processing circuit, chip, and board provided above, this disclosure provides a hardware circuit supporting data fusion operations for performing merge sorting processing on multiple data streams. In some embodiments, the data processing circuit can merge multiple ordered data streams into a single ordered fused data stream. In other embodiments, the multiple data streams to be fused can be multiple indices that correspond one-to-one with multiple associated data streams. Each index element in the index indicates the index information of the corresponding associated data element in the associated data stream. While performing merge sorting processing on the multiple indexes, the binding relationship between each index element in the fused index and the corresponding data element in the fused associated data stream is maintained. Thus, the multiple associated data streams can be merged into a single fused associated data stream according to the index order. By providing a dedicated hardware implementation of data fusion-related operations, processing can be simplified and accelerated, thereby improving machine processing efficiency. Attached Figure Description

[0011] The above and other objects, features, and advantages of exemplary embodiments of this disclosure will become readily apparent upon reading the following detailed description with reference to the accompanying drawings. In the drawings, several embodiments of this disclosure are illustrated by way of example and not limitation, and like or corresponding reference numerals denote like or corresponding parts, wherein:

[0012] Figure 1 This is a structural diagram of the board card shown in the embodiment of this disclosure;

[0013] Figure 2 This is a structural diagram illustrating the combined processing apparatus according to an embodiment of the present disclosure;

[0014] Figure 3 This is a schematic diagram illustrating the internal structure of a processor core in a single-core or multi-core computing device according to embodiments of this disclosure;

[0015] Figure 4 This is an exemplary schematic diagram illustrating data fusion processing according to embodiments of this disclosure;

[0016] Figure 5 This is a schematic diagram illustrating the structure of a data processing apparatus according to an embodiment of the present disclosure;

[0017] Figure 6This is an exemplary circuit diagram illustrating a data fusion processing according to an embodiment of this disclosure;

[0018] Figure 7 This is an exemplary circuit diagram illustrating another embodiment of data fusion processing disclosed herein; and

[0019] Figure 8 This is an exemplary flowchart illustrating a data processing method according to an embodiment of this disclosure. Detailed Implementation

[0020] The technical solutions in the embodiments of this disclosure will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this disclosure, not all of them. Based on the embodiments in this disclosure, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this disclosure.

[0021] It should be understood that the terms "first," "second," "third," and "fourth," etc., that may appear in the claims, specification, and drawings of this disclosure are used to distinguish different objects, rather than to describe a specific order. The terms "comprising" and "including" as used in the specification and claims of this disclosure indicate the presence of the described features, integrals, steps, operations, elements, and / or components, but do not exclude the presence or addition of one or more other features, integrals, steps, operations, elements, components, and / or collections thereof.

[0022] It should also be understood that the terminology used in this disclosure is for the purpose of describing particular embodiments only and is not intended to limit the disclosure. As used in this disclosure and claims, the singular forms “a,” “an,” and “the” are intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term “and / or” as used in this disclosure and claims refers to any combination and all possible combinations of one or more of the associated listed items, and includes such combinations.

[0023] As used in this specification and claims, the term "if" may be interpreted, depending on the context, as "when," "once," "in response to determination," or "in response to detection."

[0024] The specific embodiments disclosed herein will now be described in detail with reference to the accompanying drawings.

[0025] Figure 1 A schematic diagram of the structure of a board 10 according to an embodiment of this disclosure is shown. Figure 1As shown, board 10 includes chip 101, which is a system-on-chip (SoC) integrating one or more combined processing units. These combined processing units are artificial intelligence computing units used to support various deep learning and machine learning algorithms, meeting the intelligent processing needs of complex scenarios in fields such as computer vision, speech, natural language processing, and data mining. In particular, deep learning technology is widely used in cloud intelligence. A significant characteristic of cloud intelligence applications is the large volume of input data, placing high demands on the platform's storage and computing capabilities. Board 10 in this embodiment is suitable for cloud intelligence applications, possessing massive off-chip storage, on-chip storage, and powerful computing capabilities.

[0026] Chip 101 is connected to external device 103 via external interface device 102. External device 103 may be, for example, a server, computer, camera, monitor, mouse, keyboard, network card, or Wi-Fi interface. Data to be processed can be transmitted from external device 103 to chip 101 via external interface device 102. The calculation results from chip 101 can be transmitted back to external device 103 via external interface device 102. Depending on the application scenario, external interface device 102 may have different interface forms, such as a PCIe interface.

[0027] The board 10 also includes a storage device 104 for storing data, which includes one or more memory cells 105. The storage device 104 is connected to and transmits data with the controller 106 and the chip 101 via a bus. The controller 106 in the board 10 is configured to regulate the state of the chip 101. Therefore, in one application scenario, the controller 106 may include a microcontroller (MCU).

[0028] Figure 2 This is a structural diagram illustrating the combined processing device in chip 101 of this embodiment. (As shown) Figure 2 As shown, the combined processing device 20 includes a computing device 201, an interface device 202, a processing device 203, and a storage device 204.

[0029] The computing device 201 is configured to perform user-specified operations. It is mainly implemented as a single-core intelligent processor or a multi-core intelligent processor to perform deep learning or machine learning calculations. It can interact with the processing device 203 through the interface device 202 to jointly complete the user-specified operations.

[0030] Interface device 202 is used to transmit data and control commands between computing device 201 and processing device 203. For example, computing device 201 can obtain input data from processing device 203 via interface device 202 and write it to on-chip storage device of computing device 201. Further, computing device 201 can obtain control commands from processing device 203 via interface device 202 and write them to on-chip control cache of computing device 201. Alternatively or optionally, interface device 202 can also read data from storage device of computing device 201 and transmit it to processing device 203.

[0031] Processing device 203, as a general-purpose processing device, performs basic control including but not limited to data transfer, and starting and / or stopping computing device 201. Depending on the implementation, processing device 203 may be one or more types of processors, including but not limited to digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc., and their number can be determined according to actual needs. As mentioned above, computing device 201 disclosed herein can be considered as having a single-core structure or a homogeneous multi-core structure. However, when computing device 201 and processing device 203 are considered together, they are considered to form a heterogeneous multi-core structure.

[0032] Storage device 204 is used to store data to be processed. It may be DRAM or DDR memory, typically 16G or larger in size, and is used to store data of computing device 201 and / or processing device 203.

[0033] Figure 3 The diagram shows the internal structure of the processor core when the computing device 201 is a single-core or multi-core device. The computing device 301 is used to process input data such as computer vision, speech, natural language, and data mining. The computing device 301 includes three main modules: a control module 31, an arithmetic module 32, and a storage module 33.

[0034] The control module 31 coordinates and controls the operation of the computation module 32 and the storage module 33 to complete the deep learning task. It includes an instruction fetch unit (IFU) 311 and an instruction decode unit (IDU) 312. The instruction fetch unit 311 fetches instructions from the processing device 203, and the instruction decode unit 312 decodes the fetched instructions and sends the decoding result as control information to the computation module 32 and the storage module 33.

[0035] The computation module 32 includes a vector operation unit 321 and a matrix operation unit 322. The vector operation unit 321 is used to perform vector operations and can support complex operations such as vector multiplication, addition, and nonlinear transformations; the matrix operation unit 322 is responsible for the core computations of deep learning algorithms, namely matrix multiplication and convolution.

[0036] Storage module 33 is used to store or move relevant data, including neuron RAM (NRAM) 331, weight RAM (WRAM) 332, and direct memory access (DMA) module 333. NRAM 331 is used to store input neurons, output neurons, and intermediate results after computation; WRAM 332 is used to store the convolution kernels of the deep learning network, i.e., the weights; DMA 333 is connected to DRAM 204 through bus 34 and is responsible for data transfer between computing device 301 and DRAM 204.

[0037] Based on the aforementioned hardware environment, the embodiments disclosed herein provide a data processing circuit that supports data fusion operations. As mentioned in the background section, network parameter sparsification can effectively reduce the network's computational and storage requirements. However, network parameter sparsification also brings a series of impacts to subsequent processing. For example, in sparse data processing, it may be necessary to merge and sort multiple sparse data streams; or, in sparse matrix multiplication, it may be necessary to sort and accumulate vectors. Therefore, the embodiments disclosed herein provide a hardware solution for data fusion processing to simplify and accelerate such processing.

[0038] Figure 4 This diagram illustrates an exemplary principle of data fusion processing according to embodiments of this disclosure. The diagram exemplarily shows four data streams, each containing six data elements, with the data elements in each stream arranged in a first order (e.g., ascending order). After data fusion, these four data streams are merged into a single fused data stream, comprising 24 data elements, and the merged data elements are arranged in a second order (e.g., ascending order). In this merge sorting process, duplicate data elements are preserved.

[0039] Those skilled in the art will understand that the first order and the second order may be the same or different, and both may be selected from either of the following: an order from smallest to largest, or an order from largest to smallest.

[0040] Those skilled in the art will also understand that the number of data elements in each data stream may be the same or different, and this disclosure does not limit this.

[0041] In the embodiments disclosed herein, data elements are scalars. Therefore, a data path that includes several data elements can be called a data vector, and the length of the vector is equal to the number of data elements it contains.

[0042] Figure 5 A structural block diagram of a data processing circuit 500 according to an embodiment of this disclosure is shown. The data processing circuit 500 can, for example, be implemented in... Figure 2 In the computing device 201, as shown in the figure, the data processing circuit 500 may include a control circuit 510, a storage circuit 520, and an arithmetic circuit 530.

[0043] The function of control circuit 510 can be similar to Figure 3 The control module 31 may include, for example, an instruction fetching unit for fetching instructions from, for example... Figure 2 The processing device 203 has instructions and an instruction decoding unit for decoding the acquired instructions and sending the decoding results as control information to the arithmetic circuit 530 and the storage circuit 520.

[0044] In one embodiment, the control circuit 510 may be configured to control the storage circuit 520 and the arithmetic circuit 530 to perform merge sorting processing on multiple streams of data to be merged.

[0045] The storage circuit 520 can be configured to store various types of information, including at least information before and / or after the merge sort process. The storage circuit could be, for example, […]. Figure 3 WRAM332.

[0046] The arithmetic circuit 530 can be configured to merge multiple data streams into a single ordered fused data stream under the control of the control circuit.

[0047] In one embodiment, the arithmetic circuit 530 may further include an arithmetic processing circuit (not shown), which may be configured to preprocess the data before the arithmetic circuit performs the operation or postprocess the data after the operation according to the arithmetic instructions. In some application scenarios, the aforementioned preprocessing and postprocessing may include, for example, data splitting and / or data concatenation operations.

[0048] There are many ways to implement arithmetic circuits. Figure 6 An exemplary circuit diagram for data fusion processing according to one embodiment of this disclosure is shown.

[0049] As shown in the figure, in one embodiment, the storage circuit can be exemplarily divided into two parts: a first storage circuit 622 and a second storage circuit 624.

[0050] The first storage circuit 622 can be configured to store K channels of data to be merged, where K>1, and the data elements of each of the K channels are arranged in a first order. An example is shown in the figure. Figure 4 The four data streams are shown. In some embodiments, each data stream is stored contiguously, for example, as a vector, so that the data stream / vector can be accessed based on the starting address of each data stream or the starting address of the vector.

[0051] The second storage circuit 624 can be configured to store the fused data output by the arithmetic circuit, and the data elements in the output fused data are arranged in a second order. As shown in the figure, the four data streams to be fused are transformed into one fused data stream, in which the fused data elements are arranged in ascending order, and data elements of the same size are retained and repeatedly output.

[0052] In some embodiments, the arithmetic circuit may include a comparison circuit 632 and a buffer circuit 634. The comparison circuit 632 performs a comparison function, comparing the sizes of data elements in the multiple streams of data to be merged, and submitting the comparison results to the control circuit 610 for sorting. The control circuit 610 determines the insertion position of the data element in the buffer circuit 634 based on the comparison results. The buffer circuit 634 is used to buffer the compared data elements, and buffers them in order of size.

[0053] Specifically, the comparison circuit 632 can be configured to compare data elements in the data to be fused with data elements not yet output in the buffer circuit 634, and output the comparison result to the control circuit 610. The buffer circuit 634 can be configured to, according to the control of the control circuit 610, orderly store the compared data elements, and orderly output the compared data elements as fused data.

[0054] In some embodiments, buffer circuit 634 can be configured to buffer K data elements, which are sorted by size. Those skilled in the art will understand that the buffer circuit can also be configured to buffer more data elements, and the embodiments disclosed herein are not limited in this respect. Depending on the sorting method in buffer circuit 634 and the desired output sorting method, such as ascending or descending, the first or last data element in the current sequence can be output in a specified order each time. For example, in the example in the figure, buffer circuit 634 buffers data elements from left to right in descending order, outputting the rightmost data element each time, which is the smallest data element in the current sequence, such as "7".

[0055] In these embodiments, the comparison circuit 632 may include a K-1 comparator configured to compare the data element to be fused with the data element not yet output in the buffer circuit 634, that is, with the K-1 data elements remaining after the first or last data element of the current sequence is output, generate a comparison result and output it to the control circuit 610.

[0056] For example, for four data streams to be merged, a three-way comparator is shown in the figure, which compares a specified data element (9 in this case) received from the first storage circuit 622 with three data elements that are not currently output from the buffer circuit 634. The three data elements on the left in the figure are 100, 10 and 9.

[0057] In some embodiments, the comparison result of the comparator can be represented using a bitmap. For example, if the data element to be merged (e.g., 9) is greater than or equal to the data element in the buffer circuit, the comparator can output "1", otherwise, it outputs "0"; and vice versa. In the example in the figure, the comparison result of the data element to be merged (9) with the respective data elements (100, 10 and 9) in the buffer circuit is "001", which is output to the control circuit 610.

[0058] The control circuit 610 can be configured to determine the insertion position of the data element to be fused in the current sequence of the buffer circuit 634 based on the received comparison result. Specifically, the control circuit 610 can be further configured to determine the insertion position based on the change position of the bits in the bitmap. In the example shown in the figure, the comparison result is "001", indicating that the current data element to be fused is less than the first and second data elements from the left in the buffer circuit, and greater than or equal to the third data element from the left. Therefore, the insertion position is between the second and third data elements, that is, between "10" and "9".

[0059] In some embodiments, the buffer circuit 634 may be configured to insert the data element to be merged at the insertion position according to the instruction of the control circuit 610. In the example shown in the figure, the sequence after the data element is inserted into the buffer circuit 634 becomes "100,10,9,9".

[0060] Next, the buffer circuit 634 can output the rightmost data element "9". At this time, the control circuit 610 can be further configured to determine the memory access information of the next data element to be merged based on the data element output from the buffer circuit. Specifically, the control circuit retrieves the next data element to be merged from which data path the output data element belongs in the K data path and sends it to the comparison circuit 632 for comparison.

[0061] For clarity, the figure also shows the data sequence buffered in buffer circuit 634 as the sorting progresses. As shown, initially, the first data element of each of the K data paths is stored in buffer circuit 634 in descending order. In some implementations, these four data elements can be retrieved, sorted, and stored in the buffer circuit all at once. In other implementations, the data in the buffer circuit can be initialized to negative numbers, and the first data element of each path can be retrieved sequentially (e.g., from path 1 to path 4), compared with the data in the buffer circuit, and placed in the appropriate position. In this example, the first data element of all four paths is 0, so they can be arranged according to the sequence number of each path based on the order of retrieval; for example, the "0" of path 1 is placed on the far right, the "0" of path 2 is placed in the second position from the right, and so on.

[0062] Next, the rightmost "0" belonging to channel 1 in the buffer circuit is output. Based on which channel this output data element belongs to, the next data element to be merged is retrieved from that channel, namely the second data element "2" in channel 1. "2" is sent to the comparison circuit and compared with the remaining three "0"s in the buffer circuit. The comparison result is "111", which is greater than all three existing "0"s in the buffer circuit. Therefore, "2" is inserted at the end of the sequence, and the sequence in the buffer circuit becomes "2,0,0,0".

[0063] Next, the rightmost "0" belonging to the second path in the buffer circuit is output. Therefore, the second element "3" of the second path is taken out and compared with the remaining "2,0,0" in the buffer circuit. The comparison result is "111", so "3" is inserted at the end of the sequence. At this time, the sequence in the buffer circuit becomes "3,2,0,0".

[0064] Next, the rightmost "0" belonging to the third path in the buffer circuit is output. Therefore, the second element of the third path, "100", is taken out and compared with the remaining "2,0,0" in the buffer circuit. The comparison result is "111", so "100" is inserted at the end of the sequence. At this time, the sequence in the buffer circuit becomes "100,3,2,0".

[0065] Next, the rightmost "0" belonging to the 4th channel of the output buffer circuit is taken out and the second element "2" of the 4th channel is compared with the remaining "100,3,2" in the buffer circuit. The comparison result is "001", so "2" is inserted after the rightmost first element of the sequence. At this time, the sequence in the buffer circuit becomes "100,3,2,2".

[0066] Similarly, data elements in the K-channel data can be compared one by one, sorted according to size, and inserted into the appropriate positions in the buffer circuit before being output by the buffer circuit. For example, the smallest data element output by the buffer circuit each time can be stored sequentially in the second storage circuit 624. Those skilled in the art will understand that if the buffer circuit has sufficient space, the merged and sorted data elements can also be output uniformly after the sorting is completed.

[0067] As can be seen from the merged sorted data elements, when data elements of the same size exist, the merged data still retains these elements of the same size, and no deduplication operation is performed. Therefore, combining the above... Figure 6 The detailed circuit diagram describes the merge sorting scheme provided in the embodiments disclosed herein.

[0068] In some application scenarios, the multiple data streams to be merged may be multiple indexes, and these indexes correspond one-to-one with multiple associated data streams. The index element in each index indicates the index information of the corresponding associated data element in that associated data stream. For example, in a sparse vector, data elements at certain positions are retained as valid data elements, while data elements at other positions are discarded or set to zero. The position information of these valid data elements in the vector before sparsification can be identified by the index. In these application scenarios, there may be multiple sparse data streams, such as multiple sparse vectors, and it is necessary to merge these multiple sparse data streams into a single data stream, where the data elements are ordered according to the index.

[0069] At this point, in addition to merging multiple indexes into a single ordered fused index, the data processing circuit in this embodiment is also configured to merge the multiple associated data into a single ordered fused associated data, and the order of data elements in the fused associated data remains consistent with the order of data elements in the fused index. That is, after the merge sorting process, the associated data and the index always maintain a one-to-one binding relationship.

[0070] Figure 7 An exemplary circuit diagram for data fusion processing according to another embodiment of this disclosure is shown. Figure 7 The data to be merged consists of K-way indexes and their corresponding K-way associated data. Figure 7 Implementation examples and Figure 6 The difference lies in the fact that, in addition to performing merge sort on the K-way index, a similar sorting process is also performed on the associated data. Those skilled in the art will understand that... Figure 7 The K-way index in the middle is equivalent to Figure 6 K-way data in [the context]. To avoid confusion, in [the context]... Figure 7 It uses the representation of K-way indexes and K-way related data.

[0071] As shown in the figure, the first storage circuit 722 stores not only the K-way indices to be merged, but also K-way associated data corresponding one-to-one with these K-way indices. As shown in the figure, the index elements of each of these K-way indices are arranged in a first order (e.g., from smallest to largest). The index element in each index indicates the index information of the corresponding associated data element in the corresponding associated data path. The figure exemplarily shows 4-way indices and corresponding 4-way associated data. As shown in the figure, the index of the first data element D11 of the first associated data path is 0, the index of the second data element D12 is 2, the index of the third data element D13 is 5, and so on. The index of the first data element D21 of the second associated data path is 0, the index of the second data element D22 is 3, and so on. In some embodiments, each index or each associated data path is stored contiguously, for example, as an index vector or associated data vector, so that the data / vector can be accessed according to the starting address of each data path or the starting address of the vector.

[0072] To ensure that the associated data and indexes maintain a one-to-one correspondence after merge sorting, in some embodiments, the buffer circuit 734 can be further configured to: orderly store the compared index elements and their corresponding associated data elements according to the value order of the index elements. As shown in the figure, the buffer circuit 734 caches not only the index elements but also their corresponding associated data elements. Therefore, after each comparison of the index elements to determine the insertion position, the associated data element corresponding to that index element can also be inserted into the buffer circuit. Those skilled in the art will understand that the associated data element can be the associated data element itself, such as D32, D23, etc., as exemplarily shown in the figure; the associated data element can also be an address pointing to that associated data element, and the embodiments disclosed herein are not limited in this respect.

[0073] Furthermore, during ordered output, the buffer circuit 734 can be configured to output the compared index elements in order of their values ​​(e.g., from smallest to largest) as fusion indexes, and simultaneously output their corresponding associated data elements as fusion associated data. The output data is stored, for example, in the second storage circuit 724. As shown, the output data can include two vectors: a fusion index vector and a fusion associated data vector.

[0074] The merged sorted data shows that when the multi-way index contains index elements of the same size, these index elements are repeatedly output in the merged index, and the associated data elements corresponding to these index elements are synchronously output in the merged associated data. Therefore, combining the above... Figure 7 The detailed circuit diagram describes a merge sorting scheme provided in another embodiment of this disclosure.

[0075] Figure 8 An exemplary flowchart of a data processing method 800 performed using the data processing circuit described above, according to an embodiment of this disclosure, is shown.

[0076] like Figure 8 As shown, in step 810, the control circuit reads multiple data streams to be merged from the storage circuit. Next, in step 820, the arithmetic circuit merges these multiple data streams into a single, ordered merged data stream. Finally, in step 830, the arithmetic circuit outputs the merged data to the storage circuit. Although the method steps are shown sequentially in the figure, these steps cycle through each data element, and some steps occur simultaneously within these cycles. For example, while the arithmetic circuit outputs data, the control circuit simultaneously accesses the storage circuit to read the next data element to be merged.

[0077] In some embodiments, the storage circuit may include a first storage circuit and a second storage circuit. The first storage circuit is configured to store K data streams to be fused, where K>1, wherein the data elements of each of the K data streams are arranged in a first order. The second storage circuit is configured to store the fused data output by the arithmetic circuit, and the data elements of the output fused data are arranged in a second order.

[0078] In some embodiments, the first order and the second order may be the same or different, and the first order and the second order are selected from either: an ascending order or a descending order. For example, multiple streams of data arranged in ascending order can be merged into a single stream of merged data arranged in descending order, or merged into a single stream of merged data arranged in ascending order.

[0079] In some embodiments, the arithmetic circuit may include a comparison circuit and a buffer circuit. In this case, the data processing method may further include: the comparison circuit comparing data elements in the multiple streams of data to be fused with data elements not yet output in the buffer circuit, and outputting the comparison result to the control circuit; and under the control of the control circuit, the buffer circuit orderly storing the compared data elements, and orderly outputting the compared data elements as fused data.

[0080] In some embodiments, the comparison circuit may include a K-1 comparator, and the data processing method may include: the K-1 comparator comparing the data element to be fused with the K-1 data elements of the current sequence in the buffer circuit, generating a comparison result and outputting it to the control circuit.

[0081] Furthermore, the control circuit can determine the insertion position of the data element to be merged in the current sequence of the buffer circuit based on the comparison result. In one implementation, the comparison result can be represented using a bitmap, and the control circuit can determine the insertion position of the data element to be merged based on the changes in the bits in the bitmap.

[0082] Subsequently, the buffer circuit can insert the data element to be merged at the determined insertion position according to the instruction of the control circuit. Then, the buffer circuit can output the first or last data element in the current sequence in a specified order (e.g., from smallest to largest).

[0083] At this point, the control circuit can determine the memory access information of the next data element to be fused based on the data elements output from the buffer circuit, and read the data element from the storage circuit and send it to the comparison circuit for the next round of comparison.

[0084] As mentioned earlier, in some application scenarios, the multi-path data to be merged includes multiple indexes, and these indexes correspond one-to-one with multiple related data paths. The index element in each index indicates the index information of the corresponding related data element in the corresponding related data path. In this case, a similar sorting operation needs to be performed on the related data elements according to their index correspondence.

[0085] In these embodiments, the computing circuit merges these multiple associated data streams into a single ordered fused associated data stream, wherein the order of data elements in the fused associated data stream is consistent with the order of data elements (i.e., index elements) in the fused data stream (i.e., fused index).

[0086] Specifically, the buffer circuit can store the compared index elements and their corresponding associated data elements in order of their index element values; and output the compared index elements as fused data in order of their index element values, and synchronously output their corresponding associated data elements as fused associated data.

[0087] In the merge sorting process of this embodiment, when there are index elements of the same size in the multi-way index, the index elements of the same size are repeatedly output in the merged data, and the associated data elements corresponding to these index elements are simultaneously output in the merged associated data. In other words, the merge sorting process of this embodiment does not deduplicate duplicate index elements, but retains all original elements.

[0088] In some applications, the data elements in the aforementioned multi-path correlation data can be valid data elements in a sparse matrix, and the index information is used to indicate the position information of these valid data elements in the sparse matrix.

[0089] Those skilled in the art will understand that each step of the above method corresponds to the respective circuits described above in conjunction with the example circuit diagrams. Therefore, the features described above can be applied equally to the method steps, and will not be repeated here.

[0090] As described above, this disclosure provides a hardware circuit for performing data fusion operations related to merge sorting. By implementing merge sorting in hardware, processing speed can be accelerated, thereby better supporting operations related to sparsified processing, such as sparse matrix multiplication. In some embodiments, this hardware circuit can merge multiple ordered data streams into a single ordered fused data stream. In other embodiments, the multiple data streams to be merged can be multiple indices corresponding one-to-one with multiple associated data streams. Each index element in the index indicates the index information of the corresponding associated data element in the associated data stream. While performing merge sorting on the multiple indexes, the binding relationship between each index element in the merged index and the corresponding data element in the merged associated data stream is maintained. Thus, the multiple associated data streams can also be merged into a single fused associated data stream according to the order of the indices.

[0091] Depending on the application scenario, the electronic devices or apparatus disclosed herein may include servers, cloud servers, server clusters, data processing devices, robots, computers, printers, scanners, tablets, smart terminals, PC devices, IoT terminals, mobile terminals, mobile phones, dashcams, navigators, sensors, cameras, video cameras, projectors, watches, headphones, mobile storage, wearable devices, visual terminals, autonomous driving terminals, vehicles, home appliances, and / or medical devices. The vehicles include airplanes, ships, and / or vehicles; the home appliances include televisions, air conditioners, microwave ovens, refrigerators, rice cookers, humidifiers, washing machines, lights, gas stoves, and range hoods; the medical devices include MRI scanners, ultrasound machines, and / or electrocardiographs. The electronic devices or apparatus disclosed herein can also be applied in fields such as the Internet, IoT, data centers, energy, transportation, public management, manufacturing, education, power grids, telecommunications, finance, retail, construction sites, and healthcare. Furthermore, the electronic devices or apparatus disclosed herein can also be used in application scenarios related to artificial intelligence, big data, and / or cloud computing, such as cloud computing, edge computing, and terminal applications. In one or more embodiments, the high-computing-power electronic devices or apparatuses according to the present disclosure can be applied to cloud devices (e.g., cloud servers), while the low-power electronic devices or apparatuses can be applied to terminal devices and / or edge devices (e.g., smartphones or cameras). In one or more embodiments, the hardware information of the cloud devices and the hardware information of the terminal devices and / or edge devices are compatible with each other, so that suitable hardware resources can be matched from the hardware resources of the cloud devices to simulate the hardware resources of the terminal devices and / or edge devices based on the hardware information of the terminal devices and / or edge devices, so as to complete the unified management, scheduling and collaborative work of end-to-cloud or cloud-edge-end integration.

[0092] It should be noted that, for the sake of brevity, this disclosure describes some methods and their embodiments as a series of actions and combinations thereof. However, those skilled in the art will understand that the solutions disclosed herein are not limited by the order of the described actions. Therefore, based on the disclosure or teachings of this document, those skilled in the art will understand that some steps can be performed in a different order or simultaneously. Furthermore, those skilled in the art will understand that the embodiments described in this disclosure can be considered optional embodiments, that is, the actions or modules involved are not necessarily essential for the implementation of one or more solutions disclosed herein. In addition, depending on the solution, the description of some embodiments in this disclosure may have different emphases. In view of this, those skilled in the art will understand that parts not described in detail in a certain embodiment of this disclosure can also be referred to the relevant descriptions of other embodiments.

[0093] In terms of specific implementation, based on the disclosure and teachings of this document, those skilled in the art will understand that several embodiments disclosed herein can also be implemented in other ways not disclosed herein. For example, regarding the various units in the electronic device or apparatus embodiments described above, this document has decomposed them based on logical functions, but in actual implementation, there may be other decomposition methods. As another example, multiple units or components can be combined or integrated into another system, or some features or functions in a unit or component can be selectively disabled. Regarding the connection relationships between different units or components, the connections discussed above in conjunction with the accompanying drawings can be direct or indirect couplings between units or components. In some scenarios, the aforementioned direct or indirect couplings involve communication connections utilizing interfaces, where the communication interface can support electrical, optical, acoustic, magnetic, or other forms of signal transmission.

[0094] In this disclosure, the units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units. The aforementioned components or units may be located in the same location or distributed across multiple network units. Furthermore, depending on actual needs, some or all of the units can be selected to achieve the purpose of the solution described in the embodiments of this disclosure. Additionally, in some scenarios, multiple units in the embodiments of this disclosure may be integrated into one unit or each unit may exist physically independently.

[0095] In other implementation scenarios, the integrated units described above can also be implemented in hardware, i.e., as specific hardware circuits, which may include digital circuits and / or analog circuits. The physical implementation of the circuit's hardware structure may include, but is not limited to, physical devices, which may include, but are not limited to, transistors or memristors. Therefore, the various devices described herein (e.g., computing devices or other processing devices) can be implemented using appropriate hardware processors, such as central processing units, GPUs, FPGAs, DSPs, and ASICs. Furthermore, the aforementioned storage units or storage devices can be any suitable storage medium (including magnetic storage media or magneto-optical storage media), such as resistive random access memory (RRAM), dynamic random access memory (DRAM), static random access memory (SRAM), enhanced dynamic random access memory (EDRAM), high-bandwidth memory (HBM), hybrid memory cube (HMC), ROM, and RAM.

[0096] The foregoing can be better understood in accordance with the following terms:

[0097] Clause 1. A data processing circuit, comprising a control circuit, a storage circuit, and an arithmetic circuit, wherein:

[0098] The control circuit is configured to control the storage circuit and the arithmetic circuit to perform merge sorting processing on multiple streams of data to be merged.

[0099] The storage circuit is configured to store information, which includes at least information before and / or after the merge sort process; and

[0100] The computing circuit is configured to, under the control of the control circuit, merge the multiple data streams to be fused into a single ordered fused data stream.

[0101] Clause 2. The data processing circuit according to Clause 1, wherein the storage circuit includes a first storage circuit and a second storage circuit.

[0102] The first storage circuit is configured to store the K channels of data to be fused, where K>1, and the data elements of each channel of data are arranged in a first order; and

[0103] The second storage circuit is configured to store the fused data output by the arithmetic circuit, and the data elements in the output fused data are arranged in a second order.

[0104] Clause 3. The data processing circuit according to Clause 2, wherein the first order is the same as or different from the second order, and the first order and the second order are selected from either: an ascending order or an ascending order.

[0105] Clause 4. The data processing circuit according to any one of Clauses 1-3, wherein the arithmetic circuit includes a comparison circuit and a buffer circuit, wherein:

[0106] The comparison circuit is configured to compare data elements in the multi-channel data to be fused with data elements not yet output in the buffer circuit, and output the comparison result to the control circuit; and

[0107] The buffer circuit is configured to, under the control of the control circuit, orderly store the compared data elements and orderly output the compared data elements as the fused data.

[0108] Clause 5. The data processing circuit according to Clause 4, wherein the comparison circuit includes:

[0109] A K-1 comparator is configured to compare the data element to be fused with the K-1 data elements of the current sequence in the buffer circuit, generate a comparison result, and output it to the control circuit.

[0110] Clause 6. The data processing circuit according to Clause 5, wherein the control circuit is configured to determine, based on the comparison result, the insertion position of the data element to be fused in the current sequence of the buffer circuit.

[0111] Clause 7. The data processing circuit according to Clause 6, wherein the comparison result is represented using a bitmap, and the control circuit is further configured to: determine the insertion position based on the positional changes of bits in the bitmap.

[0112] Clause 8. A data processing circuit according to any one of Clauses 6-7, wherein the buffer circuit is configured to insert the data element to be merged at the insertion position according to the instruction of the control circuit.

[0113] Clause 9. The data processing circuit according to any one of Clauses 4-8, wherein the buffer circuit is further configured to output the first or last data element in the current sequence in a specified order.

[0114] Clause 10. The data processing circuit according to Clause 9, wherein the control circuit is further configured to: determine memory access information of the next data element to be fused based on the data element output from the buffer circuit.

[0115] Clause 11. The data processing circuit according to any one of Clauses 4-10, wherein the multiple data streams include multiple indexes, each index corresponding one-to-one with multiple associated data streams, each index element indicating the index information of the corresponding associated data element in the associated data stream, and the arithmetic circuit is further configured to:

[0116] The multiple associated data are merged into one ordered fused associated data, wherein the order of data elements in the fused associated data is consistent with the order of data elements in the fused data.

[0117] Clause 12. The data processing circuit according to Clause 11, wherein the buffer circuit is further configured to:

[0118] The compared index elements and their corresponding associated data elements are stored in ordered order according to the value sequence of the index elements; and

[0119] The compared index elements are output in order of their values ​​as the fused data, and their corresponding associated data elements are output synchronously as the fused associated data.

[0120] Clause 13. The data processing circuit according to Clause 12, wherein when there are index elements of the same size in the multiplexed index, the index elements of the same size are repeatedly output in the fused data, and the associated data elements corresponding to these index elements are synchronously output in the fused associated data.

[0121] Clause 14. The data processing circuit according to any one of Clauses 11-13, wherein the data elements in the multi-path associated data are valid data elements in the sparse matrix, and the index information indicates the position information of the valid data elements in the sparse matrix.

[0122] Clause 15. A chip comprising data processing circuitry as described in any one of Clauses 1-14.

[0123] Clause 16. A board including the chip described in Clause 15.

[0124] Clause 17. A method for processing data using a data processing circuit, said data processing circuit including a control circuit, a storage circuit, and an arithmetic circuit, said method comprising:

[0125] The control circuit reads multiple data streams to be merged from the storage circuit;

[0126] The computing circuit merges the multiple data streams to be fused into a single ordered fused data stream; and

[0127] The fused data is output to the storage circuit.

[0128] Clause 18. The method according to Clause 17, wherein the storage circuit includes a first storage circuit and a second storage circuit.

[0129] The first storage circuit is configured to store the K channels of data to be fused, where K>1, and the data elements of each channel of data are arranged in a first order; and

[0130] The second storage circuit is configured to store the fused data output by the arithmetic circuit, and the data elements in the output fused data are arranged in a second order.

[0131] Clause 19. The method described in Clause 18, wherein the first order is the same as or different from the second order, and the first order and the second order are selected from either: an ascending order or an ascending order.

[0132] Clause 20. The method according to any one of Clauses 17-19, wherein the arithmetic circuit includes a comparison circuit and a buffer circuit, and the method further includes:

[0133] The comparison circuit compares the data elements in the multiple data streams to be fused with the data elements not yet output in the buffer circuit, and outputs the comparison result to the control circuit; and

[0134] Under the control of the control circuit, the buffer circuit stores the compared data elements in an orderly manner and outputs the compared data elements in an orderly manner as the fused data.

[0135] Clause 21. The method according to Clause 20, wherein the comparison circuit includes a K-1 channel comparator, and the method includes:

[0136] The K-1 comparator compares the data element to be fused with the K-1 data elements of the current sequence in the buffer circuit, generates a comparison result, and outputs it to the control circuit.

[0137] Clause 22. The method described pursuant to Clause 21 further includes:

[0138] The control circuit determines the insertion position of the data element to be fused in the current sequence of the buffer circuit based on the comparison result.

[0139] Clause 23. The method according to Clause 22, wherein the comparison result is represented using a bitmap, and the method further includes: the control circuit determining the insertion position based on the changing positions of bits in the bitmap.

[0140] Clause 24. The method described under any of Clauses 22-23 further includes:

[0141] The buffer circuit inserts the data element to be merged at the insertion position according to the instruction of the control circuit.

[0142] Clause 25. The method described under any one of Clauses 20-24 further includes:

[0143] The buffer circuit outputs the first or last data element in the current sequence in a specified order.

[0144] Clause 26. The method described pursuant to Clause 25 further includes:

[0145] The control circuit determines the memory access information of the next data element to be merged based on the data elements output from the buffer circuit.

[0146] Clause 27. The method according to any one of Clauses 20-26, wherein the multi-path data includes a multi-path index, the multi-path index corresponding one-to-one with multi-path associated data, the index element in each path index indicating the index information of the corresponding associated data element in the corresponding path of associated data, and the method further includes:

[0147] The computing circuit merges the multiple associated data into one ordered fused associated data, wherein the order of data elements in the fused associated data is consistent with the order of data elements in the fused data.

[0148] Clause 28. The method described pursuant to Clause 27 further includes:

[0149] The buffer circuit stores the compared index elements and their corresponding associated data elements in an ordered manner according to the value order of the index elements; and

[0150] The compared index elements are output in order of their values ​​as the fused data, and their corresponding associated data elements are output synchronously as the fused associated data.

[0151] Clause 29. The method according to Clause 28, wherein when there are index elements of the same size in the multi-way index, the index elements of the same size are repeatedly output in the fused data, and the associated data elements corresponding to these index elements are synchronously output in the fused associated data.

[0152] Clause 30. The method according to any one of Clauses 27-29, wherein the data elements in the multipath associated data are valid data elements in the sparse matrix, and the index information indicates the position information of the valid data elements in the sparse matrix.

[0153] The embodiments of this disclosure have been described in detail above. Specific examples have been used to illustrate the principles and implementation methods of this disclosure. The descriptions of the above embodiments are only for the purpose of helping to understand the methods and core ideas of this disclosure. At the same time, for those skilled in the art, there will be changes in the specific implementation methods and application scope based on the ideas of this disclosure. Therefore, the content of this specification should not be construed as a limitation of this disclosure.

Claims

1. A data processing circuit, comprising a control circuit, a storage circuit, and an arithmetic circuit, wherein: The control circuit is configured to control the storage circuit and the arithmetic circuit to perform merge sorting processing on multiple streams of data to be merged. The storage circuit is configured to store information, which includes at least information before and / or after the merge sort process; as well as The computing circuit is configured to, under the control of the control circuit, merge the multiple data streams to be fused into a single ordered fused data stream. The multi-channel data includes multi-channel indexes, each corresponding one-to-one with multiple channels of associated data. Each index element indicates the index information of the corresponding associated data element within that channel. Furthermore, the computation circuit is configured to: The multi-way indexes are merged into one ordered fusion index, and the multi-way associated data are merged into one ordered fusion associated data. The order of data elements in the fusion associated data is consistent with the order of data elements in the fusion data. The data elements in the multi-way associated data are valid data elements in the sparse matrix, and the index information is used to indicate the position information of these valid data elements in the sparse matrix.

2. The data processing circuit according to claim 1, wherein the storage circuit includes a first storage circuit and a second storage circuit. The first storage circuit is configured to store the K channels of data to be fused, where K>1, and the data elements of each channel of data are arranged in a first order; and The second storage circuit is configured to store the fused data output by the arithmetic circuit, and the data elements in the output fused data are arranged in a second order.

3. The data processing circuit according to claim 2, wherein the first order is the same as or different from the second order, and the first order and the second order are selected from either: an order from smallest to largest, or an order from largest to smallest.

4. The data processing circuit according to any one of claims 2-3, wherein the arithmetic circuit includes a comparison circuit and a buffer circuit, wherein: The comparison circuit is configured to compare data elements in the multi-channel data to be fused with data elements not yet output in the buffer circuit, and output the comparison result to the control circuit; and The buffer circuit is configured to, under the control of the control circuit, orderly store the compared data elements and orderly output the compared data elements as the fused data.

5. The data processing circuit according to claim 4, wherein the comparison circuit comprises: A K-1 comparator is configured to compare the data element to be fused with the K-1 data elements of the current sequence in the buffer circuit, generate a comparison result, and output it to the control circuit.

6. The data processing circuit of claim 5, wherein the control circuit is configured to determine, based on the comparison result, the insertion position of the data element to be fused in the current sequence of the buffer circuit.

7. The data processing circuit of claim 6, wherein the comparison result is represented using a bitmap, and the control circuit is further configured to: determine the insertion position based on the change position of the bits in the bitmap.

8. The data processing circuit of claim 7, wherein the buffer circuit is configured to insert the data element to be merged at the insertion position according to the instruction of the control circuit.

9. The data processing circuit of claim 8, wherein the buffer circuit is further configured to output the first or last data element in the current sequence in a specified order.

10. The data processing circuit according to claim 9, wherein the control circuit is further configured to: determine the memory access information of the next data element to be fused based on the data element output in the buffer circuit.

11. The data processing circuit of claim 4, wherein the buffer circuit is further configured to: The compared index elements and their corresponding associated data elements are stored in ordered order according to the value sequence of the index elements; and The compared index elements are output in order of their values ​​as the fused data, and their corresponding associated data elements are output synchronously as the fused associated data.

12. The data processing circuit according to claim 11, wherein when there are index elements of the same size in the multi-path index, the index elements of the same size are repeatedly output in the fused data, and the associated data elements corresponding to these index elements are synchronously output in the fused associated data.

13. A chip comprising a data processing circuit according to any one of claims 1-12.

14. A circuit board comprising the chip according to claim 13.

15. A method for processing data using a data processing circuit, the data processing circuit comprising a control circuit, a storage circuit, and an arithmetic circuit, the method comprising: The control circuit reads multiple data streams to be merged from the storage circuit; The computing circuit merges the multiple data streams to be fused into a single ordered fused data stream. as well as The fused data is output to the storage circuit. The multi-path data includes a multi-path index, which corresponds one-to-one with multiple paths of associated data. Each path index contains an index element that indicates the index information of the corresponding associated data element within that path. The method further includes: The computational circuit merges the multiple indexes into a single ordered fusion index and merges the multiple associated data into a single ordered fusion associated data. The order of data elements in the fusion associated data is consistent with the order of data elements in the fusion data. The data elements in the multiple associated data are valid data elements in the sparse matrix, and the index information indicates the position information of the valid data elements in the sparse matrix.

16. The method of claim 15, wherein the storage circuit comprises a first storage circuit and a second storage circuit. The first storage circuit is configured to store the K channels of data to be fused, where K>1, and the data elements of each channel of data are arranged in a first order; and The second storage circuit is configured to store the fused data output by the arithmetic circuit, and the data elements in the output fused data are arranged in a second order.

17. The method of claim 16, wherein the first order is the same as or different from the second order, and the first order and the second order are selected from either: an ascending order or an ascending order.

18. The method according to any one of claims 16-17, wherein the arithmetic circuit includes a comparison circuit and a buffer circuit, and the method further includes: The comparison circuit compares the data elements in the multiple data streams to be fused with the data elements that have not been output in the buffer circuit, and outputs the comparison result to the control circuit. as well as Under the control of the control circuit, the buffer circuit stores the compared data elements in an orderly manner and outputs the compared data elements in an orderly manner as the fused data.

19. The method of claim 18, wherein the comparison circuit comprises a K-1 channel comparator, and the method comprises: The K-1 comparator compares the data element to be fused with the K-1 data elements of the current sequence in the buffer circuit, generates a comparison result, and outputs it to the control circuit.

20. The method of claim 19, further comprising: The control circuit determines the insertion position of the data element to be fused in the current sequence of the buffer circuit based on the comparison result.

21. The method of claim 20, wherein the comparison result is represented using a bitmap, and the method further comprises: The control circuit determines the insertion position based on the changes in the bit positions in the bitmap.

22. The method of claim 21, further comprising: The buffer circuit inserts the data element to be merged at the insertion position according to the instruction of the control circuit.

23. The method of claim 22, further comprising: The buffer circuit outputs the first or last data element in the current sequence in a specified order.

24. The method of claim 23, further comprising: The control circuit determines the memory access information of the next data element to be merged based on the data elements output from the buffer circuit.

25. The method of claim 18, further comprising: The buffer circuit stores the compared index elements and their corresponding associated data elements in an orderly manner according to the value order of the index elements. as well as The compared index elements are output in order of their values ​​as the fused data, and their corresponding associated data elements are output synchronously as the fused associated data.

26. The method according to claim 25, wherein when there are index elements of the same size in the multi-way index, the index elements of the same size are repeatedly output in the fused data, and the associated data elements corresponding to these index elements are synchronously output in the fused associated data.

Citation Information

Patent Citations

  • Method and system for data transfer

    CN104021123A

  • Parallel sorting method of multiple groups of ordered sequences

    CN105045600A

  • Device, method and application supporting vector ordering

    CN108733352A

  • Data processing device, data processing method and related product

    CN114692840A