Data processing device, data processing method and related products
By providing a data processing device and method on a device with limited hardware resources, fusion processing of multi-channel data is achieved, the efficiency problem of sparse processing is solved, and processing efficiency and resource utilization are improved.
Patent Information
- Application Number
- CN202011563214.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-12-25
- Publication Date
- 2025-09-23
- Estimated Expiration
- 2040-12-25
AI Technical Summary
Existing hardware and/or instruction sets cannot effectively support sparsification and post-sparseness related processing, making it difficult to apply deep learning models on devices with limited hardware resources.
Provided are a data processing device and method, which implement the merge, sort, and fusion of multiple data channels through fusion instructions, use specialized hardware circuits to merge multiple data channels into one channel of ordered fused data in index order, and use operation structures to represent data elements with the same index, thereby simplifying and accelerating processing.
It simplifies the processing flow of deep learning models after sparsification, improves the processing efficiency of devices with limited hardware resources, and supports the efficient execution of deep learning tasks on embedded devices.
Smart Images

Figure CN114692839B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure generally relates to the field of data processing, and more specifically, to a data processing device, a data processing method, a chip, and a board. Background Art
[0002] In recent years, the rapid development of deep learning has led to significant advances in algorithm performance across a range of fields, including computer vision and natural language processing. However, deep learning algorithms are computationally and storage-intensive. As information processing tasks become increasingly complex and the demands for real-time and accurate algorithms continue to rise, neural networks are often designed to be deeper, increasing the computational load and storage requirements. This makes existing deep learning-based artificial intelligence technologies difficult to directly apply to hardware-constrained mobile phones, satellites, or embedded devices.
[0003] Therefore, the compression, acceleration, and optimization of deep neural network models have become extremely important. Numerous studies have attempted to reduce the computational and storage requirements of neural networks without compromising model accuracy. This is of crucial importance for the engineering application of deep learning technology on embedded and mobile devices. Sparsification is one such method for achieving model lightweighting.
[0004] Network parameter sparsification is the process of reducing redundant components in a larger network through appropriate methods to reduce the network's computational load and storage requirements. Existing hardware and / or instruction sets cannot effectively support sparsification and / or post-sparseness related processing. Summary of the Invention
[0005] In order to at least partially solve one or more technical problems mentioned in the background technology, the solution disclosed herein provides a data processing device, a data processing method, a chip and a board.
[0006] In a first aspect, the present disclosure discloses a data processing device, comprising: a control circuit configured to parse a fusion instruction, wherein the fusion instruction instructs to perform fusion processing on multiple data to be fused; a storage circuit configured to store information before and / or after the fusion processing; and an operation circuit configured to merge the data elements in the multiple data to be fused into one ordered fusion-processed data according to their corresponding indexes according to the fusion instruction, wherein data elements with the same index are merged into an operation structure element, and the data element includes any of scalars, vectors or higher-dimensional data.
[0007] In a second aspect, the present disclosure provides a chip comprising the data processing device of any embodiment of the aforementioned first aspect.
[0008] In a third aspect, the present disclosure provides a board comprising the chip of any one of the embodiments of the second aspect.
[0009] In a fourth aspect, the present disclosure provides a data processing method, comprising: parsing a fusion instruction, the fusion instruction instructing to perform fusion processing on multiple data to be fused; according to the fusion instruction, merging the data elements in the multiple data to be fused into one channel of ordered fusion-processed data according to their corresponding indexes, wherein data elements with the same index are merged into operation structure elements, and the data elements include any of scalars, vectors or higher-dimensional data; and outputting the fusion-processed data.
[0010] Through the data processing device, data processing method, chip, and board provided above, embodiments of the present disclosure provide a fusion instruction for performing operations related to merge-sort fusion of multiple data channels. In some embodiments, the fusion instruction is a hardware instruction that implements data fusion processing via specialized hardware circuitry. In some embodiments, the data processing device can, based on the fusion instruction, merge multiple channels of ordered data into one channel of ordered fused data in index order. Data elements with the same index can be merged and represented in the form of an operation structure, thereby facilitating subsequent computational processing. In some embodiments, the fusion instruction can include an operation mode bit to indicate that the fusion instruction is a merge-sort fusion operation, or the fusion instruction itself can indicate a merge-sort fusion operation. In some embodiments, the data elements to be fused can be vectors or higher-dimensional data. For example, the data elements in the data to be fused can be valid data elements after sparsification in radar-based object detection. Thus, the data fusion instructions and fusion operations provided by embodiments of the present disclosure can support related processing in radar algorithms. By providing specialized fusion instructions to perform operations related to fusion processing of multiple channels of data, processing can be simplified. Furthermore, by providing specialized hardware implementations of data fusion-related operations, processing can be accelerated, thereby improving the processing efficiency of the machine. BRIEF DESCRIPTION OF THE DRAWINGS
[0011] The above and other objects, features and advantages of the exemplary embodiments of the present disclosure will become readily understood by reading the detailed description below with reference to the accompanying drawings. In the accompanying drawings, several embodiments of the present disclosure are shown in an illustrative and non-limiting manner, and the same or corresponding reference numerals represent the same or corresponding parts, wherein:
[0012] Figure 1 is a structural diagram showing a board according to an embodiment of the present disclosure;
[0013] Figure 2 is a structural diagram showing a combined processing device according to an embodiment of the present disclosure;
[0014] Figure 3is a schematic diagram showing the internal structure of a processor core of a single-core or multi-core computing device according to an embodiment of the present disclosure;
[0015] Figure 4 is an exemplary schematic diagram illustrating data fusion processing according to an embodiment of the present disclosure;
[0016] Figure 5 is a schematic diagram showing the structure of a data processing device according to an embodiment of the present disclosure;
[0017] Figure 6 is an exemplary circuit diagram illustrating a data fusion process according to an embodiment of the present disclosure;
[0018] Figure 7 is an example of an operation structure showing another embodiment of the present disclosure;
[0019] Figure 8 is an example of an operation structure showing yet another embodiment of the present disclosure;
[0020] Figure 9 The following example shows the content pointed to by each address in the fusion instruction;
[0021] Figure 10 A schematic diagram showing a data storage space according to an embodiment of the present disclosure;
[0022] Figure 11 A schematic diagram showing data blocks in a data storage space according to an embodiment of the present disclosure;
[0023] Figure 12 A block diagram showing the structure of a data processing device according to another embodiment of the present disclosure; and
[0024] Figure 13 An exemplary flow chart of a data processing method according to an embodiment of the present disclosure is shown. DETAILED DESCRIPTION
[0025] The following will clearly and completely describe the technical solutions in the embodiments of this disclosure in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of this disclosure, not all of them. Based on the embodiments of this disclosure, all other embodiments obtained by those skilled in the art without creative work are within the scope of protection of this disclosure.
[0026] It should be understood that the terms "first," "second," "third," and "fourth," etc., which may appear in the claims, specification, and drawings of this disclosure, are used to distinguish different objects rather than to describe a specific order. The terms "include" and "comprising" used in the specification and claims of this disclosure indicate the presence of the described features, integers, steps, operations, elements, and / or components, but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or combinations thereof.
[0027] It should also be understood that the terminology used in this disclosure is for the purpose of describing specific embodiments only and is not intended to limit the disclosure. As used in this disclosure and the claims, the singular forms "a," "an," and "the" are intended to include the plural forms unless the context clearly indicates otherwise. It should also be further understood that the term "and / or" as used in this disclosure and the claims refers to any and all possible combinations of one or more of the associated listed items, including and including these combinations.
[0028] As used in this specification and claims, the term "if" may be interpreted as "when" or "upon" or "in response to determining" or "in response to detecting," depending on the context.
[0029] The specific embodiments of the present disclosure are described in detail below with reference to the accompanying drawings.
[0030] Figure 1 FIG. 1 is a schematic diagram showing the structure of a board 10 according to an embodiment of the present disclosure. Figure 1 As shown, board 10 includes chip 101, which is a system-on-chip (SoC), or system-on-chip, integrated with one or more combined processing devices. The combined processing device is an artificial intelligence computing unit that supports various deep learning and machine learning algorithms to meet the intelligent processing needs in complex scenarios in fields such as computer vision, speech, natural language processing, and data mining. In particular, deep learning technology is widely used in the field of cloud intelligence. A notable feature of cloud-based intelligent applications is the large amount of input data, which places high demands on the platform's storage and computing capabilities. Board 10 of this embodiment is suitable for cloud-based intelligent applications and has extensive off-chip storage, on-chip storage, and powerful computing capabilities.
[0031] Chip 101 is connected to external device 103 via external interface device 102. External device 103 may be, for example, a server, computer, camera, monitor, mouse, keyboard, network card, or Wi-Fi interface. Data to be processed can be transmitted from external device 103 to chip 101 via external interface device 102. Calculation results from chip 101 can be transmitted back to external device 103 via external interface device 102. Depending on the application scenario, external interface device 102 may have different interface formats, such as a PCIe interface.
[0032] Board 10 also includes a memory device 104 for storing data, which includes one or more storage units 105. Memory device 104 is connected to a control device 106 and chip 101 via a bus for data transmission. Control device 106 in board 10 is configured to control the state of chip 101. To this end, in one application scenario, control device 106 may include a microcontroller (MCU).
[0033] Figure 2 FIG. 1 is a block diagram showing the combined processing device in the chip 101 of this embodiment. Figure 2 As shown in , the combined processing device 20 includes a computing device 201 , an interface device 202 , a processing device 203 and a storage device 204 .
[0034] The computing device 201 is configured to perform user-specified operations and is mainly implemented as a single-core intelligent processor or a multi-core intelligent processor to perform deep learning or machine learning calculations. It can interact with the processing device 203 through the interface device 202 to jointly complete the user-specified operations.
[0035] Interface device 202 is used to transmit data and control instructions between computing device 201 and processing device 203. For example, computing device 201 can obtain input data from processing device 203 via interface device 202 and write it to a storage device on-chip of computing device 201. Furthermore, computing device 201 can obtain control instructions from processing device 203 via interface device 202 and write them to a control cache on-chip of computing device 201. Alternatively or optionally, interface device 202 can also read data from the storage device of computing device 201 and transmit it to processing device 203.
[0036] The processing device 203 is a general processing device that performs basic controls including but not limited to data handling, starting and / or stopping the computing device 201. Depending on the implementation, the processing device 203 can be a central processing unit (CPU), a graphics processing unit (GPU) or one or more types of processors in other general and / or special processors, which include but are not limited to digital signal processors (DSP), application specific integrated circuits (ASIC), field programmable gate arrays (FPGA) or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components, etc., and their number can be determined according to actual needs. As mentioned above, with respect to the computing device 201 disclosed herein, it can be regarded as having a single-core structure or a homogeneous multi-core structure. However, when the computing device 201 and the processing device 203 are integrated and considered together, the two are regarded as forming a heterogeneous multi-core structure.
[0037] The storage device 204 is used to store data to be processed, which may be DRAM or DDR memory, and is typically 16G or larger in size, for storing data of the computing device 201 and / or the processing device 203 .
[0038] Figure 3 The figure shows the internal structure of the processor core when the computing device 201 is a single-core or multi-core device. The computing device 301 is used to process input data for computer vision, speech, natural language, data mining, etc. The computing device 301 includes three modules: a control module 31, a computing module 32, and a storage module 33.
[0039] The control module 31 coordinates and controls the operations of the computing module 32 and the storage module 33 to complete deep learning tasks. It includes an instruction fetch unit (IFU) 311 and an instruction decode unit (IDU) 312. The instruction fetch unit 311 retrieves instructions from the processing device 203, while the instruction decode unit 312 decodes the retrieved instructions and sends the decoded results as control information to the computing module 32 and the storage module 33.
[0040] The operation module 32 includes a vector operation unit 321 and a matrix operation unit 322. The vector operation unit 321 is used to perform vector operations, supporting complex operations such as vector multiplication, addition, and nonlinear transformations. The matrix operation unit 322 is responsible for the core calculations of the deep learning algorithm, namely matrix multiplication and convolution.
[0041] The storage module 33 is used to store or transfer relevant data and includes a neuron RAM (NRAM) 331, a weight RAM (WRAM) 332, and a direct memory access (DMA) module 333. NRAM 331 stores input neurons, output neurons, and intermediate computational results; WRAM 332 stores the convolution kernels (i.e., weights) of the deep learning network. DMA 333, connected to DRAM 204 via bus 34, is responsible for data transfer between the computing device 301 and DRAM 204.
[0042] The embodiments of the present disclosure are based on the aforementioned hardware environment and provide a data processing device that supports data fusion operations. As mentioned in the background technology, network parameter sparsification can effectively reduce the network's demand for computing power and storage space. However, the sparsification of network parameters will also have a series of impacts on subsequent processing. For example, in a sparse matrix multiplication operation, it may be necessary to sort and accumulate the vectors obtained in the middle of the operation to obtain the expected operation results. For another example, in a radar algorithm, it is necessary to perform fusion processing on the sparse data in radar-based object detection. In view of this, the embodiments of the present disclosure provide a hardware solution for data fusion processing to simplify and accelerate such processing.
[0043] Figure 4 The exemplary principle of data fusion processing according to the embodiment of the present disclosure is shown. The figure exemplarily shows 4 channels of data to be fused, and each channel of data includes 6 data elements. The data elements can be scalars, vectors or tensors of higher dimensions. The data elements in the figure are exemplarily shown as vectors, such as D11, D12, ... D46. These vectors have a uniform vector length, for example, D11 is (d1, d2, d3, ..., dn), and the length is n. Each data element has an associated index for indicating the position information of the data element in the corresponding channel of data. For example, the original channel of data may include 1,000 data elements, but only the data elements at some positions are valid. At this time, these valid elements can be extracted to form the above-mentioned data to be fused, and the indexes corresponding to these valid elements are extracted to indicate their positions in the original data. These indexes form the above-mentioned indexes to be fused.
[0044] The figure schematically shows the corresponding 4-way index to be merged, and each index corresponds to a path of data to be merged. The 1st path index is used to identify the position information of each data element in the 1st path data, the 2nd path index is used to identify the position information of each data element in the 2nd path data, and so on. Furthermore, the index elements in each path index are stored in order and correspond one-to-one with the data elements in the corresponding path data. In the example in the figure, the index elements in each path index are arranged in the first order (for example, from small to large order), and the data elements in each path data are also arranged in order according to the order of the corresponding index. For example, the 1st index element in the 1st path index indicates that the index of the 1st data element in the 1st path data is 0, that is, the first element; the 2nd index element in the 1st path index indicates that the index of the 2nd data element in the 1st path data is 2, that is, the third element; and so on.
[0045] After data fusion, these 4 data are merged into one channel of ordered fused data according to their corresponding indexes, and the data elements with the same index are merged into one fused data element. As shown in the figure, the fused index includes 16 index elements, which are arranged in a second order (for example, from small to large), wherein the repeated index elements in the index to be fused are removed, as shown by the dark squares in the figure. Correspondingly, the fused data also includes 16 data elements, which are arranged in order according to the order of the corresponding indexes, and the data elements with the same index are merged, as shown by the dark squares in the figure. Since the data elements may be vectors or tensors of higher dimensions, in some embodiments of the present disclosure, at least the merging of data elements with the same index can be represented in the form of operation structure elements.
[0046] In this example, the merging of data elements with the same index is schematically represented by an addition formula. For example, for the fused index element "0", its corresponding fused data element is the sum of the first data elements of each data path (D11+D21+D31+D41). For another example, for the fused index element "9", its corresponding fused data element is the sum of the 5th data element of the 1st path and the 3rd data element of the 4th path (D15+D43). When the data element is a vector, the fused data element is the corresponding vector sum. The specific representation of the fused data element will be described later.
[0047] Those skilled in the art will appreciate that the first and second orders mentioned above may be the same or different, and both may be selected from either of the following: an order from small to large, or an order from large to small. Those skilled in the art will also appreciate that, although the figures show that each path of data has an equal number of data elements, the number of data elements in each path of data may be the same or different, and the present disclosure is not limited in this respect.
[0048] Figure 5 FIG. 5 shows a structural block diagram of a data processing device 500 according to an embodiment of the present disclosure. The data processing device 500 may be implemented, for example, in Figure 2 As shown in the figure, the data processing device 500 may include a control circuit 510, a storage circuit 520 and an operation circuit 530.
[0049] The function of the control circuit 510 may be similar to Figure 3 The control module 31 may include, for example, an instruction fetch unit for obtaining an instruction from, for example, Figure 2 The processing device 203 receives the instruction, and the instruction decoding unit is used to decode the acquired instruction and send the decoding result as control information to the operation circuit 530 and the storage circuit 520.
[0050] In one embodiment, the control circuit 510 may be configured to parse a fusion instruction, where the fusion instruction indicates to perform fusion processing on multiple data to be fused.
[0051] The storage circuit 520 may be configured to store various information, including at least information before and / or after fusion processing. The storage circuit may be, for example, Figure 3 WRAM 332.
[0052] The operation circuit 530 can be configured to perform corresponding operations according to the fusion instruction. Specifically, the operation circuit 530 can merge the data elements in multiple channels of data to be fused into a channel of ordered fused data according to their corresponding indexes, where data elements with the same index are merged into an operation structure element. The data elements can be any of scalars, vectors, or higher-dimensional data.
[0053] In one embodiment, the operation circuit 530 may further include an operation processing circuit (not shown), which may be configured to pre-process the data before the operation circuit performs the operation or post-process the data after the operation according to the operation instruction. In some application scenarios, the aforementioned pre-processing and post-processing may, for example, include data splitting and / or data splicing operations.
[0054] There are many ways to implement the arithmetic circuit. Figure 6 An exemplary circuit diagram for data fusion processing according to one embodiment of the present disclosure is shown.
[0055] As shown in the figure, in one embodiment, the storage circuit can be exemplarily divided into two parts: a first storage circuit 622 and a second storage circuit 624 .
[0056] The first storage circuit 622 can be configured to store K paths of data to be fused and K path indexes corresponding to the K paths of data, where K>1. The index elements in the K path indexes indicate index information of corresponding data elements in the K path of data, i.e., the index elements and data elements have a one-to-one correspondence. In addition, the index elements of each path index in the K path indexes are arranged in order according to a first sequence, and the data elements of each path of data in the K path of data are arranged in order according to the order of the corresponding indexes.
[0057] The figure shows an example Figure 4 The 4-way index and the corresponding 4-way data are shown. Since the data elements may be vectors or higher-dimensional data, for the sake of simplicity, the storage addresses of these data elements can be used instead. In the figure, each data element is identified by the symbol Pt, which represents a pointer to the storage address of the corresponding data element (such as a vector, a three-dimensional tensor, etc.). It can be understood that the specific data elements will also be stored in the storage circuit, but for the sake of clarity, they are not shown in the figure. In some embodiments, each index can be stored continuously, for example, as a vector, so that the index / index vector can be accessed according to the starting address of each index or the starting address of the vector. Correspondingly, each data can also be stored continuously, for example, as a vector, so that the data / data vector can be accessed according to the starting address of each data or the starting address of the vector, but in this case, the vector element in the data vector is a pointer to the final data element.
[0058] The second storage circuit 624 can be configured to store the fused data output by the operation circuit. This data includes: a fused index obtained by sorting and merging the K-way indexes, fused data, and an operation structure. The fused index elements in the fused index are arranged in an orderly manner according to the second order, and the fused data elements in the fused data correspond one-to-one with the fused index elements and are arranged in an orderly manner according to the order of the fused indexes. The operation structure includes at least an operation structure element for representing a fused data element composed of data elements with the same index.
[0059] As can be seen from the example in the figure, the 4-way data to be fused becomes one-way fused data, and the corresponding 4-way index also becomes one-way fused index, wherein the fused index elements are arranged in order from small to large, and the index elements of the same size are removed. The corresponding fused data elements are arranged in the order of the fused index, and each fused data element can be an address pointer pointing to the corresponding final data element. In some embodiments, the storage space of each final data element can be a fixed size. Thus, when these data elements are stored continuously, the subsequent address can be determined by offsetting the first address by a fixed amount. For example, the figure exemplifies that the first fused data element can be the address base_addr. It is assumed that the address size occupied by each final data element is offset, then the second fused data element can be the address base_addr+offset, the third fused data element can be the address base_addr+2*offset, and so on.
[0060] Among the final data elements pointed to by these addresses, some do not require accumulation because their indexes are unique among the indexes to be fused, and are therefore original data elements; while some require accumulation because there are multiple original data elements with the same index. In the embodiments disclosed herein, at least for fused data elements composed of data elements with the same index, the final data element pointed to by them can be obtained without immediate calculation, but instead represented by the associated calculation structure element in the calculation structure. The calculation structure will be described in detail later in conjunction with specific circuits.
[0061] In some embodiments, the operation circuit may include a sorting circuit 632 and an output circuit 636 to collaboratively implement the sorting, accumulation, and fusion functions. Specifically, the sorting circuit 632 is configured to sort the K-way indexes according to the size of the index elements and output them in an orderly manner to the output circuit 636. The output circuit 636 may, at least when receiving the same index element from the sorting circuit, generate an operation structure element representing the accumulation operation of the data elements corresponding to the same index element and remove duplicate index elements.
[0062] In some embodiments, the sorting circuit 632 may include a comparison circuit 631 and a buffer circuit 633. The comparison circuit 631 implements a comparison function, comparing the sizes of index elements in multiple indices to be merged and submitting the comparison results to the control circuit 610 for sorting. The control circuit 610 determines the insertion position of the index element in the buffer circuit 633 based on the comparison results. The buffer circuit 633 is used to cache the compared index elements and the corresponding data elements, and caches them in order of the size of the index elements.
[0063] Specifically, the comparison circuit 631 may be configured to compare the index elements in the to-be-fused index with the index elements not outputted by the buffer circuit 633, and output the comparison result to the control circuit 610. The buffer circuit 633 may be configured to sequentially store the compared index elements and the information of the corresponding data elements under the control of the control circuit 610, and sequentially output the compared index elements and the information of the corresponding data elements.
[0064] In some embodiments, the buffer circuit 633 can be configured to cache K index elements, and these K index elements are sorted by size. It will be understood by those skilled in the art that the buffer circuit can also be configured to cache more index elements, and the disclosed embodiments are not limited in this respect. Depending on the sorting method in the buffer circuit 633 and the sorting method of the desired output, such as from small to large, or from large to small, the first index element or the last index element in the current sequence can be output in the specified order each time. For example, in the example in the figure, the buffer circuit 633 caches the index elements from left to right in order from large to small, and outputs the rightmost index element each time, which is the smallest index element in the current sequence, such as "7".
[0065] In these embodiments, the comparison circuit 631 may include K-1 comparators, configured to compare the index elements to be fused with the index elements not output in the buffer circuit 633, that is, to compare with the K-1 index elements remaining after the first or last index element is output in the current sequence, generate a comparison result and output it to the control circuit 610.
[0066] For example, for the 4-way data to be fused, the figure shows a 3-way comparator, which compares the specified index element (9 at this time) received from the first storage circuit 622 with the 3 index elements that are not currently output from the buffer circuit 633, which are the three index elements 100, 10 and 9 on the left in the figure.
[0067] In some embodiments, the comparison result of the comparator can be represented using a bitmap. For example, if the index element to be merged (e.g., 9) is greater than or equal to the index element in the buffer circuit, the comparator can output "1", otherwise, it outputs "0"; the reverse is also possible. In the example shown in the figure, the comparison result of the index element to be merged (9) and the various index elements in the buffer circuit (100, 10, and 9) is "001", which is output to the control circuit 610.
[0068] The control circuit 610 can be configured to determine the insertion position of the index element to be merged in the current sequence of the buffer circuit 633 based on the received comparison result. Specifically, the control circuit 610 can be further configured to determine the insertion position based on the changing position of the bit in the bitmap. In the example in the figure, the comparison result is "001", indicating that the current index element to be merged is smaller than the first and second index elements from the left in the buffer circuit, and greater than or equal to the third index element from the left. Therefore, the insertion position is between the second and third index elements, that is, between "10" and "9".
[0069] In some embodiments, the buffer circuit 633 may be configured to insert the index element to be merged into the insertion position according to the instruction of the control circuit 610. In the example shown in the figure, the sequence after the index element is inserted into the buffer circuit 633 becomes "100, 10, 9, 9".
[0070] In order to enable the data corresponding to the index to be obtained during the fusion process, in some embodiments, the buffer circuit 633 can be further configured to: orderly store the compared index elements and the corresponding data elements in the order of the index element values. As shown in the figure, in addition to caching the index elements, the buffer circuit 633 also caches the information of the data elements corresponding to them. Therefore, after each comparison of the index elements to determine the insertion position, the data element information corresponding to the index element can also be inserted into the cache circuit. Since data elements can be scalars, vectors, or higher-dimensional tensors, the address pointing to the data element can be used to represent the data element. For example, in the example in the figure, each element in the K-way data to be fused is an address, pointing to the corresponding data element, regardless of whether the data element is a scalar, vector, or high-dimensional tensor. In the description herein, a data element may refer to an address or to the final scalar, vector, or high-dimensional tensor data. Those skilled in the art can distinguish their meanings based on the context of the description.
[0071] Next, buffer circuit 633 can output the rightmost index element "9." At this point, control circuit 610 can be further configured to determine the memory access information for the next index element to be merged based on the index element output from the buffer circuit. Specifically, the control circuit retrieves the next index element to be merged from the K-way index based on which way the output index element belongs to, and sends it to comparison circuit 631 for comparison.
[0072] Furthermore, during ordered output, the buffer circuit 633 can be configured to sequentially output the compared index elements as fused indexes in order of index element values (e.g., from smallest to largest), and simultaneously output the corresponding data elements as fused data. The output data is provided to the output circuit 636 for further processing.
[0073] For clarity, the figure also shows the index sequence cached in the buffer circuit 633 as the sorting progresses. As shown in the figure, initially, the first index element of each way in the K-way index is stored in the buffer circuit 633 in descending order. In some implementations, these four index elements can be taken out, sorted, and stored in the buffer circuit at one time. In other implementations, the data in the buffer circuit can be initialized to a negative number, and the first index element of each way index can be taken out one by one in order (for example, in order from the 1st way to the 4th way), compared with the data in the buffer circuit, and placed in the appropriate position. In this example, the first index element of the 4-way index is 0, so it can be arranged according to the serial number of each way index based on the order of number retrieval, for example, the "0" of the 1st way is placed on the far right, the "0" of the 2nd way is placed on the second position from the right, and so on.
[0074] Next, the rightmost "0" in the buffer circuit, belonging to the first path, is output. Based on the path index to which this output index element belongs, the next index element to be merged is extracted from the corresponding path index, namely, the second index element "2" in the first path. "2" is fed into the comparator circuit and compared with the three remaining "0"s in the buffer circuit. The comparison result is "111", which is greater than the three existing "0"s in the buffer circuit. Therefore, "2" is inserted at the end of the sequence, and the sequence in the buffer circuit now becomes "2, 0, 0, 0".
[0075] Next, the rightmost "0" belonging to the second path in the buffer circuit is output, so the second element "3" of the second path is taken out and compared with the remaining "2,0,0" in the buffer circuit. The comparison result is "111", so "3" is inserted at the end of the sequence. At this time, the sequence in the buffer circuit becomes "3,2,0,0".
[0076] Next, the rightmost "0" belonging to the third path in the buffer circuit is output, so the second element "100" of the third path is taken out and compared with the remaining "2, 0, 0" in the buffer circuit. The comparison result is "111", so "100" is inserted at the end of the sequence. At this time, the sequence in the buffer circuit becomes "100, 3, 2, 0".
[0077] Next, the rightmost "0" belonging to the 4th path is output from the buffer circuit, and the second element "2" of the 4th path is taken out and compared with the remaining "100, 3, 2" in the buffer circuit. The comparison result is "001", so "2" is inserted after the first element of the rightmost sequence. At this time, the sequence in the buffer circuit becomes "100, 3, 2, 2".
[0078] By analogy, the index elements in the K-way index can be compared one by one, sorted by size, inserted into the appropriate position of the buffer circuit, and then output by the buffer circuit. For example, the smallest index element output by the buffer circuit each time can be output in sequence to the output circuit 636. Those skilled in the art will appreciate that if the buffer circuit has sufficient space, the merged and sorted elements can also be outputted uniformly after sorting is completed.
[0079] It can be seen from the merged and sorted index elements output that when there are index elements of the same size, the sorting circuit 632 still retains these index elements of the same size and does not perform deduplication operations, but provides them to the output circuit 636 for processing.
[0080] In some embodiments, the output circuit 636 may include a comparator 637 , a buffer 635 , and a structure generator 639 .
[0081] The comparator 637 may be configured to compare the index elements sequentially output from the sorting circuit 632 with the previous merged index element and output a comparison result, which may be "1" indicating the same or "0" indicating the different, or vice versa.
[0082] Buffer 635 can be configured to control the output of index elements based on the comparison result of comparator 637. In some embodiments, buffer 635 can output the current index element as a new merged index element only when the comparison result indicates a difference. In other words, when the comparison result indicates a difference, buffer 635 does not output the current index element, i.e., it discards the index element that is duplicated with the previous merged index element. As shown in the figure, the merged index in second storage circuit 624 does not contain any duplicate merged index elements.
[0083] The structure generator 639 may be configured to control the accumulation of data elements and generate corresponding operation structure elements based on the comparison result of the comparator 637. Specifically, at least when the comparison result indicates a match, an operation structure element is generated based on the data element corresponding to the index element. The operation structure element represents an accumulation operation of accumulating the current data element to the fused data element corresponding to the last fused index element.
[0084] Figure 6A method for generating a calculation structure element is shown in an embodiment. In this embodiment, in addition to generating a calculation structure element for a fused data element formed by merging data elements with the same index, a calculation structure element is also generated for other fused data elements. Specifically, in this embodiment, when the comparison result of the comparator 637 indicates that they are different, a calculation structure element is generated based on the data element corresponding to the current index element, where the calculation structure element represents an accumulation operation of adding the current data element to 0. In this way, a calculation structure element can be generated for all fused data elements, thereby unifying the expression method, simplifying the operation, and providing flexible processing for subsequent operations.
[0085] Figure 6 The calculation structure generated under the current embodiment is shown in the figure. As shown in the figure, when the output circuit 636 receives the first "0" output by the sorting circuit 632, since this is the first element, the comparator 637 will output, for example, "0" to indicate that it is different from the previous fused index element (for example, initialized to a negative number). The buffer 635 outputs the index "0" as the first fused index element. The structure generator 639 assigns an address to the new fused data element based on the result of the comparator 637. For example, the address of the first fused data element is base_addr. At this time, even if the index is considered to be different, an associated calculation structure element is generated for the fused data element. It is nothing more than an accumulation operation of adding the current data element to 0, for example, {add_zero, Pt11, base_addr}, where add_zero represents the address of the value 0, Pt11 represents the address of the current data element, and base_addr represents the storage address of the result, which corresponds to the address of the first fused data element assigned above. The generated calculation structure element can be stored in the calculation structure of the second storage circuit 624.
[0086] Each operation structure element may include three elements, indicating two addends and an addition result, for example, {src0_addr, src1_addr, dst_addr}, where src0_addr represents the address of the first addend, src1_addr represents the address of the second addend, and dst_addr represents the address of the addition sum.
[0087] Next, the sorting circuit 632 outputs the second "0". Since the previous fusion index element is "0", the comparator 637 will output, for example, "1" to indicate that it is the same as the previous fusion index element ("0"). At this time, the buffer 635 does not output. Based on the result of the comparator 637, the structure generator 639 believes that no new fusion data element is generated, so no address is allocated, that is, there is no new element in the fusion data. At this time, the structure generator 639 will also generate an operation structure element representing the accumulation of data elements, which accumulates the data elements to the fusion data elements corresponding to the previous fusion index element, such as {base_addr, Pt21, base_addr}, where base_addr represents the address of the fusion data element corresponding to the previous fusion index element, that is, the address of the first fusion data element allocated in the previous step, Pt21 represents the address of the current data element, and base_addr represents the storage address of the result, that is, the data elements are accumulated to the first fusion data element allocated above. The generated operation structure elements are also stored in the operation structure of the second storage circuit 624 .
[0088] When the sorting circuit 632 continues to output the third and fourth "0", the fusion index will not increase, and the fusion data elements in the corresponding fusion data will not increase. The structure generator will output the associated structure operation elements: {base_addr, Pt31, base_addr}, {base_addr, Pt41, base_addr}.
[0089] Next, the sorting circuit 632 outputs the first "2". Referring to the previous step, the buffer 635 outputs the index "2" as the second fused index element. The structure generator 639 assigns an address to the new fused data element based on the result of the comparator 637, for example, adding an offset to the address of the first fused data element. The offset is the storage address size of the corresponding final data element (such as a scalar, vector or high-dimensional tensor). At this time, an associated operation structure element is generated for the new fused data element, such as {add_zero, Pt12, base_addr+offset}. The generated operation structure element is stored in the operation structure of the second storage circuit 624.
[0090] When the sorting circuit 632 outputs the second "2", the fusion index will not increase, and the fusion data element in the corresponding fusion data will not increase. The structure generator will output the associated structure operation element: {base_addr+offset, Pt42, base_addr+offset}.
[0091] According to the above description, those skilled in the art can similarly deduce other fusion results, which will not be described one by one here.
[0092] Figure 7 An example of a calculation structure element generated according to another embodiment is shown. In this embodiment, for data elements whose indexes do not have duplicates in the data to be fused, the calculation structure element may not be generated, thereby avoiding redundant calculations.
[0093] Specifically, when the comparator 637 indicates that the currently output index element is the same as the previous fused index element, an operation structure element may be generated to indicate that the currently corresponding data element is added to the fused data element corresponding to the previous fused index element. When the comparator 637 indicates that the currently output index element is different from the previous fused index element, the current index element may be directly output as the new fused index element, and an address may be directly assigned to the new fused data element, where the corresponding original data element is stored.
[0094] The following combination Figure 6 Data example description Figure 7 The process of generating the elements of the operation structure in the calculation. As shown in the figure, when the output circuit 636 receives the first "0" output by the sorting circuit 632, since this is the first element, the comparator 637 will output, for example, "0" to indicate that it is different from the previous fusion index element (for example, initialized to a negative number). The buffer 635 outputs the index "0" as the first fusion index element. The structure generator 639 assigns an address to the new fusion data element based on the result of the comparator 637. For example, the address of the first fusion data element is base_addr. At this time, since the indexes are different, there is no need to accumulate, and the original data element corresponding to the index element can be directly moved to the assigned address base_addr. This method can ensure that the address of the fused data is continuous, which can facilitate subsequent access.
[0095] Next, the sorting circuit 632 outputs the second "0". Since the previous fusion index element is "0", the comparator 637 will output, for example, "1" to indicate that it is the same as the previous fusion index element ("0"). At this time, the buffer 635 does not output. Based on the result of the comparator 637, the structure generator 639 believes that no new fusion data element is generated, so no address is allocated, that is, there is no new element in the fusion data. At this time, the structure generator 639 will generate an operation structure element representing the accumulation of data elements, which accumulates the data elements to the fusion data elements corresponding to the previous fusion index element, such as {base_addr, Pt21, base_addr}, where base_addr represents the address of the fusion data element corresponding to the previous fusion index element, that is, the address of the first fusion data element allocated in the previous step. Please note that at this time, base_addr has already been placed in the data element pointed to by Pt11 in the previous step. Pt21 represents the address of the current data element, base_addr represents the storage address of the result, that is, the data element is added to the first fused data element allocated above. The generated operation structure element is stored in the operation structure of the second storage circuit.
[0096] When the sorting circuit 632 continues to output the third and fourth "0", the fusion index will not increase, and the fusion data elements in the corresponding fusion data will not increase. The structure generator will output the associated structure operation elements: {base_addr, Pt31, base_addr}, {base_addr, Pt41, base_addr}.
[0097] Next, the sorting circuit 632 outputs the first "2". Referring to the previous step, the buffer 635 outputs the index "2" as the second fused index element. The structure generator 639 assigns an address to the new fused data element based on the result of the comparator 637, for example, adding an offset to the address of the first fused data element. The offset is the storage address size of the corresponding final data element (such as a scalar, vector, or high-dimensional tensor). At this time, since the indexes are different, there is no need to accumulate, and the original data element corresponding to the index element (that is, the data element pointed to by Pt12) can be directly moved to the allocated address base_addr+offset.
[0098] When the sorting circuit 632 outputs the second "2", the fusion index will not increase, and the fusion data element in the corresponding fusion data will not increase. The structure generator will output the associated structure operation element: {base_addr+offset, Pt42, base_addr+offset}.
[0099] According to the above description, those skilled in the art can similarly deduce other fusion results. For example, for the fusion indexes "3" and "5", there is no situation where the indexes are the same. Therefore, only data transfer needs to be performed without generating operation structure elements.
[0100] Figure 7 The embodiment ensures that the addresses of the fused data elements in the fused data are continuous by performing data transfer during the fusion process, and for data elements that do not need to be accumulated, there is no need to operate the structure elements to represent them, thereby saving the subsequent possible amount of calculation.
[0101] Figure 8 An example of an operation structure element generated according to another embodiment is shown. Figure 7 Similarly to the embodiment of FIG, in this embodiment, for data elements whose indexes do not have duplicates in the data to be fused, operation structure elements may not be generated, thereby avoiding redundant calculations.
[0102] Specifically, when the comparator 637 indicates that the currently output index element is the same as the previous fused index element, an operation structure element may be generated to indicate that the currently corresponding data element is added to the fused data element corresponding to the previous fused index element. When the comparator 637 indicates that the currently output index element is different from the previous fused index element, the current index element may be directly output as the new fused index element, and the address of the corresponding data element may be retained. In other words, the fused data element is the same as the data element before fusion, and both point to the same address.
[0103] The following combination Figure 6 Data example description Figure 8 The process of generating the elements of the operation structure in the process. As shown in the figure, when the output circuit 636 receives the first "0" output by the sorting circuit 632, since this is the first element, the comparator 637 will output, for example, "0" to indicate that it is different from the previous fusion index element (for example, initialized to a negative number). The buffer 635 outputs the index "0" as the first fusion index element. Based on the result of the comparator 637, the structure generator 639 uses the data element corresponding to the current index element as the new fusion data element, for example, Pt11 at this time. At this time, since the indexes are different, accumulation is not required.
[0104] Next, the sorting circuit 632 outputs the second "0". Since the previous fusion index element is "0", the comparator 637 will output, for example, "1" to indicate that it is the same as the previous fusion index element ("0"). At this time, the buffer 635 does not output. Based on the result of the comparator 637, the structure generator 639 believes that no new fusion data element is generated. Therefore, it does not perform any operation on the fusion data element, and only generates an operation structure element representing the accumulation of data elements, which accumulates the data element to the fusion data element corresponding to the previous fusion index element, for example {Pt11, Pt21, Pt11}, where Pt11 represents the address of the fusion data element corresponding to the previous fusion index element. Pt21 represents the address of the current data element, and the accumulation result is still placed in the initial address of the fusion data element, that is, Pt11. The generated operation structure element is stored in the operation structure of the second storage circuit.
[0105] When the sorting circuit 632 continues to output the third and fourth "0", the fusion index will not increase, and the fusion data elements in the corresponding fusion data will not increase. The structure generator will output the associated structure operation elements: {Pt11, Pt31, Pt11}, {Pt11, Pt41, Pt11}.
[0106] Next, sorting circuit 632 outputs the first "2." Referring to the previous step, buffer 635 now outputs index "2" as the second fused index element. Based on the result of comparator 637, structure generator 639 uses the data element corresponding to the current index element as the new fused data element, for example, Pt12 in this case. Since the indices are different, no accumulation is required.
[0107] When the sorting circuit 632 outputs the second "2", the fusion index will not increase, and the fusion data element in the corresponding fusion data will not increase. The structure generator will output the associated structure operation element: {Pt12, Pt42, Pt12}.
[0108] According to the above description, those skilled in the art can similarly deduce other fusion results. For example, for the fusion indexes "3" and "5", there is no situation where the indexes are the same. Therefore, it is only necessary to use the corresponding data elements as new fusion data elements, such as Pt22 and Pt13, without generating operation structure elements.
[0109] Figure 8 Compared with the embodiment Figure 7 This reduces data handling during the fusion process, at the expense of discontinuous addresses of the fused data elements. Furthermore, for data elements that do not need to be accumulated, there is no need to represent them in the computational structure, saving potential subsequent computations.
[0110] Those skilled in the art will appreciate that other forms of hardware circuits may be designed to implement the above-mentioned merge sort fusion processing, and the present disclosure is not limited in this regard.
[0111] In the disclosed embodiment, the merge, sort and fusion processing of data can be implemented by calling the fusion instruction using the exemplary hardware circuit described above. At this time, the operation objects of the fusion instruction include the input K-way data to be fused, the K-way index corresponding to the K-way data, the size of the K-way data, and the output one-way fusion index, one-way fusion data and the associated operation structure, K>1. Among these objects, the index element in the K-way index indicates the index information of the corresponding data element in the K-way data; the index elements of each-way index in the K-way index are arranged in order according to the first order; the data elements of each-way data in the K-way data are arranged in order according to the order of the corresponding index; the fusion index elements in the fusion index are arranged in order according to the second order; the fusion data elements in the fusion data correspond one to one to the fusion index elements; and the fusion data elements composed of at least the data elements with the same index are represented by the associated operation structure elements in the operation structure.
[0112] In some embodiments, the operation object of the fusion instruction may also include the total number of output fusion index elements, indicating the number of index elements in the output fusion index. It is understood that since fusion data and fusion indexes have a one-to-one correspondence, this operation object is also simply the number of data elements in the output fusion data.
[0113] In some embodiments, the operation object of the fusion instruction may further include the total number of output operation structure elements, which is used to indicate the total number of operation structure elements in the operation structure.
[0114] As mentioned above, the first order and the second order may be the same or different, and the first order and the second order may be selected from any of the following: an order from small to large, or an order from large to small.
[0115] In some embodiments, at least one operand of the fused instruction may be represented by an address.
[0116] Figure 9 The contents pointed to by each address in the fused instruction are shown as an example.
[0117] For example, the input K-way data can be indicated by a first address, which includes K elements, where the i-th element represents the address of the i-th data, and each data element in the i-th data is also an address, pointing to a vector of a predetermined length, where 0<i≤K.
[0118] In some embodiments, the first address can be labeled data_addr. The number of elements in this address is K, indicating a fusion operation of K paths of data. data_addr is a three-level pointer, in which the K elements represent the starting addresses of the K paths of data (e.g., vectors) to be fused. Furthermore, each element in the K paths of data (vectors) is also an address, pointing to a vector of size, for example, offset.
[0119] The input K-way index may be indicated by a second address, the second address including K elements, wherein the i-th element represents the address of the i-th way index associated with the i-th way data, wherein 0<i≤K.
[0120] As mentioned above, the K-way data to be fused can have a one-to-one corresponding K-way index, so the second address can be marked as index_addr. The number of elements in this address is K, which means that the fusion is based on this K-way index. Similarly, index_addr is a secondary pointer, where K elements represent the starting address of the K-way index (e.g., vector) to be fused.
[0121] The size of the input K-way data can be indicated by the third address. The third address is a first-level pointer, which can be labeled size_addr. It also includes K elements, where the i-th element represents the number of data elements in the i-way data, where 0 < i ≤ K. Since there is a one-to-one correspondence between K-way data and K-way indexes, the i-th element in size_addr also represents the number of index elements in the i-way index.
[0122] In some embodiments, the index elements in the input K-way index are arranged in order, for example, from small to large, and the elements of the final output one-way fused index can also be arranged in order from small to large. In the merge sort fusion process disclosed herein, when there are duplicate indexes, the indexes are deduplicated and the corresponding data elements are fused into several operation structure elements to form a fused data element.
[0123] The output fused data can be stored in a fourth address, which is indicated by the fourth address in the fused instruction. The fourth address can be marked as out_data_addr, which is the address of the output fused data. The fourth address is a secondary pointer that includes L elements, where the jth element represents the jth fused data element in this fused data. Each fused data element is an address that points to a vector of a predetermined length (e.g., offset). L represents the total number of fused data elements, where L>1 and 0<j≤L.
[0124] The output fused index can be stored in the fifth address, which is indicated by the fifth address in the fusion instruction. The fifth address can be marked as out_index_addr, which is the address of the output fused index. The fifth address is a first-level pointer that includes L elements, where the jth element represents the jth fused index element in this fused index. L represents the total number of fused index elements, where L>1 and 0<j≤L. The output fused data and the output fused index also have a one-to-one correspondence.
[0125] The output operation structure can be stored in the sixth address, that is, indicated by the sixth address in the fused instruction. The sixth address can be labeled as op_addr, which is the address of the output operation structure. In some embodiments, the sixth address is a first-level pointer that includes M elements, each of which is an operation structure element.
[0126] Each operation structure element includes three sub-elements, each sub-element is an address, which is the address of two addends and the address of the addition result. In some embodiments, an operation structure element may include three 32-bit elements for representing three addresses.
[0127] Alternatively or additionally, in some embodiments, the fusion instruction's operand may also include the total number of fused data elements output. For example, after the fusion process is complete, the total number of fused data elements output will be returned, indicating the number of data elements in the output fused data path, i.e., the L value described above. This data can be written back to the parameter gpr_id0, for example.
[0128] Alternatively or additionally, in some embodiments, the fusion instruction's operands may also include the total number of output computation structure elements. For example, after the fusion process completes, the total number of output computation structure elements is returned, indicating the number of computation structure elements in the output computation structure, i.e., the value M described above. This data may be written back to the parameter gpr_id1, for example.
[0129] With the development of artificial intelligence technology, in tasks such as image processing and pattern recognition, the operands are often of the data type of multidimensional vectors (i.e., tensor data). Using only scalar or vector operations cannot enable the hardware to efficiently complete the computing tasks. Therefore, in some embodiments of the present disclosure, fusion instructions involving tensor data are also provided. At least one operation object of the fusion instruction includes tensor data, and the tensor data is indicated by at least one descriptor. Specifically, the descriptor can indicate at least one of the following information: shape information of tensor data, spatial information of tensor data. The shape information of tensor data can be used to determine the data address of the tensor data corresponding to the operand in the data storage space. The spatial information of tensor data can be used to determine the dependency between instructions, and then determine, for example, the execution order of instructions.
[0130] In one possible implementation, the spatial information of tensor data can be indicated by a spatial identifier (ID). The spatial ID can also be called a spatial alias, which refers to a spatial region used to store the corresponding tensor data. The spatial region can be a continuous space or multiple spaces. This disclosure does not limit the specific composition of the spatial region. Different spatial IDs indicate that there is no dependency between the spatial regions pointed to.
[0131] Various possible implementations of the shape information of tensor data will be described in detail below with reference to the accompanying drawings.
[0132] Tensors can contain a variety of data structures. Tensors can be of different dimensions. For example, a scalar can be considered a 0-dimensional tensor, a vector can be considered a 1-dimensional tensor, and a matrix can be a 2-dimensional or higher tensor. The shape of a tensor includes information such as the dimensions of the tensor and the size of each dimension of the tensor. For example, for a 3D tensor:
[0133] x3=[[[1,2,3],[4,5,6]];[[7,8,9],[10,11,12]]]
[0134] The shape or dimensions of this tensor can be expressed as X3 = (2, 2, 3), which means that the three parameters indicate that the tensor is a three-dimensional tensor with a first dimension of 2, a second dimension of 2, and a third dimension of 3. When storing tensor data in memory, the shape of the tensor data cannot be determined based on its data address (or storage area), and thus, related information such as the relationship between multiple tensor data cannot be determined, resulting in low processor access efficiency for tensor data.
[0135] In one possible implementation, a descriptor can be used to indicate the shape of N-dimensional tensor data, where N is a positive integer, such as N=1, 2, or 3, or zero. The three-dimensional tensor in the above example can be represented by a descriptor as (2, 2, 3). It should be noted that this disclosure does not limit the manner in which a descriptor indicates the shape of a tensor.
[0136] In one possible implementation, the value of N can be determined according to the dimension (also called order) of the tensor data, or it can be set according to the usage requirements of the tensor data. For example, when the value of N is 3, the tensor data is three-dimensional tensor data, and the descriptor can be used to indicate the shape of the three-dimensional tensor data in three dimensions (such as offset, size, etc.). It should be understood that those skilled in the art can set the value of N according to actual needs, and this disclosure does not limit this.
[0137] Although tensor data can be multi-dimensional, because the memory layout is always one-dimensional, there is a correspondence between tensors and storage on the memory. Tensor data is usually allocated in a continuous memory space, that is, tensor data can be expanded one-dimensionally (for example, row-major) and stored in the memory.
[0138] This relationship between a tensor and its underlying storage can be expressed through dimensions such as offset, size, and stride. A dimension's offset refers to the offset relative to a reference position in that dimension. A dimension's size refers to the size of that dimension, or the number of elements in that dimension. A dimension's stride refers to the spacing between adjacent elements in that dimension. For example, the stride of the three-dimensional tensor above is (6, 3, 1), meaning the stride of the first dimension is 6, the stride of the second dimension is 3, and the stride of the third dimension is 1.
[0139] Figure 10 Schematic diagram showing the data storage space according to the embodiment of the present disclosure. Figure 10 As shown, data storage space 101 stores two-dimensional data in a row-major manner, which can be represented by (x, y) (where the X axis is horizontally to the right and the Y axis is vertically downward). The size in the X-axis direction (the size of each row, or the total number of columns) is ori_x (not shown in the figure), and the size in the Y-axis direction (the total number of rows) is ori_y (not shown in the figure). The starting address PA_start (base address) of data storage space 101 is the physical address of the first data block 102. Data block 103 is a portion of the data in data storage space 101. Its offset 105 in the X-axis direction is represented by offset_x, its offset 104 in the Y-axis direction is represented by offset_y, its size in the X-axis direction is represented by size_x, and its size in the Y-axis direction is represented by size_y.
[0140] In one possible implementation, when a descriptor is used to define data block 103, the data reference point of the descriptor may be the first data block of data storage space 101, and the reference address of the descriptor may be agreed to be the starting address PA_start of data storage space 101. The content of the descriptor for data block 103 may then be determined by combining the size ori_x on the X axis and the size ori_y on the Y axis of data storage space 101, as well as the offset offset_y in the Y axis direction, the offset offset_x in the X axis direction, the size size_x in the X axis direction, and the size size_y in the Y axis direction of data block 103.
[0141] In a possible implementation, the following formula (1) can be used to express the content of the descriptor:
[0142]
[0143] It should be understood that although in the above examples, the content of the descriptor represents a two-dimensional space, those skilled in the art can set the specific dimension represented by the content of the descriptor according to actual conditions, and this disclosure does not limit this.
[0144] In one possible implementation, the base address of the data reference point of the descriptor in the data storage space can be agreed upon. Based on the base address, the content of the descriptor of the tensor data is determined according to the positions of at least two vertices at diagonal positions in N dimensional directions relative to the data reference point.
[0145] For example, the data reference point of the descriptor can be agreed to be the reference address PA_base in the data storage space. For example, a data (e.g., data at position (2, 2)) can be selected in the data storage space 81 as the data reference point, and the physical address of the data in the data storage space can be used as the reference address PA_base. The position of the two vertices at the diagonal position relative to the data reference point can be used to determine the reference address PA_base. Figure 10 The content of the descriptor of the data block 103 in the data block 103 is determined. First, the positions of at least two vertices at the diagonal positions of the data block 103 relative to the data reference point are determined. For example, the positions of the diagonal vertices from the upper left to the lower right relative to the data reference point are used, where the relative position of the upper left vertex is (x_min, y_min) and the relative position of the lower right vertex is (x_max, y_max). Then, the content of the descriptor of the data block 103 can be determined based on the reference address PA_base, the relative position of the upper left vertex (x_min, y_min), and the relative position of the lower right vertex (x_max, y_max).
[0146] In a possible implementation, the following formula (2) can be used to express the content of the descriptor (the base address is PA_base):
[0147]
[0148] It should be understood that although the vertices at the upper left corner and the lower right corner are used in the above example to determine the content of the descriptor, those skilled in the art can set the specific vertices of at least two diagonal positions according to actual needs, and this disclosure does not limit this.
[0149] In one possible implementation, the content of the tensor data descriptor can be determined based on the reference address of the descriptor's data reference point in the data storage space and the mapping relationship between the data description position and the data address of the tensor data indicated by the descriptor. The mapping relationship between the data description position and the data address can be set according to actual needs. For example, when the tensor data indicated by the descriptor is three-dimensional spatial data, the function f(x, y, z) can be used to define the mapping relationship between the data description position and the data address.
[0150] In a possible implementation, the following formula (3) can be used to express the content of the descriptor:
[0151]
[0152] In one possible implementation, the descriptor is further used to indicate the address of N-dimensional tensor data, wherein the content of the descriptor further includes at least one address parameter representing the address of the tensor data. For example, the content of the descriptor may be the following formula (4):
[0153]
[0154] Where PA is the address parameter. The address parameter can be a logical address or a physical address. When parsing the descriptor, PA can be used as any vertex, midpoint, or preset point of the vector shape, combined with the shape parameters in the X and Y directions to obtain the corresponding data address.
[0155] In a possible implementation, the address parameter of the tensor data includes a reference address of a data reference point of the descriptor in the data storage space of the tensor data, and the reference address includes a starting address of the data storage space.
[0156] In a possible implementation, the descriptor may further include at least one address parameter representing the address of the tensor data. For example, the content of the descriptor may be the following formula (5):
[0157]
[0158] PA_start is the base address parameter and will not be described in detail.
[0159] It should be understood that those skilled in the art can set the mapping relationship between the data description location and the data address according to actual conditions, and this disclosure does not limit this.
[0160] In one possible implementation, a predetermined reference address can be set within a task. All descriptors in instructions within this task use this reference address, and the descriptor content can include shape parameters based on this reference address. This reference address can be determined by setting the environment parameters for this task. A description of the reference address and its use can be found in the above embodiments. In this implementation, the descriptor content can be mapped to data addresses more quickly.
[0161] In one possible implementation, the base address can be included in the content of each descriptor, so that the base address of each descriptor can be different. Compared with the method of using environmental parameters to set a common base address, each descriptor in this method can describe data more flexibly and use a larger data address space.
[0162] In one possible implementation, the data address of the data corresponding to the operand of the processing instruction in the data storage space can be determined based on the content of the descriptor. The data address is calculated automatically by hardware, and the calculation method of the data address varies depending on the representation of the descriptor content. This disclosure does not limit the specific method for calculating the data address.
[0163] For example, the content of the descriptor in the operand is expressed using formula (1). The offsets of the tensor data indicated by the descriptor in the data storage space are offset_x and offset_y, and the size is size_x*size_y. Then, the starting data address PA1 of the tensor data indicated by the descriptor in the data storage space is (x,y) It can be determined using the following formula (6):
[0164] PA1 (x,y) =PA_start+(offset_y-1)*ori_x+offset_x (6)
[0165] The data starting address PA1 is determined according to the above formula (6) (x,y) , combined with the offsets offset_x and offset_y, and the sizes size_x and size_y of the storage area, the storage area of the tensor data indicated by the descriptor in the data storage space can be determined.
[0166] In one possible implementation, when the operand also includes a data description location for a descriptor, the data address of the data corresponding to the operand in the data storage space can be determined based on the content of the descriptor and the data description location. In this way, partial data (e.g., one or more data) in the tensor data indicated by the descriptor can be processed.
[0167] For example, the content of the descriptor in the operand is expressed using formula (2). The offsets of the tensor data indicated by the descriptor in the data storage space are offset_x and offset_y respectively, and the size is size_x*size_y. The data description position for the descriptor included in the operand is (x q ,y q ), then the data address PA2 of the tensor data indicated by the descriptor in the data storage space (x,y) It can be determined using the following formula (7):
[0168] PA2 (x,y) =PA_start+(offset_y+y q -1)*ori_x+(offset_x+x q ) (7)
[0169] In one possible implementation, the descriptor can indicate data blocks. Data blocks can effectively speed up operations and improve processing efficiency in many applications. For example, in graphics processing, convolution operations often use data blocks for fast processing.
[0170] Figure 11 Schematic diagram showing data blocks in data storage space according to an embodiment of the present disclosure. Figure 11 As shown, the data storage space 1100 also uses a row-first approach to store two-dimensional data, which can be represented by (x, y) (where the X axis is horizontal to the right and the Y axis is vertically downward). The size in the X-axis direction (the size of each row, or the total number of columns) is ori_x (not shown in the figure), and the size in the Y-axis direction (the total number of rows) is ori_y (not shown in the figure). Figure 10 Tensor data, Figure 11 The tensor data stored in consists of multiple data blocks.
[0171] In this case, the descriptor requires more parameters to represent these data blocks. Taking the X-axis (X dimension) as an example, the following parameters may be involved: ori_x, x.tile.size (the size of the block 1102), x.tile.stride (the stride 1104 in the block, i.e., the distance between the first point of the first tile and the first point of the second tile), x.tile.num (the number of tiles, shown as 3 in the figure), x.stride (the overall stride, i.e., the distance between the first point in the first row and the first point in the second row), etc. Other dimensions can similarly include corresponding parameters.
[0172] In one possible implementation, a descriptor may include a descriptor identifier and / or descriptor content. The descriptor identifier is used to distinguish the descriptor, for example, the descriptor identifier may be a number; the descriptor content may include at least one shape parameter representing the shape of the tensor data. For example, if the tensor data is three-dimensional data, and the shape parameters of two of the three dimensions of the tensor data are fixed, the descriptor content may include the shape parameter representing the other dimension of the tensor data.
[0173] In one possible implementation, the identifier and / or content of the descriptor may be stored in a descriptor storage space (internal memory), such as a register, on-chip SRAM, or other media cache. The tensor data indicated by the descriptor may be stored in a data storage space (internal memory or external memory), such as an on-chip cache or off-chip memory. This disclosure does not limit the specific locations of the descriptor storage space and the data storage space.
[0174] In one possible implementation, the identifier, content of the descriptor, and the tensor data indicated by the descriptor can be stored in the same area of the internal memory. For example, a continuous area of the on-chip cache can be used to store the relevant content of the descriptor, and its address is ADDR0-ADDR1023. Among them, the address ADDR0-ADDR63 can be used as a descriptor storage space to store the identifier and content of the descriptor, and the address ADDR64-ADDR1023 can be used as a data storage space to store the tensor data indicated by the descriptor. In the descriptor storage space, the address ADDR0-ADDR31 can be used to store the identifier of the descriptor, and the address ADDR32-ADDR63 can be used to store the content of the descriptor. It should be understood that the address ADDR is not limited to 1 bit or 1 byte. It is used here to represent an address and is an address unit. Those skilled in the art can determine the descriptor storage space, data storage space and their specific addresses according to actual conditions, and this disclosure is not limited to this.
[0175] In one possible implementation, the descriptor identifier, content, and tensor data indicated by the descriptor can be stored in different areas of the internal memory. For example, registers can be used as descriptor storage space to store the descriptor identifier and content, and on-chip cache can be used as data storage space to store the tensor data indicated by the descriptor.
[0176] In one possible implementation, when registers are used to store the identifier and content of descriptors, the register number can be used to represent the identifier of the descriptor. For example, when the register number is 0, the identifier of the descriptor stored in it is set to 0. When the descriptor in the register is valid, an area in the cache space can be allocated to store the tensor data based on the size of the tensor data indicated by the descriptor.
[0177] In one possible implementation, the identifier and content of the descriptor may be stored in internal memory, and the tensor data indicated by the descriptor may be stored in external memory. For example, the identifier and content of the descriptor may be stored on-chip, while the tensor data indicated by the descriptor may be stored off-chip.
[0178] In one possible implementation, the data address of the data storage space corresponding to each descriptor can be a fixed address. For example, a separate data storage space can be divided for tensor data, and the starting address of each tensor data in the data storage space corresponds one-to-one to the descriptor. In this case, the circuit or module responsible for parsing the computing instruction (such as an entity outside the computing device of the present disclosure) can determine the data address of the data corresponding to the operand in the data storage space based on the descriptor.
[0179] In one possible implementation, when the data address of the data storage space corresponding to the descriptor is a variable address, the descriptor can also be used to indicate the address of N-dimensional tensor data, wherein the content of the descriptor can also include at least one address parameter representing the address of the tensor data. For example, the tensor data is 3-dimensional data. When the descriptor points to the address of the tensor data, the content of the descriptor may include an address parameter representing the address of the tensor data, such as the starting physical address of the tensor data, or may include multiple address parameters of the address of the tensor data, such as the starting address + address offset of the tensor data, or the address parameters of the tensor data based on each dimension. Those skilled in the art can set the address parameters according to actual needs, and this disclosure does not limit this.
[0180] In one possible implementation, the address parameter of the tensor data may include the reference address of the descriptor's data reference point in the data storage space of the tensor data. The reference address may vary depending on the data reference point. This disclosure does not limit the selection of the data reference point.
[0181] In one possible implementation, the reference address may include the starting address of the data storage space. When the data reference point of the descriptor is the first data block in the data storage space, the reference address of the descriptor is the starting address of the data storage space. When the data reference point of the descriptor is data other than the first data block in the data storage space, the reference address of the descriptor is the address of the data block in the data storage space.
[0182] In one possible implementation, the shape parameters of the tensor data include at least one of the following: the size of the data storage space in at least one direction of the N-dimensional directions, the size of the storage area in at least one direction of the N-dimensional directions, the offset of the storage area in at least one direction of the N-dimensional directions, the positions of at least two vertices at diagonal positions in the N-dimensional directions relative to the data reference point, and the mapping relationship between the data description position of the tensor data indicated by the descriptor and the data address. The data description position is the mapping position of the point or area in the tensor data indicated by the descriptor. For example, when the tensor data is 3D data, the descriptor can use three-dimensional space coordinates (x, y, z) to represent the shape of the tensor data, and the data description position of the tensor data can be the position of the point or area mapped in the three-dimensional space represented by the three-dimensional space coordinates (x, y, z).
[0183] It should be understood that those skilled in the art can select shape parameters representing tensor data according to actual circumstances, and this disclosure does not limit this. By using descriptors in the data access process, associations between data can be established, thereby reducing the complexity of data access and improving instruction processing efficiency.
[0184] Figure 12 FIG. 1 is a block diagram showing a data processing device 1200 according to another embodiment of the present disclosure. The data processing device 1200 may be implemented, for example, in Figure 2 In the computing device 201. Figure 12 The data processing device 1200 and Figure 5 The difference is, Figure 12 The data processing device 1200 further includes a tensor interface circuit 1212 for implementing functions related to the descriptors of tensor data. Similarly, the data processing device 1200 may also include a control circuit 1210, a storage circuit 1220 and an operation circuit 1230. The specific functions and implementations of these circuits are similar to those of the Figure 5 The above are similar to those of , so they will not be repeated here.
[0185] The tensor interface unit (TIU) 1212 can be configured to implement operations associated with descriptors under the control of the control circuit 1210. These operations may include, but are not limited to, registering, modifying, deregistering, and parsing descriptors; reading and writing descriptor contents, etc. This disclosure does not limit the specific hardware type of the tensor interface unit. In this way, operations associated with descriptors can be implemented through dedicated hardware, further improving the access efficiency of tensor data.
[0186] In some embodiments, the tensor interface circuit 1212 may be configured to parse shape information of tensor data included in an operand of an instruction to determine a data address of data corresponding to the operand in the data storage space.
[0187] Optionally or additionally, in some further embodiments, the tensor interface circuit 1212 can be configured to compare the spatial information (e.g., spatial ID) of the tensor data included in the operands of two instructions to determine the dependency relationship between the two instructions, and further determine the out-of-order execution, synchronization, and other operations of the instructions.
[0188] Despite Figure 12 The control circuit 1210 and the tensor interface circuit 1212 are shown as two separate modules, but those skilled in the art will appreciate that these two circuits may also be implemented as one module or more modules, and the present disclosure is not limited in this regard.
[0189] There may be various operations related to data fusion, such as merge sort processing, sort accumulation, sort fusion processing, etc. Various instruction schemes may be designed to implement operations related to data fusion.
[0190] In one solution, a fused instruction may be designed, and the instruction may include an operation mode bit to indicate different operation modes of the fused instruction, thereby performing different operations.
[0191] In another approach, multiple fused instructions can be designed, each corresponding to one or more different operating modes, thereby performing different operations. In one implementation, a corresponding fused instruction can be designed for each operating mode. In another implementation, the operating modes can be categorized based on their characteristics, with a fused instruction designed for each type of operating mode. Furthermore, when a certain type of operating mode includes multiple operating modes, an operating mode bit can be included in the fused instruction to indicate the corresponding operating mode.
[0192] Regardless of the solution adopted, the fused instruction can indicate its corresponding operation mode through the operation mode bit and / or the instruction itself.
[0193] In the context of the present disclosure, the aforementioned fusion instruction may be a microinstruction or control signal running within one or more multi-stage operation pipelines, which may include (or indicate) one or more operations to be performed by the multi-stage operation pipeline. Depending on the operation scenario, the operation may include but is not limited to arithmetic operations such as convolution operations and matrix multiplication operations, logical operations such as AND operations, XOR operations, and OR operations, shift operations, or any combination of the aforementioned operations.
[0194] Figure 13 FIG. 13 is a flow chart illustrating an exemplary data processing method 1300 according to an embodiment of the present disclosure.
[0195] like Figure 13 As shown, in step 1310, the fusion instruction is parsed, and the fusion instruction instructs to perform fusion processing on multiple channels of data to be fused. This step can be performed by Figure 5 The control circuit 510 or Figure 12 Then, in step 1320, according to the fusion instruction, the data elements in the multiple channels of data to be fused are merged into one channel of ordered fused data according to their corresponding indexes, wherein the data elements with the same index are merged into the operation structure elements, wherein the data elements may include any of scalars, vectors or higher-dimensional data. Finally, in step 1330, the fused data is output. Steps 1320 and 1330 may be performed by, for example, Figure 5 The operation circuit 530 or Figure 12 The operation circuit 1230 is used to execute.
[0196] Those skilled in the art will appreciate that the various steps of the above method correspond to the various circuits described above in conjunction with the example circuit diagram, and therefore the features described above are equally applicable to the method steps and will not be repeated here.
[0197] As can be seen from the above description, the disclosed embodiment provides a fusion instruction for performing fusion processing of multiple channels of data to be fused. In some embodiments, the fusion instruction is a hardware instruction, which implements data fusion processing through a dedicated hardware circuit, which can speed up the processing speed, thereby better supporting operations related to post-sparse processing, such as supporting operations in radar algorithms. In some embodiments, the fusion instruction can merge multiple channels of ordered data into one channel of ordered fused data, and data with the same index can be merged and represented in the form of an operation structure, thereby facilitating subsequent calculation processing. In some embodiments, the fusion instruction may include an operation mode bit to indicate that the fusion instruction is a merge sort fusion processing operation, or the fusion instruction itself may indicate a merge sort fusion processing operation. By providing a dedicated fusion instruction to perform operations related to the fusion processing of multiple channels of data, the processing can be simplified.
[0198] According to different application scenarios, the electronic devices or devices disclosed herein may include servers, cloud servers, server clusters, data processing devices, robots, computers, printers, scanners, tablet computers, smart terminals, PC devices, Internet of Things terminals, mobile terminals, mobile phones, driving recorders, navigators, sensors, cameras, cameras, video cameras, projectors, watches, headphones, mobile storage, wearable devices, visual terminals, automatic driving terminals, vehicles, household appliances, and / or medical equipment. The vehicles include airplanes, ships and / or vehicles; the household appliances include televisions, air conditioners, microwave ovens, refrigerators, rice cookers, humidifiers, washing machines, electric lights, gas stoves, and range hoods; the medical equipment includes magnetic resonance imaging (MRI), ultrasound machines and / or electrocardiographs. The electronic devices or devices disclosed herein may also be applied to the Internet, Internet of Things, data centers, energy, transportation, public administration, manufacturing, education, power grids, telecommunications, finance, retail, construction sites, medical care and other fields. Furthermore, the electronic devices or devices disclosed herein may also be used in cloud, edge, terminal and other application scenarios related to artificial intelligence, big data and / or cloud computing. In one or more embodiments, electronic devices or apparatuses with high computing power according to the disclosed solution can be applied to cloud devices (such as cloud servers), while electronic devices or apparatuses with low power consumption can be applied to terminal devices and / or edge devices (such as smartphones or cameras). In one or more embodiments, the hardware information of the cloud device and the hardware information of the terminal device and / or edge device are compatible with each other, so that according to the hardware information of the terminal device and / or edge device, appropriate hardware resources can be matched from the hardware resources of the cloud device to simulate the hardware resources of the terminal device and / or edge device, so as to complete the unified management, scheduling and collaborative work of end-to-end or cloud-edge-to-end.
[0199] It should be noted that, for the purpose of simplicity, the present disclosure describes some methods and embodiments thereof as a series of actions and combinations thereof, but those skilled in the art will understand that the scheme of the present disclosure is not limited by the order of the actions described. Therefore, based on the disclosure or teachings of the present disclosure, those skilled in the art will understand that some of the steps therein can be performed in other orders or simultaneously. Further, those skilled in the art will understand that the embodiments described in the present disclosure can be regarded as optional embodiments, that is, the actions or modules involved therein are not necessarily necessary for the implementation of one or more schemes of the present disclosure. In addition, depending on the different schemes, the description of some embodiments of the present disclosure also has different emphases. In view of this, those skilled in the art will understand that the parts that are not described in detail in a certain embodiment of the present disclosure may also refer to the relevant descriptions of other embodiments.
[0200] In terms of specific implementation, based on the disclosure and teachings of this disclosure, those skilled in the art can understand that several embodiments disclosed in this disclosure can also be implemented in other ways not disclosed herein. For example, with respect to the various units in the electronic device or device embodiments described above, this article splits them based on the consideration of logical functions, and there may be other ways of splitting them in actual implementation. For another example, multiple units or components can be combined or integrated into another system, or some features or functions in the units or components can be selectively disabled. With respect to the connection relationship between different units or components, the connection discussed above in conjunction with the accompanying drawings can be a direct or indirect coupling between units or components. In some scenarios, the aforementioned direct or indirect coupling involves a communication connection using an interface, wherein the communication interface can support electrical, optical, acoustic, magnetic or other forms of signal transmission.
[0201] In this disclosure, the units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units. The aforementioned components or units may be located in the same location or distributed across multiple network elements. In addition, according to actual needs, some or all of the units may be selected to achieve the purpose of the solution described in the embodiments of this disclosure. In addition, in some scenarios, multiple units in the embodiments of this disclosure may be integrated into one unit or each unit may exist physically separately.
[0202] In some other implementation scenarios, the above-mentioned integrated units can also be implemented in the form of hardware, that is, as specific hardware circuits, which may include digital circuits and / or analog circuits, etc. The physical implementation of the hardware structure of the circuit may include but is not limited to physical devices, and the physical devices may include but are not limited to devices such as transistors or memristors. In view of this, the various devices described herein (such as computing devices or other processing devices) can be implemented by appropriate hardware processors, such as central processing units, GPUs, FPGAs, DSPs, and ASICs. Furthermore, the aforementioned storage unit or storage device can be any appropriate storage medium (including magnetic storage media or magneto-optical storage media, etc.), which can be, for example, resistive random access memory (RRAM), dynamic random access memory (DRAM), static random access memory (SRAM), enhanced dynamic random access memory (EDRAM), high bandwidth memory (HBM), hybrid memory cube (HMC), ROM and RAM, etc.
[0203] The foregoing content can be better understood in accordance with the following terms:
[0204] Clause 1. A data processing apparatus comprising:
[0205] a control circuit configured to parse a fusion instruction, wherein the fusion instruction instructs to perform fusion processing on multiple channels of data to be fused;
[0206] a storage circuit configured to store information before and / or after fusion processing; and
[0207] An operation circuit is configured to merge the data elements in the multiple data to be fused into one channel of ordered fused data according to their corresponding indexes according to the fusion instruction, wherein data elements with the same index are merged into operation structure elements, and the data elements include any of scalars, vectors or higher-dimensional data.
[0208] Clause 2. The data processing device according to Clause 1, wherein the operation objects of the fusion instruction include input K-way data to be fused, K-way indexes corresponding to the K-way data, the size of the K-way data, and output one-way fusion index, one-way fusion data and associated operation structures, K>1, wherein:
[0209] The index elements in the K-way index indicate index information of corresponding data elements in the K-way data;
[0210] The index elements of each index in the K-way index are arranged in order according to the first order;
[0211] The data elements of each path of data in the K paths of data are arranged in order according to the order of corresponding indexes;
[0212] The fusion index elements in the fusion index are arranged in order according to the second order;
[0213] The fused data elements in the fused data correspond one-to-one to the fused index elements; and
[0214] A fused data element composed of at least data elements of the same index is represented by associated operation structure elements in the operation structure.
[0215] Clause 3. A data processing apparatus according to clause 2, wherein the first order is the same as or different from the second order, and the first order and the second order are selected from any one of the following: an order from small to large, or an order from large to small.
[0216] Clause 4. The data processing device according to any one of clauses 2-3, wherein the operation circuit includes a sorting circuit and an output circuit, wherein
[0217] The sorting circuit is configured to sort the K-way indexes according to the size of the index elements and output them to the output circuit in order; and
[0218] The output circuit is configured to generate, at least when receiving the same index element from the sorting circuit, an operation structure element representing an accumulation operation of data elements corresponding to the same index element, and remove duplicate index elements.
[0219] Clause 5. The data processing apparatus according to clause 4, wherein the sorting circuit comprises a comparison circuit and a buffer circuit, wherein:
[0220] The comparison circuit is configured to compare the index elements to be sorted in the K-way index with the index elements not outputted in the buffer circuit, and output a comparison result to the control circuit; and
[0221] The buffer circuit is configured to sequentially store information of the compared index elements and the corresponding data elements under the control of the control circuit, and sequentially output information of the compared index elements and the corresponding data elements.
[0222] Clause 6. The data processing apparatus according to clause 5, wherein the comparison circuit comprises:
[0223] The K-1 comparators are configured to compare the index elements to be sorted in the K-way index with the K-1 index elements of the current sequence in the buffer circuit, generate a comparison result, and output it to the control circuit.
[0224] Clause 7. The data processing device according to Clause 6, wherein the control circuit is configured to determine an insertion position of the index element to be sorted in the current sequence in the buffer circuit based on the comparison result.
[0225] Clause 8. The data processing apparatus according to clause 7, wherein the comparison result is represented by a bitmap, and the control circuit is further configured to determine the insertion position according to a changed position of a bit in the bitmap.
[0226] Clause 9. A data processing device according to any one of clauses 7-8, wherein the buffer circuit is configured to insert the index elements to be sorted and the information of the data elements corresponding thereto into the insertion position according to the instruction of the control circuit.
[0227] Clause 10. A data processing device according to any one of clauses 5 to 9, wherein the buffer circuit is further configured to output information of the first or last index element in the current sequence and the data element corresponding thereto in a specified order.
[0228] Item 11. The data processing device according to Item 10, wherein the control circuit is further configured to: determine the memory access information of the next index element to be sorted based on the index element output from the buffer circuit.
[0229] Clause 12. The data processing apparatus according to any one of clauses 4 to 11, wherein the output circuit comprises a comparator, a buffer, and a structure generator, wherein:
[0230] The comparator is configured to compare the index element output from the sorting circuit with the last fused index element and output a comparison result;
[0231] The buffer is configured to output the index element as a new fused index element only when the comparison result indicates a difference; and
[0232] The structure generator is configured to generate an operation structure element based on the data element corresponding to the index element when the comparison result indicates the same, and the operation structure element represents an accumulation operation of adding the data element to the fused data element corresponding to the last fused index element.
[0233] Clause 13. A data processing device according to Clause 12, wherein the structure generator is further configured to, when the comparison result indicates inequality, generate an operation structure element based on the data element corresponding to the index element, wherein the operation structure element represents an addition operation of adding the data element to 0.
[0234] Clause 14. A data processing device according to any one of clauses 12-13, wherein each operation structure element includes three elements, indicating two addends and an addition result respectively.
[0235] Clause 15. A data processing device according to any one of clauses 1 to 14, wherein the data elements in the data to be fused are valid data elements after sparsification in radar-based object detection, and the index indicates the position information of the valid data elements in the data before sparsification.
[0236] Clause 16. A data processing device according to any one of clauses 2-15, wherein the operation object of the fusion instruction also includes the total number of output fusion index elements, which is used to indicate the total number of index elements in the output fusion index.
[0237] Clause 17. A data processing device according to any one of clauses 2-16, wherein the operation object of the fusion instruction also includes the total number of output operation structure elements, which is used to indicate the total number of operation structure elements in the operation structure.
[0238] Clause 18. A data processing apparatus according to any one of clauses 2 to 17, wherein
[0239] The input K-way data is indicated by a first address, which includes K elements, the i-th element represents the address of the i-th data, and each data element in the i-th data is an address pointing to a vector of a predetermined length, where 0<i≤K.
[0240] Clause 19. A data processing apparatus according to any one of clauses 2 to 18, wherein
[0241] The K-way index is indicated by a second address, the second address includes K elements, the i-th element represents the address of the i-th way index associated with the i-th way data, wherein 0<i≤K.
[0242] Clause 20. A data processing apparatus according to any one of clauses 2 to 19, wherein
[0243] The size of the K-way data is indicated by a third address, the third address includes K elements, the i-th element represents the number of data elements in the i-th way data, wherein 0<i≤K.
[0244] Clause 21. A data processing apparatus according to any one of clauses 2 to 20, wherein
[0245] The fused data is indicated by a fourth address, which includes L elements, where the j-th element represents the j-th data element in the fused data, each data element is an address pointing to a vector of a predetermined length, and L represents the total number of data elements in the fused data, where L≥1 and 0<j≤L.
[0246] Clause 22. A data processing apparatus according to any one of clauses 2 to 21, wherein
[0247] The one-way fusion index is indicated by the fifth address, and the fifth address includes L elements, the jth element represents the jth index element in the one-way fusion index, and L represents the total number of index elements in the fusion index, L≥1, 0<j≤L.
[0248] Clause 23. A data processing apparatus according to any one of clauses 2 to 22, wherein
[0249] The operation structure is indicated by the sixth address, which includes M elements, each element being a structure.
[0250] Clause 24. A data processing apparatus according to clause 23, wherein
[0251] The structure includes three sub-elements, each of which is an address pointing to two addends and an addition result respectively.
[0252] Clause 25. A data processing apparatus according to any one of clauses 2 to 17, wherein at least one of the operation objects comprises tensor data, and the tensor data is indicated by at least one descriptor, the descriptor indicating at least one of the following information: shape information of the tensor data and spatial information of the tensor data; and
[0253] The data processing device further includes a tensor interface circuit configured to parse the descriptor to obtain the tensor data.
[0254] Clause 26. The data processing apparatus of clause 25, wherein the tensor interface circuit is further configured to:
[0255] Determining a data address of the tensor data in a data storage space according to the shape information; and / or
[0256] Dependencies between instructions are determined based on the spatial information.
[0257] Clause 27. A data processing apparatus according to any one of Clauses 25-26, wherein the shape information of the tensor data includes at least one shape parameter representing the shape of N-dimensional tensor data, where N is a positive integer, and the shape parameter of the tensor data includes at least one of the following:
[0258] The size of the data storage space where the tensor data is located in at least one direction of the N dimensions, the size of the storage area of the tensor data in at least one direction of the N dimensions, the offset of the storage area in at least one direction of the N dimensions, the positions of at least two vertices at diagonal positions in the N dimensions relative to the data reference point, and the mapping relationship between the data description position of the tensor data and the data address.
[0259] Clause 28. The data processing apparatus according to any one of clauses 25-26, wherein the shape information of the tensor data indicates at least one shape parameter of the shape of N-dimensional tensor data including a plurality of data blocks, where N is a positive integer, and the shape parameter comprises at least one of the following:
[0260] The size of the data storage space where the tensor data is located in at least one direction of the N dimensions, the size of the storage area of a single data block in at least one direction of the N dimensions, the block step of the data block in at least one direction of the N dimensions, the number of data blocks in at least one direction of the N dimensions, and the overall step of the data block in at least one direction of the N dimensions.
[0261] Clause 29. A data processing apparatus according to any one of clauses 1 to 28, wherein
[0262] The fused instruction includes an operation mode bit to indicate the fused processing operation of the fused instruction, or the fused instruction itself indicates the fused processing operation.
[0263] Clause 30. A chip comprising the data processing device according to any one of clauses 1 to 29.
[0264] Clause 31. A board comprising the chip according to clause 30.
[0265] Article 32. A data processing method comprising:
[0266] parsing a fusion instruction, wherein the fusion instruction instructs to perform fusion processing on multiple channels of data to be fused;
[0267] According to the fusion instruction, data elements in the multiple channels of data to be fused are merged into one channel of ordered fused data according to their corresponding indexes, wherein data elements with the same index are merged into an operation structure element, wherein the data element includes any of scalars, vectors, or higher-dimensional data; and
[0268] The fused data is output.
[0269] Clause 33. The method according to clause 32, wherein the operation objects of the fusion instruction include input K-way data to be fused, K-way indexes corresponding to the K-way data, the size of the K-way data, and output one-way fusion index, one-way fusion data and associated operation structure, K>1, wherein:
[0270] The index elements in the K-way index indicate index information of corresponding data elements in the K-way data;
[0271] The index elements of each index in the K-way index are arranged in order according to the first order;
[0272] The data elements of each path of data in the K paths of data are arranged in order according to the order of corresponding indexes;
[0273] The fusion index elements in the fusion index are arranged in order according to the second order;
[0274] The fused data elements in the fused data correspond one-to-one to the fused index elements; and
[0275] A fused data element composed of at least data elements of the same index is represented by associated operation structure elements in the operation structure.
[0276] Clause 34. A data processing method according to Clause 33, wherein the first order is the same as or different from the second order, and the first order and the second order are selected from any of the following: an order from small to large, or an order from large to small.
[0277] Clause 35. The data processing method according to any one of Clauses 33-34 further comprises:
[0278] The sorting circuit sorts the K-way indexes according to the size of the index elements and outputs them to the output circuit in order; and
[0279] At least when receiving the same index element from the sorting circuit, the output circuit generates an operation structure element representing an accumulation operation of the data elements corresponding to the same index element, and removes duplicate index elements.
[0280] Clause 36. The data processing method according to clause 35, wherein the sorting circuit comprises a comparison circuit and a buffer circuit, and the method further comprises:
[0281] The comparison circuit compares the index elements to be sorted in the K-way index with the index elements not outputted in the buffer circuit, and outputs the comparison result to the control circuit; and
[0282] The buffer circuit sequentially stores the information of the compared index elements and the corresponding data elements under the control of the control circuit, and sequentially outputs the information of the compared index elements and the corresponding data elements.
[0283] Clause 37. The data processing method according to clause 36, wherein the comparison circuit comprises a K-1 comparator, and the method comprises:
[0284] The K-1-way comparator compares the index elements to be sorted in the K-way index with the K-1 index elements of the current sequence in the buffer circuit respectively, generates a comparison result and outputs it to the control circuit.
[0285] Clause 38. The data processing method according to Clause 37 further comprises:
[0286] The control circuit determines an insertion position of the index element to be sorted in a current sequence in the buffer circuit according to the comparison result.
[0287] Clause 39. The data processing method according to Clause 38, wherein the comparison result is represented by a bitmap, and the method further comprises: the control circuit determines the insertion position according to the changed position of the bit in the bitmap.
[0288] Clause 40. The data processing method according to any one of Clauses 38-39 further comprises:
[0289] The buffer circuit inserts the index elements to be sorted and the information of the data elements corresponding thereto into the insertion position according to the instruction of the control circuit.
[0290] Clause 41. A data processing method according to any one of Clauses 36-40, further comprising:
[0291] The buffer circuit outputs information of the first or last index element and the corresponding data element in the current sequence in a specified order.
[0292] Clause 42. The data processing method according to Clause 41 further comprises:
[0293] The control circuit determines the memory access information of the next index element to be sorted according to the index element output from the buffer circuit.
[0294] Clause 43. A data processing method according to any one of clauses 35 to 42, wherein the output circuit comprises a comparator, a buffer, and a structure generator, and the method comprises:
[0295] The comparator compares the index element output from the sorting circuit with the previous fused index element and outputs a comparison result;
[0296] The buffer outputs the index element as a new fused index element only when the comparison result indicates that the index element is not identical; and
[0297] When the comparison result indicates the same, the structure generator generates an operation structure element based on the data element corresponding to the index element, and the operation structure element represents an accumulation operation of adding the data element to the fused data element corresponding to the last fused index element.
[0298] Clause 44. The data processing method according to Clause 43 further comprises:
[0299] When the comparison result indicates that the data elements are not identical, the structure generator generates an operation structure element based on the data element corresponding to the index element, wherein the operation structure element represents an accumulation operation of adding the data element to 0.
[0300] Clause 45. A data processing method according to any one of clauses 43-44, wherein each operation structure element includes three elements, indicating two addends and an addition result respectively.
[0301] Clause 46. A data processing method according to any one of clauses 32-45, wherein the data elements in the data to be fused are valid data elements after sparsification in radar-based object detection, and the index indicates the position information of the valid data elements in the data before sparsification.
[0302] Clause 47. A data processing method according to any one of clauses 33-46, wherein the operation object of the fusion instruction also includes the total number of output fusion index elements, which is used to indicate the total number of index elements in the output fusion index.
[0303] Clause 48. A data processing method according to any one of clauses 33-47, wherein the operation object of the fusion instruction also includes the total number of output operation structure elements, which is used to indicate the total number of operation structure elements in the operation structure.
[0304] Clause 49. A data processing method according to any one of clauses 33 to 48, wherein
[0305] The input K-way data is indicated by a first address, which includes K elements, the i-th element represents the address of the i-th data, and each data element in the i-th data is an address pointing to a vector of a predetermined length, where 0<i≤K.
[0306] Clause 50. A data processing method according to any one of clauses 33 to 49, wherein
[0307] The K-way index is indicated by a second address, the second address includes K elements, the i-th element represents the address of the i-th way index associated with the i-th way data, wherein 0<i≤K.
[0308] Clause 51. A data processing method according to any one of clauses 33 to 50, wherein
[0309] The size of the K-way data is indicated by a third address, the third address includes K elements, the i-th element represents the number of data elements in the i-th way data, wherein 0<i≤K.
[0310] Clause 52. A data processing method according to any one of Clauses 33 to 51, wherein
[0311] The fused data is indicated by a fourth address, which includes L elements, where the j-th element represents the j-th data element in the fused data, each data element is an address pointing to a vector of a predetermined length, and L represents the total number of data elements in the fused data, where L≥1 and 0<j≤L.
[0312] Clause 53. A data processing method according to any one of Clauses 33 to 52, wherein
[0313] The one-way fusion index is indicated by the fifth address, and the fifth address includes L elements, the jth element represents the jth index element in the one-way fusion index, and L represents the total number of index elements in the fusion index, L≥1, 0<j≤L.
[0314] Clause 54. A data processing method according to any one of Clauses 33 to 53, wherein
[0315] The operation structure is indicated by the sixth address, which includes M elements, each element being a structure.
[0316] Clause 55. The data processing method according to clause 54, wherein
[0317] The structure includes three sub-elements, each of which is an address pointing to two addends and an addition result respectively.
[0318] Clause 56. A data processing method according to any one of clauses 33 to 48, wherein at least one of the operation objects includes tensor data, and the tensor data is indicated by at least one descriptor, the descriptor indicating at least one of the following information: shape information of the tensor data and spatial information of the tensor data; and the method further comprises:
[0319] The descriptor is parsed to obtain the tensor data.
[0320] Clause 57. The data processing method according to clause 56, wherein parsing the descriptor comprises:
[0321] Determining a data address of the tensor data in a data storage space according to the shape information; and / or
[0322] Dependencies between instructions are determined based on the spatial information.
[0323] Clause 58. The data processing method according to any one of clauses 56-57, wherein the shape information of the tensor data includes at least one shape parameter representing the shape of N-dimensional tensor data, N is a positive integer, and the shape parameter of the tensor data includes at least one of the following:
[0324] The size of the data storage space where the tensor data is located in at least one direction of the N dimensions, the size of the storage area of the tensor data in at least one direction of the N dimensions, the offset of the storage area in at least one direction of the N dimensions, the positions of at least two vertices at diagonal positions in the N dimensions relative to the data reference point, and the mapping relationship between the data description position of the tensor data and the data address.
[0325] Clause 59. The data processing method according to any one of clauses 56-57, wherein the shape information of the tensor data indicates at least one shape parameter of the shape of N-dimensional tensor data including a plurality of data blocks, N being a positive integer, and the shape parameter comprising at least one of the following:
[0326] The size of the data storage space where the tensor data is located in at least one direction of the N dimensions, the size of the storage area of a single data block in at least one direction of the N dimensions, the block step of the data block in at least one direction of the N dimensions, the number of data blocks in at least one direction of the N dimensions, and the overall step of the data block in at least one direction of the N dimensions.
[0327] Clause 60. A data processing method according to any one of clauses 32 to 59, wherein
[0328] The fused instruction includes an operation mode bit to indicate the fused processing operation of the fused instruction, or the fused instruction itself indicates the fused processing operation.
[0329] The above is a detailed introduction to the embodiments of the present disclosure. Specific examples are used herein to illustrate the principles and implementation methods of the present disclosure. The description of the above embodiments is only used to help understand the method and core ideas of the present disclosure. At the same time, for those skilled in the art, based on the ideas of the present disclosure, there may be changes in the specific implementation methods and application scopes. In summary, the content of this specification should not be understood as a limitation on the present disclosure.
Claims
1. A data processing device, comprising: a control circuit configured to parse a fusion instruction, wherein the fusion instruction instructs to perform fusion processing on multiple channels of data to be fused; a storage circuit configured to store information before and / or after fusion processing; as well as an operation circuit configured to, according to the fusion instruction, merge the data elements in the multiple paths of data to be fused into one path of ordered fused data according to their corresponding indexes, wherein data elements with the same index are merged into an operation structure element, wherein the data element includes any of scalars, vectors, or higher-dimensional data; The operation objects of the fusion instruction include the input K-way data to be fused, the K-way index corresponding to the K-way data, the size of the K-way data, and the output one-way fusion index, one-way fusion data and the associated operation structure, K>1, wherein: The index elements in the K-way index indicate index information of corresponding data elements in the K-way data; The index elements of each index in the K-way index are arranged in order according to the first order; The data elements of each path of data in the K paths of data are arranged in order according to the order of corresponding indexes; The fusion index elements in the fusion index are arranged in order according to the second order; The fused data elements in the fused data correspond one-to-one to the fused index elements; and A fused data element composed of at least data elements of the same index is represented by associated operation structure elements in the operation structure. 2 . The data processing apparatus according to claim 1 , wherein the first order is the same as or different from the second order, and the first order and the second order are selected from either the following: an order from small to large, or an order from large to small.
3. The data processing device according to claim 1 , wherein the operation circuit includes a sorting circuit and an output circuit, wherein The sorting circuit is configured to sort the K-way indexes according to the size of the index elements and output them to the output circuit in order; and The output circuit is configured to generate, at least when receiving the same index element from the sorting circuit, an operation structure element representing an accumulation operation of data elements corresponding to the same index element, and remove duplicate index elements.
4. The data processing apparatus according to claim 3, wherein the sorting circuit comprises a comparison circuit and a buffer circuit, wherein: The comparison circuit is configured to compare the index elements to be sorted in the K-way index with the index elements not outputted in the buffer circuit, and output a comparison result to the control circuit; and The buffer circuit is configured to sequentially store information of the compared index elements and the corresponding data elements under the control of the control circuit, and sequentially output information of the compared index elements and the corresponding data elements.
5. The data processing apparatus according to claim 4 , wherein the comparison circuit comprises: The K-1 comparators are configured to compare the index elements to be sorted in the K-way index with the K-1 index elements of the current sequence in the buffer circuit, generate a comparison result, and output it to the control circuit. 6 . The data processing apparatus according to claim 5 , wherein the control circuit is configured to determine an insertion position of the index element to be sorted in a current sequence in the buffer circuit according to the comparison result. 7 . The data processing apparatus according to claim 6 , wherein the comparison result is represented by a bitmap, and the control circuit is further configured to determine the insertion position according to a changed position of a bit in the bitmap. 8 . The data processing apparatus according to claim 6 , wherein the buffer circuit is configured to insert the information of the index elements to be sorted and the data elements corresponding thereto into the insertion position according to an instruction of the control circuit.
9. The data processing apparatus according to claim 4, wherein the buffer circuit is further configured to output information of the first or last index element and the corresponding data element in a current sequence in a specified order. 10 . The data processing device according to claim 9 , wherein the control circuit is further configured to: determine memory access information of a next index element to be sorted according to the index element output from the buffer circuit.
11. The data processing device according to claim 3, wherein the output circuit comprises a comparator, a buffer, and a structure generator, wherein: The comparator is configured to compare the index element output from the sorting circuit with the last fused index element and output a comparison result; The buffer is configured to output the index element as a new fused index element only when the comparison result indicates a difference; and The structure generator is configured to generate an operation structure element based on the data element corresponding to the index element when the comparison result indicates the same, and the operation structure element represents an accumulation operation of adding the data element to the fused data element corresponding to the last fused index element.
12. The data processing device according to claim 11, wherein the structure generator is further configured to, when the comparison result indicates inequality, generate an operation structure element based on the data element corresponding to the index element, wherein the operation structure element represents an accumulation operation of adding the data element to 0. 13 . The data processing apparatus according to claim 11 , wherein each operation structure element comprises three elements, respectively indicating two addends and an addition result.
14. The data processing device according to any one of claims 1 to 13, wherein the data elements in the data to be fused are valid data elements after sparsification in radar-based object detection, and the index indicates position information of the valid data elements in the data before sparsification.
15. The data processing device according to any one of claims 1 to 13, wherein the operation object of the fusion instruction also includes the total number of output fusion index elements, which is used to indicate the total number of index elements in the output fusion index.
16. The data processing device according to any one of claims 1 to 13, wherein the operation object of the fusion instruction further includes the total number of output operation structure elements, which is used to indicate the total number of operation structure elements in the operation structure.
17. The data processing device according to any one of claims 1 to 13, wherein The input K-way data is indicated by the first address, which includes K elements. The i-th element represents the address of the i-way data. Each data element in the i-way data is an address pointing to a vector of a predetermined length, where 0<i≤K.
18. The data processing device according to any one of claims 1 to 13, wherein The K-way index is indicated by a second address, the second address includes K elements, the i-th element represents the address of the i-th way index associated with the i-th way data, wherein 0<i≤K.
19. The data processing device according to any one of claims 1 to 13, wherein The size of the K-way data is indicated by a third address, the third address includes K elements, the i-th element represents the number of data elements in the i-th way data, wherein 0<i≤K.
20. The data processing device according to any one of claims 1 to 13, wherein The fused data is indicated by a fourth address, the fourth address includes L elements, the j-th element represents the j-th data element in the fused data, each data element is an address pointing to a vector of a predetermined length, L represents the total number of data elements in the fused data, L≥1, 0<j≤L.
21. The data processing device according to any one of claims 1 to 13, wherein The one-way fusion index is indicated by the fifth address, and the fifth address includes L elements, the jth element represents the jth index element in the one-way fusion index, and L represents the total number of index elements in the fusion index, L≥1, 0<j≤L.
22. The data processing device according to any one of claims 1 to 13, wherein The operation structure is indicated by a sixth address, which includes M elements, each of which is an operation structure element.
23. The data processing apparatus according to claim 22, wherein The operation structure element includes three sub-elements, each of which is an address pointing to two addends and an addition result respectively.
24. The data processing device according to any one of claims 1 to 13, wherein: At least one of the operation objects includes tensor data, and the tensor data is indicated by at least one descriptor, the descriptor indicating at least one of the following information: shape information of the tensor data and spatial information of the tensor data; and The data processing device further includes a tensor interface circuit configured to parse the descriptor to obtain the tensor data.
25. The data processing apparatus according to claim 24, wherein the tensor interface circuit is further configured to: Determining a data address of the tensor data in a data storage space according to the shape information; and / or Dependencies between instructions are determined based on the spatial information.
26. The data processing apparatus according to claim 24, wherein the shape information of the tensor data comprises at least one shape parameter representing the shape of N-dimensional tensor data, N being a positive integer, and the shape parameter of the tensor data comprises at least one of the following: The size of the data storage space where the tensor data is located in at least one direction of the N dimensions, the size of the storage area of the tensor data in at least one direction of the N dimensions, the offset of the storage area in at least one direction of the N dimensions, the positions of at least two vertices at diagonal positions in the N dimensions relative to the data reference point, and the mapping relationship between the data description position of the tensor data and the data address.
27. The data processing apparatus according to claim 24, wherein the shape information of the tensor data indicates at least one shape parameter of a shape of N-dimensional tensor data including a plurality of data blocks, N being a positive integer, and the shape parameter comprises at least one of the following: The size of the data storage space where the tensor data is located in at least one direction of the N dimensions, the size of the storage area of a single data block in at least one direction of the N dimensions, the block step of the data block in at least one direction of the N dimensions, the number of data blocks in at least one direction of the N dimensions, and the overall step of the data block in at least one direction of the N dimensions.
28. The data processing apparatus according to any one of claims 1 to 13, wherein The fused instruction includes an operation mode bit to indicate the fused processing operation of the fused instruction, or the fused instruction itself indicates the fused processing operation.
29. A chip comprising the data processing device according to any one of claims 1 to 28.
30. A board comprising the chip according to claim 29.
31. A data processing method comprising: parsing a fusion instruction, wherein the fusion instruction instructs to perform fusion processing on multiple channels of data to be fused; According to the fusion instruction, data elements in the multiple channels of data to be fused are merged into one channel of ordered fused data according to their corresponding indexes, wherein data elements with the same index are merged into an operation structure element, wherein the data element includes any of scalars, vectors, or higher-dimensional data; and Outputting the fused data; The operation objects of the fusion instruction include the input K-way data to be fused, the K-way index corresponding to the K-way data, the size of the K-way data, and the output one-way fusion index, one-way fusion data and the associated operation structure, K>1, wherein: The index elements in the K-way index indicate index information of corresponding data elements in the K-way data; The index elements of each index in the K-way index are arranged in order according to the first order; The data elements of each path of data in the K paths of data are arranged in order according to the order of corresponding indexes; The fusion index elements in the fusion index are arranged in order according to the second order; The fused data elements in the fused data correspond one-to-one to the fused index elements; and A fused data element consisting of at least data elements with the same index is represented by associated operation structure elements in the operation structure; The data elements in the data to be fused are valid data elements after sparseness in radar-based object detection, and the index indicates position information of the valid data elements in the data before sparseness.
32. The data processing method according to claim 31, wherein the first order is the same as or different from the second order, and the first order and the second order are selected from any one of the following: an order from small to large, or an order from large to small.
33. The data processing method according to claim 31, further comprising: The sorting circuit sorts the K-way indexes according to the size of the index elements and outputs them to the output circuit in order; as well as At least when receiving the same index element from the sorting circuit, the output circuit generates an operation structure element representing an accumulation operation of the data elements corresponding to the same index element, and removes duplicate index elements.
34. The data processing method according to claim 33, wherein the sorting circuit comprises a comparison circuit and a buffer circuit, and the method further comprises: The comparison circuit compares the index elements to be sorted in the K-way index with the index elements not outputted in the buffer circuit, and outputs the comparison result to the control circuit; as well as The buffer circuit sequentially stores the information of the compared index elements and the corresponding data elements under the control of the control circuit, and sequentially outputs the information of the compared index elements and the corresponding data elements.
35. The data processing method according to claim 34, wherein the comparison circuit comprises a K-1 comparator, and the method comprises: The K-1-way comparator compares the index elements to be sorted in the K-way index with the K-1 index elements of the current sequence in the buffer circuit respectively, generates a comparison result and outputs it to the control circuit.
36. The data processing method according to claim 35, further comprising: The control circuit determines an insertion position of the index element to be sorted in a current sequence in the buffer circuit according to the comparison result.
37. The data processing method according to claim 36, wherein the comparison result is represented by a bitmap, and the method further comprises: The control circuit determines the insertion position according to the changed position of the bit in the bitmap.
38. The data processing method according to claim 36, further comprising: The buffer circuit inserts the index elements to be sorted and the information of the data elements corresponding thereto into the insertion position according to the instruction of the control circuit.
39. The data processing method according to claim 34, further comprising: The buffer circuit outputs information of the first or last index element and the corresponding data element in the current sequence in a specified order.
40. The data processing method according to claim 39, further comprising: The control circuit determines the memory access information of the next index element to be sorted according to the index element output from the buffer circuit.
41. The data processing method according to claim 33, wherein the output circuit comprises a comparator, a buffer and a structure generator, and the method comprises: The comparator compares the index element output from the sorting circuit with the previous fused index element and outputs a comparison result; The buffer outputs the index element as a new fused index element only when the comparison result indicates that the index element is not identical; as well as When the comparison result indicates the same, the structure generator generates an operation structure element based on the data element corresponding to the index element, and the operation structure element represents an accumulation operation of adding the data element to the fused data element corresponding to the last fused index element.
42. The data processing method according to claim 41, further comprising: When the comparison result indicates that the data elements are not identical, the structure generator generates an operation structure element based on the data element corresponding to the index element, wherein the operation structure element represents an accumulation operation of adding the data element to 0.
43. The data processing method according to claim 41, wherein each operation structure element includes three elements, respectively indicating two addends and an addition result.
44. The data processing method according to any one of claims 31-43, wherein the operation object of the fusion instruction also includes the total number of output fusion index elements, which is used to indicate the total number of index elements in the output fusion index.
45. The data processing method according to any one of claims 31-43, wherein the operation object of the fusion instruction also includes the total number of output operation structure elements, which is used to indicate the total number of operation structure elements in the operation structure.
46. The data processing method according to any one of claims 31 to 43, wherein The input K-way data is indicated by the first address, which includes K elements. The i-th element represents the address of the i-way data. Each data element in the i-way data is an address pointing to a vector of a predetermined length, where 0<i≤K.
47. The data processing method according to any one of claims 31 to 43, wherein The K-way index is indicated by a second address, the second address includes K elements, the i-th element represents the address of the i-th way index associated with the i-th way data, wherein 0<i≤K.
48. The data processing method according to any one of claims 31 to 43, wherein The size of the K-way data is indicated by a third address, the third address includes K elements, the i-th element represents the number of data elements in the i-th way data, wherein 0<i≤K.
49. The data processing method according to any one of claims 31 to 43, wherein The fused data is indicated by a fourth address, the fourth address includes L elements, the j-th element represents the j-th data element in the fused data, each data element is an address pointing to a vector of a predetermined length, L represents the total number of data elements in the fused data, L≥1, 0<j≤L.
50. The data processing method according to any one of claims 31 to 43, wherein The one-way fusion index is indicated by the fifth address, and the fifth address includes L elements, the jth element represents the jth index element in the one-way fusion index, and L represents the total number of index elements in the fusion index, L≥1, 0<j≤L.
51. The data processing method according to any one of claims 31 to 43, wherein The operation structure is indicated by a sixth address, which includes M elements, each of which is an operation structure element.
52. The data processing method according to claim 51, wherein The operation structure element includes three sub-elements, each of which is an address pointing to two addends and an addition result respectively.
53. The data processing method according to any one of claims 31 to 43, wherein: At least one of the operation objects includes tensor data, and the tensor data is indicated by at least one descriptor, the descriptor indicating at least one of the following information: shape information of the tensor data and spatial information of the tensor data; And the method further includes: The descriptor is parsed to obtain the tensor data.
54. The data processing method according to claim 53, wherein parsing the descriptor comprises: Determining a data address of the tensor data in a data storage space according to the shape information; and / or Dependencies between instructions are determined based on the spatial information.
55. The data processing method according to claim 53, wherein the shape information of the tensor data comprises at least one shape parameter representing the shape of N-dimensional tensor data, N is a positive integer, and the shape parameter of the tensor data comprises at least one of the following: The size of the data storage space where the tensor data is located in at least one direction of the N dimensions, the size of the storage area of the tensor data in at least one direction of the N dimensions, the offset of the storage area in at least one direction of the N dimensions, the positions of at least two vertices at diagonal positions in the N dimensions relative to the data reference point, and the mapping relationship between the data description position of the tensor data and the data address.
56. The data processing method according to claim 53, wherein the shape information of the tensor data indicates at least one shape parameter of the shape of N-dimensional tensor data including a plurality of data blocks, N is a positive integer, and the shape parameter includes at least one of the following: The size of the data storage space where the tensor data is located in at least one direction of the N dimensions, the size of the storage area of a single data block in at least one direction of the N dimensions, the block step of the data block in at least one direction of the N dimensions, the number of data blocks in at least one direction of the N dimensions, and the overall step of the data block in at least one direction of the N dimensions.
57. The data processing method according to any one of claims 31 to 43, wherein The fused instruction includes an operation mode bit to indicate the fused processing operation of the fused instruction, or the fused instruction itself indicates the fused processing operation.
Citation Information
Patent Citations
Data storage device, and method for controlling the same
JP2008299513A
Merging entries in a deduplciation index
US20130325821A1