Data processing devices, data processing methods and related products
By providing dedicated fusion instructions and hardware circuits on devices with limited hardware resources, the system enables the merging and sorting of multiple data streams, solving the efficiency problem of sparse processing, improving processing efficiency, and supporting the application of deep learning technology in embedded and mobile devices.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-12-25
- Publication Date
- 2026-04-03
AI Technical Summary
Existing hardware and/or instruction sets cannot effectively support sparsification and related processing, making it difficult to apply deep learning technology on devices with limited hardware resources.
A data processing apparatus and method are provided, which realize the merging and sorting of multiple data streams through specialized fusion instructions and hardware circuits, including control circuits, storage circuits and arithmetic circuits, supporting data fusion processing and simplifying and accelerating the processing flow.
Data fusion processing is achieved through specialized fusion instructions and hardware circuits, which improves processing efficiency. It is suitable for devices with limited hardware resources and supports the application of deep learning technology in embedded and mobile devices.
Smart Images

Figure CN114692840B_ABST
Abstract
Description
Technical Field
[0001] This disclosure generally relates to the field of processors. More specifically, this disclosure relates to data processing apparatus, data processing methods, chips, and circuit boards. Background Technology
[0002] In recent years, the rapid development of deep learning has led to leaps in the performance of algorithms in fields such as computer vision and natural language processing. However, deep learning algorithms are computationally and storage-intensive tools. As information processing tasks become increasingly complex and the requirements for real-time performance and accuracy of algorithms continue to rise, neural networks are often designed to be deeper and deeper, resulting in ever-increasing computational and storage demands. This makes it difficult to directly apply existing deep learning-based artificial intelligence technologies to devices with limited hardware resources, such as mobile phones, satellites, or embedded devices.
[0003] Therefore, the compression, acceleration, and optimization of deep neural network models have become extremely important. Numerous studies have attempted to reduce the computational and storage requirements of neural networks without compromising model accuracy, which is of great significance for the engineering application of deep learning technology in embedded and mobile devices. Sparsity is one such method for lightweighting models.
[0004] Network parameter sparsification reduces redundant components in large networks through appropriate methods, thereby lowering the network's computational and storage requirements. Existing hardware and / or instruction sets cannot effectively support sparsification processing and / or related post-sparsing processing. Summary of the Invention
[0005] In order to at least partially solve one or more of the technical problems mentioned in the background art, the present disclosure provides a data processing apparatus, a data processing method, a chip, and a board.
[0006] In a first aspect, this disclosure discloses a data processing apparatus, comprising: a control circuit configured to parse a fusion instruction, the fusion instruction indicating that multiple data streams to be fused should be merged and sorted; a storage circuit configured to store information before and / or after the merge and sorting process; and a processing circuit configured to, according to the fusion instruction, merge the multiple data streams to be fused into a single fused data stream and output the fused data stream in an orderly manner.
[0007] In a second aspect, this disclosure provides a chip that includes the data processing apparatus of any of the embodiments of the first aspect.
[0008] In a third aspect, this disclosure provides a board including the chip of any of the embodiments of the second aspect above.
[0009] In the fourth aspect, this disclosure provides a data processing method, which includes: parsing a fusion instruction, the fusion instruction indicating that multiple data to be fused should be merged and sorted; merging the multiple data to be fused into one fused data according to the fusion instruction; and outputting the fused data in an orderly manner.
[0010] Using the data processing apparatus, data processing method, chip, and board provided above, this disclosure provides a fusion instruction for performing operations related to the merge sorting of multi-channel data. In some embodiments, the fusion instruction is a hardware instruction, implemented through dedicated hardware circuitry for data fusion processing. In some embodiments, the fusion instruction may include an operation mode bit to indicate that the fusion instruction is a merge sorting operation, or the fusion instruction itself may indicate a merge sorting operation. By providing a dedicated fusion instruction to perform operations related to the fusion processing of multi-channel data, processing can be simplified. Furthermore, by providing a dedicated hardware implementation of data fusion-related operations, processing can be accelerated, thereby improving machine processing efficiency. Attached Figure Description
[0011] The above and other objects, features, and advantages of exemplary embodiments of this disclosure will become readily apparent upon reading the following detailed description with reference to the accompanying drawings. In the drawings, several embodiments of this disclosure are illustrated by way of example and not limitation, and like or corresponding reference numerals denote like or corresponding parts, wherein:
[0012] Figure 1 This diagram shows the structure of the board card according to an embodiment of this disclosure;
[0013] Figure 2 This diagram illustrates the structure of the combined processing apparatus according to an embodiment of the present disclosure.
[0014] Figure 3 A schematic diagram illustrating the internal structure of a processor core in a single-core or multi-core computing device according to embodiments of the present disclosure;
[0015] Figure 4 This illustrates an exemplary principle of data fusion processing according to embodiments of this disclosure;
[0016] Figure 5 A structural block diagram of a data processing circuit according to an embodiment of this disclosure is shown;
[0017] Figure 6 An exemplary circuit diagram for data fusion processing according to one embodiment of this disclosure is shown;
[0018] Figure 7 An exemplary circuit diagram for data fusion processing according to another embodiment of this disclosure is shown;
[0019] Figure 8 The example illustrates the contents pointed to by each address in the fusion instruction;
[0020] Figure 9 A schematic diagram of a data storage space according to an embodiment of this disclosure is shown;
[0021] Figure 10 A schematic diagram of data blocks in a data storage space according to an embodiment of this disclosure is shown;
[0022] Figure 11 A structural block diagram of a data processing apparatus according to another embodiment of this disclosure is shown; and
[0023] Figure 12 An exemplary flowchart of a data processing method according to an embodiment of this disclosure is shown. Detailed Implementation
[0024] The technical solutions in the embodiments of this disclosure will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this disclosure, not all of them. Based on the embodiments in this disclosure, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this disclosure.
[0025] It should be understood that the terms "first," "second," "third," and "fourth," etc., in the claims, specification, and drawings of this disclosure are used to distinguish different objects, not to describe a specific order. The terms "comprising" and "including" as used in the specification and claims of this disclosure indicate the presence of the described features, integrals, steps, operations, elements, and / or components, but do not exclude the presence or addition of one or more other features, integrals, steps, operations, elements, components, and / or collections thereof.
[0026] It should also be understood that the terminology used in this disclosure is for the purpose of describing particular embodiments only and is not intended to limit the disclosure. As used in this disclosure and claims, the singular forms “a,” “an,” and “the” are intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term “and / or” as used in this disclosure and claims refers to any combination and all possible combinations of one or more of the associated listed items, and includes such combinations.
[0027] As used in this specification and claims, the term "if" may be interpreted, depending on the context, as "when," "once," "in response to determination," or "in response to detection."
[0028] The specific embodiments disclosed herein will now be described in detail with reference to the accompanying drawings.
[0029] Figure 1 A schematic diagram of the structure of a board 10 according to an embodiment of this disclosure is shown. Figure 1 As shown, board 10 includes chip 101, which is a system-on-chip (SoC) integrating one or more combined processing units. These combined processing units are artificial intelligence computing units used to support various deep learning and machine learning algorithms, meeting the intelligent processing needs of complex scenarios in fields such as computer vision, speech, natural language processing, and data mining. In particular, deep learning technology is widely used in cloud intelligence. A significant characteristic of cloud intelligence applications is the large volume of input data, placing high demands on the platform's storage and computing capabilities. Board 10 in this embodiment is suitable for cloud intelligence applications, possessing massive off-chip storage, on-chip storage, and powerful computing capabilities.
[0030] Chip 101 is connected to external device 103 via external interface device 102. External device 103 may be, for example, a server, computer, camera, monitor, mouse, keyboard, network card, or Wi-Fi interface. Data to be processed can be transmitted from external device 103 to chip 101 via external interface device 102. The calculation results from chip 101 can be transmitted back to external device 103 via external interface device 102. Depending on the application scenario, external interface device 102 may have different interface forms, such as a PCIe interface.
[0031] The board 10 also includes a storage device 104 for storing data, which includes one or more memory cells 105. The storage device 104 is connected to and transmits data with the controller 106 and the chip 101 via a bus. The controller 106 in the board 10 is configured to regulate the state of the chip 101. Therefore, in one application scenario, the controller 106 may include a microcontroller (MCU).
[0032] Figure 2 This is a structural diagram illustrating the combined processing device in chip 101 of this embodiment. (As shown) Figure 2 As shown, the combined processing device 20 includes a computing device 201, an interface device 202, a processing device 203, and a storage device 204.
[0033] The computing device 201 is configured to perform user-specified operations. It is mainly implemented as a single-core intelligent processor or a multi-core intelligent processor to perform deep learning or machine learning calculations. It can interact with the processing device 203 through the interface device 202 to jointly complete the user-specified operations.
[0034] Interface device 202 is used to transmit data and control commands between computing device 201 and processing device 203. For example, computing device 201 can obtain input data from processing device 203 via interface device 202 and write it to on-chip storage device of computing device 201. Further, computing device 201 can obtain control commands from processing device 203 via interface device 202 and write them to on-chip control cache of computing device 201. Alternatively or optionally, interface device 202 can also read data from storage device of computing device 201 and transmit it to processing device 203.
[0035] Processing device 203, as a general-purpose processing device, performs basic control including but not limited to data transfer, and starting and / or stopping computing device 201. Depending on the implementation, processing device 203 may be one or more types of processors, including but not limited to digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc., and their number can be determined according to actual needs. As mentioned above, computing device 201 disclosed herein can be considered as having a single-core structure or a homogeneous multi-core structure. However, when computing device 201 and processing device 203 are considered together, they are considered to form a heterogeneous multi-core structure.
[0036] Storage device 204 is used to store data to be processed. It may be DRAM or DDR memory, typically 16G or larger in size, and is used to store data of computing device 201 and / or processing device 203.
[0037] Figure 3 The diagram shows the internal structure of the processor core when the computing device 201 is a single-core or multi-core device. The computing device 301 is used to process input data such as computer vision, speech, natural language, and data mining. The computing device 301 includes three main modules: a control module 31, an arithmetic module 32, and a storage module 33.
[0038] The control module 31 coordinates and controls the operation of the computation module 32 and the storage module 33 to complete the deep learning task. It includes an instruction fetch unit (IFU) 311 and an instruction decode unit (IDU) 312. The instruction fetch unit 311 fetches instructions from the processing device 203, and the instruction decode unit 312 decodes the fetched instructions and sends the decoding result as control information to the computation module 32 and the storage module 33.
[0039] The computation module 32 includes a vector operation unit 321 and a matrix operation unit 322. The vector operation unit 321 is used to perform vector operations and can support complex operations such as vector multiplication, addition, and nonlinear transformations; the matrix operation unit 322 is responsible for the core computations of deep learning algorithms, namely matrix multiplication and convolution.
[0040] Storage module 33 is used to store or move relevant data, including neuron RAM (NRAM) 331, weight RAM (WRAM) 332, and direct memory access (DMA) module 333. NRAM 331 is used to store input neurons, output neurons, and intermediate results after computation; WRAM 332 is used to store the convolution kernels of the deep learning network, i.e., the weights; DMA 333 is connected to DRAM 204 through bus 34 and is responsible for data transfer between computing device 301 and DRAM 204.
[0041] Based on the aforementioned hardware environment, the embodiments disclosed herein provide a data processing scheme that executes operations related to the fusion of multi-channel data according to specialized fusion instructions. As mentioned in the background art, network parameter sparsification can effectively reduce the network's computational and storage requirements. However, network parameter sparsification also brings a series of impacts to subsequent processing. For example, in sparse data processing, it may be necessary to merge and sort multiple sparse data streams; or, in sparse matrix multiplication, it may be necessary to sort and accumulate vectors. In view of this, the embodiments disclosed herein provide a specialized fusion instruction to support data fusion processing. Furthermore, this instruction can be implemented in conjunction with a specialized data fusion processing hardware scheme to simplify and accelerate such processing.
[0042] Figure 4This diagram illustrates an exemplary principle of data fusion processing according to embodiments of this disclosure. The diagram exemplarily shows four data streams, each containing six data elements, with the data elements in each stream arranged in a first order (e.g., ascending order). After data fusion, these four data streams are merged into a single fused data stream, comprising 24 data elements, and the merged data elements are arranged in a second order (e.g., ascending order). In this merge sorting process, duplicate data elements are preserved.
[0043] Those skilled in the art will understand that the first order and the second order may be the same or different, and both may be selected from either: an ascending order or a descending order. Those skilled in the art will also understand that the number of data elements in each data stream may be the same or different, and this disclosure does not limit this.
[0044] In the embodiments disclosed herein, data elements are scalars. Therefore, a data path that includes several data elements can be called a data vector, and the length of the vector is equal to the number of data elements it contains.
[0045] Figure 5 A structural block diagram of a data processing circuit 500 according to an embodiment of this disclosure is shown. The data processing circuit 500 can, for example, be implemented in... Figure 2 In the computing device 201, as shown in the figure, the data processing circuit 500 may include a control circuit 510, a storage circuit 520, and an arithmetic circuit 530.
[0046] The function of control circuit 510 can be similar to Figure 3 The control module 31 may include, for example, an instruction fetching unit for fetching instructions from, for example... Figure 2 The processing device 203 has instructions and an instruction decoding unit for decoding the acquired instructions and sending the decoding results as control information to the arithmetic circuit 530 and the storage circuit 520.
[0047] In one embodiment, the control circuit 510 may be configured to parse a fusion instruction, wherein the fusion instruction instructs the merging and sorting of multiple streams of data to be merged.
[0048] The storage circuit 520 can be configured to store various types of information, including at least information before and / or after the merge sort process. The storage circuit could be, for example, […]. Figure 3 WRAM332.
[0049] The arithmetic circuit 530 can be configured to perform corresponding operations according to the fusion command. Specifically, the arithmetic circuit 530 can merge multiple data streams to be fused into a single fused data stream and output the fused data in an orderly manner.
[0050] In one embodiment, the arithmetic circuit 530 may further include an arithmetic processing circuit (not shown), which may be configured to preprocess the data before the arithmetic circuit performs the operation or postprocess the data after the operation according to the arithmetic instructions. In some application scenarios, the aforementioned preprocessing and postprocessing may include, for example, data splitting and / or data concatenation operations.
[0051] There are many ways to implement arithmetic circuits. Figure 6 An exemplary circuit diagram for data fusion processing according to one embodiment of this disclosure is shown.
[0052] As shown in the figure, in one embodiment, the storage circuit can be exemplarily divided into two parts: a first storage circuit 622 and a second storage circuit 624.
[0053] The first storage circuit 622 can be configured to store K channels of data to be merged, where K>1, and the data elements of each of the K channels are arranged in a first order. An example is shown in the figure. Figure 4 The four data streams are shown. In some embodiments, each data stream is stored contiguously, for example, as a vector, so that the data stream / vector can be accessed based on the starting address of each data stream or the starting address of the vector.
[0054] The second storage circuit 624 can be configured to store the fused data output by the arithmetic circuit, and the data elements in the output fused data are arranged in a second order. As shown in the figure, the four data streams to be fused are transformed into one fused data stream, in which the fused data elements are arranged in ascending order, and data elements of the same size are retained and repeatedly output.
[0055] In some embodiments, the arithmetic circuit may include a comparison circuit 632 and a buffer circuit 634. The comparison circuit 632 performs a comparison function, comparing the sizes of data elements in the multiple streams of data to be merged, and submitting the comparison results to the control circuit 610 for sorting. The control circuit 610 determines the insertion position of the data element in the buffer circuit 634 based on the comparison results. The buffer circuit 634 is used to buffer the compared data elements, and buffers them in order of size.
[0056] Specifically, the comparison circuit 632 can be configured to compare data elements in the data to be fused with data elements not yet output in the buffer circuit 634, and output the comparison result to the control circuit 610. The buffer circuit 634 can be configured to, according to the control of the control circuit 610, orderly store the compared data elements, and orderly output the compared data elements as fused data.
[0057] In some embodiments, buffer circuit 634 can be configured to buffer K data elements, which are sorted by size. Those skilled in the art will understand that the buffer circuit can also be configured to buffer more data elements, and the embodiments disclosed herein are not limited in this respect. Depending on the sorting method in buffer circuit 634 and the desired output sorting method, such as ascending or descending, the first or last data element in the current sequence can be output in a specified order each time. For example, in the example in the figure, buffer circuit 634 buffers data elements from left to right in descending order, outputting the rightmost data element each time, which is the smallest data element in the current sequence, such as "7".
[0058] In these embodiments, the comparison circuit 632 may include a K-1 comparator configured to compare the data element to be fused with the data element not yet output in the buffer circuit 634, that is, with the K-1 data elements remaining after the first or last data element of the current sequence is output, generate a comparison result and output it to the control circuit 610.
[0059] For example, for four data streams to be merged, a three-way comparator is shown in the figure, which compares a specified data element (9 in this case) received from the first storage circuit 622 with three data elements that are not currently output from the buffer circuit 634. The three data elements on the left in the figure are 100, 10 and 9.
[0060] In some embodiments, the comparison result of the comparator can be represented using a bitmap. For example, if the data element to be merged (e.g., 9) is greater than or equal to the data element in the buffer circuit, the comparator can output "1", otherwise, it outputs "0"; and vice versa. In the example in the figure, the comparison result of the data element to be merged (9) with the respective data elements (100, 10 and 9) in the buffer circuit is "001", which is output to the control circuit 610.
[0061] The control circuit 610 can be configured to determine the insertion position of the data element to be fused in the current sequence of the buffer circuit 634 based on the received comparison result. Specifically, the control circuit 610 can be further configured to determine the insertion position based on the change position of the bits in the bitmap. In the example shown in the figure, the comparison result is "001", indicating that the current data element to be fused is less than the first and second data elements from the left in the buffer circuit, and greater than or equal to the third data element from the left. Therefore, the insertion position is between the second and third data elements, that is, between "10" and "9".
[0062] In some embodiments, the buffer circuit 634 may be configured to insert the data element to be merged at the insertion position according to the instruction of the control circuit 610. In the example shown in the figure, the sequence after the data element is inserted into the buffer circuit 634 becomes "100,10,9,9".
[0063] Next, the buffer circuit 634 can output the rightmost data element "9". At this time, the control circuit 610 can be further configured to determine the memory access information of the next data element to be merged based on the data element output from the buffer circuit. Specifically, the control circuit retrieves the next data element to be merged from which data path the output data element belongs in the K data path and sends it to the comparison circuit 632 for comparison.
[0064] For clarity, the figure also shows the data sequence buffered in buffer circuit 634 as the sorting progresses. As shown, initially, the first data element of each of the K data paths is stored in buffer circuit 634 in descending order. In some implementations, these four data elements can be retrieved, sorted, and stored in the buffer circuit all at once. In other implementations, the data in the buffer circuit can be initialized to negative numbers, and the first data element of each path can be retrieved sequentially (e.g., from path 1 to path 4), compared with the data in the buffer circuit, and placed in the appropriate position. In this example, the first data element of all four paths is 0, so they can be arranged according to the sequence number of each path based on the order of retrieval; for example, the "0" of path 1 is placed on the far right, the "0" of path 2 is placed in the second position from the right, and so on.
[0065] Next, the rightmost "0" belonging to channel 1 in the buffer circuit is output. Based on which channel this output data element belongs to, the next data element to be merged is retrieved from that channel, namely the second data element "2" in channel 1. "2" is sent to the comparison circuit and compared with the remaining three "0"s in the buffer circuit. The comparison result is "111", which is greater than all three existing "0"s in the buffer circuit. Therefore, "2" is inserted at the end of the sequence, and the sequence in the buffer circuit becomes "2,0,0,0".
[0066] Next, the rightmost "0" belonging to the second path in the buffer circuit is output. Therefore, the second element "3" of the second path is taken out and compared with the remaining "2,0,0" in the buffer circuit. The comparison result is "111", so "3" is inserted at the end of the sequence. At this time, the sequence in the buffer circuit becomes "3,2,0,0".
[0067] Next, the rightmost "0" belonging to the third path in the buffer circuit is output. Therefore, the second element of the third path, "100", is taken out and compared with the remaining "2,0,0" in the buffer circuit. The comparison result is "111", so "100" is inserted at the end of the sequence. At this time, the sequence in the buffer circuit becomes "100,3,2,0".
[0068] Next, the rightmost "0" belonging to the 4th channel of the output buffer circuit is taken out and the second element "2" of the 4th channel is compared with the remaining "100,3,2" in the buffer circuit. The comparison result is "001", so "2" is inserted after the rightmost first element of the sequence. At this time, the sequence in the buffer circuit becomes "100,3,2,2".
[0069] Similarly, data elements in the K-channel data can be compared one by one, sorted according to size, and inserted into the appropriate positions in the buffer circuit before being output by the buffer circuit. For example, the smallest data element output by the buffer circuit each time can be stored sequentially in the second storage circuit 624. Those skilled in the art will understand that if the buffer circuit has sufficient space, the merged and sorted data elements can also be output uniformly after the sorting is completed.
[0070] As can be seen from the merged sorted data elements, when data elements of the same size exist, the merged data still retains these elements of the same size, and no deduplication operation is performed. Therefore, combining the above... Figure 6 The detailed circuit diagram describes the merge sorting scheme provided in the embodiments disclosed herein.
[0071] In some application scenarios, the multiple data streams to be merged may be multiple indexes, and these indexes correspond one-to-one with multiple associated data streams. The index element in each index indicates the index information of the corresponding associated data element in that associated data stream. For example, in a sparse vector, data elements at certain positions are retained as valid data elements, while data elements at other positions are discarded or set to zero. The position information of these valid data elements in the vector before sparsification can be identified by the index. In these application scenarios, there may be multiple sparse data streams, such as multiple sparse vectors, and it is necessary to merge these multiple sparse data streams into a single data stream, where the data elements are ordered according to the index.
[0072] At this point, in addition to merging multiple indexes into a single ordered fused index, the data processing circuit in this embodiment is also configured to merge the multiple associated data into a single ordered fused associated data, and the order of data elements in the fused associated data remains consistent with the order of data elements in the fused index. That is, after the merge sorting process, the associated data and the index always maintain a one-to-one binding relationship.
[0073] Figure 7 An exemplary circuit diagram for data fusion processing according to another embodiment of this disclosure is shown. Figure 7 The data to be merged consists of K-way indexes and their corresponding K-way associated data. Figure 7 Implementation examples and Figure 6 The difference lies in the fact that, in addition to performing merge sort on the K-way index, a similar sorting process is also performed on the associated data. Those skilled in the art will understand that... Figure 7 The K-way index in the middle is equivalent to Figure 6 K-way data in [the context]. To avoid confusion, in [the context]... Figure 7 It uses the representation of K-way indexes and K-way related data.
[0074] As shown in the figure, the first storage circuit 722 stores not only the K-way indices to be merged, but also K-way associated data corresponding one-to-one with these K-way indices. As shown in the figure, the index elements of each of these K-way indices are arranged in a first order (e.g., from smallest to largest). The index element in each index indicates the index information of the corresponding associated data element in the corresponding associated data path. The figure exemplarily shows 4-way indices and corresponding 4-way associated data. As shown in the figure, the index of the first data element D11 of the first associated data path is 0, the index of the second data element D12 is 2, the index of the third data element D13 is 5, and so on. The index of the first data element D21 of the second associated data path is 0, the index of the second data element D22 is 3, and so on. In some embodiments, each index or each associated data path is stored contiguously, for example, as an index vector or associated data vector, so that the data / vector can be accessed according to the starting address of each data path or the starting address of the vector.
[0075] To ensure that the associated data and indexes maintain a one-to-one correspondence after merge sorting, in some embodiments, the buffer circuit 734 can be further configured to: orderly store the compared index elements and their corresponding associated data elements according to the value order of the index elements. As shown in the figure, the buffer circuit 734 caches not only the index elements but also their corresponding associated data elements. Therefore, after each comparison of the index elements to determine the insertion position, the associated data element corresponding to that index element can also be inserted into the buffer circuit. Those skilled in the art will understand that the associated data element can be the associated data element itself, such as D32, D23, etc., as exemplarily shown in the figure; the associated data element can also be an address pointing to that associated data element, and the embodiments disclosed herein are not limited in this respect.
[0076] Furthermore, during ordered output, the buffer circuit 734 can be configured to output the compared index elements in order of their values (e.g., from smallest to largest) as fusion indexes, and simultaneously output their corresponding associated data elements as fusion associated data. The output data is stored, for example, in the second storage circuit 724. As shown, the output data can include two vectors: a fusion index vector and a fusion associated data vector.
[0077] The merged sorted data shows that when the multi-way index contains index elements of the same size, these index elements are repeatedly output in the merged index, and the associated data elements corresponding to these index elements are synchronously output in the merged associated data. Therefore, combining the above... Figure 7 The detailed circuit diagram describes a merge sorting scheme provided in another embodiment of this disclosure.
[0078] Those skilled in the art will understand that other forms of hardware circuits can be designed to implement the above-described merge sorting process, and this disclosure is not limited in this respect.
[0079] In this disclosed embodiment, data merging and sorting processing can be implemented using the exemplary hardware circuit described above by invoking a fusion instruction. The fusion instruction operates on K input data streams to be merged, the size of the K data streams, and one output fused data stream, where K > 1. The data elements of each of the K data streams are arranged in a first order, and the data elements of the output fused data stream are arranged in a second order. In some embodiments, the fusion instruction may further operate on the total number of output fused data elements, indicating the number of data elements in the output fused data stream.
[0080] As mentioned above, the first order and the second order can be the same or different, and the first order and the second order can be selected from either the following: an order from smallest to largest, or an order from largest to smallest.
[0081] In some embodiments, at least one operand of a fusion instruction can be characterized using an address.
[0082] Figure 8 The example illustrates the contents pointed to by each address in the fusion instruction.
[0083] For example, the input K-way data can be indicated by a first address, which includes K elements, where the i-th element represents the starting address of the i-th data, and 0 < i ≤ K.
[0084] As mentioned earlier, in some application scenarios, the K-way data to be merged includes K-way indices, so the first address can be marked as index_addr. This address contains K elements, representing the merging operation of the K-way indices. index_addr is a double pointer where the K elements represent the starting address of the K-way indices to be merged (e.g., vectors).
[0085] The size of the K input data paths can be indicated by a second address. This second address is a first-level pointer, denoted as `size_addr`, which also contains K elements. The i-th element represents the number of data elements in the i-th data path, where 0 < i ≤ K. In the above application scenario, the i-th element in `size_addr` represents the number of index elements in the i-th index path.
[0086] In some embodiments, the index elements in the input K-way index are arranged in ascending order, for example, and the final output merged index elements can also be arranged in ascending order. In the merge sort process disclosed herein, when there are duplicate indices, the same index will be output repeatedly.
[0087] The output fused data can be stored in the third address, which is indicated in the fusion instruction. In the above application scenario, the third address can be marked as out_index_addr, which is the address of the output fused index. The third address is a first-level pointer containing L elements, where the j-th element represents the j-th fused index element in this fused index, and L represents the total number of fused data elements, L>1, 0<j≤L.
[0088] Optionally or additionally, in some embodiments, the operation object of the fusion instruction may also include the total number of output fused data elements. For example, after the fusion process is completed, the total number of output fused data elements is returned to indicate the number of data elements in this path of fused data. This data may, for example, be written back to the parameter gpr_id0.
[0089] With the development of artificial intelligence technology, in tasks such as image processing and pattern recognition, the operands are often multi-dimensional vectors (i.e., tensor data). Using only scalar or vector operations cannot enable hardware to efficiently complete computational tasks. Therefore, some embodiments disclosed herein also provide fusion instructions involving tensor data. At least one operand of this fusion instruction includes tensor data, which is indicated by at least one descriptor. Specifically, the descriptor may indicate at least one of the following: shape information of the tensor data, and spatial information of the tensor data. The shape information of the tensor data can be used to determine the data address of the tensor data corresponding to the operand in the data storage space. The spatial information of the tensor data can be used to determine the dependencies between instructions, and thus determine, for example, the execution order of the instructions.
[0090] In one possible implementation, the spatial information of tensor data can be indicated by a spatial identifier (ID). A spatial ID, also known as a spatial alias, refers to a spatial region used to store the corresponding tensor data. This spatial region can be a continuous space or multiple segments; this disclosure does not restrict the specific composition of the spatial region. Different spatial IDs indicate that the spatial regions they point to are independent of each other.
[0091] The following section will describe in detail, with reference to the accompanying drawings, various possible ways to implement the shape information of tensor data.
[0092] Tensors can contain various forms of data composition. Tensors can be of different dimensions; for example, a scalar can be considered a 0-dimensional tensor, a vector a 1-dimensional tensor, and a matrix a 2-dimensional or higher tensor. The shape of a tensor includes information such as its dimensions and the size of each dimension. For example, for a 3D tensor:
[0093] x3=[[[1,2,3],[4,5,6]];[[7,8,9],[10,11,12]]]
[0094] The shape or dimensions of this tensor can be represented as X3 = (2, 2, 3), meaning that the tensor is a three-dimensional tensor indicated by three parameters, with the first dimension having a size of 2, the second dimension having a size of 2, and the third dimension having a size of 3. When storing tensor data in memory, the shape of the tensor data cannot be determined based on its data address (or storage area), and consequently, the relationships between multiple tensor data cannot be determined, resulting in low processor efficiency in accessing tensor data.
[0095] In one possible implementation, a descriptor can be used to indicate the shape of N-dimensional tensor data, where N is a positive integer, such as N = 1, 2, or 3, or zero. The three-dimensional tensor in the example above can be represented by the descriptor (2, 2, 3). It should be noted that this disclosure does not impose any restrictions on how the descriptor indicates the shape of the tensor.
[0096] In one possible implementation, the value of N can be determined based on the dimension (also known as the order) of the tensor data, or it can be set according to the needs of using the tensor data. For example, when N is 3, the tensor data is three-dimensional, and the descriptor can be used to indicate the shape (e.g., offset, size, etc.) of the three-dimensional tensor data in the three dimensions. It should be understood that those skilled in the art can set the value of N according to actual needs, and this disclosure does not limit this.
[0097] Although tensor data can be multidimensional, because the layout of memory is always one-dimensional, there is a correspondence between tensors and their storage in memory. Tensor data is typically allocated in contiguous storage space, meaning that tensor data can be expanded in one dimension (e.g., row-major order) and stored in memory.
[0098] The relationship between a tensor and its underlying storage can be represented by the dimension offset, dimension size, and dimension stride. The dimension offset refers to the offset relative to a reference position within that dimension. The dimension size refers to the number of elements in that dimension. The dimension stride refers to the interval between adjacent elements within that dimension. For example, the stride of the three-dimensional tensor above is (6, 3, 1), meaning the stride of the first dimension is 6, the stride of the second dimension is 3, and the stride of the third dimension is 1.
[0099] Figure 9 A schematic diagram of a data storage space according to an embodiment of this disclosure is shown. Figure 9 As shown, data storage space 91 stores two-dimensional data in row-major order, which can be represented by (x, y) (where the X-axis is horizontal to the right and the Y-axis is vertical to the bottom). The size in the X-axis direction (the size of each row, or the total number of columns) is ori_x (not shown in the figure), and the size in the Y-axis direction (the total number of rows) is ori_y (not shown in the figure). The starting address PA_start (base address) of data storage space 91 is the physical address of the first data block 92. Data block 93 is a portion of the data in data storage space 91. Its offset 95 in the X-axis direction is denoted as offset_x, its offset 94 in the Y-axis direction is denoted as offset_y, its size in the X-axis direction is denoted as size_x, and its size in the Y-axis direction is denoted as size_y.
[0100] In one possible implementation, when using a descriptor to define data block 93, the data reference point of the descriptor can use the first data block of data storage space 91, and the reference address of the descriptor can be agreed to be the starting address PA_start of data storage space 91. Then, the content of the descriptor of data block 93 can be determined by combining the size ori_x of data storage space 91 on the X-axis, the size ori_y of data storage space 91 on the Y-axis, and the offsets or offsets or offsets or sizes of data block 93 on the Y-axis, X-axis, and Y-axis.
[0101] In one possible implementation, the content of the descriptor can be represented using the following formula (1):
[0102]
[0103] It should be understood that although the content of the descriptor in the above example represents a two-dimensional space, those skilled in the art can set the specific dimension represented by the content of the descriptor according to the actual situation, and this disclosure does not limit this.
[0104] In one possible implementation, the reference address of the data reference point of the descriptor in the data storage space can be agreed upon. Based on the reference address, the content of the tensor data descriptor is determined according to the positions of at least two vertices at diagonal positions in N dimensional directions relative to the data reference point.
[0105] For example, the base address PA_base of the data reference point in the data storage space can be agreed upon. For instance, a piece of data (e.g., data at position (2, 2)) can be selected in data storage space 91 as the data reference point, and the physical address of that data in the data storage space can be used as the base address PA_base. The position of the two diagonally opposite vertices relative to the data reference point can be used to determine... Figure 9 The contents of the descriptor for data block 93 are determined first. First, the positions of at least two diagonal vertices of data block 93 relative to the data reference point are determined. For example, the positions of the diagonal vertices from the top left to the bottom right relative to the data reference point are used, where the relative positions of the top left vertex are (x_min, y_min) and the relative positions of the bottom right vertex are (x_max, y_max). Then, the contents of the descriptor for data block 63 can be determined based on the reference address PA_base, the relative positions of the top left vertex (x_min, y_min), and the bottom right vertex (x_max, y_max).
[0106] In one possible implementation, the contents of the descriptor (based on the base address PA_base) can be represented using the following formula (2):
[0107]
[0108] It should be understood that although the above example uses the top left and bottom right corners as the two diagonal vertices to determine the content of the descriptor, those skilled in the art can set the specific vertices of at least two diagonal vertices according to actual needs, and this disclosure does not limit this.
[0109] In one possible implementation, the content of the tensor data descriptor can be determined based on the reference address of the descriptor's data reference point in the data storage space, and the mapping relationship between the data description location and the data address of the tensor data indicated by the descriptor. The mapping relationship between the data description location and the data address can be set according to actual needs. For example, when the tensor data indicated by the descriptor is three-dimensional spatial data, the function f(x, y, z) can be used to define the mapping relationship between the data description location and the data address.
[0110] In one possible implementation, the content of the descriptor can be represented using the following formula (3):
[0111]
[0112] In one possible implementation, the descriptor is also used to indicate the address of N-dimensional tensor data, wherein the content of the descriptor also includes at least one address parameter representing the address of the tensor data, for example, the content of the descriptor can be the following equation (4):
[0113]
[0114] PA is the address parameter. The address parameter can be a logical address or a physical address. When resolving the descriptor, PA can be any one of the vertices, midpoints, or preset points of the vector shape, combined with the shape parameters in the X and Y directions to obtain the corresponding data address.
[0115] In one possible implementation, the address parameter of the tensor data includes the reference address of the data reference point of the descriptor in the data storage space of the tensor data, and the reference address includes the starting address of the data storage space.
[0116] In one possible implementation, the descriptor may also include at least one address parameter representing the address of the tensor data, for example, the content of the descriptor may be the following equation (5):
[0117]
[0118] PA_start is the base address parameter, which will not be elaborated further.
[0119] It should be understood that those skilled in the art can set the mapping relationship between data description location and data address according to the actual situation, and this disclosure does not impose any restrictions on this.
[0120] In one possible implementation, a pre-defined base address can be set within a task. All descriptors in instructions within this task use this base address, and the descriptor content can include shape parameters based on this base address. This base address can be determined by setting environment parameters for this task. For a description and usage of the base address, please refer to the above embodiments. In this implementation, the descriptor content can be mapped to data addresses more quickly.
[0121] In one possible implementation, the base address can be included in the content of each descriptor, allowing each descriptor to have a different base address. Compared to using environment parameters to set a common base address, this approach allows descriptors to describe data more flexibly and utilize a larger data address space.
[0122] In one possible implementation, the data address in the data storage space corresponding to the operand of the processing instruction can be determined based on the content of the descriptor. The calculation of the data address is automatically performed by the hardware, and the calculation method will differ depending on the representation of the descriptor content. This disclosure does not limit the specific calculation method of the data address.
[0123] For example, if the content of the descriptor in the operand is represented using formula (1), and the offsets of the tensor data indicated by the descriptor in the data storage space are offset_x and offset_y respectively, and the size is size_x*size_y, then the starting data address PA1 of the tensor data indicated by the descriptor in the data storage space is... (x,y) The following formula (6) can be used to determine it:
[0124] PA1 (x,y) =PA_start+(offset_y-1)*ori_x+offset_x (6)
[0125] The data starting address PA1 is determined according to the above formula (6). (x,y) By combining the offsets offset_x and offset_y, as well as the storage area sizes size_x and size_y, the storage area of the tensor data indicated by the descriptor in the data storage space can be determined.
[0126] In one possible implementation, when the operand also includes a data description location for the descriptor, the data address of the corresponding data in the data storage space can be determined based on the content of the descriptor and the data description location. In this way, a portion (e.g., one or more data items) of the tensor data indicated by the descriptor can be processed.
[0127] For example, the content of the descriptor in the operand is represented using formula (2), the tensor data indicated by the descriptor has offsets of offset_x and offset_y in the data storage space, and the size is size_x*size_y. The data description position of the descriptor included in the operand is (x q y q If the tensor data indicated by the descriptor is located at the data storage address PA2, then... (x,y) The following formula (7) can be used to determine it:
[0128] PA2 (x,y) =PA_start+(offset_y+y q -1)*ori_x+(offset_x+x q (7)
[0129] In one possible implementation, a descriptor can indicate chunks of data. Data chunking can effectively speed up computation and improve processing efficiency in many applications. For example, in graphics processing, convolution operations often use data chunking for fast processing.
[0130] Figure 10 This diagram illustrates data blocks in a data storage space according to embodiments of this disclosure. Figure 10 As shown, data storage space 1000 also uses a row-major approach to store two-dimensional data, which can be represented by (x, y) (where the X-axis is horizontal to the right and the Y-axis is vertical downwards). The dimension along the X-axis (the dimension of each row, or the total number of columns) is ori_x (not shown in the figure), and the dimension along the Y-axis (the total number of rows) is ori_y (not shown in the figure). Unlike Figure 9 tensor data, Figure 10 The tensor data stored in it consists of multiple data blocks.
[0131] In this case, the descriptor needs more parameters to represent these data tiles. Taking the X-axis (X dimension) as an example, it can involve parameters such as: `ori_x`, `x.tile.size` (size of the tile, 1002), `x.tile.stride` (stride of the tile, 1004, i.e., the distance between the first point of the first tile and the first point of the second tile), `x.tile.num` (number of tiles, shown as 3 tiles in the figure), `x.stride` (overall stride, i.e., the distance between the first point of the first row and the first point of the second row), etc. Other dimensions can similarly include corresponding parameters.
[0132] In one possible implementation, the descriptor may include an identifier and / or content. The identifier is used to distinguish descriptors; for example, the identifier may be a number. The content may include at least one shape parameter representing the shape of the tensor data. For example, if the tensor data is 3-dimensional, and the shape parameters of two of its three dimensions are fixed, the content of its descriptor may include the shape parameter representing the other dimension of the tensor data.
[0133] In one possible implementation, the identifier and / or content of the descriptor can be stored in descriptor storage space (internal memory), such as registers, on-chip SRAM, or other media caches. The tensor data indicated by the descriptor can be stored in data storage space (internal or external memory), such as on-chip cache or under-chip memory. This disclosure does not limit the specific location of the descriptor storage space or the data storage space.
[0134] In one possible implementation, the descriptor's identifier, content, and the tensor data indicated by the descriptor can be stored in the same area of internal memory. For example, a contiguous area of on-chip cache, with addresses ADDR0-ADDR1023, can be used to store the descriptor's content. Addresses ADDR0-ADDR63 can be used as the descriptor storage space to store the descriptor's identifier and content, while addresses ADDR64-ADDR1023 can be used as the data storage space to store the tensor data indicated by the descriptor. Within the descriptor storage space, addresses ADDR0-ADDR31 can be used to store the descriptor's identifier, and addresses ADDR32-ADDR63 can be used to store the descriptor's content. It should be understood that address ADDR is not limited to one bit or one byte; it is used here to represent an address, a unit of address. Those skilled in the art can determine the descriptor storage space, data storage space, and their specific addresses based on actual circumstances, and this disclosure does not impose any limitations in this regard.
[0135] In one possible implementation, the descriptor's identifier, content, and the tensor data it points to can be stored in different areas of internal memory. For example, registers can be used as descriptor storage space, storing the descriptor's identifier and content, while on-chip caches can be used as data storage space, storing the tensor data pointed to by the descriptor.
[0136] In one possible implementation, when using registers to store the identifier and content of descriptors, the register number can be used to represent the descriptor's identifier. For example, when the register number is 0, the identifier of the descriptor it stores is set to 0. When the descriptor in the register is valid, a region in the cache space can be allocated to store the tensor data according to the size of the tensor data indicated by the descriptor.
[0137] In one possible implementation, the identifier and content of the descriptor can be stored in internal memory, while the tensor data pointed to by the descriptor can be stored in external memory. For example, the identifier and content of the descriptor can be stored on-chip, and the tensor data pointed to by the descriptor can be stored off-chip.
[0138] In one possible implementation, the data address of the data storage space corresponding to each descriptor can be a fixed address. For example, a separate data storage space can be allocated for tensor data, and the starting address of each tensor data in the data storage space corresponds one-to-one with a descriptor. In this case, the circuit or module responsible for parsing computation instructions (e.g., an entity outside the computing device disclosed herein) can determine the data address of the data corresponding to the operand in the data storage space based on the descriptor.
[0139] In one possible implementation, when the data address of the data storage space corresponding to the descriptor is a variable address, the descriptor can also be used to indicate the address of N-dimensional tensor data. The content of the descriptor may further include at least one address parameter representing the address of the tensor data. For example, if the tensor data is 3-dimensional, when the descriptor points to the address of the tensor data, the content of the descriptor may include one address parameter representing the address of the tensor data, such as the starting physical address of the tensor data, or it may include multiple address parameters representing the address of the tensor data, such as the starting address of the tensor data plus an address offset, or address parameters based on each dimension of the tensor data. Those skilled in the art can set the address parameters according to actual needs, and this disclosure does not impose any limitations on this.
[0140] In one possible implementation, the address parameter of the tensor data may include the reference address of the data reference point of the descriptor within the data storage space of the tensor data. The reference address may vary depending on the data reference point. This disclosure does not impose any restrictions on the selection of the data reference point.
[0141] In one possible implementation, the base address may include the starting address of the data storage space. When the data base point of the descriptor is the first data block in the data storage space, the base address of the descriptor is the starting address of the data storage space. When the data base point of the descriptor is data other than the first data block in the data storage space, the base address of the descriptor is the address of that data block in the data storage space.
[0142] In one possible implementation, the shape parameters of the tensor data include at least one of the following: the size of the data storage space in at least one of the N-dimensional directions, the size of the storage region in at least one of the N-dimensional directions, the offset of the storage region in at least one of the N-dimensional directions, the positions of at least two vertices at diagonal positions in the N-dimensional directions relative to the data reference point, and the mapping relationship between the data description location of the tensor data indicated by the descriptor and the data address. Here, the data description location is the mapped position of a point or region in the tensor data indicated by the descriptor. For example, when the tensor data is 3D data, the descriptor can use three-dimensional spatial coordinates (x, y, z) to represent the shape of the tensor data, and the data description location of the tensor data can be the position of a point or region mapped to the tensor data in three-dimensional space, represented by three-dimensional spatial coordinates (x, y, z).
[0143] It should be understood that those skilled in the art can choose the shape parameters representing tensor data according to the actual situation, and this disclosure does not impose any restrictions on this. By using descriptors during data access, associations between data can be established, thereby reducing the complexity of data access and improving instruction processing efficiency.
[0144] Figure 11 A structural block diagram of a data processing apparatus 1100 according to another embodiment of this disclosure is shown. The data processing apparatus 1100 can, for example, be implemented in... Figure 2 In the computing device 201. Figure 11 Data processing device 1100 and Figure 5 The difference is that, Figure 11 The data processing apparatus 1100 also includes a tensor interface circuit 1112 for implementing functions related to the descriptor of tensor data. Similarly, the data processing apparatus 1100 may also include a control circuit 1110, a storage circuit 1120, and an arithmetic circuit 1130, the specific functions and implementation of which are related to... Figure 5 Those similar to those will not be repeated here.
[0145] The Tensor Interface Unit (TIU) 1112 can be configured to perform operations associated with descriptors under the control of the control circuit 1110. These operations may include, but are not limited to, registering, modifying, deregistering, and resolving descriptors; reading and writing descriptor contents, etc. This disclosure does not limit the specific hardware type of the Tensor Interface Unit. In this way, operations associated with descriptors can be implemented through dedicated hardware, further improving the efficiency of tensor data access.
[0146] In some embodiments, the tensor interface circuit 1112 may be configured to parse the shape information of the tensor data included in the operand of the instruction in order to determine the data address of the data corresponding to the operand in the data storage space.
[0147] Optionally or additionally, in some other embodiments, the tensor interface circuit 1112 may be configured to compare the spatial information (e.g., spatial ID) of the tensor data included in the operands of two instructions to determine the dependency relationship between the two instructions, and thereby determine out-of-order execution, synchronization, or other operations of the instructions.
[0148] Despite Figure 11 The control circuit 1110 and the tensor interface circuit 1112 are shown as two separate modules, but those skilled in the art will understand that these two circuits can also be implemented as one or more modules, and this disclosure is not limited in this respect.
[0149] There can be various operations related to data fusion, such as merge sorting and sorting accumulation. Multiple instruction schemes can be designed to implement these operations.
[0150] In one approach, a fusion instruction can be designed, which may include an operation mode bit to indicate different operation modes of the fusion instruction, thereby performing different operations.
[0151] In another approach, multiple fusion instructions can be designed, each corresponding to one or more different operating modes to perform different operations. In one implementation, a corresponding fusion instruction can be designed for each operating mode. In another implementation, operating modes can be categorized based on their characteristics, and a fusion instruction can be designed for each category of operating modes. Furthermore, when a certain category of operating modes includes multiple operating modes, an operating mode bit can be included in the fusion instruction to indicate the corresponding operating mode.
[0152] Regardless of the approach taken, fusion instructions can indicate their corresponding operating mode through the operation mode bit and / or the instruction itself.
[0153] In the context of this disclosure, the aforementioned fusion instruction may be a microinstruction or control signal that runs within one or more multi-stage computational pipelines, and may include (or instruct) one or more computational operations that need to be executed by the multi-stage computational pipelines. Depending on the different computational operation scenarios, the computational operations may include, but are not limited to, arithmetic operations such as convolution and matrix multiplication, logical operations such as AND, XOR, and OR, shift operations, or any combination of the aforementioned types of computational operations.
[0154] Figure 12 An exemplary flowchart of a data processing method 1000 according to an embodiment of this disclosure is shown.
[0155] like Figure 12 As shown, in step 1210, a fusion instruction is parsed, which instructs the merging and sorting of multiple streams of data to be merged. This step can, for example, be performed by... Figure 5 Control circuit 510 or Figure 11 The control circuit 1110 performs the operation. Next, in step 1220, according to the fusion instruction, the multiple data streams to be fused are merged into a single fused data stream. Finally, in step 1230, the fused data is output in an ordered manner. Steps 1220 and 1230 can, for example, be performed by... Figure 5 530 or Figure 11 The operation circuit 1130 is used to execute it.
[0156] Those skilled in the art will understand that each step of the above method corresponds to the respective circuits described above in conjunction with the example circuit diagrams. Therefore, the features described above can be applied equally to the method steps, and will not be repeated here.
[0157] As described above, this disclosure provides a fusion instruction for performing fusion processing of multiple streams of data to be fused. In some embodiments, the fusion instruction is a hardware instruction, implemented through dedicated hardware circuitry. In some embodiments, the fusion instruction may include an operation mode bit to indicate that the fusion instruction is a merge sorting operation, or the fusion instruction itself may indicate a merge sorting operation. By providing a dedicated fusion instruction to perform operations related to the fusion processing of multiple streams of data, the processing can be simplified. Furthermore, by providing a dedicated hardware implementation of data fusion-related operations, the processing can be accelerated, thereby improving the machine's processing efficiency.
[0158] Depending on the application scenario, the electronic devices or apparatus disclosed herein may include servers, cloud servers, server clusters, data processing devices, robots, computers, printers, scanners, tablets, smart terminals, PC devices, IoT terminals, mobile terminals, mobile phones, dashcams, navigators, sensors, cameras, video cameras, projectors, watches, headphones, mobile storage, wearable devices, visual terminals, autonomous driving terminals, vehicles, home appliances, and / or medical devices. The vehicles include airplanes, ships, and / or vehicles; the home appliances include televisions, air conditioners, microwave ovens, refrigerators, rice cookers, humidifiers, washing machines, lights, gas stoves, and range hoods; the medical devices include MRI scanners, ultrasound machines, and / or electrocardiographs. The electronic devices or apparatus disclosed herein can also be applied in fields such as the Internet, IoT, data centers, energy, transportation, public management, manufacturing, education, power grids, telecommunications, finance, retail, construction sites, and healthcare. Furthermore, the electronic devices or apparatus disclosed herein can also be used in application scenarios related to artificial intelligence, big data, and / or cloud computing, such as cloud computing, edge computing, and terminal applications. In one or more embodiments, the high-computing-power electronic devices or apparatuses according to the present disclosure can be applied to cloud devices (e.g., cloud servers), while the low-power electronic devices or apparatuses can be applied to terminal devices and / or edge devices (e.g., smartphones or cameras). In one or more embodiments, the hardware information of the cloud devices and the hardware information of the terminal devices and / or edge devices are compatible with each other, so that suitable hardware resources can be matched from the hardware resources of the cloud devices to simulate the hardware resources of the terminal devices and / or edge devices based on the hardware information of the terminal devices and / or edge devices, so as to complete the unified management, scheduling and collaborative work of end-to-cloud or cloud-edge-end integration.
[0159] It should be noted that, for the sake of brevity, this disclosure describes some methods and their embodiments as a series of actions and combinations thereof. However, those skilled in the art will understand that the solutions disclosed herein are not limited by the order of the described actions. Therefore, based on the disclosure or teachings of this document, those skilled in the art will understand that some steps can be performed in a different order or simultaneously. Furthermore, those skilled in the art will understand that the embodiments described in this disclosure can be considered optional embodiments, that is, the actions or modules involved are not necessarily essential for the implementation of one or more solutions disclosed herein. In addition, depending on the solution, the description of some embodiments in this disclosure may have different emphases. In view of this, those skilled in the art will understand that parts not described in detail in a certain embodiment of this disclosure can also be referred to the relevant descriptions of other embodiments.
[0160] In terms of specific implementation, based on the disclosure and teachings of this document, those skilled in the art will understand that several embodiments disclosed herein can also be implemented in other ways not disclosed herein. For example, regarding the various units in the electronic device or apparatus embodiments described above, this document has divided them based on logical functions, but in actual implementation, there may be other ways of division. As another example, multiple units or components can be combined or integrated into another system, or some features or functions in a unit or component can be selectively disabled. Regarding the connection relationships between different units or components, the connections discussed above in conjunction with the accompanying drawings can be direct or indirect couplings between units or components. In some scenarios, the aforementioned direct or indirect couplings involve communication connections utilizing interfaces, where the communication interface can support electrical, optical, acoustic, magnetic, or other forms of signal transmission.
[0161] In this disclosure, the units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units. The aforementioned components or units may be located in the same location or distributed across multiple network units. Furthermore, depending on actual needs, some or all of the units can be selected to achieve the purpose of the solution described in the embodiments of this disclosure. Additionally, in some scenarios, multiple units in the embodiments of this disclosure may be integrated into one unit or each unit may exist physically independently.
[0162] In other implementation scenarios, the integrated units described above can also be implemented in hardware, i.e., as specific hardware circuits, which may include digital circuits and / or analog circuits. The physical implementation of the circuit's hardware structure may include, but is not limited to, physical devices, which may include, but are not limited to, transistors or memristors. Therefore, the various devices described herein (e.g., computing devices or other processing devices) can be implemented using appropriate hardware processors, such as central processing units, GPUs, FPGAs, DSPs, and ASICs. Furthermore, the aforementioned storage units or storage devices can be any suitable storage medium (including magnetic storage media or magneto-optical storage media), such as resistive random access memory (RRAM), dynamic random access memory (DRAM), static random access memory (SRAM), enhanced dynamic random access memory (EDRAM), high-bandwidth memory (HBM), hybrid memory cube (HMC), ROM, and RAM.
[0163] The foregoing can be better understood in accordance with the following terms:
[0164] Clause 1. A data processing apparatus, comprising:
[0165] A control circuit configured to parse a fusion instruction that instructs the merging and sorting of multiple streams of data to be merged.
[0166] Storage circuitry configured to store information before and / or after merge sort processing; and
[0167] The processing circuit is configured to merge the multiple data streams to be merged into one fused data stream according to the fusion instruction, and output the fused data stream in an orderly manner.
[0168] Clause 2. The data processing apparatus according to Clause 1, wherein the operation object of the fusion instruction includes input K-channel data to be fused, the size of the K-channel data, and output one channel of fused data, K>1, wherein the data elements of each channel of the K-channel data are arranged in a first order, and the data elements of the output one channel of fused data are arranged in a second order.
[0169] Clause 3. The data processing apparatus according to Clause 2, wherein the first order is the same as or different from the second order, and the first order and the second order are selected from either: an ascending order or an ascending order.
[0170] Clause 4. The data processing circuit according to any one of Clauses 2-3, wherein the arithmetic circuit includes a comparison circuit and a buffer circuit, wherein:
[0171] The comparison circuit is configured to compare data elements in the multi-channel data to be fused with data elements not yet output in the buffer circuit, and output the comparison result to the control circuit; and
[0172] The buffer circuit is configured to, under the control of the control circuit, orderly store the compared data elements and orderly output the compared data elements as the fused data.
[0173] Clause 5. The data processing circuit according to Clause 4, wherein the comparison circuit includes:
[0174] A K-1 comparator is configured to compare the data element to be fused with the K-1 data elements of the current sequence in the buffer circuit, generate a comparison result, and output it to the control circuit.
[0175] Clause 6. The data processing circuit according to Clause 5, wherein the control circuit is configured to determine, based on the comparison result, the insertion position of the data element to be fused in the current sequence of the buffer circuit.
[0176] Clause 7. The data processing circuit according to Clause 6, wherein the comparison result is represented using a bitmap, and the control circuit is further configured to: determine the insertion position based on the positional changes of bits in the bitmap.
[0177] Clause 8. A data processing circuit according to any one of Clauses 6-7, wherein the buffer circuit is configured to insert the data element to be merged at the insertion position according to the instruction of the control circuit.
[0178] Clause 9. The data processing circuit according to any one of Clauses 4-8, wherein the buffer circuit is further configured to output the first or last data element in the current sequence in a specified order.
[0179] Clause 10. The data processing circuit according to Clause 9, wherein the control circuit is further configured to: determine memory access information of the next data element to be fused based on the data element output from the buffer circuit.
[0180] Clause 11. The data processing circuit according to any one of Clauses 4-10, wherein the multiple data streams include multiple indexes, each index corresponding one-to-one with multiple associated data streams, each index element indicating the index information of the corresponding associated data element in the associated data stream, and the arithmetic circuit is further configured to:
[0181] The multiple associated data are merged into one ordered fused associated data, wherein the order of data elements in the fused associated data is consistent with the order of data elements in the fused data.
[0182] Clause 12. The data processing circuit according to Clause 11, wherein the buffer circuit is further configured to:
[0183] The compared index elements and their corresponding associated data elements are stored in ordered order according to the value sequence of the index elements; and
[0184] The compared index elements are output in order of their values as the fused data, and their corresponding associated data elements are output synchronously as the fused associated data.
[0185] Clause 13. The data processing circuit according to Clause 12, wherein when there are index elements of the same size in the multiplexed index, the index elements of the same size are repeatedly output in the fused data, and the associated data elements corresponding to these index elements are synchronously output in the fused associated data.
[0186] Clause 14. The data processing circuit according to any one of Clauses 11-13, wherein the data elements in the multi-path associated data are valid data elements in the sparse matrix, and the index information indicates the position information of the valid data elements in the sparse matrix.
[0187] Clause 15. The data processing apparatus according to any one of Clauses 2-14, wherein the operation object of the fusion instruction further includes the total number of output fused data elements, used to indicate the number of data elements in one output fused data path.
[0188] Clause 16. The data processing apparatus according to any one of Clauses 2-15, wherein
[0189] The input K-channel data is indicated by a first address, which includes K elements, where the i-th element represents the starting address of the i-th channel data, and 0 < i ≤ K.
[0190] Clause 17. The data processing apparatus according to Clause 16, wherein
[0191] The size of the K data streams is indicated by a second address, which includes K elements, where the i-th element represents the number of data elements in the i-th data stream, and 0 < i ≤ K.
[0192] Clause 18. The data processing apparatus as described in Clause 17, wherein
[0193] The output fused data is indicated by a third address, which includes L elements. The j-th element represents the j-th data element in the fused data, and L represents the total number of fused data elements, where L>1 and 0<j≤L.
[0194] Clause 19. A data processing apparatus according to any one of Clauses 2-15, wherein at least one of the operated objects comprises tensor data, and the tensor data is indicated by at least one descriptor, the descriptor indicating at least one of the following: shape information and spatial information of the tensor data; and
[0195] The data processing device further includes a tensor interface circuit configured to parse the descriptor in order to obtain the tensor data.
[0196] Clause 20. The data processing apparatus according to Clause 19, wherein the tensor interface circuitry is further configured for:
[0197] Based on the shape information, determine the data address of the tensor data in the data storage space; and / or
[0198] Based on the spatial information, the dependencies between instructions are determined.
[0199] Clause 21. The data processing apparatus according to any one of Clauses 19-20, wherein the shape information of the tensor data includes at least one shape parameter representing the shape of the N-dimensional tensor data, where N is a positive integer, and the shape parameter of the tensor data includes at least one of the following:
[0200] The dimensions of the data storage space containing the tensor data in at least one of the N dimensions, the dimensions of the storage area of the tensor data in at least one of the N dimensions, the offset of the storage area in at least one of the N dimensions, the positions of at least two vertices at diagonal positions in the N dimensions relative to the data reference point, and the mapping relationship between the data description location and the data address of the tensor data.
[0201] Clause 22. The data processing apparatus according to any one of Clauses 19-20, wherein the shape information of the tensor data indicates at least one shape parameter of the shape of N-dimensional tensor data comprising a plurality of data blocks, where N is a positive integer, and the shape parameter includes at least one of the following:
[0202] The dimensions of the data storage space containing the tensor data in at least one of the N dimensions, the dimensions of the storage area of a single data block in at least one of the N dimensions, the block step size of the data block in at least one of the N dimensions, the number of data blocks in at least one of the N dimensions, and the overall step size of the data block in at least one of the N dimensions.
[0203] Clause 23. The data processing apparatus according to any one of Clauses 1-22, wherein
[0204] The merging instruction includes an operation mode bit to indicate that the merging instruction is a merge sorting operation, or the merging instruction itself indicates the merge sorting operation.
[0205] Clause 24. A chip comprising a data processing apparatus according to any one of Clauses 1-23.
[0206] Clause 25. A board including the chip described in Clause 24.
[0207] Clause 26. A data processing method, comprising:
[0208] Parse the fusion instruction, which instructs to perform merge sorting processing on multiple data streams to be merged;
[0209] According to the fusion instruction, the multiple data streams to be fused are merged into one fused data stream; and
[0210] The fused data is output in an orderly manner.
[0211] Clause 27. The data processing method according to Clause 26, wherein the operation object of the fusion instruction includes the input K-channel data to be fused, the size of the K-channel data, and the output fused data, K>1, wherein the data elements of each channel of the K-channel data are arranged in a first order, and the data elements of the output fused data are arranged in a second order.
[0212] Clause 28. The data processing method according to Clause 27, wherein the first order is the same as or different from the second order, and the first order and the second order are selected from either: an ascending order or an ascending order.
[0213] Clause 29. The data processing method according to any one of Clauses 27-28 further includes:
[0214] The comparison circuit compares the data elements in the multiple data streams to be fused with the data elements not yet output in the buffer circuit, and outputs the comparison result to the control circuit; and
[0215] The buffer circuit, under the control of the control circuit, stores the compared data elements in an orderly manner and outputs the compared data elements in an orderly manner as the fused data.
[0216] Clause 30. The data processing method according to Clause 29, wherein the comparison further includes:
[0217] The K-1 comparator compares the data element to be fused with the K-1 data elements of the current sequence in the buffer circuit, generates a comparison result, and outputs it to the control circuit.
[0218] Clause 31. The data processing method described in Clause 30 further includes:
[0219] The control circuit determines the insertion position of the data element to be fused in the current sequence of the buffer circuit based on the comparison result.
[0220] Clause 32. The data processing method according to Clause 31, wherein the comparison result is represented using a bitmap, and the control circuit determines the insertion position based on the changing positions of bits in the bitmap.
[0221] Clause 33. The data processing method according to any one of Clauses 31-32 further includes:
[0222] The buffer circuit inserts the data element to be merged at the insertion position according to the instruction of the control circuit.
[0223] Clause 34. The data processing method according to any one of Clauses 29-33 further includes:
[0224] The buffer circuit outputs the first or last data element in the current sequence in a specified order.
[0225] Clause 35. The data processing method described in Clause 34 further includes:
[0226] The control circuit determines the memory access information of the next data element to be merged based on the data elements output from the buffer circuit.
[0227] Clause 36. The data processing method according to any one of Clauses 29-35, wherein the multi-channel data includes a multi-channel index, the multi-channel index corresponding one-to-one with multi-channel associated data, the index element in each channel index indicating the index information of the corresponding associated data element in the corresponding channel of associated data, and the method further includes:
[0228] The multiple associated data are merged into one ordered fused associated data, wherein the order of data elements in the fused associated data is consistent with the order of data elements in the fused data.
[0229] Clause 37. The data processing method described in Clause 36 further includes:
[0230] The compared index elements and their corresponding associated data elements are stored in ordered order according to the value sequence of the index elements; and
[0231] The compared index elements are output in order of their values as the fused data, and their corresponding associated data elements are output synchronously as the fused associated data.
[0232] Clause 38. The data processing method according to Clause 37, wherein when there are index elements of the same size in the multi-way index, the index elements of the same size are repeatedly output in the fused data, and the associated data elements corresponding to these index elements are synchronously output in the fused associated data.
[0233] Clause 39. The data processing method according to any one of Clauses 37-38, wherein the data elements in the multi-path associated data are valid data elements in the sparse matrix, and the index information indicates the position information of the valid data elements in the sparse matrix.
[0234] Clause 40. In any of the data processing methods described in Clauses 27-39, the operation object of the fusion instruction further includes the total number of output fused data elements, used to indicate the number of data elements in one output fused data path.
[0235] Clause 41. The data processing method according to any one of Clauses 27-40, wherein
[0236] The input K-channel data is indicated by a first address, which includes K elements, where the i-th element represents the starting address of the i-th channel data, and 0 < i ≤ K.
[0237] Clause 42. The data processing method described in Clause 41, wherein
[0238] The size of the K data streams is indicated by a second address, which includes K elements, where the i-th element represents the number of data elements in the i-th data stream, and 0 < i ≤ K.
[0239] Clause 43. The data processing method described in Clause 42, wherein
[0240] The output fused data is indicated by a third address, which includes L elements. The j-th element represents the j-th data element in the fused data, and L represents the total number of fused data elements, where L>1 and 0<j≤L.
[0241] Clause 44. A data processing method according to any one of Clauses 27-40, wherein at least one of the operation objects comprises tensor data, and the tensor data is indicated by at least one descriptor, the descriptor indicating at least one of the following: shape information and spatial information of the tensor data; and the method further comprises:
[0242] The descriptor is parsed to obtain the tensor data.
[0243] Clause 45. The data processing method according to Clause 44, wherein parsing the descriptor includes:
[0244] Based on the shape information, determine the data address of the tensor data in the data storage space; and / or
[0245] Based on the spatial information, the dependencies between instructions are determined.
[0246] Clause 46. In any of the data processing methods described in Clauses 44-45, the shape information of the tensor data includes at least one shape parameter representing the shape of the N-dimensional tensor data, where N is a positive integer, and the shape parameter of the tensor data includes at least one of the following:
[0247] The dimensions of the data storage space containing the tensor data in at least one of the N dimensions, the dimensions of the storage area of the tensor data in at least one of the N dimensions, the offset of the storage area in at least one of the N dimensions, the positions of at least two vertices at diagonal positions in the N dimensions relative to the data reference point, and the mapping relationship between the data description location and the data address of the tensor data.
[0248] Clause 47. In any of the data processing methods described in Clauses 44-45, the shape information of the tensor data indicates at least one shape parameter of the shape of an N-dimensional tensor data comprising multiple data blocks, where N is a positive integer, and the shape parameter includes at least one of the following:
[0249] The dimensions of the data storage space containing the tensor data in at least one of the N dimensions, the dimensions of the storage area of a single data block in at least one of the N dimensions, the block step size of the data block in at least one of the N dimensions, the number of data blocks in at least one of the N dimensions, and the overall step size of the data block in at least one of the N dimensions.
[0250] Clause 48. The data processing method according to any one of Clauses 26-47, wherein
[0251] The merging instruction includes an operation mode bit to indicate that the merging instruction is a merge sorting operation, or the merging instruction itself indicates the merge sorting operation.
[0252] The embodiments of this disclosure have been described in detail above. Specific examples have been used to illustrate the principles and implementation methods of this disclosure. The descriptions of the above embodiments are only for the purpose of helping to understand the methods and core ideas of this disclosure. At the same time, for those skilled in the art, there will be changes in the specific implementation methods and application scope based on the ideas of this disclosure. Therefore, the content of this specification should not be construed as a limitation of this disclosure.
Claims
1. A data processing apparatus, comprising: A control circuit configured to parse a fusion instruction that instructs the merging and sorting of multiple streams of data to be merged. A storage circuit configured to store information before and / or after the merge sort process; as well as The processing circuit is configured to, according to the fusion instruction, merge the multiple data streams to be fused into one fused data stream, and output the fused data stream in an orderly manner; The multi-channel data includes multi-channel indexes, each corresponding one-to-one with multiple channels of associated data. Each index element indicates the index information of the corresponding associated data element within that channel. Furthermore, the computation circuit is configured to: The multi-way indexes are merged into a single ordered fusion index, and the multi-way associated data is merged into a single ordered fusion associated data, wherein the order of data elements in the fusion associated data is consistent with the order of data elements in the fusion data.
2. The data processing apparatus according to claim 1, wherein the operation object of the fusion instruction includes the input K-channel data to be fused, the size of the K-channel data, and the output fused data, K>1, wherein the data elements of each channel of the K-channel data are arranged in a first order, and the data elements of the output fused data are arranged in a second order.
3. The data processing apparatus according to claim 2, wherein the first order is the same as or different from the second order, and the first order and the second order are selected from either: an order from smallest to largest, or an order from largest to smallest.
4. The data processing apparatus according to any one of claims 2-3, wherein the arithmetic circuit includes a comparison circuit and a buffer circuit, wherein: The comparison circuit is configured to compare data elements in the multi-channel data to be fused with data elements not yet output in the buffer circuit, and output the comparison result to the control circuit; and The buffer circuit is configured to, under the control of the control circuit, orderly store the compared data elements and orderly output the compared data elements as the fused data.
5. The data processing apparatus according to claim 4, wherein the comparison circuit comprises: A K-1 comparator is configured to compare the data element to be fused with the K-1 data elements of the current sequence in the buffer circuit, generate a comparison result, and output it to the control circuit.
6. The data processing apparatus of claim 5, wherein the control circuit is configured to determine, based on the comparison result, the insertion position of the data element to be fused in the current sequence of the buffer circuit.
7. The data processing apparatus of claim 6, wherein the comparison result is represented using a bitmap, and the control circuit is further configured to: determine the insertion position based on the change position of bits in the bitmap.
8. The data processing apparatus according to any one of claims 6-7, wherein the buffer circuit is configured to insert the data element to be merged at the insertion position according to the instruction of the control circuit.
9. The data processing apparatus of claim 4, wherein the buffer circuit is further configured to output the first or last data element in the current sequence in a specified order.
10. The data processing apparatus according to claim 9, wherein the control circuit is further configured to: determine the memory access information of the next data element to be fused based on the data element output from the buffer circuit.
11. The data processing apparatus of claim 4, wherein the buffer circuit is further configured to: The compared index elements and their corresponding associated data elements are stored in ordered order according to the value sequence of the index elements; and The compared index elements are output in order of their values as the fused data, and their corresponding associated data elements are output synchronously as the fused associated data.
12. The data processing apparatus according to claim 11, wherein when there are index elements of the same size in the multi-path index, the index elements of the same size are repeatedly output in the fused data, and the associated data elements corresponding to these index elements are synchronously output in the fused associated data.
13. The data processing apparatus according to claim 11, wherein the data elements in the multi-path associated data are valid data elements in the sparse matrix, and the index information indicates the position information of the valid data elements in the sparse matrix.
14. The data processing apparatus according to any one of claims 2-3, wherein the operation object of the fusion instruction further includes the total number of output fused data elements, used to indicate the number of data elements in one output fused data channel.
15. The data processing apparatus according to any one of claims 2-3, wherein The input K-channel data is indicated by a first address, which includes K elements, where the i-th element represents the starting address of the i-th channel data, and 0 < i ≤ K.
16. The data processing apparatus according to any one of claims 2-3, wherein The size of the K data streams is indicated by a second address, which includes K elements, where the i-th element represents the number of data elements in the i-th data stream, and 0 < i ≤ K.
17. The data processing apparatus according to any one of claims 2-3, wherein The output fused data is indicated by a third address, which includes L elements. The j-th element represents the j-th data element in the fused data, and L represents the total number of fused data elements, where L>1 and 0<j≤L.
18. The data processing apparatus according to any one of claims 2-3, wherein, At least one of the operation objects includes tensor data, and the tensor data is indicated by at least one descriptor, the descriptor indicating at least one of the following: shape information and spatial information of the tensor data; and The data processing device further includes a tensor interface circuit configured to parse the descriptor in order to obtain the tensor data.
19. The data processing apparatus of claim 18, wherein the tensor interface circuit is further configured to: Based on the shape information, determine the data address of the tensor data in the data storage space; and / or Based on the spatial information, the dependencies between instructions are determined.
20. The data processing apparatus of claim 18, wherein the shape information of the tensor data includes at least one shape parameter representing the shape of the N-dimensional tensor data, where N is a positive integer, and the shape parameter of the tensor data includes at least one of the following: The dimensions of the data storage space containing the tensor data in at least one of the N dimensions, the dimensions of the storage area of the tensor data in at least one of the N dimensions, the offset of the storage area in at least one of the N dimensions, the positions of at least two vertices at diagonal positions in the N dimensions relative to the data reference point, and the mapping relationship between the data description location and the data address of the tensor data.
21. The data processing apparatus of claim 18, wherein the shape information of the tensor data indicates at least one shape parameter of the shape of N-dimensional tensor data comprising a plurality of data blocks, where N is a positive integer, and the shape parameter includes at least one of the following: The dimensions of the data storage space containing the tensor data in at least one of the N dimensions, the dimensions of the storage area of a single data block in at least one of the N dimensions, the block step size of the data block in at least one of the N dimensions, the number of data blocks in at least one of the N dimensions, and the overall step size of the data block in at least one of the N dimensions.
22. The data processing apparatus according to any one of claims 1-3, wherein The merging instruction includes an operation mode bit to indicate that the merging instruction is a merge sorting operation, or the merging instruction itself indicates the merge sorting operation.
23. A chip comprising a data processing apparatus according to any one of claims 1-22.
24. A board comprising the chip according to claim 23.
25. A data processing method, comprising: Parse the fusion instruction, which instructs to perform merge sorting processing on multiple data streams to be merged; According to the fusion instruction, the multiple data streams to be fused are merged into one fused data stream; as well as The fused data is output in an orderly manner; The multi-path data includes a multi-path index, which corresponds one-to-one with multiple paths of associated data. Each path index contains an index element that indicates the index information of the corresponding associated data element within that path. The method further includes: The multi-way indexes are merged into a single ordered fusion index, and the multi-way associated data is merged into a single ordered fusion associated data, wherein the order of data elements in the fusion associated data is consistent with the order of data elements in the fusion data.
26. The data processing method according to claim 25, wherein the operation object of the fusion instruction includes the input K-channel data to be fused, the size of the K-channel data, and the output fused data, K>1, wherein the data elements of each channel of the K-channel data are arranged in a first order, and the data elements of the output fused data are arranged in a second order.
27. The data processing method of claim 26, wherein the first order is the same as or different from the second order, and the first order and the second order are selected from either: an order from smallest to largest, or an order from largest to smallest.
28. The data processing method according to any one of claims 26-27, further comprising: The comparison circuit compares the data elements in the multiple data to be fused with the data elements that have not been output in the buffer circuit, and outputs the comparison result to the control circuit. as well as The buffer circuit, under the control of the control circuit, stores the compared data elements in an orderly manner and outputs the compared data elements in an orderly manner as the fused data.
29. The data processing method of claim 28, wherein the comparison further comprises: The K-1 comparator compares the data element to be fused with the K-1 data elements of the current sequence in the buffer circuit, generates a comparison result, and outputs it to the control circuit.
30. The data processing method according to claim 29, further comprising: The control circuit determines the insertion position of the data element to be fused in the current sequence of the buffer circuit based on the comparison result.
31. The data processing method according to claim 30, wherein the comparison result is represented using a bitmap, and the control circuit determines the insertion position based on the changing positions of bits in the bitmap.
32. The data processing method according to any one of claims 30-31, further comprising: The buffer circuit inserts the data element to be merged at the insertion position according to the instruction of the control circuit.
33. The data processing method according to claim 28, further comprising: The buffer circuit outputs the first or last data element in the current sequence in a specified order.
34. The data processing method according to claim 33, further comprising: The control circuit determines the memory access information of the next data element to be merged based on the data elements output from the buffer circuit.
35. The data processing method according to claim 28, further comprising: The compared index elements and their corresponding associated data elements are stored in an ordered manner according to the value order of the index elements. as well as The compared index elements are output in order of their values as the fused data, and their corresponding associated data elements are output synchronously as the fused associated data.
36. The data processing method according to claim 35, wherein when there are index elements of the same size in the multi-way index, the index elements of the same size are repeatedly output in the fused data, and the associated data elements corresponding to these index elements are synchronously output in the fused associated data.
37. The data processing method according to claim 35, wherein the data elements in the multi-path associated data are valid data elements in the sparse matrix, and the index information indicates the position information of the valid data elements in the sparse matrix.
38. The data processing method according to any one of claims 26-27, wherein the operation object of the fusion instruction further includes the total number of output fused data elements, used to indicate the number of data elements in one output fused data channel.
39. The data processing method according to any one of claims 26-27, wherein The input K-channel data is indicated by a first address, which includes K elements, where the i-th element represents the starting address of the i-th channel data, and 0 < i ≤ K.
40. The data processing method according to claim 38, wherein The size of the K data streams is indicated by a second address, which includes K elements, where the i-th element represents the number of data elements in the i-th data stream, and 0 < i ≤ K.
41. The data processing method according to claim 40, wherein The output fused data is indicated by a third address, which includes L elements. The j-th element represents the j-th data element in the fused data, and L represents the total number of fused data elements, where L>1 and 0<j≤L.
42. The data processing method according to any one of claims 26-27, wherein, At least one of the operation objects includes tensor data, and the tensor data is indicated by at least one descriptor, the descriptor indicating at least one of the following: shape information of the tensor data and spatial information of the tensor data; The method also includes: The descriptor is parsed to obtain the tensor data.
43. The data processing method according to claim 42, wherein parsing the descriptor includes: Based on the shape information, determine the data address of the tensor data in the data storage space; and / or Based on the spatial information, the dependencies between instructions are determined.
44. The data processing method according to claim 42, wherein the shape information of the tensor data includes at least one shape parameter representing the shape of the N-dimensional tensor data, where N is a positive integer, and the shape parameter of the tensor data includes at least one of the following: The dimensions of the data storage space containing the tensor data in at least one of the N dimensions, the dimensions of the storage area of the tensor data in at least one of the N dimensions, the offset of the storage area in at least one of the N dimensions, the positions of at least two vertices at diagonal positions in the N dimensions relative to the data reference point, and the mapping relationship between the data description location and the data address of the tensor data.
45. The data processing method of claim 42, wherein the shape information of the tensor data indicates at least one shape parameter of the shape of N-dimensional tensor data comprising multiple data blocks, where N is a positive integer, and the shape parameter includes at least one of the following: The dimensions of the data storage space containing the tensor data in at least one of the N dimensions, the dimensions of the storage area of a single data block in at least one of the N dimensions, the block step size of the data block in at least one of the N dimensions, the number of data blocks in at least one of the N dimensions, and the overall step size of the data block in at least one of the N dimensions.
46. The data processing method according to any one of claims 25-27, wherein The merging instruction includes an operation mode bit to indicate that the merging instruction is a merge sorting operation, or the merging instruction itself indicates the merge sorting operation.
Citation Information
Patent Citations
Method and system for data transfer
CN104021123A
Parallel sorting method of multiple groups of ordered sequences
CN105045600A
Device, method and application supporting vector ordering
CN108733352A
Merge instruction processing method and device, electronic equipment and storage medium
CN110851787A