Data processing apparatus, data processing method and electronic device
By establishing direct transmission channels between computing unit groups, the problem of data interaction between computing units relying on caching in existing technologies is solved, thereby improving the computing efficiency and bandwidth utilization efficiency of parallel processors.
Patent Information
- Application Number
- CN202211547138.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-12-05
- Publication Date
- 2025-10-31
- Estimated Expiration
- 2042-12-05
AI Technical Summary
In existing parallel processors, data interaction between computing units relies on caching, resulting in low bandwidth utilization efficiency and an inability to effectively improve computing efficiency.
Establish direct transmission channels between computing unit groups, including combinations of inter-unit and inter-group request-feedback channels, to form request and feedback links and reduce reliance on caching.
By using direct transmission channels, the consumption of cache bandwidth is reduced, and the data transmission efficiency and computing performance between computing units are improved.
Smart Images

Figure CN115934622B_ABST
Abstract
Description
Technical Field
[0001] Embodiments of this disclosure relate to a data processing apparatus, a data processing method, and an electronic device. Background Technology
[0002] With the development of artificial intelligence technology, higher demands are being placed on the bandwidth and computing power of parallel processors. However, due to limitations in current processes and materials, bandwidth and computing power have significant physical limits. For tasks with high data repetition rates, such as matrix operations, increasing data reuse can effectively reduce bandwidth usage and thus significantly improve computational efficiency.
[0003] Typical parallel processors consist of multiple computing units, each containing multiple processing elements. Each processing element performs computations on data, and these elements interact with each other via vector memory. However, there is no direct data exchange path between computing units. Therefore, if interaction is required, one computing unit must first store the data in a cache, and then the other unit must read it from the cache. The cache, as the primary medium for interaction, bears a significant portion of the data transfer workload. Therefore, if data exchange between computing units could be achieved without going through the cache—for example, if the vector memory in different computing units stores data from the same workgroup, making it visible to all processing elements running operations within that workgroup, and allowing these elements to freely access the data in the vector memory—then cache bandwidth usage could be greatly reduced. Summary of the Invention
[0004] At least one embodiment of this disclosure provides a data processing apparatus, including a plurality of computing units, which are divided into a plurality of computing unit groups. Each computing unit group includes at least two computing units, and each computing unit includes a vector memory. Two adjacent computing unit groups are coupled together through at least one corresponding first transmission channel to connect the plurality of computing unit groups in series, and information is transmitted between the vector memories included in different computing units in two adjacent computing unit groups.
[0005] For example, in the data processing apparatus provided in at least one embodiment of this disclosure, each computing unit group includes at least two computing units coupled through at least one corresponding second transmission channel to transmit information between vector memories included in the at least two computing units respectively.
[0006] For example, in a data processing apparatus provided in at least one embodiment of this disclosure, at least one second transmission channel includes at least one inter-unit request-feedback channel combination, the inter-unit request-feedback channel combination including a second request channel from the starting calculation unit to the pointing calculation unit and a second feedback channel from the pointing calculation unit to the starting calculation unit.
[0007] For example, in the data processing apparatus provided in at least one embodiment of this disclosure, each computing unit group includes at least two computing units, including a first computing unit and a second computing unit, and at least one second transmission channel includes a first inter-unit request-feedback channel combination with the first computing unit as the starting computing unit and a second inter-unit request-feedback channel combination with the second computing unit as the starting computing unit.
[0008] For example, in a data processing apparatus provided in at least one embodiment of this disclosure, at least one first transmission channel includes at least one inter-group request-feedback channel combination, each of the at least one inter-group request-feedback channel combination including a first request channel from the starting computing unit group to the pointing computing unit group and a first feedback channel from the pointing computing unit group to the starting computing unit group.
[0009] For example, in the data processing apparatus provided in at least one embodiment of this disclosure, each computing unit group includes at least two computing units, including a first computing unit and a second computing unit, and at least one inter-group request-feedback channel combination includes a first inter-group request-feedback channel combination and a second inter-group request-feedback channel combination. The first request channel and the first feedback channel of the first inter-group request-feedback channel combination are disposed between the vector memories of the two first computing units respectively included in two adjacent computing unit groups, and the first request channel and the first feedback channel of the second inter-group request-feedback channel combination are disposed between the vector memories of the two second computing units respectively included in two adjacent computing unit groups.
[0010] For example, in the data processing apparatus provided in at least one embodiment of this disclosure, a first request channel between each first computing unit in a plurality of computing unit groups is connected in series to form a first request link, a first request channel between each second computing unit in a plurality of computing unit groups is connected in series to form a second request link, a first feedback channel between each first computing unit in a plurality of computing unit groups is connected in series to form a first feedback link, and a first feedback channel between each second computing unit in a plurality of computing unit groups is connected in series to form a second feedback link.
[0011] For example, in the data processing apparatus provided in at least one embodiment of this disclosure, the first request link and the second request link are unidirectional links; the first feedback link and the second feedback link are unidirectional links.
[0012] For example, in at least one embodiment of the data processing apparatus provided in this disclosure, the data processing apparatus includes a request network and a feedback network. The request network includes a second request channel and a first request link and a second request link in each of a plurality of computing unit groups. The feedback network includes a second feedback channel and a first feedback link and a second feedback link in each of a plurality of computing unit groups.
[0013] For example, in the data processing apparatus provided in at least one embodiment of this disclosure, each of the plurality of computing units further includes a plurality of processing elements and a command queue, and the vector memory includes a vector memory body and a vector memory node. The vector memory body is used to receive stimuli from the command queue and the plurality of processing elements and to interact with the vector memory node. The vector memory node is used to interact with the vector memory in the adjacent computing unit through a combination of inter-unit request-feedback channels or a combination of inter-group request-feedback channels.
[0014] At least one embodiment of this disclosure provides a data processing method for a data processing apparatus provided in at least one embodiment of this disclosure. Multiple computing unit groups include a first computing unit group and a second computing unit group. The data processing method includes: a vector memory of a source computing unit in the first computing unit group transmitting a read request to a vector memory of a destination computing unit in the second computing unit group via at least one first transmission channel; and the vector memory of the destination computing unit in the second computing unit group transmitting the requested data to the vector memory of the source computing unit in the first computing unit group via at least one first transmission channel.
[0015] At least one embodiment of this disclosure provides a data processing method for a data processing apparatus provided in at least one embodiment of this disclosure. Multiple computing unit groups include a first computing unit group and a second computing unit group. The data processing method includes: a vector memory of a source computing unit in the first computing unit group transmitting a read request to a vector memory of a destination computing unit in the second computing unit group via a request network; and the vector memory of the destination computing unit in the second computing unit group transmitting the requested data to the vector memory of the source computing unit in the first computing unit group via a feedback network.
[0016] For example, in a data processing method provided in at least one embodiment of this disclosure, the vector memory includes a vector memory node, and the processing method further includes: using the vector memory node to interact with the vector memory in an adjacent computing unit through an inter-cell request-feedback channel combination or an inter-group request-feedback channel combination.
[0017] For example, in a data processing method provided in at least one embodiment of this disclosure, the vector memory of the source computing unit in the first computing unit group transmits a read request to the vector memory of the destination computing unit in the second computing unit group through a request network. This includes: during the transmission process from the source computing unit in the first computing unit group to the destination computing unit in the second computing unit group, determining the transmission path of the read request based on the identifier of the current computing unit and the identifier of the destination computing unit using a shortest path determination algorithm, wherein the computing units traversed in the transmission path are referred to as nodes in the transmission path.
[0018] For example, in the data processing method provided in at least one embodiment of this disclosure, determining the transmission path of a read request using the shortest path determination algorithm includes: when the identifier of the current computing unit is equal to the identifier of the destination computing unit, for the current node: determining whether there are sufficient processing resources to process the read request; if there are sufficient processing resources, storing the read requests in a cache for processing operations according to priority order, and notifying the previous node to input the information corresponding to the read request, parsing and processing the read requests in order; or if there are insufficient processing resources, processing read requests with a first priority in order, and notifying the previous node that it needs to wait for read requests with a second priority, where the first priority is higher than the second priority.
[0019] For example, in the data processing method provided in at least one embodiment of this disclosure, determining whether there are sufficient processing resources to process the read request includes: determining the remaining resources of the current node for processing the read request and whether there are sufficient resources in the vector memory of the current node.
[0020] For example, in the data processing method provided in at least one embodiment of this disclosure, determining the transmission path of a read request through the shortest path judgment algorithm further includes: when the identifier of the current computing unit is not equal to the identifier of the destination computing unit, for the current node: determining whether there are enough forwarding resources to forward the read request; if there are enough forwarding resources, storing the read request in a cache for forwarding operations according to priority order, and notifying the previous node to input the information corresponding to the read request, and forwarding the read request to the next node in order; or if there are not enough forwarding resources, storing the read request with third priority in a cache for forwarding operations, and notifying the node corresponding to the read request with third priority to output the read request; for the read request with fourth priority, notifying the node corresponding to the read request with fourth priority to wait, with third priority being higher than fourth priority.
[0021] For example, in the data processing method provided in at least one embodiment of this disclosure, determining whether there are sufficient forwarding resources to forward the read request includes: determining the resource status of the current node in forwarding the read request and whether the next node can receive the read request.
[0022] For example, at least one embodiment of the data processing method provided in this disclosure further includes: during the transmission of a read request, adjacent nodes establish a connection through a handshake mechanism, and transmit instructions or data through a combination of inter-unit request-feedback channels or a combination of inter-group request-feedback channels between nodes.
[0023] For example, in the data processing method provided in at least one embodiment of this disclosure, the adjacent nodes include a master node and a slave node, and the handshake mechanism includes: the master node prepares data and sends a ready signal to enable, and outputs first data to the slave node; the slave node determines whether it can receive the first data, and if it can receive the first data, the slave node will receive the ready signal to enable and receive the first data.
[0024] For example, in the data processing method provided in at least one embodiment of this disclosure, the handshake mechanism further includes: if the master node detects that the receive ready signal is enabled, maintaining the transmit ready signal enabled and outputting the second data located after the first data to the slave node; or if the master node does not detect that the receive ready signal is enabled, the master node maintains the status quo.
[0025] At least one embodiment of this disclosure provides an electronic device, including a data processing device provided in at least one embodiment of this disclosure. Attached Figure Description
[0026] To more clearly illustrate the technical solutions of the embodiments of this disclosure, the accompanying drawings of the embodiments will be briefly described below. Obviously, the drawings described below only relate to some embodiments of this disclosure and are not intended to limit this disclosure.
[0027] Figure 1 A schematic diagram of a data processing device for parallel computing is shown.
[0028] Figure 2 A method for Figure 1 A schematic diagram of the internal structure of the computing unit of a data processing device;
[0029] Figure 3 A schematic diagram of a data processing apparatus provided in at least one embodiment of the present disclosure is shown;
[0030] Figure 4 A schematic diagram of the structure of a request network provided in at least one embodiment of this disclosure is shown;
[0031] Figure 5A schematic diagram of the structure of a feedback network provided in at least one embodiment of this disclosure is shown;
[0032] Figure 6 This illustration shows a schematic diagram of the structure of nodes in a request network and a feedback network provided in at least one embodiment of the present disclosure;
[0033] Figure 7 A schematic flowchart of a data processing method provided in at least one embodiment of the present disclosure is shown;
[0034] Figure 8 A schematic flowchart of another data processing method provided by at least one embodiment of the present disclosure is shown. Detailed Implementation
[0035] To make the objectives, technical solutions, and advantages of the embodiments of this disclosure clearer, the technical solutions of the embodiments of this disclosure will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this disclosure. All other embodiments obtained by those skilled in the art based on the described embodiments of this disclosure without creative effort are within the scope of protection of this disclosure.
[0036] Unless otherwise defined, the technical or scientific terms used in this disclosure shall have the ordinary meaning understood by one of ordinary skill in the art to which this disclosure pertains. The terms “first,” “second,” and similar terms used in this disclosure do not indicate any order, quantity, or importance, but are merely used to distinguish different components. Similarly, the terms “an,” “a,” or “the,” and similar terms do not indicate a quantity limitation, but rather indicate the presence of at least one. The terms “including,” “comprising,” or “containing,” and similar terms mean that the element or object preceding the word encompasses the elements or objects listed following the word and their equivalents, without excluding other elements or objects. The terms “connected,” “linked,” or similar terms are not limited to physical or mechanical connections, but can include electrical connections, whether direct or indirect. The terms “upper,” “lower,” “left,” and “right,” etc., are used only to indicate relative positional relationships, and these relative positional relationships may change accordingly when the absolute position of the described objects changes.
[0037] Figure 1 A schematic diagram of a data processing device for parallel computing is shown. For example, this data processing device is a parallel processor, such as a general-purpose graphics processing unit (GPGPU).
[0038] like Figure 1As shown, the data processing device includes multiple computing units (CUs) CU0 to CU15 and a cache (not shown). Each computing unit interacts with the cache via a cache input / output interface (cache_IO) (e.g., CU0 interacts with the cache via cache_IO0). In this data processing device, there is no direct data interaction path between one computing unit and another. Therefore, if interaction is required, one computing unit must first store the data in the cache, and then the other computing unit reads it from the cache. In a GPGPU, the smallest unit for a predetermined task to be run by a computing unit is called a thread bundle. For example, each thread bundle contains 32 or 64 threads, and the thread bundles running within a computing unit are independent of each other.
[0039] Figure 2 A method for Figure 1 A schematic diagram of the internal structure of the computing unit 200 of the data processing device. (See diagram below.) Figure 2 As shown, the computing unit 200 includes processing elements 201 and 202, a vector memory 203, and a command queue 204. Processing elements 201 and 202 each perform calculations on vector data, and each contains numerous vector registers for calculations (e.g., floating-point multiplication and division, complex calculations using trigonometric functions such as cosine and sin). The vector memory 203 handles data storage and interaction between processing elements 201 and 202, for example, through interfaces PE0 idx / data and PE1 idx / data. The command queue 204 controls the operation of processing elements 201 and 202 and the vector memory 203 based on specific programs, for example, through interfaces PE0_cmd, VM_cmd, and PE1_cmd to facilitate interaction and communication between the command queue 204 and corresponding devices. The vector data of the computing unit 200 interacts with the outside world through cache_IO.
[0040] Figure 1 The structure of the data processing device shown can handle the computation of independent vector data very well. In this data processing device, the computation between computing units has a high degree of independence, but it is obviously disadvantageous for data sharing within the parallel processor.
[0041] Another data processing device includes, for example, 16 computing units CU0 to CU15. In this device, adjacent odd and even computing units are connected, dividing the 16 computing units CU0 to CU15 into 8 computing unit groups. Each computing unit group includes two computing units. For example, computing units CU0 and CU1 form one computing unit group; or, for example, computing units CU2 and CU3 form one computing unit group, and so on, with odd-numbered computing units and even-numbered computing units forming one computing unit group. The two computing units within each computing unit group are coupled through an instruction transmission channel (message) and a data address transmission channel (idx / data) to transmit instructions, addresses, and data. The vector data of the computing units in the data processing device interacts with the outside world through cache_IO. This data processing device can achieve data sharing between vector memories within the computing unit group, but for data sharing between a larger range of computing units, caching is still required.
[0042] At least one embodiment of this disclosure provides a data processing apparatus including a plurality of computing units, which are divided into a plurality of computing unit groups. Each computing unit group includes at least two computing units, and each computing unit includes a vector memory. Two adjacent computing unit groups are coupled together through at least one corresponding first transmission channel to connect the plurality of computing unit groups in series, and information is transmitted between the vector memories included in different computing units in two adjacent computing unit groups.
[0043] The data processing apparatus provided in the above embodiments of this disclosure can achieve data sharing between vector memories of computing units at a relatively low cost, effectively save bandwidth consumption, improve the performance of parallel processors, effectively shorten the data transmission latency between computing units, and improve the efficiency of data transmission between computing units.
[0044] At least one embodiment of this disclosure also provides a data processing method for the above-described data processing apparatus and an electronic device including the above-described data processing apparatus.
[0045] The embodiments of this disclosure will now be described in detail with reference to the accompanying drawings, but this disclosure is not limited to these specific embodiments.
[0046] Figure 3 A schematic diagram of a data processing apparatus 400 provided in at least one embodiment of the present disclosure is shown. For example, the data processing apparatus is a parallel processor, such as a general-purpose graphics processing unit (GPGPU), a data processor, a tensor processor, etc.
[0047] like Figure 3As shown, the data processing device 400 includes multiple computing units CU0 to CU15, which are divided into multiple computing unit groups. Each computing unit group includes at least two computing units, for example, two computing units. For example, computing unit CU0 and computing unit CU1 form one computing unit group, and computing unit CU2 and computing unit CU3 form another computing unit group, and so on, with odd-numbered computing units and even-numbered computing units forming one computing unit group. Each computing unit includes a vector memory, multiple processing elements, and a command queue. For example, computing unit CU0 includes a vector memory VM0, processing elements PE0_0 and PE0_1, and a command queue CQ0.
[0048] Two adjacent computing unit groups are coupled through at least one corresponding first transmission channel to connect multiple computing unit groups in series, and to transmit information between the vector memories included in different computing units within the two adjacent computing unit groups. For example, a computing unit group consisting of computing units CU0 and CU1 and a computing unit group consisting of computing units CU2 and CU3 are two adjacent computing unit groups. These two computing unit groups are coupled through first transmission channels (including channels req_ch0_2, ret_ch2_0, req_ch3_1, and ret_ch1_3) to transmit information between the vector memory VM0 included in computing unit CU0 and the vector memory VM2 included in computing unit CU2, and between the vector memory VM1 included in computing unit CU1 and the vector memory VM3 included in computing unit CU3. The situation between other adjacent computing unit groups is the same as described above and will not be repeated.
[0049] For example, in some embodiments of this disclosure, each computing unit group includes at least two computing units coupled through at least one corresponding second transmission channel to transmit information between vector memories included in the at least two computing units respectively.
[0050] For example, such as Figure 3 As shown in the diagram, the computing unit group at the top, comprising computing unit CU0 and computing unit CU1, is coupled via a second transmission channel (including channels req_ch1_0, ret_ch0_1, req_ch0_1, and ret_ch1_0) to transmit information between the vector memory VM0 and vector memory VM1 included in computing unit CU0 and computing unit CU1, respectively. The situation for other computing unit groups is the same as described above and will not be repeated.
[0051] For example, in some embodiments of this disclosure, at least one second transmission channel includes at least one inter-unit request-feedback channel combination, which includes a second request channel from the starting computing unit to the pointing computing unit and a second feedback channel from the pointing computing unit to the starting computing unit.
[0052] For example, such as Figure 3 As shown, between computing unit CU0 and computing unit CU1, the second transmission channel (including req_ch1_0, ret_ch0_1, req_ch0_1, and ret_ch1_0) includes two inter-unit request-feedback channel combinations. One inter-unit request-feedback channel combination includes a second request channel req_ch1_0 from the starting computing unit CU1 (hereinafter referred to as the "starting computing unit") to the pointed computing unit CU0 (hereinafter referred to as the "pointing computing unit") and a second feedback channel ret_ch0_1 from the pointing computing unit CU0 to the starting computing unit CU1. The other inter-unit request-feedback channel combination includes a second request channel req_ch0_1 from the starting computing unit CU0 to the pointing computing unit CU1 and a second feedback channel ret_ch1_0 from the pointing computing unit CU1 to the starting computing unit CU0.
[0053] For example, such as Figure 3 As shown, between computing unit CU2 and computing unit CU3, the second transmission channel (including req_ch3_2, ret_ch2_3, req_ch2_3, and ret_ch3_2) includes two inter-unit request-feedback channel combinations. One inter-unit request-feedback channel combination includes a second request channel req_ch3_2 from the starting computing unit CU3 to the computing unit CU2 and a second feedback channel ret_ch2_3 from the computing unit CU2 to the starting computing unit CU3. The other inter-unit request-feedback channel combination includes a second request channel req_ch2_3 from the starting computing unit CU2 to the computing unit CU3 and a second feedback channel ret_ch3_2 from the computing unit CU3 to the starting computing unit CU2.
[0054] For example, in some embodiments of this disclosure, each computing unit group includes at least two computing units, including a first computing unit and a second computing unit, and at least one second transmission channel includes a first inter-unit request-feedback channel combination with the first computing unit as the starting computing unit and a second inter-unit request-feedback channel combination with the second computing unit as the starting computing unit.
[0055] For example, such as Figure 3As shown, channels req_ch0_1 and ret_ch1_0 form a first inter-unit request-feedback channel combination with the first computing unit CU0 as the starting computing unit, and channels req_ch1_0 and ret_ch0_1 form a second inter-unit request-feedback channel combination with the second computing unit CU1 as the starting computing unit.
[0056] Within each computing unit group, two inter-unit request-feedback channels are provided to enhance the data interaction capabilities within the computing unit group. It should be noted that the directions of the request and feedback channels between computing unit groups are not necessarily opposite, and this disclosure does not impose any restrictions on this.
[0057] For example, in some embodiments of this disclosure, at least one first transmission channel between two adjacent computing unit groups includes at least one inter-group request-feedback channel combination, each of the at least one inter-group request-feedback channel combination including a first request channel from the starting computing unit group to the computing unit group and a first feedback channel from the computing unit group to the starting computing unit group.
[0058] For example, such as Figure 3 As shown, the first transmission channels req_ch0_2 and ret_ch2_0 form an inter-group request-feedback channel combination. This inter-group request-feedback channel combination includes a first request channel req_ch0_2 from the starting computing unit group (here, the computing unit group composed of computing unit CU0 and computing unit CU1) (hereinafter referred to as the "starting computing unit group") to the pointed computing unit group (here, the computing unit group composed of computing unit CU2 and computing unit CU3) (hereinafter referred to as the "pointing computing unit group") and a first feedback channel ret_ch2_0 from the pointing computing unit group to the starting computing unit group.
[0059] For example, in some embodiments of this disclosure, each computing unit group includes at least two computing units, including a first computing unit and a second computing unit, and at least one inter-group request-feedback channel combination includes a first inter-group request-feedback channel combination and a second inter-group request-feedback channel combination. The first request channel and the first feedback channel of the first inter-group request-feedback channel combination are disposed between the vector memories of the two first computing units respectively included in two adjacent computing unit groups, and the first request channel and the first feedback channel of the second inter-group request-feedback channel combination are disposed between the vector memories of the two second computing units respectively included in two adjacent computing unit groups.
[0060] For example, such as Figure 3As shown, the computing unit group composed of computing units CU0 and CU1 and the computing unit group composed of computing units CU2 and CU3 include a first inter-group request-feedback channel combination and a second inter-group request-feedback channel combination. The first request channel req_ch0_2 and the first feedback channel ret_ch2_0 of the first inter-group request-feedback channel combination are located between the vector memories VM0 and VM2 of the two first computing units CU0 and CU2. The first request channel req_ch3_1 and the first feedback channel ret_ch1_3 of the second inter-group request-feedback channel combination are located between the vector memories VM1 and VM3 of the two second computing units CU1 and CU3.
[0061] For example, in some embodiments of this disclosure, first request channels between each first computing unit in a plurality of computing unit groups are connected in series to form a first request link, first request channels between each second computing unit in a plurality of computing unit groups are connected in series to form a second request link, first feedback channels between each first computing unit in a plurality of computing unit groups are connected in series to form a first feedback link, and first feedback channels between each second computing unit in a plurality of computing unit groups are connected in series to form a second feedback link.
[0062] For example, such as Figure 3 As shown, the first request channels (req_ch0_2, req_ch2_4, ..., req_ch12_14) between each first computing unit (CU0, CU2, ..., CU14) are as follows: some first request channels are... Figure 3 (Not shown in the image) are connected in series to form a first request link, and the first request channels (req_ch3_1, req_ch5_3, ..., req_ch15_13, some of the first request channels are in series between the various second computing units (CU1, CU3, ..., CU15). Figure 3 (Not shown in the image) are connected in series to form a second request link. The first feedback channels (ret_ch2_0, ret_ch4_2, ..., ret_ch14_12, some of the first feedback channels are in series) between the various first computing units (CU0, CU2, ..., CU14). Figure 3 (Not shown in the diagram) are connected in series to form a first feedback link. The first feedback channels (ret_ch1_3, ret_ch3_5, ..., ret_ch13_15) between each of the second computing units (CU1, CU3, ..., CU15) are some of the first feedback channels in the diagram. Figure 3 (Not shown in the image) are connected in series to form a second feedback link.
[0063] For example, in some embodiments of this disclosure, the first request link and the second request link are unidirectional links; the first feedback link and the second feedback link are unidirectional links.
[0064] For example, such as Figure 3 As shown, the first request link is a unidirectional link with the direction CU0->CU2->……->CU14, and the second request link is a unidirectional link with the direction CU15->CU13->……->CU1; the first feedback link is a unidirectional link with the direction CU14->CU12->……->CU0, and the second feedback link is a unidirectional link with the direction CU1->CU3->……->CU15.
[0065] It should be noted that, Figure 3 The directions of the request and feedback links shown are merely an example. This disclosure does not impose any restrictions as long as both the request and feedback links are unidirectional and can traverse all computational units.
[0066] For example, in some embodiments of this disclosure, the data processing apparatus includes a request network and a feedback network. The request network includes a second request channel and a first request link and a second request link for each of a plurality of computing unit groups. The feedback network includes a second feedback channel and a first feedback link and a second feedback link for each of a plurality of computing unit groups.
[0067] Figure 4 A schematic diagram of the structure of a request network provided in at least one embodiment of this disclosure is shown.
[0068] like Figure 4 As shown, corresponding to Figure 3 In this case, the request network includes the second request channels req_ch1_0, req_ch0_1, req_ch3_2, ..., req_ch15_14, req_ch14_15 in each of the multiple computing unit groups, as well as the first request link (the first request link composed of the first request channels req_ch0_2, req_ch2_4, ..., req_ch12_14) and the second request link (the second request link composed of the first request channels req_ch3_1, req_ch5_3, ..., req_ch15_13).
[0069] Within the computing unit group, the request network is bidirectional to ensure high data interoperability under high spatial correlation.
[0070] Figure 5 A schematic diagram of the structure of a feedback network provided in at least one embodiment of this disclosure is shown.
[0071] like Figure 5 As shown, corresponding to Figure 3In this case, the feedback network includes the second feedback channels ret_ch1_0, ret_ch0_1, ret_ch3_2, ..., ret_ch15_14, ret_ch14_15 in each of the multiple computing unit groups, as well as the first feedback link (the first feedback link composed of the first feedback channels ret_ch2_0, ret_ch4_2, ..., ret_ch14_12) and the second feedback link (the second feedback link composed of the first feedback channels ret_ch1_3, ret_ch3_5, ..., ret_ch13_15).
[0072] In the above embodiments of this disclosure, the transmission of request data and feedback data is organized in the form of a network, which eliminates the need for overall arbitration, effectively shortens the latency of data transmission between various computing units, and improves the efficiency of data transmission between computing units.
[0073] For example, in some embodiments of this disclosure, the vector memory in the computing unit includes a vector memory body and vector memory nodes, wherein the vector memory body is used to receive commands from a command queue and stimuli from multiple processing elements and to interact with the vector memory nodes, and the vector memory nodes are used to interact with the vector memory in adjacent computing units through inter-unit request-feedback channel combinations or inter-group request-feedback channel combinations.
[0074] Figure 6 A schematic diagram of the structure of nodes in a request network and a feedback network provided in at least one embodiment of the present disclosure is shown.
[0075] like Figure 6 As shown, for a node CUm in the request network and feedback network, it includes a command queue CQ, processing elements PE0 and PE1, and a vector memory VMm. The vector memory VMm includes a vector memory body VMMm and a vector memory node VMm_node.
[0076] The vector memory body VMMm communicates with the command queue CQ, processing elements PE0 and PE1 in the same computing unit. For example, there is an interface / channel PE0_VM_idx / data for transmitting addresses and data and an interface / channel VM_PE0_data for transmitting data between the processing element PE0 and the vector memory body VMMm; there is an interface / channel CQ_VM_cmd for transmitting commands and an interface / channel VM_CQ_rddone for transmitting response signals between the command queue CQ and the vector memory body VMMm.
[0077] The vector memory body VMMm communicates with the vector memory node VMm_node. For example, there are interfaces / channels between VMMm and VMm_node for transmitting request data (req_self), feedback data (ret_self), data sent to the processing element (node_PE_data), converting received read requests into instructions (node_CQ_cmd), converting received read requests into addresses / data (node_VMM_idx / data), and transmitting response signals (node_CQ_rddone). Before data transmission occurs between interfaces / channels req_self and ret_self, VMMm and VMm_node connect via a handshake mechanism. Interfaces / channels req_self_rts, req_self_rtr, ret_self_rts, and ret_self_rtr are used to transmit handshake signals. The handshake mechanism will be explained below and will not be repeated here. The interface / channel credit is used to transmit data related to the credit mechanism. This data represents the amount of read requests that can be stored internally. When a new read request is received, the value of this data is decremented by 1. The associated nodes of node CUM are CUn, CUk, and CUj. Since the relevant signals are directly connected to the vector memory, vector memories VMn, VMk, and VMj are used instead of CUn, CUk, and CUj for explanation. VMn is the vector memory of computing unit CUn, which is in the same computing unit group as computing unit CUM. VMk and VMj are the vector memories of computing units CUk and CUj, which are adjacent to computing unit CUM. Between vector memories VMm and VMn are two inter-unit request-feedback channel combinations (one combination includes a request channel req_n_m and a feedback channel ret_m_n, and the other includes a request channel req_m_n and a feedback channel ret_n_m). The vector memory VMm and VMj, as well as VMm and VMk, include inter-group request-feedback channel combinations. The inter-group request-feedback channel combination between VMm and VMj includes a request channel req_m_j and a feedback channel ret_j_m. The inter-group request-feedback channel combination between VMm and VMk includes a request channel req_k_m and a feedback channel ret_m_k.
[0078] Figure 7A schematic flowchart illustrating a data processing method provided in at least one embodiment of this disclosure is shown. This data processing method is applied, for example, to the aforementioned data processing apparatus. For instance, the plurality of computing unit groups in the data processing apparatus include a first computing unit group and a second computing unit group. The terms "first computing unit group" and "second computing unit group" are used for convenience in describing the data processing method and do not specifically refer to any two computing unit groups, but rather to represent any two computing unit groups involved in the operation.
[0079] like Figure 7 As shown, the data processing method includes the following steps S801 to S802.
[0080] Step S801: The vector memory of the source computing unit in the first computing unit group transmits the read request to the vector memory of the destination computing unit in the second computing unit group through at least one first transmission channel.
[0081] Step S802: The vector memory of the destination computing unit in the second computing unit group transmits the requested data to the vector memory of the source computing unit in the first computing unit group through at least one first transmission channel.
[0082] For example, when a source computing unit in the first computing unit group needs to read data stored in the vector memory of a destination computing unit in the second computing unit group, the read request is transmitted to the destination computing unit through at least one first transmission channel. The vector memory in the destination computing unit transmits the requested data and a completion indication signal to the vector memory of the source computing unit through at least one first transmission channel. The requested data is stored in the corresponding processing element of the source computing unit, and the completion indication signal is received by the command queue to mark the completion of the entire operation.
[0083] Figure 8 A schematic flowchart illustrating another data processing method provided by at least one embodiment of this disclosure is shown. This data processing method is applied, for example, to the data processing apparatus described above. For instance, the plurality of computing unit groups in the data processing apparatus include a first computing unit group and a second computing unit group. Similarly, "first computing unit group" and "second computing unit group" are used to refer to any two computing unit groups involved in the operation.
[0084] like Figure 8 As shown, the data processing method includes the following steps S901 to S902.
[0085] Step S901: The vector memory of the source computing unit in the first computing unit group transmits the read request to the vector memory of the destination computing unit in the second computing unit group through the request network.
[0086] For instructions on requesting a network, please refer to, for example, [link to relevant documentation]. Figure 4 Related descriptions.
[0087] Step S902: The vector memory of the destination computing unit in the second computing unit group transmits the requested data to the vector memory of the source computing unit in the first computing unit group through a feedback network.
[0088] For an explanation of feedback networks, please refer to, for example, [link to relevant documentation]. Figure 5 Related descriptions.
[0089] For example, such as Figure 3 As shown, in one example, the first computing unit group is a computing unit group consisting of computing unit CU0 and computing unit CU1, and the second computing unit group is a computing unit group consisting of computing unit CU6 and computing unit CU7. The source computing unit is computing unit CU0, the destination computing unit is computing unit CU7, computing unit CU0 performs a read operation on computing unit CU7, the vector memory VM0 of computing unit CU0 transmits the read request to the vector memory VM7 of destination computing unit CU7 through the request network, and the vector memory VM7 transmits the requested data to the vector memory VM0 through the feedback network.
[0090] For example, in some embodiments of this disclosure, as described above, the vector memory includes vector memory nodes, and the processing method may further include: using the vector memory nodes to interact with vector memories in adjacent computing units via inter-cell request-feedback channel combinations or inter-group request-feedback channel combinations.
[0091] For example, in some embodiments of this disclosure, step S901 may include step S911: during the transmission process from the source computing unit in the first computing unit group to the destination computing unit in the second computing unit group, based on the identifier of the current computing unit and the identifier of the destination computing unit, the transmission path of the read request is determined by the shortest path judgment algorithm, and the computing units passed through in the transmission path are called nodes in the transmission path.
[0092] For node-to-node access in a network, a multipath problem exists, requiring the selection of a preferred path from multiple possible paths. To minimize the number of nodes in the access path and reduce the impact on other data transmissions, at least one embodiment of this disclosure provides a shortest path determination method. For each node in the request network and the feedback network, there are two transmission directions, one of which is transmission within the computing unit group (see reference). Figure 3 This can be considered as a left-right direction, with one direction being the transfer between computational unit groups (see reference). Figure 3This can be viewed as an up-down direction. Based on the identifier Cur_id of the current computing unit and the identifier Tag_id of the destination computing unit, the direction selection for each node can be quickly determined. Specifically, for the request network, the direction selection scheme is shown in Table 1 below.
[0093] Table 1: Path selection scheme for network requests
[0094]
[0095] In this context, "updown" indicates the transmission direction between computing unit groups, while "leftright" indicates the transmission direction within a computing unit group. CQ&VM indicates that when a read request arrives at the destination computing unit, it enters the command queue for instruction parsing and then enters the vector memory for address and data processing. CQ&PE indicates that when the requested data arrives at the source computing unit, it enters the command queue for completion indication signal calibration and then enters the processing element for data storage. During the transmission of a read request, the request carries the identifiers of the source and destination computing units. The destination computing unit identifier indicates the target node the read request needs to reach, and the source computing unit identifier indicates the original node that issued the read request. During the transmission of the read request, the destination computing unit identifier is used as the Tag_id, and the transmission direction of the read request is determined by comparing Cur_id and Tag_id. Similarly, during the return transmission of the requested data, the source computing unit identifier is used as the Tag_id, and the transmission direction of the requested data is determined by comparing Cur_id and Tag_id.
[0096] For example, such as Figure 3 As shown in the example, the source computing unit is CU0, the destination computing unit is CU7, computing unit CU0 performs a read operation on computing unit CU7, and requests the access order of each node in the network as CU0->CU2->CU4->CU6->CU7, and the feedback access order of each node in the network is CU7->CU6->CU4->CU2->CU0.
[0097] For example, in some embodiments of this disclosure, determining the transmission path of a read request using a shortest path algorithm may include: if the identifier of the current computing unit is equal to the identifier of the destination computing unit, for the current node: determining whether there are sufficient processing resources to process the read request; if there are sufficient processing resources, storing the read requests in priority order into a cache for processing operations, and notifying the previous node to input the information corresponding to the read request, parsing and processing the read requests in sequence; or if there are insufficient processing resources, processing read requests with a first priority in sequence, and notifying the previous node that it needs to wait for read requests with a second priority. For example, the first priority is higher than the second priority.
[0098] For example, in some embodiments of this disclosure, determining whether there are sufficient processing resources to process a read request may include: determining the remaining resources of the current node for processing the read request and whether there are sufficient resources in the vector memory of the current node.
[0099] For example, in some embodiments of this disclosure, the shortest path determination method may further include: when the identifier of the current computing unit is not equal to the identifier of the destination computing unit, for the current node: determining whether there are sufficient forwarding resources to forward the read request; if there are sufficient forwarding resources, storing the read requests in a cache for forwarding operations according to priority, and notifying the previous node to input the information corresponding to the read request, and forwarding the read request to the next node in sequence; or if there are insufficient forwarding resources, storing the read request with the third priority in a cache for forwarding operations, and notifying the node corresponding to the read request with the third priority to output the read request; for the read request with the fourth priority, notifying the node corresponding to the read request with the fourth priority that it needs to wait. For example, the third priority is higher than the fourth priority.
[0100] For example, in some embodiments of this disclosure, determining whether there are sufficient forwarding resources to forward a read request may include: determining the resource availability of the current node for forwarding read requests and whether the next node can receive read requests.
[0101] For example, the data processing method provided in this disclosure embodiment may further include: during the transmission of a read request, adjacent nodes establish a connection through a handshake mechanism, and transmit instructions or data through a combination of inter-unit request-feedback channels or a combination of inter-group request-feedback channels between nodes.
[0102] For example, in some embodiments of this disclosure, adjacent nodes include a master node and a slave node, and a handshake mechanism is used: the master node prepares the data and enables the sending of a ready signal, and outputs the first data to the slave node; the slave node determines whether it can receive the first data, and if it can receive the first data, the slave node enables the receiving of the ready signal and receives the first data. Here, the master node is the node that sends the data, and the slave node is the node that receives the data. Each node may act as a master node or a slave node in different transactions.
[0103] For example, in some embodiments of this disclosure, the handshake mechanism further includes: if the master node detects that the receive ready signal is enabled, maintaining the transmit ready signal enabled and outputting the second data following the first data to the slave node; or if the master node does not detect that the receive ready signal is enabled, the master node maintains the status quo.
[0104] For example, such as Figure 6 As shown, for adjacent nodes VMm and VMk, node VMm acts as a slave node and node VMk acts as a master node. Master node VMk prepares data and enables the transmit ready signal (req_k_m_rts = 1), then outputs the first data to slave node VMm. Slave node VMm determines whether it can receive the first data. If it can, slave node VMk enables the receive ready signal (req_m_k_rtr = 1) and receives the first data. If master node VMk detects the receive ready signal (req_m_k_rtr = 1), it maintains the transmit ready signal enabled (req_k_m_rts = 1) and outputs the second data following the first data to slave node VMm; otherwise, if master node VMk does not detect the receive ready signal (req_m_k_rtr = 0), master node VMk maintains its current state.
[0105] It should be noted that the embodiments of this disclosure do not limit the first data and the second data. For example, the first data may be an instruction and the second data may be an address, or the first data and the second data may be transmitted in the form of data packets.
[0106] The following is combined with Figure 6 Table 1 details the workflow of the data processing method provided by at least one embodiment of this disclosure. The specific workflow includes steps S1 to S7:
[0107] Step S1: When node Cum receives a request from the previous node, it uses the shortest path determination scheme shown in Table 1 to determine whether the target node has been reached based on the identifier Tag_id of the destination computing unit and the identifier Cur_id of the current computing unit. If the target node has not been reached, it determines the direction of transmission. If the direction of transmission is up-down or left-right, it jumps to step S2; if the target node has been reached, it jumps to step S6.
[0108] Step S2: Consider the following in both the up-down and left-right directions: 1) the read request's enabling status, 2) the current node's resource availability for forwarding read requests, and 3) whether the next node can receive read requests. If resources are sufficient, proceed to step S3; if resources are insufficient, proceed to step S4.
[0109] Step S3: When resources are sufficient, read requests are stored in the cache for forwarding operations according to priority. The previous node is notified to input the information corresponding to the read request by receiving a ready signal, and the read request is forwarded to the next node in sequence.
[0110] For example, the cache used for forwarding operations is a first-in-first-out (FIFO) queue, managed through a credit mechanism.
[0111] Step S4: When resources are insufficient, store high-priority request information into a cache for forwarding operations according to priority, and notify the previous node of the high-priority request to output a read request by enabling the receive ready signal. For the remaining low-priority read requests, notify the corresponding node that it needs to wait by disabling the receive ready signal.
[0112] Step S5: When the read request has reached the destination computing unit, simultaneously consider: 1) the enabled status of the received read request, 2) the remaining resources of the current node for processing the read request, and 3) whether there are sufficient resources in the vector memory of the current node. If there are sufficient resources, proceed to step S6; if there are insufficient resources, proceed to step S7.
[0113] It should be noted that if a read request has reached the destination compute unit, it needs to be stored in the cache for processing operations and, at the appropriate time, converted into the instruction `node_CQ_cmd` and address / data `node_VMM_idx / data`, and then sent to the vector memory of the destination compute unit. The vector memory requires a certain amount of time to process each instruction. Operations on the node itself may cause the vector memory to be fully loaded. In this case, the vector memory cannot receive new `node_CQ_cmd` and `node_VMM_idx / data`, so it is necessary to determine the resource status of the vector memory, based on whether the receive ready signal is enabled.
[0114] Step S6: Store the read requests in the cache for processing operations according to their priority, and notify the previous node to input the information corresponding to the read request by receiving the ready signal. Then, parse and process the read requests in order.
[0115] Step S7: Process high-priority read requests as in step S6. For low-priority read requests, notify the corresponding node that it needs to wait by disabling the readiness signal.
[0116] The technical effects of the above data processing methods and Figure 3 The data processing device shown has the same technical effect, and will not be described again here.
[0117] According to at least one embodiment of this disclosure, an electronic device is also provided, which includes the aforementioned data processing device. For example, the data processing device is a parallel processor, such as a general-purpose graphics processing unit (GPGPU), a data processor, or a tensor processor. For example, the electronic device may also include, for example, a central processing unit (CPU), which, together with the data processing device, forms a heterogeneous computing system via a system bus. For example, the bus may be a Peripheral Component Interconnect Standard (PCI) bus or an Extended Industry Standard Architecture (EISA) bus. The communication bus may be divided into an address bus, a data bus, a control bus, etc.
[0118] For example, in at least one embodiment of this disclosure, the electronic device may further include input devices such as a touchscreen, touchpad, keyboard, mouse, camera, microphone, accelerometer, gyroscope, etc.; output devices such as a liquid crystal display, speaker, vibrator, etc.; storage devices such as magnetic tape, hard disk (HDD or SDD), etc.; and communication devices such as network interface cards such as LAN cards, modems, etc. The communication device can allow the electronic device to communicate wirelessly or wiredly with other devices to exchange data and perform communication processing via a network such as the Internet. A drive is connected to the I / O interface as needed. A removable storage medium, such as a disk, optical disk, magneto-optical disk, semiconductor memory, etc., is installed on the drive as needed so that computer programs read from it can be installed into the storage device as needed.
[0119] For example, the electronic device may further include peripheral interfaces. These peripheral interfaces can be various types of interfaces, such as USB interfaces, Lightning interfaces, etc. The communication device can communicate wirelessly with networks and other devices, such as the Internet, intranets and / or wireless networks such as cellular telephone networks, wireless local area networks (LANs) and / or metropolitan area networks (MANs). Wireless communication can use any of a variety of communication standards, protocols, and technologies, including but not limited to Global System for Mobile Communications (GSM), Enhanced Data GSM Environment (EDGE), Wideband Code Division Multiple Access (W-CDMA), Code Division Multiple Access (CDMA), Time Division Multiple Access (TDMA), Bluetooth, Wi-Fi (e.g., based on IEEE 802.11a, IEEE 802.11b, IEEE 802.11g, and / or IEEE 802.11n standards), Voice over Internet Protocol (VoIP), Wi-MAX, protocols for email, instant messaging, and / or Short Message Service (SMS), or any other suitable communication protocol.
[0120] The electronic device may be, for example, a system-on-a-chip (SOC) or a device that includes the SOC. For example, it may be any device such as a mobile phone, tablet computer, laptop computer, e-book reader, game console, television, digital photo frame, navigator, home appliance, communication base station, industrial controller, server, etc. It may also be any combination of data processing device and hardware. The embodiments disclosed herein do not limit this.
[0121] The following points should be noted regarding this disclosure:
[0122] (1) The accompanying drawings of the embodiments of this disclosure only involve the structures involved in the embodiments of this disclosure. Other structures can be referred to the general design.
[0123] (2) Where there is no conflict, the embodiments of this disclosure and the features in the embodiments can be combined with each other to obtain new embodiments.
[0124] The above description is only a specific embodiment of this disclosure, but the protection scope of this disclosure is not limited thereto. The protection scope of this disclosure should be determined by the protection scope of the claims.
Claims
1. A data processing apparatus, wherein, The data processing device is a parallel processor, comprising multiple computing units, wherein the multiple computing units are divided into multiple computing unit groups, each computing unit group comprising at least two computing units, each computing unit comprising a vector memory, multiple processing elements, and a command queue, each processing element being used to perform calculations on vector data, and the vector memory being used to perform data storage and interaction among the multiple processing elements. Two adjacent computing unit groups are coupled through at least one corresponding first transmission channel to connect the plurality of computing unit groups in series, and information is transmitted between vector memories included in different computing units in the two adjacent computing unit groups, wherein the at least one first transmission channel includes at least one inter-group request-feedback channel combination. Wherein, each of the computing unit groups includes at least two computing units, including a first computing unit and a second computing unit. The at least one inter-group request-feedback channel combination includes a first inter-group request-feedback channel combination and a second inter-group request-feedback channel combination. The first request channel and the first feedback channel of the first inter-group request-feedback channel combination are located between the vector memories of the two first computing units included in the two adjacent computing unit groups, respectively. The first request channel and the first feedback channel of the second group request-feedback channel combination are located between the vector memory of the two second computing units included in the two adjacent computing unit groups, respectively.
2. The data processing apparatus according to claim 1, wherein, Each computing unit group includes at least two computing units coupled through at least one corresponding second transmission channel to transmit information between the vector memories included in each of the at least two computing units.
3. The data processing apparatus according to claim 2, wherein, The at least one second transmission channel includes at least one inter-unit request-feedback channel combination. The inter-unit request-feedback channel combination includes a second request channel from the starting computing unit to the pointing computing unit and a second feedback channel from the pointing computing unit to the starting computing unit.
4. The data processing apparatus according to claim 3, wherein, Each group of computing units includes at least two computing units, including a first computing unit and a second computing unit. The at least one second transmission channel includes a first inter-unit request-feedback channel combination with the first computing unit as the starting computing unit and a second inter-unit request-feedback channel combination with the second computing unit as the starting computing unit.
5. The data processing apparatus according to claim 3, wherein, Each of the at least one inter-group request-feedback channel combinations includes a first request channel from the starting computing unit group to the pointing computing unit group and a first feedback channel from the pointing computing unit group to the starting computing unit group.
6. The data processing apparatus according to claim 5, wherein, The first request channels among the first computing units in the plurality of computing unit groups are connected in series to form a first request link. The first request channels between the various second computing units in the multiple computing unit groups are connected in series to form a second request link. The first feedback channels among the first computing units in the plurality of computing unit groups are connected in series to form a first feedback link. The first feedback channels of each second computing unit in the plurality of computing unit groups are connected in series to form a second feedback link.
7. The data processing apparatus according to claim 6, wherein, The first request link and the second request link are unidirectional links; The first feedback link and the second feedback link are unidirectional links.
8. The data processing apparatus according to claim 6 or 7, wherein, The data processing device includes a request network and a feedback network. The request network includes the second request channel within each of the plurality of computing unit groups, as well as the first request link and the second request link. The feedback network includes a second feedback channel within each of the plurality of computing unit groups, as well as the first feedback link and the second feedback link.
9. The data processing apparatus according to claim 5, wherein, Each of the plurality of computing units further includes multiple processing elements and command queues. The vector memory includes a vector memory body and vector memory nodes. The vector memory core is used to receive stimuli from the command queue and the plurality of processing elements and to interact with the vector memory nodes. The vector memory node is used to interact with the vector memory in the adjacent computing unit through the inter-unit request-feedback channel combination or the inter-group request-feedback channel combination.
10. A data processing method of the data processing apparatus according to claim 1, wherein, The plurality of computing unit groups includes a first computing unit group and a second computing unit group, and the data processing method includes: The vector memory of the source computing unit in the first computing unit group transmits a read request to the vector memory of the destination computing unit in the second computing unit group through the at least one first transmission channel; and The vector memory of the destination computing unit in the second computing unit group transmits the requested data to the vector memory of the source computing unit in the first computing unit group through the at least one first transmission channel.
11. A data processing method of the data processing apparatus according to claim 8, wherein, The plurality of computing unit groups includes a first computing unit group and a second computing unit group, and the data processing method includes: The vector memory of the source computing unit in the first computing unit group transmits a read request to the vector memory of the destination computing unit in the second computing unit group through the request network; The vector memory of the destination computing unit in the second computing unit group transmits the requested data to the vector memory of the source computing unit in the first computing unit group through the feedback network.
12. The data processing method according to claim 11, wherein, The vector memory includes vector memory nodes. The processing method further includes: The vector memory node interacts with the vector memory in adjacent computing units through the inter-unit request-feedback channel combination or the inter-group request-feedback channel combination.
13. The data processing method according to claim 11, wherein, The vector memory of the source computing unit in the first computing unit group transmits the read request to the vector memory of the destination computing unit in the second computing unit group through the request network, including: During the transmission process from the source computing unit in the first computing unit group to the destination computing unit in the second computing unit group, the transmission path of the read request is determined by the shortest path determination algorithm based on the identifier of the current computing unit and the identifier of the destination computing unit, wherein the computing units traversed in the transmission path are referred to as nodes in the transmission path.
14. The data processing method according to claim 13, wherein, Determining the transmission path of the read request using the shortest path determination algorithm includes: If the identifier of the current computing unit is equal to the identifier of the destination computing unit, the current node is used to process the read request. For the current node: Determine if there are sufficient processing resources to process the read request. If sufficient processing resources are available, the read requests are stored in a cache for processing operations in order of priority, and the previous node is notified to input the information corresponding to the read request. The read requests are then parsed and processed sequentially. In the absence of sufficient processing resources, read requests with the first priority are processed sequentially, while read requests with the second priority are notified to wait. The first priority is higher than the second priority.
15. The data processing method according to claim 14, wherein, Determining whether there are sufficient processing resources to process the read request includes: Determine the remaining resources of the current node in processing the read request and whether there are sufficient resources in the vector memory of the current node.
16. The data processing method according to claim 14, wherein, Determining the transmission path of the read request using the shortest path determination algorithm also includes: If the identifier of the current computing unit is not equal to the identifier of the destination computing unit, the current node is used to forward the read request. For the current node: Determine if there are sufficient forwarding resources to forward the read request. If sufficient forwarding resources are available, the read requests are stored in a cache for forwarding operations in priority order, and the previous node is notified to input the information corresponding to the read request. The read requests are then forwarded to the next node in sequence. In the absence of sufficient forwarding resources, read requests with the third priority are stored in the cache used for forwarding operations, and the node corresponding to the read request with the third priority is notified to output the read request. For read requests with the fourth priority, the node corresponding to the read request with the fourth priority is notified to wait, wherein the third priority is higher than the fourth priority.
17. The data processing method according to claim 16, wherein, Determining whether there are sufficient forwarding resources to forward the read request includes: Determine the resource status of the current node forwarding the read request and whether the next node can receive the read request.
18. The data processing method according to any one of claims 11 to 17, further comprising: During the transmission of the read request, adjacent nodes establish a connection through a handshake mechanism and transmit instructions or data through the inter-unit request-feedback channel combination or the inter-group request-feedback channel combination between the nodes.
19. The data processing method according to claim 18, wherein, The adjacent nodes include master nodes and slave nodes. The handshake mechanism includes: The master node prepares the data and sends a ready signal to enable it, and outputs the first data to the slave node; The slave node determines whether it can receive the first data. If it can receive the first data, the slave node will receive a ready signal to enable and receive the first data.
20. The data processing method according to claim 19, wherein, The handshake mechanism also includes: If the master node detects that the receive ready signal is enabled, it maintains the transmit ready signal enabled and outputs the second data following the first data to the slave node; or If the master node does not detect the receive ready signal enable, the master node remains unchanged.
21. An electronic device comprising: The data processing apparatus according to any one of claims 1 to 9.
Citation Information
Patent Citations
Multipath computer system
CN106095720A
Processing circuit and neural network operating method thereof
CN108470009A
Matrix convolution calculation method, interface, coprocessor and system based on RISC-V architecture
CN109857460A
Multiprocessor data processing system
CN1637734A