Data transmission method and device, equipment, storage medium and product
By adding routing units to the streaming processor cluster, the problem of large transmission delay between calculation units is solved, and the optimization and efficiency improvement of data transmission paths are achieved.
Patent Information
- Application Number
- CN202510461903.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-11
- Publication Date
- 2025-07-18
AI Technical Summary
In the existing streaming processor cluster architecture, the data transmission delay between computing units is large and needs to be transferred through unrelated computing units, resulting in low transmission efficiency.
A routing unit is added in the streaming processor cluster, and the routing unit receives transmission requests and directly or indirectly sends data to the target computing unit to avoid redirection of irrelevant computing units and optimizes the path length.
Significantly reduce data transmission delay, improve the overall transmission efficiency of streaming processor clusters, reduce the transmission pressure of routing units, and achieve efficient and stable data transmission.
Smart Images

Figure CN120336248A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of data transmission, and in particular, to a data transmission method, apparatus, device, storage medium, and product. Background Art
[0002] As Figure 1 is the existing Stream Processor Cluster (SPC) architecture, which includes 4 SPCs, and each SPC contains 8 Compute Units (CUs) and a Cluster Bus Interface (CBI). However, the existing SPC architecture has the following problems:
[0003] (1) Only data can be transmitted between CUs in the same SPC through CUs, and it needs to be relayed through irrelevant CUs in the middle, resulting in large transmission latency. For example, to transmit the data of CU0 in SPC0 to CU7 in SPC0, it needs to be relayed through 6 irrelevant CUs such as CU1, CU2,..., CU6 in the middle, significantly increasing the transmission latency.
[0004] (2) Data is transmitted between different SPCs through CBI, and it needs to be relayed through irrelevant CUs in the middle, resulting in large transmission latency. For example, to transmit the data of CU1 in SPC0 to CU1 in SPC1, the data needs to be transmitted to CU2, CU3,..., CU7 in SPC0 in sequence, then transmitted from CU7 in SPC0 to CBI1 in SPC1 through CBI0 in SPC0, and finally transmitted from CBI1 in SPC1 to CU1 in SPC1 through CU0 in SPC1, significantly increasing the transmission latency. Summary of the Invention
[0005] This application provides a data transmission method, apparatus, device, storage medium, and product to solve the problem of large transmission latency in the prior art.
[0006] To achieve the above object, an embodiment of this application provides a data transmission method, which is applied to a routing unit newly added in a stream processor cluster, where the stream processor cluster is located in a data processor, and the data transmission method includes:
[0007] Receiving a transmission request initiated by a first computing unit; wherein, the transmission request includes: data to be transmitted and a destination node message, and the destination node message is used to represent one or more second computing units to which the data to be transmitted needs to be transmitted; the first computing unit is located in the stream processor cluster or other stream processor clusters, and the second computing unit is located in the stream processor cluster or other stream processor clusters;
[0008] Send the data to be transmitted to one or more of the second computing units.
[0009] As an improvement to the above solution, the step of sending the data to be transmitted to one or more of the second computing units includes:
[0010] For each first candidate computing unit, send the data to be transmitted to the first candidate computing unit; the first candidate computing unit is the second computing unit located in the streaming processor cluster;
[0011] For each second candidate computing unit, send the data to be transmitted to the target routing unit, and the target routing unit sends the data to be transmitted to the second candidate computing unit; the second candidate computing unit is the second computing unit located in other streaming processor clusters, and the second candidate computing unit and the target routing unit are located in the same streaming processor cluster.
[0012] As an improvement to the above solution, the step of, for each first candidate computing unit, sending the data to be transmitted to the first candidate computing unit includes:
[0013] For each of the first candidate computing units, send the data to be transmitted to each of the first candidate computing units that meet the preset conditions, and the first candidate computing units that meet the preset conditions send the data to be transmitted to the first candidate computing units that do not meet the preset conditions.
[0014] As an improvement to the above solution, the transmission request includes a first transmission request initiated by the first computing unit and directly sent to the second computing unit;
[0015] The step of receiving the transmission request initiated by the first computing unit includes:
[0016] When the first computing unit and the second computing unit are located in the streaming processor cluster and the path length between the first computing unit and the second computing unit is greater than the preset length, receive the first transmission request;
[0017] When the first computing unit is located in the streaming processor cluster and the second computing unit is located in other streaming processor clusters, receive the first transmission request.
[0018] As an improvement to the above solution, the transmission request includes a second transmission request initiated by the first computing unit and sent through the routing unit in the streaming processor cluster where it is located to the second computing unit;
[0019] The step of receiving the transmission request initiated by the first computing unit includes:
[0020] In the case where the first computing unit is located in other streaming processor clusters, and the first computing unit and the second computing unit are located in different streaming processor clusters, receive the second transmission request.
[0021] As an improvement to the above solution, the transmission request further includes: a broadcast identifier; the destination node message is a mask, and each element of the mask is used to indicate whether the data to be transmitted needs to be sent to the second computing unit corresponding to the element.
[0022] To achieve the above object, an embodiment of the present application further provides a data transmission device, which is applied to a routing unit newly added in a streaming processor cluster, and the streaming processor cluster is located in a data processor. The data transmission device includes:
[0023] A receiving subunit, configured to receive a transmission request initiated by a first computing unit; wherein the transmission request includes: data to be transmitted and a destination node message, and the destination node message is used to indicate one or more second computing units to which the data to be transmitted needs to be transmitted; the first computing unit is located in the streaming processor cluster or other streaming processor clusters, and the second computing unit is located in the streaming processor cluster or other streaming processor clusters;
[0024] A sending subunit, configured to send the data to be transmitted to one or more of the second computing units.
[0025] To achieve the above object, an embodiment of the present application further provides a data transmission device, including a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor. When the processor executes the computer program, the data transmission method as described above is implemented.
[0026] To achieve the above object, an embodiment of the present application further provides a computer-readable storage medium, which includes a stored computer program; wherein, when the computer program runs, it controls the device where the computer-readable storage medium is located to execute the data transmission method as described above.
[0027] To achieve the above object, an embodiment of the present application further provides a computer program product, including a computer program / instructions, and when the computer program / instructions are executed by a processor, the data transmission method as described above is implemented.
[0028] Compared with the prior art, a data transmission method, apparatus, device, storage medium, and product provided by an embodiment of the present application add a hardware structure in advance in a streaming processor cluster: a routing unit. The routing unit receives a transmission request initiated by a first computing unit and sends the data to be transmitted in the transmission request to one or more second computing units, avoiding the need for irrelevant computing units to transfer in the middle, effectively shortening the path length between the transmission source node, i.e., the first computing unit, and the transmission destination node, i.e., the second computing unit, thereby significantly reducing the data transmission delay and improving the overall transmission efficiency of the streaming processor cluster. BRIEF DESCRIPTION OF THE DRAWINGS
[0029] Figure 1 is an architecture diagram of a prior art streaming processor cluster;
[0030] Figure 2 is a flowchart of a data transmission method provided by an embodiment of the present application;
[0031] Figure 3 is a partial diagram of an architecture of a streaming processor cluster provided by an embodiment of the present application;
[0032] Figure 4 is a block diagram of a structure of a data processor provided by an embodiment of the present application;
[0033] Figure 5 is a block diagram of a structure of a data transmission apparatus provided by an embodiment of the present application;
[0034] Figure 6 is a block diagram of a structure of a data transmission device provided by an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0035] The following will clearly and completely describe the technical solutions in the embodiments of the present application with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without creative efforts shall fall within the protection scope of the present application.
[0036] Refer to Figure 2 , Figure 2 which is a flowchart of a data transmission method provided by an embodiment of the present application. The data transmission method is applied to a routing unit added in advance in a streaming processor cluster. The streaming processor cluster is located in a data processor. The data transmission method includes:
[0037] S1. Receive a transmission request initiated by a first computing unit; wherein, the transmission request includes: data to be transmitted and a destination node message, and the destination node message is used to represent one or more second computing units to which the data to be transmitted needs to be transmitted; the first computing unit is located in the streaming processor cluster or another streaming processor cluster, and the second computing unit is located in the streaming processor cluster or another streaming processor cluster;
[0038] S2. Send the data to be transmitted to one or more of the second computing units.
[0039] It should be noted that the embodiments of the present application are applicable not only to the point-to-point data transmission scenario, but also to the broadcast scenario, and can transmit data to multiple second computing units simultaneously. Refer to Figure 3 , in the embodiments of the present application, a hardware structure: a routing unit is added to the SPC in advance. The routing unit is configured to be able to interact not only with each computing unit in the same SPC, but also with the routing units in other SPCs, so as to realize data transmission within the SPC and across SPCs, avoid the need for intermediate transfer through irrelevant computing units, effectively shorten the path length between the transmission source node, that is, the first computing unit, and the transmission destination node, that is, the second computing unit, thereby significantly reducing the data transmission delay and improving the overall transmission efficiency of the streaming processor cluster.
[0040] It should be noted that the first computing unit and the second computing unit may be located in the same streaming processor cluster of the same data processor, or in different streaming processor clusters of the same data processor, or in different streaming processor clusters of different data processors. A data processor includes one or more streaming processor clusters, such as Figure 4 , this data processor includes N SPCs. Among them, the data processor may be an artificial intelligence chip such as a Graphics Processing Unit (GPU) or a General-Purpose Graphics Processing Unit (GPGPU).
[0041] In an alternative embodiment, the sending the data to be transmitted to one or more of the second computing units includes:
[0042] For each first candidate computing unit, send the data to be transmitted to the first candidate computing unit; the first candidate computing unit is the second computing unit located in the streaming processor cluster;
[0043] For each second candidate computing unit, send the data to be transmitted to the target routing unit, and the target routing unit sends the data to be transmitted to the second candidate computing unit; the second candidate computing unit is the second computing unit located in other streaming processor clusters, and the second candidate computing unit and the target routing unit are located in the same streaming processor cluster.
[0044] In the embodiment of the present application, for the second computing unit located in the streaming processor cluster, that is, the first candidate computing unit, the data to be transmitted is directly and / or indirectly sent from the first computing unit to the first candidate computing unit through the routing unit in the streaming processor cluster.
[0045] For the second computing unit that is not located in the streaming processor cluster, that is, the second candidate computing unit, first, the data to be transmitted is sent from the first computing unit to the target routing unit through the routing unit in the streaming processor cluster, and then the target routing unit sends the data to be transmitted to the second candidate computing unit; the target routing unit is the routing unit in the streaming processor cluster where the second candidate computing unit is located.
[0046] In an alternative embodiment, the step of sending the data to be transmitted to each first candidate computing unit includes:
[0047] For each of the first candidate computing units, send the data to be transmitted to each of the first candidate computing units that meet the preset conditions, and the first candidate computing units that meet the preset conditions send the data to be transmitted to the first candidate computing units that do not meet the preset conditions.
[0048] It should be noted that when the data needs to be transmitted to multiple first candidate computing units simultaneously, the routing unit will face serious transmission pressure. In the embodiment of the present application, the data to be transmitted is first sent to specific computing units, that is, the first candidate computing units that meet the preset conditions, and then these specific computing units send the data to the remaining computing units, that is, the first candidate computing units that do not meet the preset conditions, which can greatly reduce the transmission pressure of the routing unit, improve the propagation speed and efficiency of the broadcast data, and in addition, by reasonably allocating the broadcast tasks, the balance of the data transmission load is achieved, avoiding overload of some nodes due to processing a large amount of broadcast data, and ensuring the stable operation of the entire streaming processor cluster.
[0049] The embodiments of the present application do not specifically limit the preset conditions. Optionally, if the first candidate computing units that meet the preset conditions are the first candidate computing units with even indices, then the first candidate computing units with even indices send the data to be transmitted to the adjacent first candidate computing units with odd indices; if the first candidate computing units that meet the preset conditions are the first candidate computing units with odd indices, then the first candidate computing units with odd indices send the data to be transmitted to the adjacent first candidate computing units with even indices.
[0050] For example, the first computing unit is CU1 of SPC1, and the first candidate computing units include: CU3 - CU7 of SPC1. The first candidate computing units with even indices, namely CU2, CU4, and CU6 of SPC1, are used as secondary forwarding nodes. Then, the routing unit of SPC1 first receives the data to be transmitted sent by CU1 of SPC1 and sends the data to be transmitted to CU2, CU4, and CU6 of SPC1. Then, CU2 of SPC1 sends it to the adjacent CU3, CU4 of SPC1 sends it to the adjacent CU5, and CU6 of SPC1 sends it to the adjacent CU7, finally realizing the transmission of the data of CU1 of SPC1 to CU3 - CU7 of SPC1.
[0051] This hierarchical cascading broadcast mechanism improves the performance of the streaming processor cluster in the following two aspects: First, it distributes the centralized broadcast load to multiple computing units, significantly reducing the instantaneous bandwidth pressure on the routing unit; second, by using the parallel transmission characteristics between computing units, the data reception delay of the first candidate computing units that do not meet the preset conditions only increases by a single-hop delay compared to the first candidate computing units that meet the preset conditions, thereby achieving a broadcast efficiency close to that of a fully connected architecture.
[0052] In an alternative embodiment, the transmission request includes a first transmission request initiated by the first computing unit and directly sent to the second computing unit that needs to be transmitted.
[0053] Receiving the transmission request initiated by the first computing unit includes:
[0054] When the first computing unit and the second computing unit are located in the streaming processor cluster and the path length between the first computing unit and the second computing unit is greater than the preset length, receiving the first transmission request;
[0055] When the first computing unit is located in the streaming processor cluster and the second computing unit is located in another streaming processor cluster, receiving the first transmission request.
[0056] In the embodiment of the present application, in order to reduce the transmission pressure of the routing unit, when the first computing unit determines that the first computing unit and the second computing unit are located in the streaming processor cluster, and the path length between the first computing unit and the second computing unit is greater than a preset length, or the first computing unit is located in the streaming processor cluster and the second computing unit is located in another streaming processor cluster, the first transmission request is directly sent to the routing unit of the streaming processor cluster, and the routing unit of the streaming processor cluster performs data forwarding, which can avoid all data being forwarded through the routing unit and significantly reduce the transmission pressure of the routing unit. Wherein, the first transmission request is a transmission request initiated by the first computing unit and directly sent to the routing unit.
[0057] For example, when CU0 of SPC1 sends data to CU3 of SPC1, if the data is transmitted along the direct connection line between the computing units, the path length will increase to 3; while if it is selected to be relayed through the routing unit, the path length can be shortened to 2, which can significantly reduce the delay of data transmission and improve the data transmission efficiency. When CU0 of SPC1 sends data to CU3 of SPC2, it is also relayed by the routing unit, significantly reducing the delay of data transmission and improving the data transmission efficiency.
[0058] Correspondingly, when the first computing unit determines that the first computing unit and the second computing unit are located in the streaming processor cluster and the path length between the first computing unit and the second computing unit is less than or equal to the preset length, the first computing unit directly sends the data to be transmitted to the second computing unit. In the embodiment of the present application, data is transmitted through the direct connection line between the computing units without additional relaying, avoiding the additional delay and complexity that the routing unit may bring.
[0059] Exemplarily, the preset length is 1 unit length. When CU1 of SPC1 needs to broadcast data to CU0, CU2 - CU7 of SPC1, the following operations will be performed:
[0060] Direct connection transmission: CU1 of SPC1 directly sends data to CU0 and CU2 of SPC1 with a path length of 1;
[0061] Routing relay: CU1 of SPC1 sends a transmission request to the routing unit of SPC1;
[0062] Routing processing: The routing unit of SPC1 only forwards data to the eligible even - numbered CUs (CU2, CU4, CU6 of SPC1), and then CU2, CU4, CU6 of SPC1 respectively pass the data to CU3, CU5, CU7 of SPC1.
[0063] Therefore, when a computing unit needs to transfer data to other computing units in the same streaming processor cluster, a decision is made based on the path length between the two computing units. When the path length is greater than the preset length, the data will be relayed through the routing unit of the same streaming processor cluster to improve the transmission efficiency and path optimization. Conversely, if the path length is less than or equal to the preset length, the data will be transmitted through the direct connection line between the computing units, avoiding the additional latency and complexity that the routing unit may bring.
[0064] The computing unit will split the transmission requests for different paths and send them simultaneously. Therefore, when the computing unit is concurrently processing multiple data streams, it must strictly maintain the transmission order at the single-CU granularity. This mechanism is achieved through the following constraints:
[0065] Transmission requests with a path length less than or equal to the preset length only activate the direct connection channels between the computing units;
[0066] Transmission requests with a path length greater than the preset length are forced to be redirected to the routing unit.
[0067] This constraint fundamentally eliminates the risk of data interleaving that may be caused by cross-path transmission, ensuring deterministic latency under a broadcast / unicast mixed load.
[0068] In an alternative embodiment, the transmission request includes a second transmission request initiated by the first computing unit and sent through the routing unit in the streaming processor cluster where it is located, and needs to be transmitted to the second computing unit;
[0069] Receiving the transmission request initiated by the first computing unit includes:
[0070] In the case where the first computing unit is located in another streaming processor cluster and the first computing unit and the second computing unit are located in different streaming processor clusters, receiving the second transmission request.
[0071] In order to reduce the transmission pressure on the routing unit in the embodiments of the present application, when the first computing unit determines that the first computing unit is located in another streaming processor cluster and the first computing unit and the second computing unit are located in different streaming processor clusters, the second transmission request is sent to the routing unit of the streaming processor cluster, and the routing unit of the streaming processor cluster forwards the data, which can avoid all data being forwarded through the routing unit and significantly reduce the transmission pressure on the routing unit. Among them, the second transmission request is a transmission request initiated by the first computing unit and indirectly sent to the routing unit.
[0072] Therefore, when a computing unit in a certain streaming processor cluster needs to transfer data to a computing unit in another streaming processor cluster, it can directly forward the data between the streaming processor clusters through the routing unit. For example, when CU0 of SPC0 needs to transfer data to CU1 of SPC1, the data is first sent from CU0 of SPC0 to the routing unit of SPC0, then forwarded from the routing unit of SPC0 to the routing unit of SPC1, and finally the routing unit of SPC1 transfers the data to CU1 of SPC1; when CU0 of SPC0 needs to transfer data to CU2 of SPC2, the data can be forwarded from the routing unit of SPC0 to the routing unit of SPC1, and then forwarded through the routing unit of SPC1 to the routing unit of SPC2, and finally the routing unit of SPC2 transfers the data to CU2 of SPC2. This data transfer method can greatly shorten the data latency and improve the data transfer efficiency.
[0073] In an alternative embodiment, the transmission request further includes: a broadcast identifier; the destination node message is a mask, and each element of the mask is used to indicate whether the data to be transmitted needs to be sent to the second computing unit corresponding to the element.
[0074] It should be noted that when the first computing unit initiates a broadcast, the first computing unit will create a transmission request, which includes the data to be transmitted, a broadcast identifier, and a mask.
[0075] The broadcast identifier is used to indicate that the data to be transmitted needs to be broadcast. When the routing unit detects this broadcast identifier, it means that the routing unit has the ability to broadcast the data to all the computing units marked in the mask.
[0076] Each element of the mask is used to indicate whether the data to be transmitted needs to be sent to the second computing unit corresponding to the element. That is to say, this mask is used to identify which CUs are the destination nodes for the broadcast data. Specifically, the element is a binary element. For example, if CU0 of SPC0 needs to broadcast data to CU1 to CU7 of SPC0, it will generate a 32-bit bit mask, and its hexadecimal representation is 0xFE. Each bit element in the mask corresponds to a CU, and a value of 1 indicates that the CU needs to receive the data, and this CU is the second computing unit. When a certain CU receives the broadcast data to be transmitted, it will reset the element corresponding to itself in the mask and decide whether to continue forwarding the data to other CUs according to the updated mask. This mechanism ensures that the broadcast data can be efficiently and accurately delivered to all the second computing units marked in the mask.
[0077] A data transmission method provided by an embodiment of the present application adds a hardware structure in advance in a streaming processor cluster: a routing unit. The routing unit receives a transmission request initiated by a first computing unit and sends the data to be transmitted in the transmission request to one or more second computing units, avoiding the need for irrelevant computing units to transfer in the middle, effectively shortening the path length between the transmission source node, that is, the first computing unit, and the transmission destination node, that is, the second computing unit, thereby significantly reducing the data transmission delay and improving the overall transmission efficiency of the system.
[0078] See Figure 5 , Figure 5 is a structural block diagram of a data transmission device 10 provided by an embodiment of the present application. The data transmission device 10 includes:
[0079] A receiving subunit 11, configured to receive a transmission request initiated by a first computing unit; wherein, the transmission request includes: data to be transmitted and a destination node message, and the destination node message is used to represent one or more second computing units to which the data to be transmitted needs to be transmitted; the first computing unit is located in the streaming processor cluster or other streaming processor clusters, and the second computing unit is located in the streaming processor cluster or other streaming processor clusters;
[0080] A sending subunit 12, configured to send the data to be transmitted to one or more of the second computing units.
[0081] Optionally, the sending subunit 12 specifically includes:
[0082] A first sending subunit, configured to send the data to be transmitted to each first candidate computing unit; the first candidate computing unit is the second computing unit located in the streaming processor cluster;
[0083] A second sending subunit, configured to send the data to be transmitted to a target routing unit for each second candidate computing unit, and the target routing unit sends the data to be transmitted to the second candidate computing unit; the second candidate computing unit is the second computing unit located in other streaming processor clusters, and the second candidate computing unit and the target routing unit are located in the same streaming processor cluster.
[0084] Optionally, the second sending subunit is specifically configured to:
[0085] For each of the first candidate computing units, send the data to be transmitted to each of the first candidate computing units that meet the preset conditions, and the first candidate computing units that meet the preset conditions send the data to be transmitted to the first candidate computing units that do not meet the preset conditions.
[0086] Optionally, the transmission request includes a first transmission request initiated by the first computing unit and directly sent, which needs to be transmitted to the second computing unit;
[0087] The receiving subunit 11 includes:
[0088] A first receiving subunit, configured to receive the first transmission request when the first computing unit and the second computing unit are located in the streaming processor cluster and the path length between the first computing unit and the second computing unit is greater than a preset length; and to receive the first transmission request when the first computing unit is located in the streaming processor cluster and the second computing unit is located in another streaming processor cluster.
[0089] Optionally, the transmission request includes a second transmission request initiated by the first computing unit and sent through a routing unit in the streaming processor cluster where it is located, which needs to be transmitted to the second computing unit;
[0090] The receiving subunit 11 includes:
[0091] A second receiving subunit, configured to receive the second transmission request when the first computing unit is located in another streaming processor cluster and the first computing unit and the second computing unit are located in different streaming processor clusters.
[0092] Optionally, the transmission request further includes: a broadcast identifier; the destination node message is a mask, and each element of the mask is used to indicate whether the data to be transmitted needs to be sent to the second computing unit corresponding to the element.
[0093] It should be noted that the working processes of the various modules in the data transmission device 10 according to the embodiments of the present application may refer to the working process of the data transmission method described in the above embodiments, and will not be elaborated here.
[0094] A data transmission device 10 provided by an embodiment of the present application, by pre - adding a routing unit in the streaming processor cluster, the routing unit receives a transmission request initiated by a first computing unit and sends the data to be transmitted in the transmission request to one or more second computing units, avoiding the need for irrelevant computing units to transfer in the middle, effectively shortening the path length between the transmission source node, i.e., the first computing unit, and the transmission destination node, i.e., the second computing unit, thereby significantly reducing the data transmission delay and improving the overall transmission efficiency of the system.
[0095] In addition, an embodiment of the present application further provides a computer - readable storage medium, where the computer - readable storage medium includes a stored computer program; wherein, when the computer program runs, it controls the device where the computer - readable storage medium is located to execute the data transmission method described in any of the above embodiments.
[0096] In addition, an embodiment of the present application further provides a computer program product, including a computer program / instructions. When the computer program / instructions are executed by a processor, the data transmission method described in any of the above embodiments is implemented.
[0097] See Figure 6 , Figure 6 FIG. is a structural block diagram of a data transmission device 20 provided by an embodiment of the present application. The data transmission device 20 includes: a processor 21, a memory 22, and a computer program stored in the memory 22 and executable on the processor 21. When the processor 21 executes the computer program, the steps in the above data transmission method embodiment are implemented. Alternatively, when the processor 21 executes the computer program, the functions of each module / unit in the above device embodiments are implemented.
[0098] Exemplarily, the computer program may be divided into one or more modules / units. The one or more modules / units are stored in the memory 22 and executed by the processor 21 to complete the present application. The one or more modules / units may be a series of computer program instruction segments capable of performing specific functions, and the instruction segments are used to describe the execution process of the computer program in the data transmission device 20.
[0099] The data transmission device 20 may include, but is not limited to, a processor 21 and a memory 22. Those skilled in the art can understand that the schematic diagram is only an example of the data transmission device 20, and does not constitute a limitation on the data transmission device 20. It may include more or fewer components than shown, or combine certain components, or different components. For example, the data transmission device 20 may further include an input / output device, a network access device, a bus, etc.
[0100] The processor 21 may be a central processing unit (CPU), or may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor, or the processor may also be any conventional processor, etc. The processor 21 is the control center of the data transmission device 20, and connects various parts of the entire data transmission device 20 through various interfaces and lines.
[0101] The memory 22 can be used to store the computer programs and / or modules. By running or executing the computer programs and / or modules stored in the memory 22, and by invoking the data stored in the memory 22, the processor 21 realizes various functions of the data transmission device 20. The memory 22 mainly includes a program storage area and a data storage area. Among them, the program storage area can store an operating system, application programs required for at least one function (such as a sound playback function, an image playback function, etc.); the data storage area can store data created according to the use of the mobile phone (such as audio data, phone book, etc.). In addition, the memory 22 can include a high-speed random access memory, and can also include a non-volatile memory, such as a hard disk, a memory, a plug-in hard disk, a Smart Media Card (SMC), a Secure Digital (SD) card, a Flash Card, at least one magnetic disk storage device, a flash memory device, or other volatile solid-state storage devices.
[0102] Among them, if the modules / units integrated in the data transmission device 20 are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on such an understanding, to implement all or part of the processes in the above method embodiments of the present application, it can also be completed by instructing relevant hardware through a computer program. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by the processor 21, the steps of the above method embodiments can be realized. Among them, the computer program includes computer program code, and the computer program code can be in the form of source code, object code, executable file or some intermediate form, etc. The computer-readable medium can include: any entity or device capable of carrying the computer program code, a recording medium, a USB flash drive, a mobile hard disk, a magnetic disk, an optical disc, a computer memory, a read-only memory (ROM), a random access memory (RAM), an electrical carrier signal, a telecommunication signal, and a software distribution medium, etc.
[0103] It should be noted that the device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. In addition, in the attached drawings of the device embodiments provided in this application, the connection relationship between the modules indicates that they have a communication connection, which can be specifically implemented as one or more communication buses or signal lines. Those of ordinary skill in the art can understand and implement it without creative efforts.
[0104] The above is the preferred implementation manner of this application. It should be noted that for those of ordinary skill in the art in this technical field, several improvements and refinements can be made without departing from the principle of this application, and these improvements and refinements are also regarded as the protection scope of this application.
Claims
1. A data transmission method, characterized in that, Applied to a pre-added routing unit in a streaming processor cluster, where the streaming processor cluster is located in a data processor, the data transmission method includes: Receiving a transmission request initiated by a first computing unit; wherein, the transmission request includes: data to be transmitted and a destination node message, and the destination node message is used to represent one or more second computing units to which the data to be transmitted needs to be transmitted; the first computing unit is located in the streaming processor cluster or another streaming processor cluster, and the second computing unit is located in the streaming processor cluster or another streaming processor cluster; Sending the data to be transmitted to one or more of the second computing units.
2. The data transmission method according to claim 1, wherein The sending the data to be transmitted to one or more of the second computing units includes: For each first candidate computing unit, sending the data to be transmitted to the first candidate computing unit; the first candidate computing unit is a second computing unit located in the streaming processor cluster; For each second candidate computing unit, sending the data to be transmitted to a target routing unit, and the target routing unit sends the data to be transmitted to the second candidate computing unit; the second candidate computing unit is a second computing unit located in another streaming processor cluster, and the second candidate computing unit and the target routing unit are located in the same streaming processor cluster.
3. The data transmission method according to claim 2, wherein The for each first candidate computing unit, sending the data to be transmitted to the first candidate computing unit includes: For each of the first candidate computing units, sending the data to be transmitted to each of the first candidate computing units that meet a preset condition, and the first candidate computing unit that meets the preset condition sends the data to be transmitted to the first candidate computing unit that does not meet the preset condition.
4. The data transmission method according to claim 1, wherein The transmission request includes a first transmission request initiated by the first computing unit and directly sent to the second computing unit; The receiving the transmission request initiated by the first computing unit includes: When the first computing unit and the second computing unit are located in the streaming processor cluster and the path length between the first computing unit and the second computing unit is greater than a preset length, receiving the first transmission request; When the first computing unit is located in the streaming processor cluster and the second computing unit is located in another streaming processor cluster, receiving the first transmission request.
5. The data transmission method according to claim 1, characterized in that The transmission request includes a second transmission request initiated by the first computing unit and sent through a routing unit in the streaming processor cluster where it is located to the second computing unit; The receiving the transmission request initiated by the first computing unit includes: When the first computing unit is located in another streaming processor cluster and the first computing unit and the second computing unit are located in different streaming processor clusters, receiving the second transmission request.
6. The data transmission method according to claim 1, wherein The transmission request further includes: a broadcast identifier; the destination node message is a mask, and each element of the mask is used to represent whether the data to be transmitted needs to be sent to the second computing unit corresponding to the element.
7. A data transmission device, characterized in that, Applied to a pre-added routing unit in a streaming processor cluster, the streaming processor cluster is located in a data processor, and the data transmission device includes: A receiving subunit, configured to receive a transmission request initiated by a first computing unit; wherein, the transmission request includes: data to be transmitted and a destination node message, and the destination node message is used to characterize one or more second computing units to which the data to be transmitted needs to be transmitted; the first computing unit is located in the streaming processor cluster or another streaming processor cluster, and the second computing unit is located in the streaming processor cluster or another streaming processor cluster; A sending subunit, configured to send the data to be transmitted to one or more of the second computing units.
8. A data transmission device, characterized in that, Comprising a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor, the processor implements the data transmission method according to any one of claims 1 to 6 when executing the computer program.
9. A computer-readable storage medium, characterized in that, The computer-readable storage medium includes a stored computer program; wherein, the computer program controls the device where the computer-readable storage medium is located to execute the data transmission method according to any one of claims 1 to 6 when running.
10. A computer program product, characterized in that, Comprising computer program / instructions, which implement the data transmission method according to any one of claims 1 to 6 when executed by a processor.