A computing cluster, in-network computing method and related device

CN122824735APending Publication Date: 2026-09-25海光信息技术(成都)有限公司
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610727419.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-05-25
Publication Date
2026-09-25

AI Technical Summary

Benefits of technology

[0017]可以看出,本申请实施例提供的计算集群,包括:至少一个发起方计算节点和多个对端计算节点;网络设备,所述发起方计算节点与对端计算节点通过所述网络设备互联,所述网络设备用于接收所述发起方计算节点发送的任务请求报文,所述任务请求报文至少包括任务操作指令和组播ID项,以及基于所述任务操作指令确定当前任务,如果当前任务为在网计算任务,基于所述组播ID项,将所述任务请求报文转发至对端计算节点,以获取各个对端计算节点返回的计算数据报文并执行所述在网计算任务。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122824735A_ABST
    Figure CN122824735A_ABST
Patent Text Reader

Abstract

The embodiment of the present application provides a kind of computing cluster, in-network computing method and related equipment, the computing cluster includes: at least one initiator computing node and multiple opposite end computing nodes;Network equipment, the initiator computing node is interconnected with opposite end computing node by the network equipment, the network equipment is used to receive the task request message sent by the initiator computing node, the task request message at least includes task operation instruction and multicast ID item, determines current task based on the task operation instruction, if current task is in-network computing task, based on the multicast ID item, the task request message is forwarded to opposite end computing node, to obtain the computing data message returned by each opposite end computing node and executes the in-network computing task.The computing cluster provided by the embodiment of the present application can effectively realize in-network computing while improving bandwidth utilization efficiency, simplifies model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer technology, specifically to a computing cluster, an online computing method, and related equipment. Background Technology

[0002] In-Network Computing (INC) is a technology that offloads some or all of the computing tasks to network devices (such as switches and network interface cards) to reduce communication latency, improve data transmission efficiency, and reduce the computing burden on terminal devices. It is of great significance for further promoting the improvement of computing efficiency.

[0003] Against this backdrop, how to provide a new type of on-network computing solution has become a technical problem that urgently needs to be solved by those skilled in the art. Summary of the Invention

[0004] In view of this, embodiments of this application provide a computing cluster, an on-network computing method, and related equipment, which can effectively realize on-network computing while improving bandwidth utilization efficiency and simplifying the model.

[0005] To achieve the above objectives, the embodiments of this application provide the following technical solutions.

[0006] In a first aspect, embodiments of this application provide a computing cluster, including: at least one initiating computing node and multiple peer computing nodes; a network device, wherein the initiating computing node and the peer computing nodes are interconnected through the network device, and the network device is used to receive a task request message sent by the initiating computing node, wherein the task request message includes at least a task operation instruction and a multicast ID item, and determines the current task based on the task operation instruction; if the current task is an on-network computing task, forwards the task request message to the peer computing nodes based on the multicast ID item, so as to obtain the computing data messages returned by each peer computing node and execute the on-network computing task.

[0007] Optionally, each compute node includes: At least one network transmission device is configured to transmit the task request message to the network device, or to receive a task request message forwarded by the network device.

[0008] Optionally, the network transmission device is configured with a multicast forwarding table, which includes a multicast ID field. The multicast ID field is used to indicate the multicast communication relationship between the initiating computing node and the peer computing node. The multicast ID field includes at least the multicast IP, the virtual address base address of the computing node's memory, and its size.

[0009] Optionally, each computing node further includes: The data structure unit is used to store data within each computing node and to forward the data, task operation instructions, and the multicast ID item to the network transmission device.

[0010] Optionally, the network device includes: Multiple ports are used to connect to the network transmission devices in each computing node, with one port corresponding to one computing node.

[0011] Optionally, the network device further includes: The multicast protocol processing module is used to parse the task operation instructions in the received task request message to obtain the current task content and transmit the task request message based on the current task. The computing unit is used to execute the on-network computing task after obtaining the computing data packets returned by each peer computing node.

[0012] Secondly, embodiments of this application provide an on-network computing method applied to the computing cluster described in the first aspect above, comprising: initiating a task request and transmitting a task request message to a network device, the task request message including at least a task operation instruction and a multicast ID item, the task operation instruction indicating the current task content, and the multicast ID item indicating the multicast communication relationship between the initiating computing node and the peer computing node; determining the current task content based on the task operation instruction; if the current task is an on-network computing task, forwarding the task request message to the peer computing node based on the multicast ID item to obtain computing data messages returned by each peer computing node and execute the on-network computing task.

[0013] Thirdly, embodiments of this application provide an on-network computing device applied to the computing cluster described in the first aspect above, comprising: a task request initiation module, configured to initiate a task request and transmit a task request message to a network device, the task request message including at least a task operation instruction and a multicast ID item, the task operation instruction indicating the current task content, and the multicast ID item indicating the multicast communication relationship between the initiating computing node and the peer computing node; and an on-network computing task execution module, configured to determine the current task content based on the task operation instruction, and if the current task is an on-network computing task, forward the task request message to the peer computing node based on the multicast ID item to obtain computing data messages returned by each peer computing node and execute the on-network computing task.

[0014] Fourthly, embodiments of this application provide an electronic device, characterized in that it includes at least one memory and at least one processor, wherein the memory stores one or more computer-executable instructions, and the processor invokes the one or more computer-executable instructions to execute the online computing method as described in the first aspect above.

[0015] Fifthly, embodiments of this application provide a storage medium that stores one or more computer-executable instructions, which, when executed, implement the online computing method as described in the first aspect above.

[0016] In a sixth aspect, embodiments of this application provide one or more computer-executable instructions, which, when executed, implement the online computing method as described in the first aspect above.

[0017] As can be seen, the computing cluster provided in this application embodiment includes: at least one initiating computing node and multiple peer computing nodes; a network device, wherein the initiating computing node and the peer computing nodes are interconnected through the network device, and the network device is used to receive a task request message sent by the initiating computing node. The task request message includes at least a task operation instruction and a multicast ID item, and determines the current task based on the task operation instruction. If the current task is an on-network computing task, the task request message is forwarded to the peer computing nodes based on the multicast ID item to obtain the computing data messages returned by each peer computing node and execute the on-network computing task.

[0018] The computing cluster provided in this application embodiment includes at least one initiating computing node and multiple peer computing nodes. The initiating computing node and peer computing nodes can be interconnected through network devices (e.g., switches). The network devices are used to receive task request messages sent by the initiating computing node. The task request message includes at least a task operation instruction and a multicast ID. The network devices can then determine the current task based on the task operation instruction. After determining that the current task is an on-network computing task, the network devices can forward the task request message to the peer computing nodes based on the multicast ID to obtain the computing data messages returned by each peer computing node and execute the on-network computing task (e.g., data aggregation). Thus, the computing cluster provided in this application embodiment can effectively realize on-network computing tasks. At the same time, compared with the programming model in the existing on-network computing scheme, the computing cluster model provided in this application embodiment for realizing on-network computing tasks is simple, reduces complexity, and improves bandwidth utilization efficiency. Attached Figure Description

[0019] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of this application. For those skilled in the art, other drawings can be obtained based on the provided drawings without creative effort.

[0020] Figure 1 This is an example diagram of an optional structure of a computing cluster provided in an embodiment of this application; Figure 2 This is an example diagram of another optional structure of the computing cluster provided in the embodiments of this application; Figure 3 This is a flowchart illustrating the online computing method provided in an embodiment of this application; Figure 4 This is an example diagram of an optional structure of the online computing device provided in the embodiments of this application. Detailed Implementation

[0021] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0022] In modern high-performance computing (HPC) and artificial intelligence (AI) training scenarios, large-scale distributed computing clusters typically rely on aggregate communication operations (such as AllReduce, Broadcast, and AllGather) to achieve data synchronization among multiple nodes. Traditional aggregate communication relies on software algorithms (such as Ring AllReduce and Tree AllReduce) to perform data aggregation on the host side (CPU / GPU), resulting in high communication latency and high CPU (Central Processing Unit) / GPU (Graphics Processing Unit) computing resource consumption.

[0023] In-Network Computing (INC) is an emerging computing paradigm that aims to reduce latency and improve overall computing efficiency by utilizing programmable network devices (such as switches and smart NICs) to perform computations directly during data transmission. Its core idea is to shift computing power from server accelerators to network devices to address the increased network speed and the performance of GPGPUs (General-purpose computing on graphics processing units). GPGPUs utilize graphics processors to perform general-purpose computing tasks that were originally handled by the central processing unit. Key technologies include programmable switching technology, smart NICs, software-defined networking (SDN), and protocol-independent switching architecture (PISA), which enable network devices to support flexible packet processing logic. Implementation methods encompass communication offloading, intra-network caching and analysis, and low-latency computing, making it suitable for high-performance computing, artificial intelligence, and edge computing scenarios. In the future, INC will evolve towards a convergence of programmability and flexibility, standardization and industrialization, further driving improvements in computing efficiency.

[0024] One on-network computing solution is the SHARP (Scalable Hierarchical Aggregation and Reduction Protocol) scheme based on switch integration. Specifically, it performs data aggregation and reduction operations at the switching layer through scale-up topology. Scale-up refers to improving performance by adding hardware resources to individual servers or devices, thereby significantly improving communication efficiency in multi-GPU systems. This technology allows these operations to be processed directly within the switch, rather than transmitting large amounts of intermediate data between GPUs, thus reducing overall data transmission volume, lowering latency, and alleviating the burden on GPUs. This solution is beneficial for applications requiring large-scale parallel computing and frequent inter-node communication, such as deep learning training and complex scientific simulations, providing key support for improving overall system performance and scalability.

[0025] However, while the aforementioned switch-based SHARP technology has shown significant effectiveness in improving the efficiency of aggregated communication between multi-GPU systems and reducing latency, it also has some limitations and drawbacks, such as high cost, software ecosystem dependence, and closed-source protocol.

[0026] Another on-network computing solution is the SHARP scheme using InfiniBand (IB) network interface cards (NICs). IB NICs are network adapters based on the InfiniBand protocol, primarily used for data transmission between different nodes. This solution significantly optimizes communication efficiency in distributed computing (such as AI training) through in-network computation. Its core principle is to offload aggregate communication operations (such as gradient aggregation in AllReduce) from the terminal device to the dedicated hardware engine of the InfiniBand switch, directly performing data reduction (such as summation and maximization) at the network layer, avoiding multiple data transmissions between nodes. Specifically, each node only needs to send data to the switch once. The switch performs aggregation through its built-in Arithmetic Logic Unit (ALU) and broadcasts the result back to all nodes, simplifying traditional multi-hop communication into a single-hop operation, reducing bandwidth requirements by more than 50%, and reducing latency to sub-microsecond levels. This technology is particularly well-suited for large-scale GPU clusters, supporting hierarchical topologies and multi-task concurrency, making it a key acceleration solution for high-performance computing and AI training.

[0027] However, the SHARP scheme for the aforementioned IB network cards has the following drawbacks: First, IB network cards are a scale-out technology that requires dedicated hardware support, such as InfiniBand network cards and switches. These devices are expensive, and the deployment cost is 5-10 times that of traditional Ethernet (TCP / IP). Even if an Ethernet-based RoCE (RDMA over Converged Ethernet) scheme is adopted, network cards and switches that support RDMA (Remote Direct Memory Access) are still required, resulting in a large overall investment. Second, the IB network card programming model is relatively complex. Finally, IB network cards use proprietary protocols and are closed-source.

[0028] In view of this, embodiments of this application provide a computing cluster that can effectively achieve on-network computing while improving bandwidth utilization efficiency and simplifying the model.

[0029] Figure 1 This is an example diagram of an optional structure of a computing cluster provided in an embodiment of this application, such as... Figure 1As shown, the computing cluster includes: at least one initiating computing node and multiple peer computing nodes; a network device, wherein the initiating computing node and the peer computing nodes are interconnected through the network device, and the network device is used to receive task request messages sent by the initiating computing node. The task request message includes at least a task operation instruction and a multicast ID field, and determines the current task based on the task operation instruction. If the current task is an on-network computing task, the task request message is forwarded to the peer computing nodes based on the multicast ID field to obtain the computing data messages returned by each peer computing node and execute the on-network computing task.

[0030] In this embodiment of the application, the computing cluster consists of a large number of computing nodes and network devices connecting these computing nodes, which together complete large-scale computing tasks. Among them, the computing nodes are responsible for executing the main and complex computing tasks and are the fundamental source of computing power. The network devices not only transmit data but also complete part of the computing tasks during the transmission process, thereby reducing latency and improving overall computing efficiency.

[0031] In an optional embodiment, the computing node may include any of the following acceleration devices: GPGPU (General-purpose computing on graphics processing units), TPU (Tensor Processing Unit), APU (Accelerated Processing Unit), DSP (Digital Signal Processor), or FPGA (Field-Programmable Gate Array).

[0032] In this embodiment of the application, multiple computing nodes are interconnected through the network device. The multiple computing nodes may include an initiating computing node as the task initiator and a peer computing node that receives the task request. A task request initiated by an initiating computing node can be transmitted to multiple other computing nodes, i.e., multiple peer computing nodes, through the network device.

[0033] Reference Figure 1In this embodiment of the application, four computing nodes are used as an example, and the computing nodes are GPGPUs. The four computing nodes may include GPGPU0, GPGPU1, GPGPU2, and GPGPU3. GPGPU0, as the initiator of the task request, needs to send the task request to other computing nodes, including GPGPU1, GPGPU2, and GPGPU3, through the network device. GPGPU1 to GPGPU3 are the peer computing nodes.

[0034] Continue to refer to Figure 1 In an optional implementation, each computing node includes at least one network transmission device, such as OE (Over Ethernet). OE is a hardware device capable of network transmission, which can be used to transmit task request messages to the network device or receive task request messages forwarded by the network device.

[0035] It should be noted that, for the sake of simplicity, the diagram only shows the internal composition of computing nodes GPGPU0 and GPGPU1, as well as the communication relationship between the two computing nodes. Computing nodes GPGPU2 and GPGPU3 are the same as GPGPU1 and are not shown in the diagram.

[0036] In an optional embodiment, the task request message can be a UDP (User Datagram Protocol) format message. UDP messages have the advantages of being simple and fast. Therefore, within the computing node, the data, instructions, content, etc. that the computing node needs to transmit to the network device can be quickly encapsulated into UDP message format and sent to the network device.

[0037] In other embodiments, the task request message may also be any other message format.

[0038] Continue to refer to Figure 1 In this embodiment of the application, the network transmission device is further configured with a multicast forwarding table, which includes a multicast ID field, which is used to indicate the multicast communication relationship between the initiating computing node and the peer computing node.

[0039] In this embodiment of the application, a multicast task refers to a collaborative computing task in which the same data or instruction is simultaneously distributed to multiple computing nodes in the computing cluster. In on-network computing, it is necessary to establish multicast communication relationships between various computing nodes. Based on the multicast communication relationships, data or instructions can be distributed to various computing nodes to collect data and complete on-network computing.

[0040] Reference Figure 1In the specific implementation, before performing on-network computing, the computing nodes (including GPGPU0, GPGPU1, GPGPU2, and GPGPU3) can agree on the starting address and size of the device memory corresponding to different multicasts through software. In the figure, computing nodes GPGPU0 and GPGPU1 are marked as S0, indicating that computing nodes GPGPU0 and GPGPU1 agree on the starting address and size of the device memory corresponding to different multicasts through software. GPGPU2 and GPGPU3 are the same, but are not marked in the figure. Among them, the starting address of the memory is different, but the size can be the same.

[0041] In an optional implementation, the memory can be any storage medium, such as HBM (High Bandwidth Memory) or GDDR (Graphics Double Data Rate).

[0042] Furthermore, construct a multicast forwarding table, such as Figure 1 As shown, the multicast forwarding table includes a multicast ID field (MCid). Different multicasts correspond to different multicast ID fields, such as MCid1, MCid2, MCid3, MCid4, ..., MCidn.

[0043] In an optional implementation, the multicast ID item includes at least the multicast IP, the virtual address base address (VA) of the compute node's memory, and its size.

[0044] In an optional implementation, the multicast ID item may also include a process ID (HPI). HPI (Hardware process ID) refers to the hardware's ability to index the corresponding page table base address based on the HPI, and can be called a process ID or process number.

[0045] like Figure 1 As shown in the multicast forwarding table, a multicast ID can be represented as "MCid: multicast ip HPI\VA\size", where the multicast ip is the IP address corresponding to the multicast. In on-network computing, task request messages can be sent to network devices based on the multicast ip address.

[0046] VA (Virtual Address) represents the base address of the memory of a compute node. The base address of the memory of a compute node is mapped as VA->MCID+VA. Specifically, the memory mapping address of each compute node can be determined based on the multicast ID item and the virtual address base address of the memory of the peer compute node. The memory mapping address is the virtual address base address of the memory of each compute node.

[0047] HPI is used by subsequent peer computing nodes to provide a basis for converting virtual addresses to physical addresses (PAs) after receiving task request messages.

[0048] In this embodiment of the application, the multicast forwarding table may include, but is not limited to, fields such as multicast IP, HPI, and VA, which can be added according to the functional requirements.

[0049] In the optional implementation, HPI, VA, and PA have no bit width limit and can be any bit width.

[0050] Continue to refer to Figure 1 Each computing node may further include a data fabric unit for storing data within each computing node and for forwarding the data, task operation instructions, and the multicast ID item to the network transmission device.

[0051] Reference Figure 1 In this embodiment of the application, the initiating computing node further includes a kernel, which can initiate task requests. The GPGPU hardware executes the instructions compiled by the kernel to implement the task requests.

[0052] In a specific implementation, the initiating computing node initiates a computing task through the kernel and instructions. First, it converts the virtual address (VA) into a physical address so that the hardware data structure unit can discover the relevant fields of the multicast ID item and then forward the data, task operation instructions (OP), and multicast ID item to the network transmission device (OE).

[0053] In this embodiment of the application, the task operation instruction OP is used to indicate the current task content. The task operation instruction needs to be sent to the network device, which will parse and execute the current task. The current task may include multicast tasks and on-network computing tasks.

[0054] After receiving the kernel compilation instructions, the network transmission device encapsulates the task operation instructions, multicast ID items, and data into a UDP network message, namely the task request message, and then transmits the task request message to the network device (as shown in S11 in the figure) based on the multicast IP.

[0055] In this embodiment of the application, the network device can be a switch, which can be any programmable network device. The switch needs to have programmability so that it can recognize the format of the received task request message.

[0056] The network device can be used to receive task request messages sent by the initiating computing node. The task request message includes at least a task operation instruction (OP) and a multicast ID.

[0057] In an optional implementation, the network device includes multiple ports for connecting to network transmission devices in each computing node, wherein one port corresponds to one computing node.

[0058] like Figure 1 As shown, the network device includes multiple ports (port 0, port 1, port 2, and port 3), with each port corresponding to a computing node.

[0059] After receiving the task request message, the network device can return an ACK instruction to the initiating computing node GPGPU0.

[0060] In this embodiment of the application, the network device may further include: a multicast protocol processing module, used to parse the task operation instructions in the received task request message to obtain the current task content and transmit the task request message based on the current task.

[0061] In an optional implementation, if the current task operation instruction is parsed and it is found that the current task is an on-network computing task, the task request message can be forwarded to the peer computing node (not shown in the figure) based on the multicast ID item, in order to wait for the computing data messages returned by each peer computing node.

[0062] In a specific implementation, when the network transmission device OE of the peer computing node (e.g., GPGPU1) receives a task request message, it can obtain the process ID (HPI) and the virtual address base address VA of the memory of the initiating computing node by parsing the multicast ID item and task operation instructions in the task request message.

[0063] In this embodiment of the application, the task request message may also include an address offset (OFFSET). During data access, the address offset is used to locate the specific address or address range of the data for access.

[0064] In a specific implementation, the corresponding page table can be determined based on the process ID, the target virtual address can be calculated based on the virtual address base address and the address offset, and the target virtual address can be converted into a physical address PA through an address translation unit; then data access can be realized based on the physical address.

[0065] In this embodiment, the current task is an on-network computing task, which may include at least a data aggregation task. This application takes a data aggregation task as an example. Assume that the current task is to read the data in each computing node and perform a summation calculation. For example, the data read from GPGPU1, GPGPU2, and GPGPU3 are 2, 3, and 4, respectively. Each peer computing node will re-encapsulate the data in its node into a UDP packet (i.e., the computing data packet) and transmit it to the network device (as shown in S12 in the figure). After receiving the UDP packet, the network device returns an ACK instruction to the peer computing node GPGPU1.

[0066] In an optional implementation, the network device may further include a computing unit, configured to execute the on-network computing task after receiving the computing data packets returned by each peer computing node.

[0067] Specifically, the network device can perform computing tasks through the computing unit. For example, it can sum the data returned by each peer computing node with the data transmitted in computing node GPGPU0, such as 1. After the calculation is completed, the network computing result is obtained. The network computing result can be returned to the initiator (e.g., GPGPU0) or notified to each computing node by broadcasting.

[0068] Furthermore, in this embodiment of the application, when the network device receives the task request message from the initiator, it parses the task operation instruction, and the current task may also be a multicast task.

[0069] Figure 2 This is an example diagram of another optional structure of the computing cluster provided in the embodiments of this application; wherein, Figure 1 and Figure 2 The architecture of the computing cluster is the same, and the similarities will not be repeated here. The difference between this embodiment and the previous embodiments is that the task operation instruction indicates that the current task is a multicast task.

[0070] Specifically, refer to Figure 2 The network device parses the current task operation instruction and obtains that the current task is a multicast task. Then, based on the multicast ID, it forwards the task request message to the peer computing node (marked as S13 in the figure). After the peer computing node completes the data access, it returns an ACK instruction to the network device. In other words, if the current task is a multicast task, it is only necessary to transmit the data of the initiator to each peer computing node. There is no need to perform computing tasks or modify the data content in the meantime.

[0071] As can be seen, the computing cluster provided in this application embodiment includes at least one initiating computing node and multiple peer computing nodes. The initiating computing node and peer computing nodes can be interconnected through network devices (e.g., switches). The network device is used to receive task request messages sent by the initiating computing node. The task request message includes at least a task operation instruction and a multicast ID item. The network device can then determine the current task based on the task operation instruction. After determining that the current task is an on-network computing task, it can forward the task request message to the peer computing nodes based on the multicast ID item to obtain the computing data messages returned by each peer computing node and execute the on-network computing task (e.g., data aggregation). Thus, the computing cluster provided in this application embodiment can effectively realize on-network computing tasks. At the same time, compared with the programming model in the existing on-network computing scheme, the computing cluster model provided in this application embodiment for realizing on-network computing tasks is simple, reduces complexity, and improves bandwidth utilization efficiency.

[0072] In a further optional implementation, based on the computing cluster provided in the embodiments of this application, the embodiments of this application also provide an on-network computing method. The on-network computing method described below can be referred to in correspondence with the computing cluster content described above. In an optional implementation, Figure 3 This is a flowchart illustrating the online computing method provided in the embodiments of this application, with reference to... Figure 3 The on-network computing method is applied to the computing cluster described in the foregoing embodiments, and the method may include the following steps.

[0073] Step S301: Initiate a task request and transmit the task request message to the network device. The task request message includes at least a task operation instruction and a multicast ID item. The task operation instruction is used to indicate the current task content, and the multicast ID item is used to indicate the multicast communication relationship between the initiating computing node and the peer computing node.

[0074] In this embodiment of the application, before initiating a task request, the method may further include: determining the starting address and address size of the memory corresponding to each computing node through software.

[0075] Combination Figure 1 As shown, this embodiment of the application takes four computing nodes, and the computing nodes are GPGPUs as an example. The four computing nodes may include GPGPU0, GPGPU1, GPGPU2, and GPGPU3. GPGPU0, as the initiator of the task request, needs to send the task request to other computing nodes, including GPGPU1, GPGPU2, and GPGPU3, through the network device. GPGPU1 to GPGPU3 are the peer computing nodes.

[0076] It should be noted that, for the sake of simplicity, the diagram only shows the communication relationship between the two computing nodes GPGPU0 and GPGPU1. The computing nodes GPGPU2 and GPGPU3 are the same as GPGPU1 and are not shown in the diagram.

[0077] As shown in Figure S0, compute node GPGPU0 and compute node GPGPU1 can agree on the starting address and size of the device memory corresponding to different multicasts through software. GPGPU2 and GPGPU3 are the same, but are not shown in the figure. Among them, the starting address of the memory is different, but the size can be the same.

[0078] In an optional implementation, the memory can be any storage medium, such as HBM (High Bandwidth Memory) or GDDR (Graphics Double Data Rate).

[0079] In addition, a multicast forwarding table is constructed, which includes a multicast ID field, which includes at least the multicast IP, the virtual address base address (VA) of the compute node's memory, and the size.

[0080] In the process of constructing the multicast forwarding table, the method for determining the virtual address base address of the memory of each computing node may include: determining the memory mapping address of each computing node based on the multicast ID item and the virtual address base address of the memory of the peer computing node, wherein the memory mapping address is the virtual address base address of the memory of each computing node.

[0081] like Figure 1 As shown, the multicast forwarding table includes a multicast ID field (MCid). Different multicasts correspond to different multicast ID fields, such as MCid1, MCid2, MCid3, MCid4, ..., MCidn.

[0082] In an optional implementation, the multicast ID item may also include a process ID (HPI). HPI (Hardware process ID) refers to the hardware's ability to index the corresponding page table base address based on the HPI, and can be called a process ID or process number.

[0083] like Figure 1 As shown in the multicast forwarding table, a multicast ID can be represented as "MCid: multicast ip HPI\VA\size", where the multicast ip is the multicast IP, i.e., the IP address corresponding to the multicast. In on-network computing, task request messages can be sent to network devices based on the multicast IP address.

[0084] VA (Virtual Address) represents the base address of the memory of a compute node. The base address of the memory of a compute node is mapped as VA->MCID+VA. Specifically, the memory mapping address of each compute node can be determined based on the multicast ID item and the virtual address base address of the memory of the peer compute node. The memory mapping address is the virtual address base address of the memory of each compute node.

[0085] HPI is used by subsequent peer computing nodes to provide a basis for converting virtual addresses to physical addresses (PAs) after receiving task request messages.

[0086] In this embodiment of the application, the multicast forwarding table may include, but is not limited to, fields such as multicast IP, HPI, and VA, which can be added according to the functional requirements.

[0087] In the optional implementation, HPI, VA, and PA have no bit width limit and can be any bit width.

[0088] Furthermore, the initiating computing node can initiate a task request through the kernel, and the GPGPU hardware executes the instructions compiled by the kernel to implement the task request.

[0089] In a specific implementation, the initiating computing node initiates a task request through the kernel and instructions. It first converts the virtual address (VA) into a physical address so that the hardware data structure unit within the computing node can discover the relevant fields of the multicast ID item and then forward the data, task operation instructions (OP), and multicast ID item to the network transmission device (OE).

[0090] In this embodiment of the application, the task operation instruction OP is used to indicate the current task content. The task operation instruction needs to be sent to the network device, which will parse and execute the current task. The current task is at least a network computing task.

[0091] After receiving the kernel compilation instructions, the network transmission device can encapsulate the task operation instructions, multicast ID items, and data into a UDP network message, namely the task request message, and then transmit the task request message to the network device based on the multicast IP.

[0092] In an optional implementation, the task request message may further include: an address offset OFFSET; during data access, it is necessary to locate the specific address or address range of the data based on the address offset.

[0093] Step S302: Based on the task operation instruction, determine the current task content. If the current task is an on-network computing task, based on the multicast ID item, forward the task request message to the peer computing node to obtain the computing data messages returned by each peer computing node and execute the on-network computing task.

[0094] In an optional implementation, determining the current task content based on the task operation instruction includes: receiving and parsing the operation instruction OP to obtain the current task content, wherein the current task content indicates the current task to be executed, and the current task includes at least a network computing task.

[0095] After the network device receives the task request message transmitted by the computing node initiator, it can parse the task operation instructions in the task request message through a reliable multicast protocol to confirm the current task content, which includes at least a network computing task.

[0096] If the current task is an on-network computing task, the task request message can be forwarded to the peer computing node based on the multicast ID item to obtain the computing data messages returned by each peer computing node and execute the on-network computing task.

[0097] Furthermore, after forwarding the task request message to the peer computing node, the method further includes: parsing the multicast ID item MCid, address offset OFFSET and operation instruction OP in the task request message to obtain the process ID (HPI) and the virtual address base address VA of the memory of the initiating computing node; Furthermore, based on the process ID, the corresponding page table is determined, the target virtual address can be calculated according to the virtual address base address and address offset, the target virtual address is converted into a physical address PA using the page table, and data access is performed based on the physical address.

[0098] In this embodiment, the current task is an on-network computing task, which may include at least a data aggregation task. This application takes a data aggregation task as an example. Assume that the current task is to read the data in each computing node and perform a summation calculation. For example, the data read from GPGPU1, GPGPU2, and GPGPU3 are 2, 3, and 4, respectively. Each peer computing node will re-encapsulate the data in its node into a UDP packet and transmit it to the network device (as shown in S12 in the figure). After receiving the UDP packet, the network device returns an ACK instruction to the peer computing node.

[0099] Furthermore, the online computing method may also include: after completing the online computing task, returning the online computing result to the initiating computing node, or notifying the initiating computing node of the online computing result via broadcast.

[0100] Specifically, the network device can perform computing tasks through the computing unit. For example, it can sum the data returned by each peer computing node with the data transmitted in computing node GPGPU0, such as 1. After the calculation is completed, the network computing result is obtained. The network computing result can be returned to the initiator (e.g., GPGPU0) or notified to each computing node by broadcasting.

[0101] Furthermore, in this embodiment of the application, the current task may also include a multicast task; Accordingly, the method may further include: if the current task is a multicast task, forwarding the task request message to the peer computing node based on the multicast ID item, in order to wait for the confirmation instructions returned by each peer computing node.

[0102] In other words, if the current task is a multicast task, it is only necessary to transmit the data of the initiating computing node to each peer computing node, without performing computing tasks or modifying the data content in the middle.

[0103] In a further optional implementation, based on the online computing method provided in the embodiments of this application, the embodiments of this application also provide an online computing device. The online computing device described below can be referred to in correspondence with the online computing method applied to computational centralization described above. In an optional implementation, Figure 4 This is an example diagram of an optional structure of the online computing device provided in an embodiment of this application. (Refer to...) Figure 4 The on-network computing device is applied to the computing cluster as described in the foregoing embodiments, and the device may include: The task request initiation module 401 is used to initiate a task request and transmit a task request message to the network device. The task request message includes at least a task operation instruction and a multicast ID item. The task operation instruction is used to indicate the current task content, and the multicast ID item is used to indicate the multicast communication relationship between the initiating computing node and the peer computing node. The on-net computing task execution module 402 is used to determine the current task content based on the task operation instruction. If the current task is an on-net computing task, it forwards the task request message to the peer computing node based on the multicast ID item to obtain the computing data messages returned by each peer computing node and execute the on-net computing task.

[0104] In an optional implementation, the on-network computing device further includes: The online computing result return module is used to return the online computing result to the initiating computing node after the online computing task is completed, or to notify the initiating computing node of the online computing result via broadcast.

[0105] In an optional implementation, the on-network computing task execution module 402 is used to determine the current task content based on the task operation instructions, including: The operation instruction is received and parsed to obtain the current task content, which indicates the current task to be executed, and the current task includes at least a network computing task.

[0106] In an optional implementation, the on-network computing device further includes: The preprocessing module is used to determine the starting address and size of the memory corresponding to each computing node by software before initiating a task request; Construct a multicast forwarding table, which includes a multicast ID field. The multicast ID field includes at least the multicast IP address, the virtual address base address of the computing node's memory, and its size.

[0107] In an optional implementation, the preprocessing module is used to determine the virtual address base address of the memory of each computing node during the construction of the multicast forwarding table, including: Based on the multicast ID and the virtual address base address of the memory of the peer computing node, the memory mapping address of each computing node is determined, wherein the memory mapping address is the virtual address base address of the memory of each computing node.

[0108] In an optional implementation, the task request message further includes: an address offset; the multicast ID field further includes: the process ID of the current task; The on-network computing task execution module 402 is further configured to, after forwarding the task request message to the peer computing node, parse the multicast ID item, address offset and operation instructions in the task request message to obtain the process ID and the virtual address base address of the computing node's memory. The corresponding page table is determined based on the process ID, the target virtual address is calculated based on the virtual address base address and the address offset, and the target virtual address is converted into a physical address using the page table; Data access is performed based on the physical address.

[0109] In an optional implementation, the current task may also include a multicast task; The online computing device further includes: The multicast task execution module is used to forward the task request message to the peer computing node based on the multicast ID item if the current task is a multicast task, in order to wait for the confirmation instructions returned by each peer computing node.

[0110] In a further optional implementation, embodiments of this application also provide an electronic device, including at least one memory and at least one processor, wherein the memory stores one or more computer-executable instructions, and the processor invokes the one or more computer-executable instructions to execute the on-network computing method applied to a computing cluster as described in the foregoing embodiments.

[0111] In a further optional implementation, this application embodiment also provides a storage medium that stores one or more computer-executable instructions. When the one or more computer-executable instructions are executed, the on-network computing method applied to a computing cluster as described in the foregoing embodiments is implemented.

[0112] In a further optional implementation, this application embodiment also provides a computer program product, including one or more computer-executable instructions, which, when executed, implement the on-network computing method applied to a computing cluster as described in the foregoing embodiments.

[0113] The foregoing describes multiple embodiment schemes provided by the embodiments of this application. The optional methods described in each embodiment scheme can be combined and cross-referenced with each other without conflict, thereby extending to a variety of possible embodiment schemes. These can all be considered as the embodiment schemes disclosed and published by the embodiments of this application.

[0114] While the embodiments disclosed above are described in this application, this application is not limited thereto. Any person skilled in the art can make various modifications and alterations without departing from the spirit and scope of this application; therefore, the scope of protection of this application should be determined by the scope defined in the claims.

Claims

1. A computing cluster, characterized in that, include: At least one initiating computing node and multiple peer computing nodes; A network device is used to interconnect the initiating computing node and the peer computing node. The network device is used to receive a task request message sent by the initiating computing node. The task request message includes at least a task operation instruction and a multicast ID field. Based on the task operation instruction, the current task is determined. If the current task is an on-network computing task, the task request message is forwarded to the peer computing node based on the multicast ID field to obtain the computing data messages returned by each peer computing node and execute the on-network computing task.

2. The computing cluster according to claim 1, characterized in that, Each computing node includes: At least one network transmission device is configured to transmit the task request message to the network device, or to receive a task request message forwarded by the network device.

3. The computing cluster according to claim 2, characterized in that, The network transmission device is configured with a multicast forwarding table, which includes a multicast ID field. The multicast ID field is used to indicate the multicast communication relationship between the initiating computing node and the peer computing node. The multicast ID field includes at least the multicast IP, the virtual address base address of the computing node's memory, and its size.

4. The computing cluster according to claim 2, characterized in that, Each computing node also includes: The data structure unit is used to store data within each computing node and to forward the data, task operation instructions, and the multicast ID item to the network transmission device.

5. The computing cluster according to claim 2, characterized in that, The network device includes: Multiple ports are used to connect to the network transmission devices in each computing node, with one port corresponding to one computing node.

6. The computing cluster according to claim 5, characterized in that, The network device also includes: The multicast protocol processing module is used to parse the task operation instructions in the received task request message to obtain the current task content and transmit the task request message based on the current task. The computing unit is used to execute the on-network computing task after obtaining the computing data packets returned by each peer computing node.

7. An online computing method, characterized in that, Applied to the computing cluster as described in any one of claims 1-6, comprising: Initiate a task request and transmit the task request message to the network device. The task request message includes at least a task operation instruction and a multicast ID item. The task operation instruction is used to indicate the current task content, and the multicast ID item is used to indicate the multicast communication relationship between the initiating computing node and the peer computing node. Based on the task operation instructions, the current task content is determined. If the current task is an on-network computing task, the task request message is forwarded to the peer computing node based on the multicast ID item to obtain the computing data messages returned by each peer computing node and execute the on-network computing task.

8. The on-network computing method according to claim 7, characterized in that, Also includes: After completing the online computing task, the online computing results will be returned to the initiating computing node, or the online computing results will be broadcast to the initiating computing node.

9. The on-network computing method according to claim 7, characterized in that, The process of determining the current task content based on the task operation instructions includes: The operation instruction is received and parsed to obtain the current task content, which indicates the current task to be executed, and the current task includes at least a network computing task.

10. The on-network computing method according to claim 9, characterized in that, Before initiating a task request, the method further includes: The starting address and size of the memory corresponding to each computing node are determined by software. Construct a multicast forwarding table, which includes a multicast ID field. The multicast ID field includes at least the multicast IP address, the virtual address base address of the computing node's memory, and its size.

11. The on-network computing method according to claim 10, characterized in that, The method for determining the virtual address base address of each computing node's memory during the construction of the multicast forwarding table includes: Based on the multicast ID and the virtual address base address of the memory of the peer computing node, the memory mapping address of each computing node is determined, wherein the memory mapping address is the virtual address base address of the memory of each computing node.

12. The on-network computing method according to claim 11, characterized in that, The task request message also includes: address offset; the multicast ID field also includes: the process ID of the current task; After forwarding the task request message to the peer computing node, the method further includes: The multicast ID, address offset, and operation instructions in the task request message are parsed to obtain the process ID and the virtual address base address of the memory of the initiating computing node; The corresponding page table is determined based on the process ID, the target virtual address is calculated based on the virtual address base address and the address offset, and the target virtual address is converted into a physical address using the page table; Data access is performed based on the physical address.

13. The on-network computing method according to claim 9, characterized in that, The current task also includes multicast tasks; The method further includes: If the current task is a multicast task, the task request message is forwarded to the peer computing node based on the multicast ID, and the node waits for confirmation instructions from each peer computing node.

14. An online computing device, characterized in that, Applied to the computing cluster as described in any one of claims 1-6, comprising: The task request initiation module is used to initiate a task request and transmit the task request message to the network device. The task request message includes at least a task operation instruction and a multicast ID item. The task operation instruction is used to indicate the current task content, and the multicast ID item is used to indicate the multicast communication relationship between the initiating computing node and the peer computing node. The on-net computing task execution module is used to determine the current task content based on the task operation instruction. If the current task is an on-net computing task, it forwards the task request message to the peer computing node based on the multicast ID item to obtain the computing data messages returned by each peer computing node and execute the on-net computing task.

15. An electronic device, characterized in that, It includes at least one memory and at least one processor, the memory storing one or more computer-executable instructions, and the processor invoking the one or more computer-executable instructions to perform the online computing method as described in any one of claims 7-13.

16. A storage medium, characterized in that, The storage medium stores one or more computer-executable instructions, which, when executed, implement the on-network computing method as described in any one of claims 7-13.

17. A computer program product comprising one or more computer-executable instructions, wherein when the one or more computer-executable instructions are executed, they implement the online computing method as described in any one of claims 7-13.