Online computing method and device, electronic equipment, storage medium and product
By distinguishing between non-full modes and full modes on-network computing methods, reducing network transmission of loading requests and calculation results, solving network congestion problems, realizing the saving of network traffic and ensuring computing accuracy.
Patent Information
- Application Number
- CN202510387626.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-28
- Publication Date
- 2025-07-18
AI Technical Summary
There is a problem of network congestion in network computing, which leads to an increase in time delay of network equipment, especially in large-scale distributed training and data transmission.
A method of on-network calculation is proposed. By distinguishing between non-full mode and full mode, the loading request in non-full mode is not multicast to the initiating node. The computing node further calculates based on its own data and network equipment calculation results to reduce network traffic; all nodes in the full mode participate in the calculation to ensure accuracy.
Effectively slow down network congestion, save network traffic, and automatically select appropriate modes based on accuracy requirements and traffic conditions to achieve a good balance between accuracy and flow control.
Smart Images

Figure CN120342955A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer technology, and more specifically, to a method, device, electronic device, storage medium, and product for in-network computing. Background Art
[0002] In-Network Computing (INC) is a new computing mode that offloads computing tasks to network devices (such as switches or routers, etc.) for execution. Its core idea is to utilize the computing power of network devices to process data during data transmission, thereby reducing the amount of data transmission, lowering communication overhead, and enhancing overall computing efficiency.
[0003] In-Network Computing has broad application prospects in many fields such as artificial intelligence (AI), the Internet of Things, edge computing, and 5G / 6G communications. For example, in AI training, In-Network Computing can significantly reduce the communication overhead between nodes, thereby accelerating the model training process.
[0004] Currently, network congestion problems are prevalent in In-Network Computing. How to effectively alleviate or overcome network congestion is one of the focuses of concern in the industry. Summary of the Invention
[0005] The present invention provides a method, device, electronic device, storage medium, and product for in-network computing, which helps to alleviate or overcome network congestion.
[0006] The technical solutions of the embodiments of the present invention are as follows:
[0007] An in-network computing method, which is applicable to a network device, and the method includes:
[0008] Receiving a loading request from a computing node, where the loading request includes identification information of a computing node group, and the computing node group includes the computing node;
[0009] When it is determined that the device is operating in a non-full mode, based on the identification information, multicasting the loading request to the remaining computing nodes in the computing node group except the computing node;
[0010] Receiving data associated with the loading request from the remaining computing nodes;
[0011] Performing a predetermined calculation on the data to obtain a first calculation result;
[0012] Sending the first calculation result to the computing node, so that the computing node performs the calculation on the first calculation result and the data in the computing node associated with the loading request to obtain a second calculation result.
[0013] In one embodiment, after obtaining the second calculation result, the method includes:
[0014] Receiving a first multicast request from the computing node, where the first multicast request includes the identification information and the second calculation result;
[0015] Based on the identification information, multicasting the second calculation result to the remaining computing nodes in the computing node group except the computing node.
[0016] In one embodiment, the method includes:
[0017] When it is determined to work in the full mode, based on the identification information, multicasting the loading request to all the computing nodes in the computing node group;
[0018] Receiving data associated with the loading request from each of the all computing nodes;
[0019] Performing the calculation on the data received from each of the all computing nodes and associated with the loading request to obtain a third calculation result;
[0020] Sending the third calculation result to the computing node.
[0021] In one embodiment, after obtaining the third calculation result, the method includes:
[0022] Receiving a second multicast request from the computing node, where the second multicast request includes the identification information and the third calculation result;
[0023] Based on the identification information, multicasting the third calculation result to all the computing nodes in the computing node group.
[0024] In one embodiment, it includes:
[0025] Generating a mode selection signal based on a predetermined performance metric;
[0026] Based on the mode selection signal, determining that the network device works in the non-full mode or the full mode.
[0027] In one embodiment, the generating a mode selection signal based on a predetermined performance metric includes:
[0028] Monitoring the traffic metric of the network device;
[0029] When the traffic index is greater than a predetermined first threshold, generate a mode selection signal for indicating that the network device operates in the non-full mode; when the traffic index is less than or equal to the predetermined first threshold, generate a mode selection signal for indicating that the network device operates in the full mode.
[0030] In one embodiment, generating the mode selection signal based on a predetermined performance index includes:
[0031] Obtain the accuracy requirement index of the calculation;
[0032] When the accuracy requirement index is greater than a predetermined second threshold, generate a mode selection signal for indicating that the network device operates in the full mode; when the accuracy requirement index is less than or equal to the predetermined second threshold, generate a mode selection signal for indicating that the network device operates in the non-full mode.
[0033] An in-network computing device, the device being applicable to a network device, the device includes:
[0034] A first receiving module, configured to receive a loading request from a computing node, the loading request including identification information of a computing node group, and the computing node group includes the computing node;
[0035] A multicast module, configured to, when determining to operate in the non-full mode, based on the identification information, multicast the loading request to the remaining computing nodes in the computing node group except the computing node;
[0036] A second receiving module, configured to receive data associated with the loading request from the remaining computing nodes;
[0037] An in-network computing module, configured to perform a predetermined calculation on the data to obtain a first calculation result;
[0038] A sending module, configured to send the first calculation result to the computing node, so that the computing node performs the calculation on the first calculation result and the data associated with the loading request in the computing node to obtain a second calculation result.
[0039] In one embodiment, the multicast module is configured to, when determining to operate in the full mode, based on the identification information, multicast the loading request to all computing nodes in the computing node group;
[0040] The second receiving module is configured to receive data associated with the loading request from each of the all computing nodes;
[0041] The calculation module is used to perform the calculation on the data associated with the load request received from all the computing nodes respectively, so as to obtain a third calculation result;
[0042] The sending module is used to send the third calculation result to the computing node.
[0043] An electronic device, comprising:
[0044] Memory;
[0045] processor;
[0046] The memory stores an application program executable by the processor, which is used to enable the processor to execute any of the above-mentioned on-line computing methods.
[0047] A computer-readable storage medium stores computer-readable instructions, which, when executed by a processor, cause the processor to execute any of the above-described on-line computing methods.
[0048] A program product comprises a computer program, wherein when the computer program is executed by a processor, the computer program implements any of the above-mentioned on-line computing methods.
[0049] It can be seen from the above technical solution that a load request is received from a computing node, the load request includes identification information of a computing node group, and the computing node group includes computing nodes; when it is determined to work in a non-full mode, based on the identification information, the load request is multicast to the remaining computing nodes in the computing node group except the computing node; data associated with the load request is received from the remaining computing nodes; a predetermined calculation is performed on the data to obtain a first calculation result; the first calculation result is sent to the computing node, so that the computing node performs the calculation on the first calculation result and the data associated with the load request in the computing node to obtain a second calculation result. It can be seen from this that the network device does not need to send the load request to the computing node that initiates the load request, and also does not need to receive data associated with the load request from the computing node that initiates the load request, thereby saving network traffic and effectively alleviating network congestion. In addition, the second calculation result calculated by the computing node that initiates the load request is multicast to the remaining computing nodes, and does not need to be sent to the computing node that initiates the load request, further saving network traffic.
[0050] In addition, the embodiments of the present invention can also automatically select a suitable mode based on accuracy requirements or flow conditions, and can achieve a good balance between accuracy requirements and flow control. BRIEF DESCRIPTION OF THE DRAWINGS
[0051] Figure 1 It is an exemplary schematic diagram of the first process of full protocol operation based on network computing in the related art.
[0052] Figure 2 It is a schematic diagram of an exemplary second process of full reduction operation based on in-network computing in the related art.
[0053] Figure 3 It is a schematic flowchart of an in-network computing method according to an embodiment of the present invention.
[0054] Figure 4 It is a schematic structural diagram of an in-network computing system according to an embodiment of the present invention.
[0055] Figure 5 It is a schematic diagram of an exemplary first process of full reduction operation based on in-network computing according to an embodiment of the present invention.
[0056] Figure 6 It is a schematic diagram of an exemplary second process of full reduction operation based on in-network computing according to an embodiment of the present invention.
[0057] Figure 7 It is a schematic structural diagram of an in-network computing device according to an embodiment of the present invention.
[0058] Figure 8 It is a schematic structural diagram of an electronic device according to an embodiment of the present invention. Detailed Embodiments
[0059] To make the objectives, technical solutions, and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings.
[0060] For the sake of brevity and intuitiveness in description, the solutions of the present invention will be elaborated below by describing several representative embodiments. A large number of details in the embodiments are only used to help understand the solutions of the present invention. However, it is obvious that the technical solutions of the present invention can be implemented without being limited to these details. To avoid unnecessarily obscuring the solutions of the present invention, some embodiments are not described in detail but only the frameworks are given. Hereinafter, "including" means "including but not limited to", and "according to..." means "at least according to..., but not limited to only according to...". Due to the language habits of Chinese, when the quantity of a component is not specifically pointed out hereinafter, it means that the component can be one or more, or can be understood as at least one.
[0061] By processing data directly on network devices, on-line computing can improve the real-time performance of computing. However, there is a traffic management problem in on-line computing. That is, while executing computing tasks on network devices, how to efficiently manage and optimize data traffic in the network to avoid congestion, reduce latency, and improve overall network performance. In current on-line computing, network congestion is often encountered due to excessive traffic, resulting in increased time delays in network devices. For example, network congestion often occurs during distributed training and large-scale data transmission (especially in data parallelism and pipeline parallelism scenarios).
[0062] After research, it is found that there are often redundant communication processes in network computing, which will cause or aggravate network congestion. If the communication process can be optimized to minimize the traffic, network congestion can be alleviated or overcome.
[0063] The following uses the use of on-network computing in the related art to perform full protocol operations as an example to illustrate the network congestion problem in the related art.
[0064] The core functions of the full-reduce operation mainly include: (1) Data aggregation: Each computing node (such as GPU, CPU or server) independently calculates the local gradient or parameter of its data slice; then, these local data are aggregated (that is, reduced, such as summing or averaging, etc.). (2) Data broadcast: The aggregated global data is multicast to all computing nodes to ensure that each computing node obtains the same update result. Latency is one of the key indicators to measure the performance of the full-reduce, especially in large-scale distributed training, where low latency is crucial to improving the overall training efficiency. Latency is the total time required to complete a full-reduce operation, including the time overhead of steps such as data transmission, reduction calculation and broadcast.
[0065] Assume that the computing node is implemented as 4 GPUs (eg, GPU0-GPU3), the network device is implemented as a switch, and GPU0-GPU3 respectively contain 4 block data (eg, block data 0-block data 3). The full reduction operation of the related art may include a first process and a second process.
[0066] In the first process: GPU0~GPU3 respectively send their own protocol requests to the network device; the network device multicasts the protocol requests sent by each GPU to GPU0~GPU3 (including the GPU that initiates the protocol request); GPU0~GPU3 sends their own data to be negotiated associated with the protocol request to the network device; the network device performs protocol processing based on the data to be negotiated sent by each GPU, and sends the protocol result to the GPU that initiated the protocol request.
[0067] In the second process: GPU0 to GPU3 respectively send the reduction results received from the network device to the network device, so that the network device multicasts the reduction results sent by GPU0 to GPU3 to each GPU (including the GPU that initiated the reduction request).
[0068] Figure 1 It is a schematic diagram of the first process of the full reduction operation based on in-network computing in the related art. For example, taking GPU0 as an example, Figure 1 The first process shown includes the following steps:
[0069] Step S1: GPU0 sends a reduction request to the switch.
[0070] Step S2: The switch multicasts the load request sent by GPU0 to GPU0 to GPU3.
[0071] Step S3: GPU0 to GPU3 respectively send the data to be reduced (for example, their respective block data 0) associated with the reduction request sent by GPU0 to the switch.
[0072] Step S4: The switch performs reduction processing on the data to be reduced received from GPU0 to GPU3 (that is, the 4 block data 0 received from GPU0 to GPU3 respectively), and returns the reduction result to GPU0.
[0073] GPU1 to GPU3 respectively synchronously execute the above steps similar to GPU0. For example, taking GPU1 as an example, the first process includes: GPU1 sends a reduction request to the switch; the switch multicasts the reduction request sent by GPU1 to GPU0 to GPU3; GPU0 to GPU3 respectively send the data to be reduced (for example, their respective block data 1) associated with the reduction request sent by GPU1 to the switch; the switch performs reduction processing on the data to be reduced received from GPU0 to GPU3 (that is, the 4 block data 1 received from GPU0 to GPU3 respectively), and returns the reduction result to GPU1.
[0074] Similarly, GPU2 to GPU3 respectively synchronously execute the above steps to complete the first process.
[0075] After completing the first process, the second process is executed. In the second process, GPU0 to GPU3 respectively send the reduction results received from the switch to the switch, so that the switch multicasts the reduction results sent by GPU0 to GPU3 to GPU0 to GPU3.
[0076] Figure 2 It is a schematic diagram of the second process of the full reduction operation based on in-network computing in the related art. For example, taking GPU0 as an example, the second process includes the following steps:
[0077] Step S5: GPU0 sends the protocol result received from the switch in step S4 (ie, the protocol result of block data 0 in GPU0, block data 0 of GPU1, block data 0 of GPU2 and block data 0 of GPU3) to the switch.
[0078] Step S6: The switch multicasts the reduction result sent by GPU0 to GPU0 to GPU3. Therefore, GPU0 to GPU3 can obtain the reduction results of block data 0 in GPU0, block data 0 in GPU1, block data 0 in GPU2 and block data 0 in GPU3.
[0079] In the second process, GPU1~GPU3 respectively synchronously execute the above steps similar to GPU0. For example, taking GPU1 as an example, the second process includes: GPU1 sends the protocol result received from the switch (that is, the protocol result of block data 1 in GPU0, block data 1 of GPU1, block data 1 of GPU2 and block data 1 of GPU3) to the switch; the switch multicasts the protocol result sent by GPU1 to GPU0~GPU3. Therefore, GPU0~GPU3 can also obtain the protocol results of block data 1 in GPU0, block data 1 of GPU1, block data 1 of GPU2 and block data 1 of GPU3 respectively. Similarly, GPU2~GPU3 respectively synchronously execute the above steps, so that GPU0~GPU3 can also obtain the protocol results of block data 2 in GPU0, block data 2 of GPU1, block data 2 of GPU2 and block data 2 of GPU3, as well as the protocol results of block data 3 in GPU0, block data 3 of GPU1, block data 3 of GPU2 and block data 3 of GPU3. At this point, the second process is completed.
[0080] It can be seen that in the first process of the related technology: the protocol request of each GPU is multicast to all GPUs (including the GPU that initiated the protocol request), and the switch needs to receive data from all GPUs (including the GPU that initiated the protocol request). In the second process of the related technology: the switch multicasts the protocol result of each GPU to all GPUs (including the GPU that initiated the protocol request). It can be seen that to complete a full protocol operation, the network device transmits a lot of traffic, which is easy to cause network congestion, and thus has the disadvantage of large delay.
[0081] The above takes the full specification operation of specific calculations in in-network computing as an example to illustrate the scenarios that cause network congestion in related technologies. In fact, specific calculations in in-network computing can also be implemented as data aggregation and caching, content distribution and computing reuse, stateful forwarding and multicast communication, dynamic task allocation and scheduling, machine learning inference, real-time data processing and analysis, distributed computing, load balancing, edge computing and cloud-native applications, and so on. Similarly, in the above in-network computing in related technologies, there are also similar scenarios that cause network congestion due to sending a loading request to the computing node that initiated the loading request or returning the calculation result to the computing node that initiated the loading request.
[0082] In an embodiment of the present invention, a new working mode (referred to as the non-full mode) that can mitigate network congestion is proposed. In the non-full mode: the loading request is not multicast to the computing node that initiated the loading request, and there is no need to obtain the data associated with the loading request from the computing node that initiated the loading request. The network device calculates a first calculation result based on the data associated with the loading request of the remaining computing nodes in the computing node group (that is, the computing nodes in the computing node group other than the computing node that initiated the loading request), and sends the first calculation result to the computing node that initiated the loading request. The computing node that initiated the loading request calculates a second calculation result based on the data associated with the loading request and the first calculation result saved by itself. The network device multicasts the second calculation result to the remaining computing nodes in the computing node group (without multicasting to the computing node that initiated the loading request). It can be seen that in the non-full mode, the network device does not need to send the loading request to the computing node that initiated the loading request, the computing node that initiated the loading request does not need to send the data associated with the loading request in the computing node that initiated the loading request to the network device, and the network device does not need to multicast the second calculation result to the computing node that initiated the loading request, thus saving network traffic and effectively mitigating network congestion.
[0083] In an embodiment of the present invention, a working mode (referred to as the full mode) that ensures accuracy requirements is also proposed. In the full mode: the loading request sent by the computing node is multicast by the network device to all computing nodes in the computing node group, and all computing nodes will provide the network device with their respective data associated with the loading request, so the result of in-network computing has good accuracy.
[0084] Moreover, in an embodiment of the present invention, an appropriate mode can also be automatically selected based on accuracy requirements or traffic conditions to achieve a good balance between accuracy requirements and traffic control.
[0085] The above disclosure details the technical defects existing in the related art, the causes of the technical defects, and the thinking and analysis process of overcoming the technical defects. In fact, the cognition of the above technical defects is not common knowledge in the field, but a novel discovery made by the inventor in his research. In addition, the cause tracing of the technical defects and the thinking and analysis process of overcoming the technical defects are also the gradual analysis results of the inventor in the actual research process, and are not common knowledge in the field.
[0086] Figure 3 The figure is an exemplary flow chart of an on-line computing method according to an embodiment of the present invention. Figure 3 The method shown can be performed by a processor in a network device. The processor can be implemented as any one of a central processing unit (CPU), a graphics processing unit (GPU), a tensor processing unit (TPU), a neural network processor (NPU), a deep learning processor (DPU), an accelerated processing unit (APU), and a general-purpose graphics processing unit (GPGPU).
[0087] Network devices play different roles in the network architecture and can be used to implement functions such as data transmission, network connection, flow control, security protection, etc. For example, network devices can be implemented as routers, switches, network cards, hubs, repeaters, bridges, gateways, firewalls, access points, etc.
[0088] like Figure 3 As shown, the method includes:
[0089] Step 101: receiving a loading request from a computing node, the loading request including identification information of a computing node group, and the computing node group includes computing nodes.
[0090] Here, the computing node group is a combination of computing nodes that need to perform computing tasks. The computing node group includes the computing node that initiates the load request. For example, each computing node in the computing node group can be implemented as a CPU, GPU, APU or GPGPU, etc.
[0091] Step 102: When it is determined to work in a non-full mode, based on the identification information, multicast the load request to the remaining computing nodes in the computing node group except the computing node.
[0092] Step 103: Receive data associated with the loading request from the remaining computing nodes.
[0093] Step 104: Perform a predetermined calculation on the data to obtain a first calculation result.
[0094] For example, the predetermined calculation can be implemented as: reduction calculation (such as Reduce, AllReduce, etc.), data aggregation, secure calculation, cache and load balancing calculation, machine learning and deep learning calculation, complex event processing calculation, network function virtualization (NFV) calculation, industrial control calculation, optimization calculation for content distribution and streaming media, and so on.
[0095] Step 105: Send the first calculation result to the computing node so that the computing node performs the calculation on the first calculation result and the data in the computing node associated with the loading request to obtain a second calculation result.
[0096] Here, the computing node performs the same calculation as in Step 104 on the data associated with the loading request saved by itself and the first calculation result received from the network device to obtain a second calculation result. For example, assuming that the calculation in Step 104 is implemented as a reduction calculation, after the computing node obtains the first calculation result (that is, the reduction calculation result of the data of the remaining computing nodes), it obtains the data associated with the loading request from itself, and then performs a reduction calculation on the data associated with the loading request and the first calculation result to obtain the second calculation result as the final reduction result.
[0097] In the embodiment of the present invention, the working mode of the network device is divided into a non-full mode and a full mode.
[0098] In the full mode: Each computing node will receive the loading requests sent by all computing nodes from the network device; the network device executes the complete calculation process to obtain the calculation results of each loading request; each computing node will receive the calculation results of the loading requests sent by itself.
[0099] In the non-full mode: each computing node will receive a load request from the remaining computing nodes other than itself from the network device; the network device will perform a partial calculation process to obtain a first calculation result for each load request; each computing node will complete the remaining calculation process based on the first calculation result and the data associated with the load request stored by itself to obtain a second calculation result. Specifically, in the non-full mode: the load request issued by the computing node is multicast by the network device to the remaining computing nodes in the computing node group other than the computing node that initiates the load request, and the network device then performs calculations to obtain the first calculation result. Since the network device does not need to send the load request to the computing node that initiates the load request, and the computing node that initiates the load request does not need to send the data associated with the load request in the computing node that initiates the load request to the network device, network traffic is saved and network congestion can be effectively alleviated.
[0100] In one embodiment, after obtaining the second calculation result, the method includes: receiving a first multicast request from a computing node, the first multicast request including identification information and the second calculation result; based on the identification information, multicasting the second calculation result to the remaining computing nodes in the computing node group except the computing node.
[0101] It can be seen that in the non-full mode, the second calculation result calculated by the computing node that initiates the load request based on the first calculation result is multicast to the remaining computing nodes in the computing node group except the computing node (no need to be multicast to the computing node that initiates the load request), which can further save network traffic and further effectively alleviate network congestion.
[0102] In one embodiment, the method includes: when it is determined to work in full mode, based on identification information, multicasting a load request to all computing nodes in a computing node group; receiving data associated with the load request from all computing nodes respectively; performing predetermined calculations on the data associated with the load request received from all computing nodes respectively to obtain a third calculation result; and sending the third calculation result to the computing node.
[0103] Specifically, in the full mode: the load request issued by the computing node is multicasted by the network device to all computing nodes in the computing node group, the network device receives data associated with the load request from all computing nodes respectively, and performs on-network calculation on the data to obtain the third calculation result. It can be seen that in the full mode, the load request issued by the computing node is multicasted by the network device to all computing nodes in the computing node group, so all computing nodes will provide their respective data associated with the load request to the network device, so that the third calculation result has good accuracy.
[0104] In one embodiment, after obtaining the third calculation result, the method includes: receiving a second multicast request from a computing node, the second multicast request including identification information and the third calculation result; and multicasting the third calculation result to all computing nodes in the computing node group based on the identification information.
[0105] Therefore, in the full mode, the third calculation result is multicast to all computing nodes in the computing node group (including the computing node that initiated the load request), and all computing nodes in the computing node group can obtain high-precision calculation results.
[0106] The full mode has the advantage of high in-network computing accuracy, and the non-full mode has the advantage of effectively alleviating network congestion. In the embodiments of the present invention, an appropriate mode can be automatically selected based on the accuracy requirement or traffic condition, and a good balance can be achieved between the accuracy requirement and traffic control.
[0107] In one embodiment, the method includes: generating a mode selection signal based on a predetermined performance metric; and determining that the network device operates in the non-full mode or the full mode based on the mode selection signal.
[0108] In one embodiment, generating a mode selection signal based on a predetermined performance metric includes: monitoring a traffic metric of the network device; when the traffic metric is greater than a predetermined first threshold, generating a mode selection signal for indicating that the network device operates in the non-full mode; and when the traffic metric is less than or equal to the predetermined first threshold, generating a mode selection signal for indicating that the network device operates in the full mode.
[0109] The traffic metrics of network devices are important parameters for measuring network performance and traffic status. For example, the traffic metrics of network devices can include: (1) Bandwidth utilization: Bandwidth utilization refers to the ratio of the actually used bandwidth to the total bandwidth, which is used to evaluate the usage of network resources. Excessive bandwidth utilization may lead to network congestion; (2) Throughput: Throughput represents the amount of data successfully transmitted per unit time, usually measured in bits per second (bps). It reflects the actual transmission capacity of network devices; (3) Network latency: Network latency refers to the time required for a data packet to travel from the sender to the receiver, which is an important indicator for measuring network response speed. High latency may affect the user experience of real-time applications. (4) Jitter: Jitter refers to the variation in transmission delay between data packets. High jitter may cause data packets to be out of order; (5) Packet loss rate: The packet loss rate refers to the proportion of lost data packets in the total number of transmitted data packets during data transmission. An excessively high packet loss rate will result in data retransmission and increase network load; (6) Session count: The session count refers to the number of active connections in the network within a certain period of time. Excessive sessions may cause the network device to be overloaded; (7) Flow count: The flow count refers to the total number of flows exported by the network device at a given time point, which is used to evaluate the overall scale of network traffic; (8) Total traffic: Total traffic refers to the traffic passing through the network device within one second, usually expressed in bits per second or percentage; (9) Traffic size: Traffic size refers to the total amount of data passing through the network device over a period of time, usually measured in megabytes (MB); (10) Round-trip time (RTT): Round-trip time refers to the total time for a data packet to travel from the sender to the receiver and then back to the sender, which is used to measure the two-way latency of the network; (11) Utilization: Utilization refers to the actual usage rate of the network device, usually expressed as a percentage. High utilization may indicate that the network is approaching saturation; (12) Back-to-back: Back-to-back refers to the number of data packets that the network device can fully forward at the maximum rate, which reflects the network device's ability to handle bursty data; (13) Maximum new connection rate: The maximum new connection rate refers to the maximum number of connections that the network device can establish per unit time, which is used to evaluate the network device's ability to handle new connections.
[0110] For example, the above metrics can be monitored and analyzed in real time through network monitoring tools (such as Wireshark, NetFlow, PRTG, etc.).
[0111] The above exemplary description presents typical instances of the traffic metrics of network devices. Those skilled in the art can realize that such a description is merely exemplary and is not used to limit the protection scope of the embodiments of the present invention.
[0112] Depending on influencing factors such as the nature of the task, resource limitations, and performance goals, different in-network computing tasks usually have different levels of precision requirements. For example, for application scenarios with high computational precision requirements (such as high-performance computing (HPC), artificial intelligence training (AI Training), industrial automation and real-time control, and fault diagnosis in data center networks, etc.), the full mode can be selected to ensure the precision requirements. For application scenarios with low computational precision requirements (such as lightweight inference in edge computing, network traffic monitoring and analysis, or energy-driven computing, etc.), the non-full mode can be selected to alleviate network congestion.
[0113] In one embodiment, generating a mode selection signal based on a predetermined performance metric includes: obtaining the precision requirement metric of in-network computing; when the precision requirement metric is greater than a predetermined second threshold, generating a mode selection signal for instructing the network device to operate in the full mode; when the precision requirement metric is less than or equal to the predetermined second threshold, generating a mode selection signal for instructing the network device to operate in the non-full mode. For example, the precision requirement metric can include: the requirement metric for computational precision, the requirement metric for task precision, the requirement metric for network function precision, the requirement metric for application precision, and so on.
[0114] For example, computational precision can include: (1) Floating-Point Precision: Floating-point precision is usually used for tasks that require high-precision calculations, such as high-performance computing (HPC) and deep learning training. Floating-point precision (such as 32-bit floating-point) can provide high computational precision but requires more computing resources and energy consumption; (2) Fixed-Point Precision: Suitable for resource-constrained devices, such as edge computing and Internet of Things devices. By reducing the bit width (such as 8-bit, 16-bit fixed-point), a trade-off can be made between computing resources and precision.
[0115] In Figure 3 In step 102 of the method shown, the network device multicasts the load request to the corresponding computing nodes in the computing node group based on the identification information of the computing node group. Moreover, in Figure 3 In other steps of the method shown, the network device can also multicast the second calculation result or the third calculation result to the corresponding computing nodes in the computing node group based on the identification information.
[0116] In one embodiment, the above multicast process can be implemented through a multi-level look-up table process at the network device, and the multi-level look-up table process has the advantage of convenient implementation.
[0117] Considering that the implementation of multicast by means of multi-level look-up tables has the disadvantage of high memory occupancy, the embodiments of the present invention also propose a processing method for implementing multicast without the need for multi-level look-up tables. In one embodiment, the identification information (sw-mcid) of a computing node group includes: the fusion information of the computing node multicast identification information (gpu-mcid) and the computing node connection identification information (gpu-port-id). Among them: gpu-mcid is used to represent the unique identifier of the computing node group (for example, the unique identifier of computing node group 1 can be 0000, the unique identifier of computing node group 2 can be 0001, etc.); gpu-port-id is used to represent the intersection physical ports of the respective computing nodes in the computing node group.
[0118] Based on the computing node multicast identification information (gpu-mcid), perform a look-up table process at the computing node to obtain the gpu-port-id corresponding to the gpu-mcid (that is, the intersection physical ports of the respective computing nodes in the computing node group). Then, fuse the sw-mcid with the gpu-port-id to obtain the sw-mcid. Next, the computing node carries the sw-mcid in the multicast request and sends the multicast request to the network device. Then, at the network device, perform a look-up table process based on the sw-mcid to obtain the port list corresponding to the sw-mcid, and send the data to be multicast-transmitted to each port in the port list. The following is a demonstration description of the multicast process.
[0119] Assume that: the physical port of GPU0 with identification information 0 is connected to the physical port (P0) of the network device with identification information 0. The physical port of GPU0 with identification information 1 is connected to the physical port (P1) of the network device with identification information 1. The physical port of GPU1 with identification information 0 is connected to the physical port (P2) of the network device with identification information 2. The physical port of GPU1 with identification information 1 is connected to the physical port (P3) of the network device with identification information 3. The physical port of GPU2 with identification information 0 is connected to the physical port (P4) of the network device with identification information 4, and the physical port of GPU2 with identification information 1 is connected to the physical port (P5) of the network device with identification information 5. The physical port of GPU3 with identification information 0 is connected to the physical port (P6) of the switch with identification information 6, and the physical port of GPU3 with identification information 1 is connected to the physical port (P7) of the switch with identification information 7.
[0120] Table 1 is a demonstration schematic table of sw-mcid, gpu-mcid, and gpu-port-id.
[0121]
[0122] Table 1
[0123] For computing node group 1 containing computing nodes (GPU0, GPU1), query Table 1 at GPU0 or GPU1 to obtain the corresponding gpu-mcid as: 0000. The port 0 and port 1 of GPU0 are connected to the network device, and the port 0 and port 1 of GPU1 are connected to the network device. It can be seen that the intersection of the physical ports of GPU0 and GPU1 is 0 and 1. Therefore, the intersection physical ports (gpu-port-id) determined based on Table 1 are: 0 and 1. At GPU0 or GPU1, based on the gpu-mcid and gpu-port-id, in the form of string concatenation, the sw-mcid of computing node group 1 is obtained as: 00000, 00001.
[0124] For computing node group 2 containing computing nodes (GPU0, GPU2), query Table 1 at GPU0 or GPU2 to obtain the corresponding gpu-mcid as: 0001. The port 0 and port 1 of GPU0 are connected to the network device, and the port 0 and port 1 of GPU2 are connected to the network device. It can be seen that the intersection of the physical ports of GPU0 and GPU2 is 0 and 1. Therefore, the intersection physical ports (gpu-port-id) determined based on Table 1 are: 0 and 1. At GPU0 or GPU2, based on the gpu-mcid and gpu-port-id, in the form of string concatenation, the sw-mcid of computing node group 2 containing two computing nodes (GPU0, GPU2) can be obtained as: 00010, 00011.
[0125] Any computing node in any computing node group can carry the sw-mcid in the multicast request and send the multicast request to the network device. Then, at the network device, perform a table lookup process based on the sw-mcid to obtain the port list corresponding to the sw-mcid, and copy the data to be multicast to each port in the port list, or read data from each port in the list.
[0126] Table 2 is a demonstration correspondence table between sw-mcid and the physical port identification information (sw-port-id) of the network device. Among them: when the value of sw-port-id is 1, it means the port is selected; when the value of sw-port-id is 0, it means the port is not selected.
[0127]
[0128] Table 2
[0129] Preferably, Table 2 can be stored at the network device. For example, when the network device receives a multicast request containing 00000 and 00001 from a computing node, querying Table 2 gives the selected ports as: P0, P1, P2, and P3. Therefore, the network device sends multicast data to P0, P1, P2, and P3 respectively. It can be seen that performing a single table lookup at the network device can achieve multicast, effectively reducing the memory requirement.
[0130] In the above exemplary description, in the multicast requests sent by the computing node that is the multicast request initiator, the corresponding computing node groups all include the computing node that is the multicast request initiator. In fact, for the loading request and the first multicast request of the embodiments of the present invention, they do not need to be sent to the computing node that is the multicast request initiator. Therefore, when the computing node that is the loading request or the first multicast request initiator determines the sw-mcid based on Table 1, it is necessary to select and remove the computing node group of the multicast request initiator.
[0131] For example, at GPU0 of the computing node group including computing nodes (GPU0, GPU1, GPU3), query Table 1 using the remaining computing node group (i.e., (GPU1, GPU3)) composed of the remaining computing nodes after removing GPU0 from (GPU0, GPU1, GPU3) to obtain the gpu-port-id of the remaining computing node group (i.e., (GPU1, GPU3)). Then, fuse the gpu-mcid of the remaining computing node group with the gpu-port-id of the remaining computing node group to obtain the sw-mcid of the remaining computing node group. Then, GPU0 can carry this sw-mcid in the loading request. After the network device receives this loading request, it can obtain the data associated with the loading request from the physical port corresponding to the sw-mcid. Similarly, GPU0 can carry this sw-mcid and the second calculation result in the first multicast request. After the network device receives this first multicast request, it can send the second calculation result to each physical port corresponding to the sw-mcid.
[0132] Next, taking the execution of the full reduction operation using in-network computing as an example, the embodiments of the present invention will be exemplarily described. Figure 4 It is a schematic structural diagram of an in-network computing system according to an embodiment of the present invention. In Figure 4 it, the network device is implemented as a switch, and the computing node group includes 4 GPUs connected to the switch, namely GPU0 to GPU3. Each GPU contains its own data block to be reduced (for example, data block 0 to data block 3).
[0133] The full reduction operation of the embodiments of the present invention can include a first process and a second process.
[0134] In the first process: Each of GPUs 0 to 3 respectively sends its own reduction request to the switch; the switch multicasts the reduction requests sent by each GPU to the remaining GPUs in the compute node group other than the GPU that initiated the reduction request; the switch respectively receives the respective data to be reduced associated with the reduction requests from the remaining GPUs; the switch performs reduction processing based on the data to be reduced sent by the remaining GPUs and sends the reduction result to the GPU that initiated the reduction request.
[0135] In the second process: Each of GPUs 0 to 3 respectively receives from the switch the reduction result calculated by the switch, and performs reduction processing on the chunk data corresponding to its own sent reduction request and the reduction result calculated by the switch that it stores itself, to obtain the reduction result calculated by the GPU. Each of GPUs 0 to 3 sends the reduction result calculated by itself to the switch, so that the switch multicasts the reduction results sent by each GPU to the remaining GPUs among GPUs 0 to 3 (excluding the GPU that initiated the reduction request).
[0136] Figure 5 FIG. is a schematic diagram of the first process of the full reduction operation based on in-network computing according to an embodiment of the present invention.
[0137] For example, taking GPU0 as an example, Figure 5 The first process shown includes the following steps:
[0138] Step S1: GPU0 sends a reduction request to the switch, and the reduction request includes the identifier of the compute node group. This compute node group includes GPUs 0 to 3.
[0139] Step S2: The network traffic metric of the switch is greater than the first threshold, so it is determined that the switch is operating in a non-full mode. The switch operating in the non-full mode multicasts the reduction request sent by GPU0 to GPUs 1, 2, and 3, without multicasting to GPU0.
[0140] Step S3: GPUs 1, 2, and 3 respectively send the data to be reduced associated with the reduction request sent by GPU0 (for example, their respective chunk data 0) to the switch.
[0141] Step S4: The switch performs reduction processing on the data to be reduced received from GPUs 1, 2, and 3 (that is, the 3 chunk data 0 respectively received from GPUs 1, 2, and 3) and returns the reduction result to GPU0.
[0142] GPU1 to GPU3 respectively synchronously execute the above steps similar to GPU0. For example, taking GPU1 as an example, the first process includes: GPU1 sends a reduction request to the switch, and the reduction request contains the identifier of the computing node group; the switch multicasts the reduction request sent by GPU1 to GPU0, GPU2, and GPU3; GPU0, GPU2, and GPU3 respectively send the data to be reduced (such as their respective block data 1) associated with the reduction request sent by GPU1 to the switch; the switch performs reduction processing on the data to be reduced received from GPU0, GPU2, and GPU3 (that is, the 3 block data 1 respectively received from GPU0, GPU2, and GPU3), and returns the reduction result to GPU1.
[0143] Similarly, GPU2 to GPU3 respectively synchronously execute the above steps to complete the first process.
[0144] After completing the first process, the second process is executed. In the second process, GPU0 to GPU3 respectively perform reduction calculations on the reduction results of their respective sent reduction requests received from the switch and the data saved by themselves and associated with their respective sent reduction requests, so as to obtain the reduction results calculated by themselves. Moreover, GPU0 to GPU3 respectively send the reduction results calculated by themselves to the switch, so that the switch multicasts the reduction results sent by GPU0 to GPU3 to GPU0 to GPU3.
[0145] Figure 6 FIG. is a schematic diagram of the second process of the full reduction operation based on in-network computing according to an embodiment of the present invention. For example, taking GPU0 as an example, the second process includes the following steps:
[0146] Step S5: GPU0 performs a reduction operation based on the reduction result received from the switch in step S4 (that is, the reduction results of the block data 0 of GPU1, the block data 0 of GPU2, and the block data 0 of GPU3) and the block data (i.e., the block data 0 of GPU0) saved by itself and corresponding to the reduction request sent by GPU0, so as to obtain the reduction result calculated by GPU0. GPU0 sends a multicast request to the switch, and the multicast request contains the identifier information of the computing node group and the reduction result calculated by GPU0.
[0147] Step S6: The switch multicasts the reduction result calculated by GPU0 sent by GPU0 to GPU0 to GPU3. Therefore, GPU0 to GPU3 can obtain the reduction results of the block data 0 in GPU0, the block data 0 of GPU1, the block data 0 of GPU2, and the block data 0 of GPU3.
[0148] In the second process, GPU1 to GPU3 respectively and synchronously execute the above steps similar to GPU0. For example, taking GPU1 as an example, the second process includes: GPU1 performs reduction operation based on the reduction result received from the switch in step S4 (that is, the reduction result of block data 1 of GPU0, block data 1 of GPU2 and block data 1 of GPU3) and the block data corresponding to the reduction request sent by GPU0 (that is, block data 1 of GPU1) stored by itself, and obtains the reduction result calculated by GPU1. GPU1 sends a multicast request to the switch, and the multicast request contains the identification information of the computing node group and the reduction result calculated by GPU1. Similarly, GPU2-GPU3 respectively perform the above steps synchronously, so that GPU0-GPU3 can also obtain the reduction results of block data 2 in GPU0, block data 2 in GPU1, block data 2 in GPU2 and block data 2 in GPU3 (calculated by GPU2) and the reduction results of block data 3 in GPU0, block data 3 in GPU1, block data 3 in GPU2 and block data 3 in GPU3 (calculated by GPU3). So far, the second process is completed.
[0149] It can be seen that since the switch in the non-full mode does not need to send the protocol request to the GPU that initiated the protocol request, the GPU that initiated the protocol request does not need to send the data associated with the protocol request in the GPU that initiated the protocol request to the switch, thus saving network traffic and effectively alleviating network congestion. Moreover, the protocol result finally calculated by the GPU that initiated the protocol request based on the protocol result calculated by the switch is multicast to the remaining GPUs in the GPU group except the GPU that initiated the protocol request, which can further save network traffic and further effectively alleviate network congestion.
[0150] The above takes the implementation of on-line computing as a full-protocol operation as an example to illustrate the on-line computing of the implementation mode of the present invention. In fact, on-line computing can also be implemented as data aggregation and caching, content distribution and computing reuse, stateful forwarding and multicast communication, dynamic task allocation and scheduling, machine learning reasoning, real-time data processing and analysis, distributed computing, load balancing, edge computing and cloud native applications and other application environments.
[0151] Figure 7 FIG. 1 is an exemplary structural diagram of an on-line computing device according to an embodiment of the present invention. The device is applicable to network equipment. Figure 7As shown in the figure, the device includes: a first receiving module 701, configured to receive a loading request from a computing node, where the loading request includes identification information of a computing node group, and the computing node group includes computing nodes; a multicast module 702, configured to, when it is determined that the device operates in a non-full mode, based on the identification information, multicast the loading request to the remaining computing nodes in the computing node group except the computing node; a second receiving module 703, configured to receive data associated with the loading request from the remaining computing nodes; an in-network computing module 704, configured to perform a predetermined calculation on the data to obtain a first calculation result; and a sending module 705, configured to send the first calculation result to the computing node, so that the computing node performs the calculation on the first calculation result and the data in the computing node associated with the loading request to obtain a second calculation result.
[0152] In one embodiment, the device can be used as a software module or a hardware module to be integrated or embedded into a network device. For example, the network device can be implemented as a router, a switch, a network card, a hub, a repeater, a bridge, a gateway, a firewall, an access point, and the like.
[0153] In one embodiment, the multicast module 702 is configured to, when it is determined that the device operates in a full mode, based on the identification information, multicast the loading request to all computing nodes in the computing node group; the second receiving module 703 is configured to receive data associated with the loading request from all computing nodes respectively; the in-network computing module 704 is configured to perform a predetermined calculation on the data respectively received from all computing nodes and associated with the loading request to obtain a third calculation result; and the sending module 705 is configured to send the third calculation result to the computing node.
[0154] In summary, in the embodiment of the present invention, a loading request is received from a computing node, where the loading request includes identification information of a computing node group, and the computing node group includes computing nodes; when it is determined that the device operates in a non-full mode, based on the identification information, the loading request is multicast to the remaining computing nodes in the computing node group except the computing node; data associated with the loading request is received from the remaining computing nodes; a predetermined calculation is performed on the data to obtain a first calculation result; and the first calculation result is sent to the computing node, so that the computing node performs the calculation on the first calculation result and the data in the computing node associated with the loading request to obtain a second calculation result. Thus, it can be seen that the network device does not need to send the loading request to the computing node that initiates the loading request, saving network traffic and effectively alleviating network congestion. In addition, the second calculation result calculated by the computing node that initiates the loading request is multicast to the remaining computing nodes without being sent to the computing node that initiates the loading request, further saving network traffic. In addition, the embodiment of the present invention can also automatically select a suitable mode based on the accuracy requirement or traffic condition, and can achieve a good balance between the accuracy requirement and traffic control.
[0155] Embodiments of the present invention also propose an electronic device having a processor-memory architecture. Figure 8 is a structural diagram of an electronic device according to an embodiment of the present invention. As Figure 8 shown, the electronic device includes a processor 801, a memory 802, and a computer program stored on the memory 802 and executable on the processor 801. When the computer program is executed by the processor 801, it implements any of the above online computing methods. Among them, the memory 802 can be specifically implemented as various storage media such as electrically erasable programmable read-only memory (EEPROM), flash memory, programmable read-only memory (PROM), etc. The processor 801 can be implemented as including one or more central processing units or one or more field programmable gate arrays, where the field programmable gate array integrates one or more central processing unit cores. Specifically, the central processing unit or the central processing unit core can be implemented as a CPU, GPU, GPGPU, MCU, or DSP, etc.
[0156] It should be noted that not all steps and modules in the above processes and structural diagrams are necessary, and some steps or modules can be ignored according to actual needs. The execution order of each step is not fixed and can be adjusted according to needs. The division of each module is only for the convenience of description and is a functional division. In actual implementation, one module can be implemented by multiple modules, and the functions of multiple modules can also be implemented by the same module. These modules can be located in the same device or in different devices.
[0157] The hardware modules in each embodiment can be implemented mechanically or electronically. For example, a hardware module can include a specially designed permanent circuit or logic device (such as a dedicated processor, such as an FPGA or ASIC) for performing specific operations. For example, specific operations can be completed in various types of chips (such as artificial intelligence chips). The hardware module can also include a programmable logic device or circuit (such as including a general-purpose processor or other programmable processor) temporarily configured by software for performing specific operations. As for whether to specifically use a mechanical method, a dedicated permanent circuit, or a temporarily configured circuit (such as configured by software) to implement the hardware module, it can be determined according to cost and time considerations.
[0158] The present invention also provides a machine-readable storage medium storing instructions for causing a machine to execute the method as described in this application. Specifically, a system or device equipped with the storage medium can be provided. On this storage medium, software program codes for implementing the functions of any one of the above-mentioned embodiments are stored, and the computer (or CPU or MPU) of the system or device is caused to read and execute the program codes stored in the storage medium. In addition, based on the instructions of the program codes, the operating system operating on the computer can be used to complete part or all of the actual operations. The program codes read from the storage medium can also be written into the memory provided in the expansion board inserted into the computer or into the memory provided in the expansion unit connected to the computer. Subsequently, based on the instructions of the program codes, the CPU etc. installed on the expansion board or the expansion unit are caused to execute part and all of the actual operations, thereby implementing the functions of any one of the above-mentioned embodiments. The storage medium embodiments for providing the program codes include floppy disks, hard disks, magneto-optical disks, optical disks (such as CD-ROM, CD-R, CD-RW, DVD-ROM, DVD-RAM, DVD-RW, DVD+RW), magnetic tapes, non-volatile memory cards, and ROMs. Optionally, the program codes can be downloaded from a server computer or the cloud via a communication network.
[0159] In this document, "schematic" means "serving as an instance, example, or illustration", and any illustration or embodiment described as "schematic" in this document should not be construed as a more preferred or more advantageous technical solution. To make the drawings concise, only the parts related to the present invention are schematically shown in each drawing, and do not represent their actual structures as products. Additionally, to make the drawings concise and easy to understand, in some drawings, for components with the same structure or function, only one of them is schematically shown, or only one of them is labeled. In this document, "a" does not mean that the quantity of the relevant parts of the present invention is limited to "only one", and "a" does not exclude the case where the quantity of the relevant parts of the present invention is "more than one". In this document, "upper", "lower", "front", "rear", "left", "right", "inner", "outer", etc. are only used to represent the relative positional relationship between the relevant parts, rather than defining the absolute positions of these relevant parts.
[0160] The above are only the preferred embodiments of the present invention, and are not used to limit the protection scope of the present invention. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.
Claims
1. A method for in-network computing, characterized in that, The method is applicable to a network device, and the method includes: Receiving a loading request from a computing node, where the loading request includes identification information of a computing node group, and the computing node group includes the computing node; When it is determined that the device is operating in a non-full mode, based on the identification information, multicasting the loading request to the remaining computing nodes in the computing node group except the computing node; Receiving data associated with the loading request from the remaining computing nodes; Performing a predetermined calculation on the data to obtain a first calculation result; Sending the first calculation result to the computing node, so that the computing node performs the calculation on the first calculation result and the data in the computing node associated with the loading request to obtain a second calculation result.
2. The method according to claim 1, characterized in that, After obtaining the second calculation result, the method includes: Receiving a first multicast request from the computing node, where the first multicast request includes the identification information and the second calculation result; Based on the identification information, multicasting the second calculation result to the remaining computing nodes in the computing node group except the computing node.
3. The method according to claim 1, wherein Including: When it is determined that the device is operating in a full mode, based on the identification information, multicasting the loading request to all computing nodes in the computing node group; Receiving data associated with the loading request from each of the all computing nodes; Performing the calculation on the data associated with the loading request received from each of the all computing nodes to obtain a third calculation result; Sending the third calculation result to the computing node.
4. The method according to claim 3, wherein After obtaining the third calculation result, the method includes: Receiving a second multicast request from the computing node, where the second multicast request includes the identification information and the third calculation result; Based on the identification information, multicasting the third calculation result to all computing nodes in the computing node group.
5. The method according to any one of claims 1-4, characterized in that, Including: Generating a mode selection signal based on a predetermined performance metric; Based on the mode selection signal, determining that the network device is operating in the non-full mode or the full mode.
6. The method according to claim 5, characterized in that, The generating a mode selection signal based on a predetermined performance metric includes: Monitoring a traffic metric of the network device; When the traffic metric is greater than a predetermined first threshold, generating a mode selection signal for indicating that the network device is operating in the non-full mode; when the traffic metric is less than or equal to the predetermined first threshold, generating a mode selection signal for indicating that the network device is operating in the full mode.
7. The method according to claim 5, wherein The generating a mode selection signal based on a predetermined performance metric includes: Obtaining a precision requirement metric for the calculation; When the precision requirement metric is greater than a predetermined second threshold, generating a mode selection signal for indicating that the network device is operating in the full mode; when the precision requirement metric is less than or equal to the predetermined second threshold, generating a mode selection signal for indicating that the network device is operating in the non-full mode.
8. A computing device in the network, characterized in that, The apparatus is applicable to a network device, and the apparatus includes: A first receiving module, configured to receive a loading request from a computing node, where the loading request includes identification information of a computing node group, and the computing node group includes the computing node; A multicast module, configured to, when determining to operate in a non-full mode, multicast the loading request to the remaining computing nodes in the computing node group except the computing node based on the identification information; A second receiving module, configured to receive data associated with the loading request from the remaining computing nodes; An in-network computing module, configured to perform a predetermined calculation on the data to obtain a first calculation result; A sending module, configured to send the first calculation result to the computing node, so that the computing node performs the calculation on the first calculation result and the data in the computing node associated with the loading request to obtain a second calculation result.
9. The apparatus according to claim 8, wherein: The multicast module is configured to, when determining to operate in a full mode, multicast the loading request to all computing nodes in the computing node group based on the identification information; The second receiving module is configured to receive data associated with the loading request from each of the all computing nodes; The computing module is configured to perform the calculation on the data respectively received from each of the all computing nodes and associated with the loading request to obtain a third calculation result; The sending module is configured to send the third calculation result to the computing node.
10. An electronic device, characterized in that, Comprising: A memory; A processor; Wherein an application program executable by the processor is stored in the memory, and is configured to enable the processor to execute the in-network computing method according to any one of claims 1-7.
11. A computer-readable storage medium, characterized in that, Computer-readable instructions are stored on the computer-readable storage medium, and when being executed by the processor, enable the processor to execute the in-network computing method according to any one of claims 1-7.
12. A program product, comprising a computer program, characterized in that, When being executed by the processor, the computer program implements the in-network computing method according to any one of claims 1-7.
Citation Information
Cited By
Data communication method and device, electronic equipment and storage medium
CN121509528A