Online computing error recovery method, electronic device and computer program product

By switching to the alternate aggregation tree in network computing, the aggregation operation interruption caused by network device errors is solved, and more efficient distributed training is achieved.

CN120186007APending Publication Date: 2025-06-20ZTE CORP
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510257827.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-03
Publication Date
2025-06-20

AI Technical Summary

Technical Problem

In distributed computing, when an error occurs in the network device, the aggregation operation needs to be restarted, affecting the efficiency and training time of network computing.

Method used

When an error occurs in the aggregation tree, if the preset switching conditions are met, switch to the alternate aggregation tree and continue the aggregation process after the switching to avoid interruption and restart operations.

Benefits of technology

It reduces the recovery time of network computing errors, shortens the time of distributed training, and improves the efficiency of distributed training.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120186007A_ABST
    Figure CN120186007A_ABST
Patent Text Reader

Abstract

The invention provides an in-network computing error recovery method, which comprises the following steps of: in an aggregation process of executing in-network computing, switching an aggregation tree to a standby aggregation tree under the condition that network equipment in the current aggregation tree has an error and meets a preset switching condition, and continuing the aggregation process after the aggregation tree is switched to the standby aggregation tree; wherein the aggregation tree for network computing comprises a plurality of end side hosts and multiple levels of network equipment, the multiple levels of network equipment are in signal connection level by level, and the end side hosts are in signal connection with the network equipment at the lowest level. The invention further provides electronic equipment and a computer program product.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of communication technologies, and particularly to an in-network computing error recovery method, an electronic device, and a computer program product. Background Art

[0002] In-network computing is a technology that offloads the computing process of distributed training tasks to network devices, which can reduce data transmission volume, improve throughput, and reduce training time. Currently, when errors occur in distributed computing devices (network devices and end-side hosts), it is necessary to restart the aggregation operation on the control plane, which will affect the aggregation efficiency of in-network computing and ultimately lead to an extension of the training time. How to reduce the impact of errors in distributed computing devices on the training time is an urgent problem to be solved. Summary of the Invention

[0003] The present disclosure provides an in-network computing error recovery method, an electronic device, and a computer program product.

[0004] In a first aspect, an embodiment of the present disclosure provides an in-network computing error recovery method, including:

[0005] During the aggregation process of in-network computing, when an error occurs in a network device in the current aggregation tree and a preset switching condition is satisfied, switch the aggregation tree to a standby aggregation tree, and continue the aggregation process after switching to the standby aggregation tree; wherein, the aggregation tree of the in-network computing includes multiple end-side hosts and multiple levels of network devices, the multiple levels of network devices are connected by signals in sequence, and the end-side hosts are connected by signals to the network devices at the lowest level.

[0006] In a second aspect, an embodiment of the present disclosure provides an electronic device, which includes a memory and a processor; the memory stores a computer program that can be executed by the processor, and when the computer program is executed by the processor, it implements the in-network computing error recovery method provided by the embodiment of the present disclosure.

[0007] In a third aspect, an embodiment of the present disclosure provides a computer program product, which includes a computer program, and when the computer program is executed by a processor, it implements the in-network computing error recovery method provided by the embodiment of the present disclosure.

[0008] In the in-network computing error recovery method in the embodiments of the present disclosure, when an error occurs in a network device in the aggregation tree and a switching condition is satisfied, the aggregation tree is switched to a standby aggregation tree, and the aggregation process is continued after switching to the standby aggregation tree, reducing the recovery time of in-network computing errors, thereby shortening the distributed training time and improving the efficiency of distributed training. Brief Description of the Drawings

[0009] In the drawings of the embodiments of the present disclosure:

[0010] Figure 1 It is the architecture diagram of in-network computing in the related art;

[0011] Figure 2 It is the flowchart of a method for in-network computing error recovery provided by an embodiment of the present disclosure;

[0012] Figure 3 It is the schematic diagram of problems occurring in a spine network device provided by an embodiment of the present disclosure;

[0013] Figure 4 It is the format of an error indication message extended based on the UEC protocol in an embodiment of the present disclosure;

[0014] Figure 5 It is the format of an error indication message extended based on the RoCE protocol in an embodiment of the present disclosure;

[0015] Figure 6 It is the flowchart of a method for in-network computing error recovery provided by an embodiment of the present disclosure;

[0016] Figure 7 It is the flowchart of another method for in-network computing error recovery provided by an embodiment of the present disclosure;

[0017] Figure 8 It is the block diagram of the composition of an electronic device provided by an embodiment of the present disclosure. Detailed implementation manners

[0018] To enable those skilled in the art to better understand the technical solutions of the present disclosure, the embodiments of the present disclosure will be described in detail below with reference to the accompanying drawings.

[0019] The present disclosure will be described more fully hereinafter with reference to the accompanying drawings. However, the illustrated embodiments may be embodied in different forms and the present disclosure should not be construed as limited to the embodiments set forth hereinafter. On the contrary, these embodiments are provided so that this disclosure will be thorough and complete, and will fully convey the scope of the present disclosure to those skilled in the art.

[0020] The accompanying drawings of the embodiments of the present disclosure are used to provide a further understanding of the embodiments of the present disclosure, and constitute a part of the specification. Together with the detailed embodiments, they are used to explain the present disclosure and do not constitute a limitation to the present disclosure. By describing the detailed embodiments with reference to the accompanying drawings, the above and other features and advantages will become more apparent to those skilled in the art.

[0021] The present disclosure may be described with reference to the plan views and / or cross-sectional views by means of the ideal schematic diagrams of the present disclosure. Therefore, the example illustrations may be modified according to the manufacturing technology and / or tolerances.

[0022] In the case of no conflict, the embodiments of the present disclosure and the features in the embodiments may be combined with each other.

[0023] The terms used in the present disclosure are only for describing specific embodiments and are not intended to limit the present disclosure. As used in the present disclosure, the term "and / or" includes any and all combinations of one or more of the related listed items. As used in the present disclosure, the singular forms "a" and "the" are also intended to include the plural forms unless the context clearly indicates otherwise. As used in the present disclosure, the terms "comprising", "made of", specify the presence of the described features, wholes, steps, operations, elements and / or components, but do not exclude the presence or addition of one or more other features, wholes, steps, operations, elements, components and / or their groups.

[0024] Unless otherwise defined, all terms (including technical and scientific terms) used in the present disclosure have the same meaning as commonly understood by those of ordinary skill in the art. It will also be understood that terms such as those defined in common dictionaries should be interpreted as having a meaning consistent with their meaning in the context of the relevant art and the present disclosure, and will not be interpreted as having an idealized or overly formal meaning unless the present disclosure clearly so defines.

[0025] The present disclosure is not limited to the embodiments shown in the drawings, but includes modifications to the configurations formed based on the manufacturing process. Therefore, the regions illustrated in the drawings have schematic properties, and the shapes of the regions shown in the figures illustrate the specific shapes of the regions of the elements, but are not intended to be restrictive.

[0026] With the rapid development of artificial intelligence (AI for short), the requirement for computing power is getting higher and higher. Distributed in-network computing can effectively alleviate the problem of computing power shortage and reduce the occupancy of network resources.

[0027] Figure 1 It is an architecture diagram of in-network computing in the related art. As Figure 1As shown in the figure, the architecture of in-network computing includes a controller ((In-Network Computing Manager, abbreviated as IM) 10, network devices 20, and end-hosts 30. Among them, the network devices 20 are used to receive the aggregated packets sent by the end-hosts 30, complete the aggregation operation, and send the aggregated packets to the next-level network devices or return them to the end-hosts 30. The network devices 20 include multiple network devices, and the multiple network devices form the physical topology of in-network computing. In some related technologies, the multiple network devices include at least one leaf network device 21 and at least one spine network device 22. The end-hosts 30 include user applications and network cards. The user applications initiate in-network computing aggregation requests to the controller 10. After the aggregation requests are successful, the aggregated packets are sent to the leaf network device 21 and the spine network device 22 through the network cards for aggregation operations. The controller 10 is the abbreviation of the in-network computing controller, which is used to collect the physical topology of in-network computing, build an aggregation tree based on the physical topology according to the user's aggregation requests, allocate the resources required for in-network computing, and send the resource allocation results to the leaf network device 21 and the spine network device 22. The end-hosts 30, the leaf network device 21, and the spine network device 22 perform aggregation calculations according to the requirements of the controller 10 and send the aggregation results to the end-hosts 30.

[0028] In-network computing can complete end-to-end training tasks. The end-to-end training tasks include:

[0029] Step S10, the controller 10 obtains the physical topology of in-network computing.

[0030] Step S20, the end-host 30 sends an aggregation request to the controller 10.

[0031] Step S30, the controller 10 responds to the aggregation request of the end-host 30 and creates an aggregation tree based on the physical topology.

[0032] Step S40, the controller 10 returns the aggregation tree parameters and resource allocation information to the network devices. The resource allocation information includes but is not limited to forwarding table entries and resources occupied by switches.

[0033] Step S50, the end-host 30 sends aggregated packets to the leaf network device 21 and the spine network device 22.

[0034] Step S60, the leaf network device 21 and the spine network device 22 perform aggregation operations, and the spine network device 22 returns the aggregation results to the end-host 30.

[0035] In Figure 1Among them, the dashed line indicates the information transmission direction between the network device 20 and the controller 10, and between the end-side host 30 and the controller 10. The dotted line indicates the information transmission direction from bottom to top of the network device 20 and the end-side host 30. The dotted line indicates the information transmission direction from top to bottom of the network device 20 and the end-side host 30.

[0036] In the above end-to-end process, if any one of the end-side host 30, the leaf network device 21, and the spine network device 22 has an error, it will affect the in-network computing process and increase the training time. Although the draft in-network computing standard of the Ultra Ethernet Consortium (UEC) classifies errors into transient errors and non-transient errors, the draft also proposes various solutions for transient errors, such as packet retransmission or flow control at the link layer and the transport layer, so that it can be restored in a short time and has little impact on the task completion time. For non-transient errors, the draft does not have corresponding solutions, and only proposes that the aggregation tree creation of the control plane needs to be re-performed and the aggregation operation needs to be re-executed. And for the controller to reallocate resources, the aggregation operation needs to be restarted. Therefore, it seriously affects the completion time of the training task.

[0037] Therefore, the embodiments of the present disclosure provide an in-network computing error recovery method, an electronic device, a readable medium, and a program product. When errors occur in the end-side host and the network device, the aggregation operation is not interrupted, and it is not necessary to restart the aggregation operation, which can reduce the completion time of distributed training.

[0038] In a first aspect, the embodiments of the present disclosure provide an in-network computing error recovery method.

[0039] Figure 2 It is a flowchart of an in-network computing error recovery method provided by the embodiments of the present disclosure. As Figure 2 shown, the in-network computing error recovery method provided by the embodiments of the present disclosure includes:

[0040] Step S201, during the aggregation process of in-network computing, when an error occurs in a network device in the current aggregation tree and meets a preset switching condition, switch the current aggregation tree to a standby aggregation tree, and continue the aggregation process after switching to the standby aggregation tree.

[0041] Among them, the aggregation tree of in-network computing includes multiple end-side hosts and multiple levels of network devices. The multiple levels of network devices are connected by signals in sequence. The end-side host is signal-connected to the network device at the lowest level. The end-side host and the network device are used to perform the aggregation operation of the distributed training task.

[0042] It should be noted that the error here refers to an error in the aggregation module in the network device and the edge host, resulting in the inability to perform the aggregation operation. However, the communication between the edge host and the network device is not affected by the error, that is, the edge host and the network device, as well as network devices at different levels, can still communicate with each other.

[0043] The current aggregation tree is the aggregation tree currently performing the aggregation operation. Before performing the aggregation operation, the controller constructs one or more aggregation trees in response to the aggregation request of the edge host. When the controller constructs an aggregation tree, this aggregation tree is used as the current aggregation tree. When the controller constructs multiple aggregation trees, the aggregation tree with the optimal performance is selected as the current aggregation tree, and the remaining aggregation trees are used as standby aggregation trees.

[0044] It should be noted that when switching the aggregation tree, the aggregation operations of the unaffected network devices and / or edge hosts do not stop being executed.

[0045] In the in-network computing error recovery method in the embodiments of the present disclosure, when a network device in the aggregation tree has an error and meets the switching conditions, the current aggregation tree is switched to a standby aggregation tree, and the subsequent aggregation process continues after switching to the standby aggregation tree, reducing the recovery time of the in-network computing error, thereby shortening the time of distributed training and improving the efficiency of distributed training.

[0046] In some embodiments, switching the current aggregation tree to a standby aggregation tree includes: selecting, from the standby aggregation tree list, the aggregation tree that does not include the faulty network device and has the highest priority as the new aggregation tree; wherein, the standby aggregation tree list is pre-generated by the controller based on the physical topology of the in-network computing and the edge hosts participating in the in-network computing.

[0047] In some embodiments, the controller responds to the aggregation request of the edge host, generates multiple aggregation trees based on the physical topology of the in-network computing and the edge hosts participating in the in-network computing, selects the aggregation tree with the highest priority from them as the current aggregation tree, uses the other aggregation trees as standby aggregation trees, and forms a standby aggregation tree list. In the standby aggregation tree list, the standby aggregation trees are sorted according to the priority of the standby aggregation trees.

[0048] Among them, the priority of the aggregation tree is determined according to the performance of the aggregation tree. The priority of the aggregation tree can be determined according to network performance parameters such as the network delay and network bandwidth utilization rate of the aggregation tree. The embodiments of the present disclosure do not limit the determination method of the priority.

[0049] In the embodiments of the present disclosure, the physical topology is a topology structure composed of all end-side hosts and network devices. When constructing an aggregation tree, not all end-side hosts and network devices are necessarily used. Therefore, the end-side hosts and network devices in the aggregation tree may be some of the end-side hosts and network devices in the physical topology. Since a network device can be signal-connected to different end-side hosts, even if multiple constructed aggregation trees have the same end-side hosts, different network devices may be adopted. In the embodiments of the present disclosure, network devices include, but are not limited to, switches and routers.

[0050] The information of the aggregation tree includes the aggregation relationship between end-side hosts and network devices, and the computing resources required for the aggregation operation. When a network device receives the aggregation tree, it can obtain the computing resources that the network device needs to provide from the information of the aggregation tree.

[0051] In the embodiments of the present disclosure, after the controller generates the standby aggregation tree list, it can store the standby aggregation tree list by itself, or send the standby aggregation tree list to the end-side hosts, and the end-side hosts store the standby aggregation tree list.

[0052] When the standby aggregation tree list is stored in the end-side hosts, switching the current aggregation tree to the standby aggregation tree includes: the end-side hosts respond to the aggregation tree switching trigger message carried in the data-plane packet, select the aggregation tree with the highest priority that does not include faulty network devices from the standby aggregation tree list as the new aggregation tree, and continue the aggregation process after switching to the new aggregation tree.

[0053] Since the switching trigger message is sent to the network device through the data-plane packet, and the standby aggregation tree list is stored in the end-side hosts, the aggregation tree switching does not pass through the controller, which greatly reduces the error recovery time, thereby shortening the time of distributed training.

[0054] When the standby aggregation tree list is stored in the controller, switching the current aggregation tree to the standby aggregation tree includes: the controller responds to the aggregation tree switching trigger information, selects the aggregation tree with the highest priority that does not include faulty network devices from the standby aggregation tree list as the new aggregation tree; and sends the new aggregation tree to the end-side hosts, and the end-side hosts perform subsequent aggregation processes according to the new aggregation tree.

[0055] At this time, although the new aggregation tree is sent by the controller, it is pre-stored in the controller. Therefore, compared with the prior art, there is no need for the controller to recalculate the aggregation tree, saving the time for recalculating the aggregation tree and reducing the error recovery time.

[0056] In some embodiments, when the aggregation tree is switched, each of the end-side hosts of the aggregation tree is switched, that is, each end-side host in the aggregation tree is switched to the alternative aggregation tree.

[0057] In some embodiments, when the aggregation tree is switched, the edge hosts that are signal - connected to the faulty network device are switched, while the remaining edge hosts are not switched. Only the edge hosts that are signal - connected to the faulty network device are switched to the standby aggregation tree, and the edge hosts that are not signal - connected to the faulty network device are not switched and still use the original aggregation tree.

[0058] In some embodiments, the switching condition is preset. The preset switching conditions include: the computing device with an error is a network device at a level below the highest level, and the number of edge hosts connected to the network device with an error is greater than or equal to a preset first configuration value (HostUnderLeaf). The specific value of the first configuration value can be set by the user, and the embodiments of the present disclosure do not limit the specific value of the first configuration value.

[0059] As Figure 1 shown, the computing devices in the aggregation tree include four edge hosts 30 and three network devices. The three network devices form two levels, namely, the high level and the low level; among them, the low level includes two network devices, that is, two leaf network devices 21, and the high level includes one network device, that is, one spine network device 22. Each leaf network device 21 is signal - connected to two edge hosts 30, and the two leaf network devices 21 are respectively signal - connected to the spine network device 22.

[0060] When a leaf network device 21 has an error, it is determined whether the switching condition is met according to the number of edge hosts 30 connected to the leaf network device 21. When the number of edge hosts 30 connected to the faulty leaf network device 21 is greater than or equal to the preset first configuration value, the aggregation tree needs to be switched; otherwise, the aggregation packets of the faulty leaf network device are discarded and the aggregation process continues without interruption, and there is no need to restart the aggregation operation.

[0061] When the aggregation tree includes network devices at multiple levels, the network device with an error sends an error message to the upper - level network device and the edge hosts connected to it, and the upper - level network device forwards the error message to the network device at a higher level and the edge hosts connected to it until it is forwarded to the spine network device. In this way, when all network devices and edge hosts in the aggregation tree receive the error message, the edge host selects a standby aggregation tree from the standby aggregation tree list as the new aggregation tree, and the edge host starts sending aggregation packets to the standby aggregation tree from the error point. Among them, the error point is the first aggregation packet that the faulty network device cannot send, that is, the next aggregation packet of the packet sequence number (PSN) of the packet that is correctly aggregated and returns a result in the error message.

[0062] Exemplarily, Figure 3 is a schematic diagram of a problem occurring in a spine network device provided by an embodiment of the present disclosure. As Figure 3 shown, the physical topology includes six end-side hosts host1-host6 and two levels of network devices. Among them, the highest-level network devices include spine network device switch4 and spine network device switch5, and the lowest-level network devices include leaf network devices switch1-switch3. Moreover, end-side host host1 and end-side host host2 are signal-connected to leaf network device switch1, end-side host host3 and end-side host host4 are signal-connected to leaf network device switch2, and end-side host host5 and end-side host host6 are signal-connected to leaf network device switch3. Leaf network devices switch1-switch3 are both signal-connected to spine network device switch4 and signal-connected to spine network device switch5.

[0063] The controller constructs two aggregation trees according to the aggregation operation request. The first aggregation tree includes end-side hosts host1-host6, leaf network devices switch1-switch3, and spine network device switch4, and the second aggregation tree includes end-side hosts host1-host6, leaf network devices switch1-switch3, and spine network device switch5. The performance of the first aggregation tree is better than that of the second aggregation tree. Therefore, the first aggregation tree is used as the current aggregation tree, and the second aggregation tree is used as the standby aggregation tree.

[0064] During the aggregation process, if the spine network device switch4 in the current aggregation tree has an error, the aggregation packets sent by the leaf network devices switch1-switch3 to the spine network device switch4 cannot perform aggregation operations. The spine network device switch4 sends an error packet to the leaf network devices switch1-switch3, and the leaf network devices switch1-switch3 forward the error packet to the end-side hosts host1-host6. When the end-side hosts host1-host6 perform aggregation tree switching, they switch to the second aggregation tree and continue to send aggregation packets from the error point for the subsequent aggregation process.

[0065] For example, if an error occurs in the spine network device switch4, and the packet sequence number of the packet that the spine network device switch4 has correctly completed aggregation and returned the result is 150, after switching the aggregation tree to the spine network device switch5, the aggregation packet sequence number sent from the error point is 151, and the aggregation packet continues the aggregation process along the aggregation tree with the root node being the spine network device switch5.

[0066] When the network device with an error is a network device below the highest level (the lowest level or the intermediate level), there are three handling methods. First, if the number of end-side hosts connected to the network device with an error is less than a preset first configuration value, it means that the impact of this leaf network device on the training result is small, and this leaf network device can be discarded. If the number of end-side hosts connected to the network device with an error is greater than or equal to the preset first configuration value, the aggregation tree can be switched, or the aggregation operation can be restarted.

[0067] In some embodiments, when the network device with an error is a network device below the highest level and the number of end-side hosts connected to the network device with an error is less than the preset first configuration value, the aggregation packets of the network device with an error are discarded.

[0068] For example, as Figure 3 , when an error occurs in the leaf network device switch1 and the number of end-side hosts connected to the leaf network device switch1 is less than the preset first configuration value, the spine network device switch4 discards the aggregation packets of the leaf network device switch1 and only aggregates the aggregation packets of the leaf network device switch2.

[0069] In some embodiments, when a network device in the current aggregation tree has an error and meets the preset restart condition, the aggregation tree is removed; and after the network device with an error resumes normal operation, a new aggregation operation is started again, that is, a new aggregation tree is created and training is performed again. It should be noted that when removing the aggregation tree, the backup aggregation tree needs to be completely removed, and after waiting for the computing device with an error to resume normal operation, a new aggregation operation is started again.

[0070] In some embodiments, the restart condition includes: the computing device with an error is a network device at a level below the highest level and the number of end-side hosts connected to the network device with an error is greater than or equal to the preset first configuration value; and / or, the number of end-side hosts with an error is greater than the preset second configuration value (MaxHostError); and / or, the network device with an error has a communication failure.

[0071] When errors occur in network devices at the lowest level and the intermediate level, and the number of end - side hosts connected to the network device with errors is greater than or equal to a preset first configuration value, the aggregation operation can be restarted or the aggregation tree can be switched. Specifically, whether to restart the aggregation operation or switch the aggregation tree can be set by the user himself.

[0072] When a communication failure occurs in a network device, the entire aggregation operation is affected and the aggregation operation needs to be restarted.

[0073] In the embodiments of the present disclosure, when a leaf network device has an error, the leaf network device can notify the spine network device of its error through an error indication message. Or, when the spine network device cannot receive the aggregation message of the leaf network device, it determines that the leaf network device has an error. When an end - side host has an error, the end - side host can actively send an error indication message to report its error, or the controller can detect the error of the end - side host through heartbeat.

[0074] In some embodiments, when a computing device has an error, the computing device with the error can send an error indication message. The error indication message is a message used to notify network devices and end - side hosts and belongs to a data - plane message. The error indication message includes one or more of a task identifier, a group identifier, an aggregation tree identifier, a packet sequence number that has been correctly aggregated, and an identifier of the computing device where the error occurs, and these information are encapsulated in the error indication message.

[0075] Among them, the task identifier Job_id is used to identify which training task the task being processed by the computing device belongs to.

[0076] The group identifier Group_id is used to identify the aggregation group to which the computing device executing the aggregation task belongs.

[0077] The aggregation tree identifier Tree_id is used to identify the identity of the aggregation tree.

[0078] The packet sequence number PSN that has been correctly aggregated is used to identify the sequence number of the aggregation packet.

[0079] The computing device identifier is used to identify the identity of the computing device.

[0080] In some embodiments, error indication messages can adopt multiple encapsulation formats. In the embodiments of the present disclosure, error indication messages can be encapsulated based on the Ultra Ethernet Consortium (UEC) or the Remote Direct Memory Access over Converged Ethernet (RoCE) protocol. Among them, the transport layer of the UEC protocol uses the Ultra Ethernet Transport layer (Ultra Ethernet Transport layer, referred to as UET), and uses the packet delivery service Header (PDSHeader) and the Semantic sublayer Header (SES Header) of the UET for extension.

[0081] Figure 4 This is the format of the error indication message extended based on the UEC protocol in the embodiments of the present disclosure. As Figure 4 shown, the PDS layer is extended through existing control messages. The PDS Header includes:

[0082] TYPE = 11, which is used to specify the control packet of the UET.

[0083] CTL_TYPE uses any one of the reserved values 9 - 15. For example, CTL_TYPE = 9 is used to represent an error indication message.

[0084] PSN, SPDCID, and DPDCID respectively represent the packet sequence number, source PDC identifier, and destination PDC identifier. These three values still follow the original definitions. Among them, PSN is directly used to represent the message sequence number that has been correctly aggregated and the result has been returned.

[0085] Tree-id and Node-id select the original payload Payload field and, according to CTL_TYPE = 9, are used to indicate the current aggregation tree identifier and the computing device identifier where the error occurs.

[0086] Group_id and Job_id are carried through the encapsulation of the semantic SES Header. Group_id is encapsulated in the Resource Index field, and Job_id is encapsulated in the Job ID field. The other fields follow the original semantic header encapsulation, which will not be elaborated in the embodiments of the present disclosure.

[0087] The RoCE protocol is extended based on the Base Transport header (BTH).Figure 5 This is the format of the error indication message extended based on the RoCE protocol in the embodiments of the present disclosure. As Figure 5 shown, the extended message based on BTH includes:

[0088] The Opcode is specified as the control message of RoCE. The OpCode uses reserved values. Bits 7 - 5 are 110 - 111, and bits 4 - 0 are 00000 - 11111. For example, Opcode = 11000000 (binary) = 192 (decimal) can be selected. When formulating this Opcode, the error indication content follows the BTH.

[0089] The PSN is directly used to represent the message sequence number that has been correctly aggregated and returned the result. Other fields in the BTH still follow the original definitions, which will not be elaborated in the embodiments of the present disclosure.

[0090] In the extended message, new error indication Error_Indication content is added, specifically including the task identifier Job_id, the aggregation group identifier Group_id, the aggregation tree identifier Tree_id, and the computing device Node_id.

[0091] To better understand the present disclosure, taking a switch as an example below, in combination with Figure 1 and Figure 6 the in-network computing error recovery method is further introduced.

[0092] Figure 6 This is the flowchart of an in-network computing error recovery method provided by the embodiments of the present disclosure. As Figure 6 shown, the in-network computing error recovery method includes:

[0093] Step S601, the controller generates multiple aggregation trees in response to the aggregation request of the end-side host, selects the aggregation tree with the optimal performance as the current aggregation tree, and uses other aggregation trees as standby aggregation trees and arranges them in the standby aggregation tree list in the order of priority.

[0094] When generating the aggregation tree, the controller also generates the resources required for each switch in the aggregation tree to complete the aggregation operation. If the optimal aggregation tree among multiple aggregation trees is the Figure 1 aggregation tree shown, then the Figure 1 aggregation tree is used as the current aggregation tree.

[0095] Step S602, determine the type of the computing device where the error occurs. If an error occurs in the end-side host, execute step S603; if an error occurs in the leaf switch, execute step S606; if an error occurs in the spine aggregation tree, execute step S610.

[0096] The computing devices in the aggregation tree perform aggregation operations according to the commands issued by the aggregation tree controller. The lowest-level switch receives the aggregation packets from the end-side hosts connected to it, performs aggregation processing, and sends the aggregation result in the form of an aggregation packet to the switch at the upper level. The upper-level switch receives the aggregation packets from the switch at the lower level connected to it, performs aggregation processing, and sends the aggregation result in the form of an aggregation packet to the switch at the upper level, and so on, until the highest-level spine switch performs aggregation processing.

[0097] Step S603: Determine whether the number of end-side hosts with errors is greater than the second configured value. If so, execute Step S604; if not, execute Step S605.

[0098] Step S604: The controller deletes the aggregation tree. After the end-side host or leaf switch resumes normal operation, the aggregation operation is restarted according to the aggregation request of the end-side host.

[0099] When the controller determines that the aggregation operation needs to be restarted, the controller not only deletes the current aggregation tree but also deletes the standby aggregation tree. After the end-side host resumes normal operation, the aggregation operation is restarted according to the aggregation request of the end-side host.

[0100] Step S605: Discard the end-side hosts with errors and continue the aggregation process.

[0101] When the number of end-side hosts with errors is less than or equal to the second configured value, the impact on the training result is small. Therefore, the leaf switch discards the end-side hosts with errors and only aggregates the aggregation packets of the end-side hosts without errors, and the aggregation process is not interrupted.

[0102] Step S606: Determine whether the number of end-side hosts connected under the leaf switch with errors is less than the first configured value. If not, execute Step S607; if so, execute Step S608.

[0103] Step S607: Whether to report to the spine switch. If so, execute Step S604; if not, execute Step S609.

[0104] When the number of end-side hosts connected under the leaf switch with errors is greater than or equal to the first configured value, the leaf switch can report to the controller or switch the aggregation tree.

[0105] When the leaf switch selects to report to the spine switch, the spine switch waits for the leaf switch to resume normal operation and then restarts the aggregation operation. At this time, the aggregation process is interrupted.

[0106] Step S608: The spine switch discards the leaf switch with errors and only aggregates the aggregation packets of other leaf switches.

[0107] Step S609: The leaf switch sends an error indication report to the spine switch. The spine switch sends the error indication report to all end-side hosts through other leaf switches, and then proceeds to Step S611.

[0108] Step S610: The spine switch sends an error indication packet to other leaf switches in the aggregation tree, and the other leaf switches then forward it to the corresponding end-side hosts.

[0109] The error indication packet carries the current task identifier, group identifier, identifier of the aggregation tree, sequence number of the packets that have been correctly aggregated, and the identifier of the computing device where the error occurred.

[0110] Step S611: After all end-side hosts receive the error indication packet, select the next aggregation tree in the backup aggregation tree list as the new current aggregation tree, and send aggregation packets from the error point to the backup aggregation tree to complete the aggregation of the remaining gradients.

[0111] Combined Figure 3 As shown, assume that an error occurs in spine switch switch4 during the aggregation process. The aggregation packets sent by leaf switches switch1 - switch3 to spine switch switch4 cannot be aggregated. The aggregation group ID for aggregation in spine switch switch4 is 1, and the sequence number of the packets that spine switch switch4 has correctly aggregated and returned the result is 150. Then, spine switch switch4 sends an error indication packet to leaf switches switch1 - switch3. The error indication packet encapsulates Tree-id as 1, PSN as 150, and Node-id as spine switch switch4. After receiving the error indication packet, leaf switches switch1, switch2, and switch3 forward it to all end-side hosts host1 - host6. When end-side hosts host1 - host6 receive the error indication packet, they know from the Tree-id that the aggregation tree is 1, and from the Node-id that the error computing device is spine switch switch4. Then, they select the aggregation tree with Tree-id 2 in the backup aggregation tree and send aggregation packets with PSN 151 from the error point to perform the aggregation of the remaining gradients. The aggregation packets continue the aggregation process along the aggregation tree with spine switch switch5 as the root node, and this in-network computing error recovery does not interrupt the aggregation process.

[0112] It should be noted that when switching the aggregation tree, if the standby aggregation trees all include the faulty leaf switches, that is, the faulty leaf switches cannot be avoided even by switching the aggregation tree, then the aggregation tree is removed. After the leaf switches resume normal operation, the aggregation operation is restarted. This in-network computing error recovery method interrupts the aggregation process.

[0113] Figure 7 This is a flowchart of another in-network computing error recovery method provided by an embodiment of the present disclosure. As Figure 7 shown, the in-network computing error recovery method includes:

[0114] Step S701, the controller generates multiple aggregation trees in response to the aggregation request of the end-side host, selects the optimal aggregation tree as the current aggregation tree, and arranges the other aggregation trees as standby aggregation trees in the standby aggregation tree list in order of priority.

[0115] Step S702, determine the type of the faulty computing device. If the end-side host is faulty, execute step S703; if the leaf switch is faulty, execute step S706; if the spine aggregation tree is faulty, execute step S710.

[0116] The switches in the aggregation tree perform aggregation operations according to the commands issued by the aggregation tree controller. The switches at the lowest level receive the aggregation packets from the end-side hosts connected to them and perform aggregation processing, and send the aggregation results in the form of aggregation packets to the switches at the upper level. The upper-level switches receive the aggregation packets from the lower-level switches connected to them and perform aggregation processing, and send the aggregation results in the form of aggregation packets to the switches at the upper level, and so on, until the switches at the highest level perform aggregation processing.

[0117] Step S703, determine whether the number of faulty end-side hosts is greater than the second configuration value. If so, execute step S704; if not, execute step S705.

[0118] Step S704, the controller removes the aggregation tree. After the end-side host or the leaf switch resumes normal operation, the aggregation operation is restarted according to the aggregation request of the end-side host.

[0119] When the controller determines that the aggregation operation needs to be restarted, the controller not only removes the current aggregation tree but also removes the standby aggregation trees. The controller will remove the aggregation tree and restart the aggregation operation according to the aggregation request of the end-side host after the end-side host resumes normal operation. At this time, the aggregation process is interrupted.

[0120] Step S705, discard the faulty end-side host and continue the aggregation process.

[0121] When the number of faulty edge hosts is less than or equal to the second configuration value, the impact on the training result is small. Therefore, the leaf switch discards the faulty edge hosts and aggregates the aggregation packets of the non-faulty edge hosts without interrupting the aggregation process.

[0122] Step S706: Determine whether the number of edge hosts connected to the faulty leaf switch is less than the first configuration value. If not, execute step S707; if so, execute step S708.

[0123] Step S707: Send an error message to the controller, and then execute step S710.

[0124] The error indication message carries the current task identifier, group identifier, identifier of the aggregation tree, sequence number of the aggregation packet that has been correctly aggregated, and identifier of the computing device where the error occurred.

[0125] Step S708: The spine switch discards the faulty leaf switch and only aggregates the aggregation packets of other leaf switches.

[0126] The spine switch discards the aggregation packets corresponding to the faulty leaf switch and only aggregates the aggregation packets of other leaf switches. This in-network computing error recovery method does not interrupt the aggregation process.

[0127] Step S709: The spine switch sends an error indication message to all leaf switches and the controller in the aggregation tree, and the leaf switch then forwards it to the corresponding edge host.

[0128] Step S710: The controller sends the backup aggregation tree to the spine switch, leaf switch, and edge host.

[0129] Step S711: After the edge host receives the backup aggregation tree, it uses the backup aggregation tree as the current aggregation tree and sends aggregation packets from the error point to the backup aggregation tree to complete the aggregation of the remaining gradients.

[0130] In a second aspect, an embodiment of the present disclosure provides an electronic device.

[0131] Figure 8 It is a block diagram of the components of an electronic device provided by an embodiment of the present disclosure. As Figure 8 shown, an electronic device provided by an embodiment of the present disclosure includes a processor 701 and a memory 702; the memory 702 stores a computer program executable by the processor 701, and when the computer program is executed by the processor 701, it implements any one of the in-network computing error recovery methods of the embodiments of the present disclosure.

[0132] In some embodiments, the electronic device further includes an I / O interface (read / write interface) 703. The I / O interface 703 is connected between the processor 701 and the memory 702 and can implement information interaction between the memory 702 and the processor 701, including but not limited to a data bus (Bus), etc.

[0133] Among them, the processor is a device with data processing capabilities, including but not limited to a central processing unit (CPU), etc.; the memory is a device with data storage capabilities, including but not limited to a random access memory (RAM, more specifically such as SDRAM, DDR, etc.), a read-only memory (ROM), an electrically erasable programmable read-only memory (EEPROM), a flash memory (FLASH); the I / O interface (read / write interface) is connected between the processor and the memory and can implement information interaction between the memory and the processor, including but not limited to a data bus (Bus), etc.

[0134] Those of ordinary skill in the art can understand that all or some of the steps, systems, and functional modules / units in the devices disclosed above can be implemented as software, firmware, hardware, and their appropriate combinations.

[0135] The embodiments of the present disclosure also provide a computer-readable medium, on which a computer program is stored. When the computer program is executed by a processor, it implements any one of the online computing error recovery methods described in the above embodiments.

[0136] The embodiments of the present disclosure also provide a computer program product, which includes a computer program. When the computer program is executed by a processor, it implements any one of the online computing error recovery methods described in the above embodiments. In a hardware implementation, the division between the functional modules / units mentioned in the above description does not necessarily correspond to the division of physical components; for example, a physical component can have multiple functions, or a function or step can be executed by several physical components in cooperation.

[0137] Those of ordinary skill in the art can understand that all or some of the steps, systems, and functional modules / units in the devices disclosed above can be implemented as software, firmware, hardware, and their appropriate combinations.

[0138] In a hardware implementation, the division between the functional modules / units mentioned in the above description does not necessarily correspond to the division of physical components; for example, a physical component can have multiple functions, or a function or step can be executed by several physical components in cooperation.

[0139] Some physical components or all physical components may be implemented as software executed by a processor, such as a central processing unit (CPU), a digital signal processor or a microprocessor, or implemented as hardware, or implemented as an integrated circuit, such as an application-specific integrated circuit. Such software may be distributed on a computer-readable medium, which may include a computer storage medium (or non-temporary medium) and a communication medium (or temporary medium). As known to those of ordinary skill in the art, the term computer storage medium includes volatile and non-volatile, removable and non-removable media implemented in any method or technology for storing information (such as computer-readable instructions, data structures, program modules or other data). Computer storage media include, but are not limited to, random access memory (RAM, more specifically SDRAM, DDR, etc.), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory (FLASH) or other disk storage; read-only compact disk (CD-ROM), digital versatile disk (DVD) or other optical disk storage; magnetic cassettes, magnetic tapes, disk storage or other magnetic storage; any other medium that can be used to store desired information and can be accessed by a computer. Furthermore, it is well known to those skilled in the art that communication media typically embodies computer readable instructions, data structures, program modules, or other data in a modulated data signal such as a carrier wave or other transport mechanism, and may include any information delivery media.

[0140] The present disclosure has disclosed example embodiments, and although specific terms are employed, they are used and should be interpreted only in a general illustrative sense and not for limiting purposes. In some instances, it will be apparent to those skilled in the art that, unless otherwise expressly stated, features, characteristics, and / or elements described in conjunction with a particular embodiment may be used alone or in combination with features, characteristics, and / or elements described in conjunction with other embodiments. Therefore, those skilled in the art will appreciate that various changes in form and detail may be made without departing from the scope of the present disclosure as set forth in the appended claims.

Claims

1. A method for recovering from an online computing error, comprising: During the aggregation process of performing on-line computing, if an error occurs in a network device in the current aggregation tree and a preset switching condition is met, the aggregation tree is switched to a backup aggregation tree, and the aggregation process is continued after switching to the backup aggregation tree; wherein the on-line computing aggregation tree includes multiple end-side hosts and multi-level network devices, the multi-level network devices are signal-connected level by level, and the end-side host is signal-connected to the network device at the lowest level.

2. The method according to claim 1, wherein: The step of switching the aggregation tree to the standby aggregation tree includes: Select an aggregation tree that does not include an erroneous network device and has the highest priority from the backup aggregation tree list as a new aggregation tree; The standby aggregation tree list is generated in advance by the controller based on the physical topology of the in-network computing and the end-side hosts participating in the in-network computing.

3. The method according to claim 2, wherein: The step of switching the aggregation tree to the standby aggregation tree includes: The backup aggregation tree list is stored in the end-side host. In response to the aggregation tree switching trigger information carried in the data plane message, the end-side host selects an aggregation tree that does not include an erroneous network device and has the highest priority from the backup aggregation tree list as a new aggregation tree.

4. The method according to claim 2, wherein: The step of switching the aggregation tree to the standby aggregation tree includes: The backup aggregation tree list is stored in the controller. In response to the aggregation tree switching trigger information, the controller selects an aggregation tree that does not include an erroneous network device and has the highest priority from the backup aggregation tree list as a new aggregation tree; and sends the new aggregation tree to the end-side host.

5. The method according to any one of claims 1 to 4, wherein: Each of the end-side hosts performs the switching.

6. The method according to any one of claims 1 to 4, wherein: The end-side host that is aggregated and connected to the wrong network device performs the switching, while the other end-side hosts do not perform the switching.

7. The method according to any one of claims 1 to 4, wherein: The preset switching conditions include: The computing device where the error occurs is a network device at a level below the highest level, and the number of the end-side hosts connected to the network device where the error occurs is greater than or equal to a preset first configuration value.

8. The method according to claim 3, wherein: The data plane message includes an error indication message, and the error indication message includes one or more of a task identifier, a group identifier, an aggregation tree identifier, a sequence number of a message that has completed aggregation correctly, and an identifier of a network device where an error occurs.

9. The method according to claim 8, wherein: The error indication message is extended by using the data packet transmission service header and semantic sublayer header of the Hyper Ethernet transport layer, or by using the basic transport header.

10. An electronic device comprising a memory and a processor; the memory stores a computer program executable by the processor, and the computer program, when executed by the processor, implements the on-line computing error recovery method as described in any one of claims 1 to 9.

11. A computer program product, comprising a computer program, wherein when the computer program is executed by a processor, the method for recovering from an error in network computing according to any one of claims 1 to 9 is implemented.