Message transmission method and apparatus, and computing device
By enabling the target node in the ring topology network to autonomously select a backup data path to transmit messages, the service interruption and efficiency reduction problems caused by faulty nodes are solved, and efficient message transmission and network reliability are achieved.
Patent Information
- Application Number
- CN202410444393.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-04-12
- Publication Date
- 2025-10-21
AI Technical Summary
During message transmission in a ring topology network, faulty nodes can cause service interruption and reduced transmission efficiency.
When a data path failure is detected, the target node directly redetermines at least one second data path from multiple data paths for message transmission, avoiding reporting the failure information to the management node and redetermining a reachable path using existing routing information.
It avoids interruption of message transmission services, improves transmission efficiency and network reliability, and enhances tolerance to faulty nodes.
Smart Images

Figure CN120825529A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of networking technology, and in particular to a message transmission method, apparatus, and computing device. Background Art
[0002] Due to its low cost, torus topology networks are widely used in data centers. However, when a torus topology network experiences a fault, the faulty node (hereinafter referred to as the faulty node) reports the fault to the torus topology network's management node. Upon receiving the fault information, the management node re-issues routing information to the torus topology network, indicating the reachable paths between different nodes on the torus topology network. This allows the torus topology network to determine the reachable paths for packets when it needs to transmit them, using the newly issued routing information.
[0003] However, if a faulty node occurs on the data path of the message during the message transmission process, the message transmission service will be affected by the faulty node and cause service interruption, which seriously reduces the message transmission efficiency. Summary of the Invention
[0004] The present application provides a message transmission method, apparatus, and computing device, which can prevent message transmission services from being affected by faults, thereby avoiding service interruptions caused by faults, and further ensuring message transmission efficiency when faults occur.
[0005] In a first aspect, a message transmission method is provided, which is applied to a torus topology network. The torus topology network includes multiple nodes, each of the multiple nodes storing target routing information, the target routing information being used to indicate multiple data paths in the torus topology network. The method comprises: a target node among the multiple nodes determines a first data path from multiple data paths for transmitting a first message; if a fault occurs on the first data path, the target node determines at least one second data path from the multiple data paths; and the target node transmits the first message via the at least one second data path.
[0006] In this solution, the torus topology network stores target routing information, which indicates multiple data paths on the torus topology network. Based on this information, after a target node determines a first data path from the multiple data paths for transmitting a first message, if a failure occurs on the first data path, the target node can redetermine at least one second data path from the multiple data paths and transmit the first message via the redetermined at least one second data path.
[0007] When a fault occurs on the first data path, the target node directly uses the multiple data paths indicated by the currently stored target routing information to redetermine a data path (i.e., at least one second data path) that can reach the destination node of the first message, and transmits the first message through the at least one redetermined second data path. In this way, the ring topology network does not need to report fault information to the management node to trigger the management node to re-issue routing information, thereby determining the data path for transmitting the message based on the re-issued routing information. Therefore, the present application avoids the interruption of message transmission services caused by a fault in the data path of the message to be transmitted, thereby avoiding the impact of the fault on the data path of the message to be transmitted on the message transmission efficiency, and thus can ensure that the message transmission efficiency will not be seriously reduced when a fault occurs in the data path of the message to be transmitted. In addition, by determining multiple second data paths for transmitting the first message, the number of faulty nodes that can be tolerated during the message transmission process can be increased, thereby helping to improve the network reliability of the ring topology network.
[0008] In one possible implementation, the destination node is the source node of the first message, or the destination node is a node on the first data path. This not only helps to increase the probability of successful message transmission, but also helps to increase the node range applicable to the message transmission method improved in this application.
[0009] In another possible implementation, if the destination node is a node on the first data path, the destination node determines at least one second data path from multiple data paths, including: the destination node determines at least one second data path from multiple target data paths; the multiple target data paths include the data path between the destination node and the destination node of the first message. This not only helps reduce the number of hops in the second data path, but also helps expand the selection range of the second data path.
[0010] In another possible implementation, if the destination node is a node on the first data path, the destination node determines at least one second data path from multiple data paths, including: the destination node determines at least one second data path from multiple fourth data paths; the multiple fourth data paths include data paths between the source node of the first message and the destination node of the first message. This not only helps ensure the accuracy of the second data path, that is, accurately transmit the first message to the destination node, but also helps simplify the logic of the destination routing algorithm.
[0011] In another possible implementation, the target node transmits the first message through at least one second data path, including: the target node updates the first message to obtain the second message; the source node indicated by the second message is the target node, and the destination node indicated by the second message is the destination node of the first message; the target node transmits the second message through at least one second data path.
[0012] In another possible implementation, the second message includes a fault identifier that indicates the number of times a fault was detected during transmission of the first message. In this way, a node receiving the second message can determine, based on the fault identifier, whether a fault was detected during transmission of the original message of the second message. Furthermore, the node can determine whether to discard the second message based on the number of times a fault was detected, thereby helping to ensure the accuracy of the message discard operation.
[0013] In another possible implementation, the method further includes: the target node transmitting a target message to at least the first node of the second data path; the target message is used to indicate that the source node of the first message is the target node.
[0014] In another possible implementation, the target message is further used to indicate the number of times a fault is detected during the transmission of the first message.
[0015] In another possible implementation, the at least one second data path includes a target second data path, and the target second data path includes the first node. The method further includes: if the target second data path fails, the first node discards the received message. This helps avoid deadlock in a ring topology network.
[0016] In another possible implementation, the failure of the first data path includes at least one of a failure of a first port on the target node, a failure of a second port on the destination node of the first message, a failure of a second node on the first data path, and a failure of a third node on the first data path. The first port is connected to a node on the first data path, the second port is connected to a node on the first data path, the second node is the next hop node of the target node, the third node is separated from the target node by at least one node, and the at least one node is located on the first data path. This helps ensure the reliability of the first data path, thereby helping to increase the probability of successful transmission of the first message.
[0017] In another possible implementation, the at least one second data path includes multiple second data paths, and different second data paths in the multiple second data paths do not have the same nodes. This helps to increase the number of faulty nodes that can be tolerated during the transmission of the first message.
[0018] In another possible implementation, each of the at least one second data path includes a fourth node, which is a next-hop node of the target node and a non-faulty node. This helps improve the success rate of message transmission.
[0019] In another possible implementation, the method further includes: the target node receiving a target value; and the target value being used to determine the number of at least one second data path. In this way, the user can specify the number of second data paths by sending the target value to the target node.
[0020] In another possible implementation, the target routing information indicates that M data paths exist between the target node and the destination node of the first message, where M is a positive integer greater than 1. The target node determines at least one second data path from the multiple data paths, including: the target node determines K second data paths with the fewest hops from the M data paths, where K is a positive integer less than or equal to a target value. This helps improve message transmission efficiency.
[0021] In another possible implementation, the target routing information indicates that M data paths exist between the target node and the destination node of the first message, where M is a positive integer greater than 1; and the target node determines at least one second data path from the multiple data paths, including: the target node randomly determines K second data paths from the M data paths. This helps increase the diversity of selection of the second data paths.
[0022] In a second aspect, a message transmission device is provided, comprising: functional units for executing any one of the methods provided in the first aspect, wherein the actions performed by each functional unit are implemented via hardware or via hardware executing corresponding software implementations. For example, the message transmission device comprises: a routing module and a transmission module. The routing module is configured to determine, from a plurality of data paths, a first data path for transmitting a first message; the routing module is further configured to determine, from the plurality of data paths, at least one second data path if a fault occurs on the first data path; and the transmission module is configured to transmit the first message via the at least one second data path.
[0023] In a third aspect, a computing device is provided, which may include a processor, a memory, and a computer program / instructions stored in the memory; the processor executes the computer program to enable the computing device to implement the steps of any one of the methods provided in the first aspect above.
[0024] In a fourth aspect, a node of a torus topology network is provided, the node including a processor, a memory and a computer program / instructions stored in the memory; the processor executes the computer program so that the node implements the steps of any one of the methods provided in the first aspect above.
[0025] In the fifth aspect, a torus topology network is provided, which includes multiple nodes, each of which includes a processor, a memory, and a computer program / instructions stored in the memory; the processor executes the computer program to enable the torus topology network to implement the steps of any one of the methods provided in the first aspect above.
[0026] In a sixth aspect, a computer program product is provided, comprising a computer program / instruction, which implements the steps of any one of the methods provided in the first aspect when the computer program / instruction is executed by a processor.
[0027] In a seventh aspect, a computer-readable storage medium is provided, on which a computer program / instruction is stored. When the computer program / instruction is executed by a processor, the steps of any one of the methods provided in the first aspect are implemented.
[0028] Among them, the technical effects brought about by any implementation method in the second to seventh aspects can be referred to the technical effects brought about by different implementation methods in the above-mentioned first aspect, and will not be repeated here. BRIEF DESCRIPTION OF THE DRAWINGS
[0029] Figure 1 A schematic diagram of a deadlock provided for this application;
[0030] Figure 2 A schematic diagram of a storage network provided in this application;
[0031] Figure 3 A schematic diagram of a system architecture provided for this application;
[0032] Figure 4 A schematic diagram of another system architecture provided for this application;
[0033] Figure 5 A flowchart of a message transmission method provided by this application;
[0034] Figure 6 A schematic diagram of a torus topology network provided in this application;
[0035] Figure 7 A flowchart of another message transmission method provided by this application;
[0036] Figure 8 A schematic diagram of another torus topology network provided in this application;
[0037] Figure 9 A flowchart of another message transmission method provided by this application;
[0038] Figure 10 A schematic diagram of a message transmission device provided in this application. DETAILED DESCRIPTION
[0039] To facilitate understanding, the relevant terms involved in this application are first briefly introduced.
[0040] Data Paths (DP): refers to the path for transmitting messages between different components.
[0041] In this application, a data path refers to a path for transmitting messages between different nodes on a torus topology network.
[0042] Multiple Data Paths (MDP): refers to multiple paths for transmitting messages between different components.
[0043] A torus topology is a network in which multiple nodes are arranged in a regular grid. Adjacent nodes in the same row or column are connected, and the two most distant nodes in the same row or column are also connected. Each row or column in a torus topology is a ring.
[0044] Virtual channel (VC): The buffer of the network card device can be divided into multiple queues. A queue can be called a virtual channel. A queue is used to store messages received by the network card device / messages to be sent.
[0045] Deadlock: refers to a state in which multiple resources are waiting for each other and cannot continue execution.
[0046] In this application, deadlock refers to a state in which multiple nodes on a torus topology network are waiting for each other and cannot continue to transmit messages.
[0047] For example, Figure 1 As shown in the figure, nodes 1, 2, 3, and 4 are each configured with one virtual channel. Node 1's virtual channel 1 is occupied by message 1, node 2's virtual channel 1 is occupied by message 2, node 3's virtual channel 1 is occupied by message 3, and node 4's virtual channel 1 is occupied by message 4. Node 1 waits for node 2 to release virtual channel 1 so it can transmit message 1 to node 2. Node 2 waits for node 3 to release virtual channel 1 so it can transmit message 2 to node 3. Node 3 waits for node 4 to release virtual channel 1 so it can transmit message 3 to node 4. Node 4 waits for node 1 to release virtual channel 1 so it can transmit message 4 to node 1. This creates a circular dependency between nodes 1, 2, 3, and 4, preventing all messages from being transmitted and causing a deadlock.
[0048] Hop count: refers to the total number of hops required for a message to be routed from the source node to the destination node. Routing from one node to the next is called a hop.
[0049] The technical solution provided in this application will be described in detail below with reference to the accompanying drawings.
[0050] Currently, the networking forms of data centers include fat tree topology, spine-leaf topology, etc. Figure 2 As shown, a data center may include a computing network and a storage network. The computing network may include multiple computing nodes, and the storage network may include multiple storage nodes. The following uses the storage network as an example to illustrate the current networking form.
[0051] To differentiate north-south traffic from east-west traffic within the storage network, the storage network is divided into a front-end network and a back-end network. The front-end network is responsible for north-south communication between user compute nodes and storage nodes, such as transmitting read and write traffic. The back-end network is responsible for data communication between storage nodes, such as erasure codes (EC) and garbage collection (GC). In this networking configuration, when the number of storage nodes exceeds the port limit of a single switch, the storage nodes are connected to the front-end switch, which is then connected to the aggregation switch, which is then connected to the compute nodes. As can be seen, when the number of nodes is large, this networking configuration requires more switches, resulting in higher networking costs.
[0052] In order to reduce costs, related technologies have proposed a ring topology network (also known as a ring direct connection topology network). Since the ring topology network directly connects different nodes, that is, there is no need to connect different nodes through switches, the cost of networking is greatly reduced.
[0053] In a torus topology, multiple reachable paths exist between any two nodes, meaning multiple data paths exist between any two nodes. To ensure normal communication in a torus topology, a routing algorithm is used to determine the paths (i.e., data paths) used between different nodes. Currently, torus topology networks use routing algorithms that select the shortest path for communication, such as deterministic dimension-ordered routing (DOR) and shortest path first (SPF).
[0054] When a torus topology network using the above-mentioned routing algorithm experiences a fault, the faulty node will report the fault information to the torus topology network's management node. After receiving the fault information, the management node will re-issue routing information to the torus topology network, indicating the reachable paths between different nodes. In other words, the paths between different nodes do not include the faulty node. In this way, when the torus topology network needs to transmit a message, it can determine the reachable path for the message using the newly issued routing information. However, if a fault occurs on the data path used to transmit the message during the message transmission process, the torus topology network needs to wait for the faulty node to report the fault information and re-determine the reachable path based on the new routing information. Therefore, the message transmission service will be affected by the fault, resulting in service interruption, which in turn seriously reduces the message transmission efficiency.
[0055] In view of this, the present application provides a message transmission method, which is applied to a torus topology network. The torus topology network stores routing information, and the routing information indicates multiple data paths on the torus topology network. Based on this, after a destination node on the torus topology network determines a first data path from multiple data paths for transmitting a first message, if a failure occurs on the first data path, the destination node can continue to re-determine at least one second data path from the multiple data paths for transmitting the first message.
[0056] When a fault occurs on the first data path, the target node directly uses the multiple data paths indicated by the currently stored target routing information to redetermine a data path (i.e., at least one second data path) that can reach the destination node of the first message, and transmits the first message via the at least one redetermined second data path. Thus, the ring topology network does not need to report fault information to the management node to trigger the management node to re-issue routing information, thereby determining the data path for transmitting the message based on the re-issued routing information. Therefore, the message transmission method provided in this application avoids interruption of message transmission services caused by faults in the data path of the message to be transmitted, thereby preventing faults in the data path of the message to be transmitted from affecting message transmission efficiency, thereby ensuring that message transmission efficiency will not be severely reduced when a fault occurs in the data path of the message to be transmitted. Furthermore, by determining multiple second data paths for transmitting the first message, the number of faulty nodes that can be tolerated during message transmission can be increased, thereby improving the network reliability of the ring topology network.
[0057] Next, the system architecture involved in the technical solution provided in this application is further introduced with reference to the accompanying drawings.
[0058] The system architecture provided in this application is applicable to the above-mentioned message transmission method. The system architecture may include a management node and a ring topology network. The ring topology network may include multiple nodes, each of which may communicate with the management node.
[0059] Each node may include a storage medium, which may be referred to as a local storage medium of the node. The local storage medium of the node may be used to store routing information sent by the management node.
[0060] It should be noted that in this application, multiple can refer to two or more. In addition, the management node can be called the management plane or control plane, and the ring topology network can be called the data plane, which will not be repeated hereafter.
[0061] Illustratively, each node on the ring topology network may communicate with the management node through a switch, or each node on the ring topology network may also communicate directly with the management node, which is not limited in the present application.
[0062] For example, the torus topology network can be a two-dimensional (2D) network, a three-dimensional (3D) network, etc., and this application does not limit this. Each node in the 2D torus topology network can include 4 ports, and each node in the 3D torus topology network can include 6 ports, wherein the ports of a node in the torus topology network can be used to connect to other nodes in the torus topology network.
[0063] In this application, the management node can be used to manage nodes on a ring topology network, such as sending routing information to the nodes on the ring topology network. The routing information can be used to indicate multiple reachable paths (i.e., data paths) between every two nodes on the ring topology network. The nodes on the ring topology network can be used to execute the message transmission method provided in this application, thereby preventing the message transmission service from being affected by the fault, thereby avoiding service interruption caused by the fault, and ensuring that the message transmission efficiency will not be seriously reduced when a fault occurs.
[0064] Optionally, the nodes in the torus topology network may be computing devices, etc. In this application, computing devices may be storage nodes, computing nodes, etc. wherein computing nodes are used to perform computing tasks, and storage nodes are used to perform storage tasks.
[0065] In this application, a computing device may be a server or an electronic device.
[0066] A server can be a single physical or logical server, or it can be two or more physical or logical servers sharing different responsibilities and working together to implement the server's functions. For example, a server can be a blade server, a high-density server, a rack server, or a tower server. Electronic devices can include ultra-mobile personal computers (UMPCs), laptops, netbooks, desktop computers, all-in-one computers, and the like.
[0067] It should be noted that the present application does not limit the types of nodes in the torus topology network, and the above is only an exemplary description. Below, the present application is exemplified by taking the nodes in the torus topology network as computing devices as an example.
[0068] In the present application, a node on a ring topology network may include a network card device, which may be used to execute the message transmission method provided in the present application. The network card device may include a storage medium, which may be referred to as a local storage medium of the network card device. The local storage medium of the network card device may be used to store routing information issued by a management node.
[0069] Optionally, the network card device may be a programmable network card, a data processing unit (DPU) network card, or the like.
[0070] It should be noted that this application does not limit the device type of the network card device, and the above is only an exemplary description.
[0071] Optionally, the management node may be a server or an electronic device.
[0072] Figure 3 A schematic diagram of a system architecture provided for this application.
[0073] like Figure 3 As shown, the system architecture may include a computing network, a front-end network of a storage network, and a back-end network of the storage network. The computing network may communicate with storage nodes in the back-end network via the front-end network.
[0074] exist Figure 3 In the illustrated system architecture, the backend network includes multiple storage nodes, such as Node 1, ..., Node 16. These storage nodes form a torus topology, enabling multiple data paths between any two storage nodes. The frontend network includes multiple compute nodes, which can communicate with multiple storage nodes via the frontend network.
[0075] Optionally, the system architecture may further include a management node and a switch. The management node may communicate with each storage node on the ring topology network through the switch to send routing information to each storage node.
[0076] Optionally, the system architecture may further include an electronic device, which may communicate with the management node.
[0077] In this application, a user may send a target value to a management node via an electronic device. The target value may be used to specify the maximum number of faulty nodes supported by a torus topology network. The storage node may determine the number of at least one second data path based on the target value.
[0078] Need to explain, Figure 3 The backend network shown includes 16 storage nodes, which is only for exemplary purposes. This application does not limit the number of nodes in the backend network, that is, this application does not limit the number of nodes in the ring topology network.
[0079] Figure 4 A schematic diagram of another system architecture provided for this application.
[0080] like Figure 4 As shown, the system architecture may include a computing network, a front-end network of a storage network, and a back-end network of the storage network. The computing network may communicate with storage nodes in the back-end network through the front-end network.
[0081] exist Figure 4 In the illustrated system architecture, the computing network includes multiple computing nodes, such as nodes 1, ..., and 16. These computing nodes form a torus topology, enabling multiple data paths between any two computing nodes. The backend network can include multiple storage nodes, which can communicate with the computing nodes via the frontend network.
[0082] Need to explain, Figure 4 For other related descriptions of the system architecture shown, please refer to Figure 3 The description of the system architecture shown will not be repeated here.
[0083] Need to explain, Figure 3 、 Figure 4 The system architecture shown is merely an exemplary description of the system architecture applicable to the message transmission method provided in this application, and does not constitute a limitation on the system architecture for executing the message transmission method provided in this application.
[0084] The present application also provides a target routing algorithm that can be used to implement the message transmission method provided by the present application. The target routing algorithm can be run by nodes on a torus topology network. In other words, the nodes on the torus topology network can implement the message transmission method provided by the present application by running the target routing algorithm.
[0085] Exemplarily, the target routing algorithm can be run by a network card device on a node, thereby implementing the message transmission method provided in this application.
[0086] It should be noted that the system architecture and application scenarios described in this application are intended to more clearly illustrate the technical solutions of this application and do not constitute a limitation on the technical solutions provided in this application. Ordinary technicians in this field will know that with the evolution of the system architecture and the emergence of new application scenarios, the technical solutions provided in this application are also applicable to similar technical problems.
[0087] For ease of understanding, the message transmission method provided in this application is exemplarily introduced below in combination with the above system architecture and accompanying drawings.
[0088] This application will be divided into three parts to provide an exemplary introduction to the scheme of the message transmission method.
[0089] Part I: Combining Figures 5 and 6 , introduces a scheme in which the source node of the first message acts as the target node to transmit the message.
[0090] Part II, Combination Figures 7 and 8 , introducing a scheme in which a node on the first data path acts as a target node to transmit a message.
[0091] Part III: Combination Figure 9 , introduces a scheme for transmitting messages by nodes in the second data path.
[0092] The following, combined Figures 5 and 6 , introducing the first part of this application.
[0093] Figure 5 This is a flow chart of a message transmission method provided by the present application. Exemplarily, the method may include the following steps 501-505.
[0094] It should be noted that the "step" in this application can be abbreviated as "S" and will not be repeated later.
[0095] Figure 5 The message transmission method shown is suitable for Figure 3 、 Figure 4 The torus topology shown in the figure is as follows. Figure 3 The present application is exemplified by taking the torus topology network shown as an example.
[0096] In this application, the management node may send target routing information to the torus topology network, and the target routing information may be used to indicate multiple data paths on the torus topology network, wherein the multiple data paths include a data path between every two nodes on the torus topology network.
[0097] For example, Figure 3 As shown, the management node can send target routing information to each node in the torus topology network through a switch. After receiving the target routing information, each node stores the target routing information in a storage medium. For example, the management node can send the target routing information to the network card device of each node through a switch. After receiving the target routing information, the network card device of each node can store the target routing information on the node.
[0098] In one example, the target routing information may be stored in a local storage medium of each node. In another example, the target routing information may be stored in a local storage medium of a network card device of each node.
[0099] It should be noted that this application does not limit the storage location of the target routing information, and the above is only an exemplary description. In addition, for the method of generating the target routing information by the management node and the method of sending the routing information to the ring topology network, reference can be made to the methods in the relevant technology, and this application does not limit this.
[0100] Hereinafter, the source node of the first message is referred to as the first source node, and the destination node of the first message is referred to as the first destination node, to illustrate the present application. The first source node may also be referred to as the initial source node, which will not be described in detail later.
[0101] In this application, the operations performed by the first source node, such as steps 501 to 505, etc., can be performed by the network card device on the first source node and will not be described in detail later.
[0102] Step 501: A first source node determines a first data path from multiple data paths for transmitting a first message.
[0103] In the present application, after the first source node (ie, the target node) generates the first message, it can determine the first data path for transmitting the first message based on multiple data paths indicated by the target routing information stored on the node.
[0104] The operations performed by the first source node, such as steps 501 to 505, may be operations performed by the first source node by running a target routing algorithm, and will not be described in detail later.
[0105] Exemplarily, the multiple data paths indicated by the target routing information include multiple fourth data paths, and the multiple fourth data paths include data paths between the first source node and the first destination node. Based on this, the first source node can determine the first data path from the multiple fourth data paths indicated by the target routing information.
[0106] In one example, the first data path can be the data path with the fewest hops among the plurality of fourth data paths. This helps improve the transmission efficiency of the first message. In another example, the first data path can be any one of the plurality of fourth data paths. This helps increase the selection diversity of the first data path.
[0107] It should be noted that the present application does not limit the specific algorithm for the first source node to determine the first data path, and the above is only an exemplary description.
[0108] For example, Figure 6 As shown in (a), node 9 is the first source node and node 4 is the first destination node. Node 9 stores target routing information, which indicates multiple fourth data paths between node 9 and node 4. For example, the multiple fourth data paths may include data path 1, data path 2, data path 3, and data path 5. Data path 1 includes nodes 10, 11, 12, and 8; data path 2 includes nodes 5, 6, 7, and 3; data path 3 includes nodes 13, 14, 15, and 16; data path 4 includes nodes 5 and 8; and data path 5 includes nodes 13 and 1. Based on this, after node 9 generates the first message, it determines data path 1 (i.e., the first data path) from data paths 1 to 5 to transmit the first message.
[0109] Step 502: The first source node determines whether there is a fault on the first data path.
[0110] If the judgment result is no, then execute step 503. If the judgment result is yes, then execute step 504.
[0111] The first data path includes at least one of a first port on the first source node, a second port on the first destination node, a second node on the first data path, and a third node on the first data path. The first port is connected to a node on the first data path, the second port is connected to a node on the first data path, the second node is the next hop node of the first source node, the third node is separated from the first source node by at least one node, and the at least one node is located on the first data path.
[0112] In the present application, after determining the first data path, the first source node can determine whether there is a fault on the first data path. If there is no fault, the first message is transmitted through the first data path. If there is a fault, at least one second data path is re-determined from the multiple data paths indicated by the target routing information for transmitting the message.
[0113] Optionally, step 502 includes multiple implementation methods, and the following is exemplified by methods 1 to 4.
[0114] Mode 1: The fault on the first data path includes a fault on the first port. Based on this, step 502 may include: the first source node determining whether the first port has a fault. If the first port has a fault, the result of step 502 is yes. If the first port has no fault, the result of step 502 is no.
[0115] In the present application, the first source node stores status information of multiple ports, and the status information of the multiple ports includes status information of each port on the first source node. The status information of a port is used to indicate whether a port is in a fault state or a non-fault state. After the first source node determines that the first data path is used to transmit the message, it determines the port on the first source node connected to the first data path, thereby obtaining the first port. On this basis, the first source node determines the status information of the first state from the status information of the multiple ports, and judges whether the first port has a fault based on the status information of the first port. If the status information of the first port indicates a fault state, then the first port has a fault; if the status information of the first port indicates a non-fault state, then the first port does not have a fault.
[0116] For example, Figure 6 As shown in (b) of FIG1 , node 9 stores status information for multiple ports, including the status information for each port of node 9. After node 9 determines that data path 1 is used to transmit the first message, it determines that port 1 (i.e., the first port) is connected to node 10. Node 9 then determines the status information for port 1 from the status information for multiple ports and, based on the status information for port 1, determines whether port 1 has a fault.
[0117] In this application, different ports have different port identifiers, and the state information of multiple ports is stored in association with the port identifiers of the multiple ports. Based on this, the first source node can filter the state information of the first port from the state information of the multiple ports according to the port identifier of the first port.
[0118] In this implementation, when a fault occurs at the first port, the first data path is determined to be faulty, thereby re-determining the data path for transmitting the first message. This helps to improve the coverage of the solution of the present application, thereby helping to improve the message transmission success rate.
[0119] Mode 2: The fault on the first data path includes a fault on the second port. Based on this, step 502 may include: the first source node determining whether the second port has a fault. If the second port has a fault, the determination result of step 502 is yes. If the second port has no fault, the determination result of step 502 is no.
[0120] It should be noted that for other related instructions of Method 2, please refer to the instructions of Method 1 and will not be repeated here.
[0121] In this implementation, when a fault occurs on the second port, it is determined that a fault exists on the first data path, thereby redetermining the data path for transmitting the first message. This not only helps to improve the coverage of the solution of the present application, thereby helping to improve the message transmission success rate, but also helps to redetermine the data path for transmitting the message as soon as possible when a fault occurs on the first data path, thereby improving the message transmission success rate.
[0122] Mode 3: The fault on the first data path includes a fault on the second node. Based on this, step 502 may include: the first source node determines whether the second node has a fault.
[0123] If the second node has a fault, the judgment result of step 502 is yes. If the second node has no fault, the judgment result of step 502 is no. If the second node has a fault, the second node can be considered a faulty node. If the second node has no fault, the second node can be considered a normal node.
[0124] In this method, when a fault occurs in the next-hop node of the first source node, it is determined that a fault exists on the first data path. In this way, at least one second data path can be determined by the previous-hop node of the faulty node to continue transmitting the message through the at least one second data path, thereby helping to reduce the total number of hops when the at least one second data path transmits the first message, that is, the sum of the hops of different second data paths to transmit the first message, thereby helping to reduce the number of nodes occupied when transmitting the first message and improving the network performance of the ring topology network.
[0125] In this application, after the first source node determines the first data path, it determines the next hop node from at least one node on the first data path and determines whether the next hop node has a fault. Figure 6 As shown in (b) of FIG, after node 9 determines that data path 1 is used to transmit the first message, it determines the next hop node to be node 10 from among nodes 10, 11, 12, and 8, and determines whether node 10 is faulty. Because node 10 is faulty, node 9's judgment result is yes.
[0126] The following is an illustrative introduction to the process of determining whether a fault occurs at the first source node through two implementation methods.
[0127] Optionally, a first implementation manner in which the first source node determines whether a fault exists may include S1 to S2.
[0128] S1: The first source node sends a detection signal to the second node. The detection signal is used to detect whether the second node has a fault.
[0129] In this application, the first source node determines the first data path, and after determining that the second node on the first data path is the next hop node, sends a detection signal to the second node to determine whether the second node has a fault. Figure 6 As shown in (b) in FIG, node 9 sends a detection signal to node 10 to detect whether node 10 has a fault.
[0130] S2: If no response signal is received from the second node within the first time period, the first source node determines that a fault occurs in the second node.
[0131] In this application, after a first source node sends a detection signal to a second node, the second node can receive the detection signal and return a response signal to the first source node. Based on this, after the first source node sends a detection signal, it can determine whether a response signal from the second node is received within a first duration. If no response signal is received within the first duration, it indicates that the second node is unable to respond to the first source node normally, and a fault is determined in the second node. Conversely, if a response signal is received within the first duration, it is determined that the fourth node is not faulty.
[0132] Optionally, the second node includes a third port, and the first source node and the second node are connected via a first link. The second node failure may include at least one of a third port failure, a first link failure, and a software / hardware failure of the second node.
[0133] In one example, a second node receives a detection signal sent by a first source node via a first link via a third port. If a fault occurs on the third port or the first link, the second node cannot receive the detection signal. In this case, the second node will not return a response signal to the first source node. As a result, the first source node will not be able to receive a response signal from the second node within a first time period, and the second node will be determined to be faulty.
[0134] In another example, the second node can successfully receive the detection signal sent by the first source node. However, due to a software or hardware failure, the second node cannot return a response signal to the first source node. Based on this, the first source node will not be able to receive the response signal returned by the second node within the first time period, and the second node will be determined to be faulty.
[0135] It should be noted that this application does not limit the fault type of the second node, and the above is only an exemplary description.
[0136] For example, Figure 6 As shown in (b), after node 9 sends a detection signal to node 10, node 10 is unable to return a response signal to node 9 due to a fault. Based on this, node 9 is determined to have a fault because it did not receive a response signal returned by node 10 within the first time period.
[0137] It should be noted that this application does not restrict the types of detection signals and response signals. Reference can be made to the types of signals used to detect faulty nodes in related technologies, which will not be detailed here. Furthermore, this application does not restrict the specific value of the first duration, which can be set based on the actual scenario. For example, it can be set to a few milliseconds, tens of milliseconds, etc.
[0138] In the above implementation, the first source node determines in real time whether the second node has a fault by detecting the signal. Since the second node has a fault in real time, it helps to ensure the accuracy of the detection result.
[0139] Optionally, the second implementation manner in which the first source node determines whether a fault exists may include S3 to S4.
[0140] In the present application, a first source node stores status information of multiple nodes on a torus topology network. The status information of a node is used to indicate a fault state or a non-fault state. Exemplarily, the status information of the multiple nodes may be stored in a storage medium of the first source node or in a storage medium of a network card device of the first source node.
[0141] S3: The first source node determines the status information of the second node from the status information of multiple nodes.
[0142] In the present application, after the first source node determines the first data path, it filters out the status information of the second node from the status information of multiple nodes.
[0143] For example, Figure 6 As shown in (b), node 9 stores the status information of node 1, ..., and node 16. After node 9 determines the first data path and determines that the next hop node is node 10, it filters out the status information of node 10 from the status information of node 1, ..., and node 16.
[0144] In this application, different nodes in a torus topology network have different node identifiers, and the state information of multiple nodes can be stored in association with the node identifiers of the multiple nodes. Based on this, the first source node can filter out the state information of node 10 from the state information of multiple nodes based on the node identifier of node 10.
[0145] S4: The first source node determines that a fault exists in the second node according to the fault status indicated by the status information of the second node.
[0146] In the present application, if the status information of the second node indicates a fault state, the first source node determines that the second node has a fault. If the status information of the second node indicates a non-fault state, the first source node determines that the second node has no fault.
[0147] For example, Figure 6 As shown in (b) in FIG, the status information of the node 10 indicates a fault state, and the node 9 determines that the node 10 has a fault.
[0148] In the above implementation, the first source node determines whether the second node is faulty based on the status information of multiple nodes, which helps to improve the efficiency of determining the faulty node.
[0149] It should be noted that the present application does not limit the method by which the first source node determines whether a fault exists, and the above is merely an exemplary description.
[0150] In this application, the initial state indicated by the state information of each node in the plurality of nodes is a non-fault state. Figure 6 Nodes 9 and 10 shown in (b) in FIG. 1 are used to exemplify the process of updating the status information of multiple nodes.
[0151] Exemplarily, among the status information of multiple nodes stored on node 9, the status information of node 10 indicates an initial state that is a non-faulty state. Based on this, node 9 sends a detection signal to node 10 according to a preset period to determine whether node 10 is a faulty node. After node 9 sends a detection signal to node 10 every period, if no response signal is received from node 10 within a first time period, node 9 determines that node 10 is in a faulty state. Based on this, node 9 updates the status information of node 10 stored in the node to obtain updated status information of node 10, and the updated status information of node 10 indicates a faulty state.
[0152] In addition, node 9 may also send updated status information of node 10 to nodes 1-8 and nodes 11-16. After receiving the updated status information of node 10, nodes 1-8 and nodes 11-16 store the updated status information of node 10. For example, the updated status information of node 10 may be used to overwrite the historical status information of node 10.
[0153] It should be noted that this application does not limit the method of storing update status information, and the above is only an exemplary description.
[0154] In addition, the present application does not limit the duration of the preset period, and can be set according to the actual scenario. For example, the preset period can be a few milliseconds, tens of milliseconds, etc.
[0155] It should be noted that for other related instructions of Method 3, please refer to the instructions of Method 1 to Method 2, which will not be repeated here.
[0156] Mode 4: The fault on the first data path includes a fault on the third node. Based on this, step 502 may include: the first source node determines whether the third node has a fault.
[0157] If the third node has a fault, the judgment result of step 502 is yes. If the third node has no fault, the judgment result of step 502 is no. If the third node has a fault, the third node can be considered a faulty node. If the third node has no fault, the third node can be considered a normal node.
[0158] In this application, the third node refers to a type of node, which includes other nodes on the first data path except the second node. Based on this, the first node may include at least one node. Figure 6 As shown in (a), the third node includes node 11, node 11 and node 8.
[0159] In this application, after the first source node determines the first data path, it determines the third node from multiple nodes on the first data path and determines whether the third node has a fault. Figure 6 As shown in (c) of FIG5 , after node 9 determines that data path 1 is used to transmit the first message, it can determine whether nodes 11, 12, and 8 are faulty. If none of nodes 11, 12, and 8 are faulty, the result of step 502 is no. Conversely, if at least one of nodes 11, 12, and 8 is faulty, the result of step 502 is yes.
[0160] In this manner, when a fault occurs in the third node on the first data path, it is determined that a fault occurs in the first data path. This not only helps to improve the coverage of the solution of the present application, thereby helping to improve the message transmission success rate, but also helps to re-determine the data path for transmitting the message as soon as possible when a fault occurs in the first data path, thereby further improving the message transmission success rate.
[0161] Optionally, the first source node may determine whether the third node is faulty based on the status information of the third node.
[0162] In the present application, after the first source node determines the first data path, it filters the status information of the third node on the first data path from the status information of multiple nodes, and determines whether the third node on the first data path is faulty based on the status information of the third node on the first data path. If the status information of at least some of the third nodes indicates a faulty state, it is determined that the third node is faulty. If the status information of all the third nodes indicates a non-faulty state, it is determined that the third node is not faulty.
[0163] For example, Figure 6As shown in (c) in FIG, node 9 stores the status information of nodes 1, ..., and 16. After determining the first data path, node 9 filters the status information of nodes 11, 12, and 8 from the status information of nodes 1, ..., and 16. The status information of node 12 indicates a fault state, while the status information of nodes 11 and 8 indicates a normal state. Based on this, node 9 determines that a fault exists on the first data path.
[0164] It should be noted that for other related descriptions of Method 4, reference can be made to the descriptions of Methods 1 to 3 above, which will not be repeated here. In addition, the above Methods 1 to 4 can be used one by one, or they can be used in combination, and this application does not limit this.
[0165] Step 503: The first source node transmits a first message through a first data path.
[0166] In this application, if the judgment result of step 502 is no, the first source node transmits the first message to the next hop node (ie, the fourth node). Figure 6 As shown in (a) in FIG, node 9 transmits a first message to node 10.
[0167] Step 504: The first source node determines at least one second data path from the plurality of data paths.
[0168] The second data path and the first data path are different data paths.
[0169] It should be noted that the data paths are named as the first data path, the second data path, etc. in order to distinguish the data paths in different scenarios, which will not be further described later.
[0170] Optionally, the second data path does not include the faulty node on the first data path. This helps to increase the number of second data paths that can successfully transmit messages, thereby helping to improve the transmission success rate of the first message.
[0171] It should be noted that for the relevant description of the first source node determining the second data path, reference may be made to the description of the first source node determining the first data path, which will not be repeated here.
[0172] In the present application, after the first source node detects a fault on the first data path, it re-determines at least one second data path from multiple fourth data paths between the first source node and the first destination node for transmitting the first message.
[0173] In one example, the first source node may determine a second data path, so as to prevent the first message from occupying too many data paths, thereby helping to avoid deadlock in the ring topology network.
[0174] In another example, the first source node may determine multiple second data paths. In this way, the first message can be transmitted through the multiple data paths, thereby helping to increase the number of faulty nodes that can be tolerated during the transmission of the first message, thereby helping to increase the probability of successful transmission of the first message to the destination node.
[0175] The following is an exemplary introduction to the number of faulty nodes that the message transmission method provided in this application can tolerate.
[0176] For example, in a 2D torus topology network, each node includes four ports. Based on this, if a fault occurs on the first data path, the first source node can determine up to three disjoint second data paths, such as second data path 1, second data path 2, and second data path 3. If a faulty node exists on two of the three second data paths, such as second data path 1 and second data path 2, the first message can be transmitted via second data path 3.
[0177] From this, it can be seen that for a 2D torus topology network, the message transmission method provided by this application can tolerate at least three faulty nodes during message transmission, such as a faulty node on the first data path, a faulty node on the second data path 1, and a faulty node on the second data path 2. In other words, even if there are three faulty nodes during message transmission, the message can still be successfully transmitted to the destination node.
[0178] On this basis, if there is still a faulty node in the ring topology network and the faulty node is not on the second data path 3, then even if there are 4 faulty nodes in the ring topology network, the first message can still be transmitted through the second data path 3.
[0179] Based on the above, when transmitting messages based on the message transmission method provided in this application, for 2D torus topology networking, it supports any 3 node failures and partial 4 node failures, and for 3D torus topology networking, it supports any 5 node failures and partial 6 node failures.
[0180] The following is an exemplary introduction to an implementation method of determining the number of at least one second data path.
[0181] Exemplarily, the target routing information indicates M data paths between the first source node and the first destination node. That is, there are M data paths between the first source node and the first destination node, where M is a positive integer greater than 1.
[0182] Optionally, a target value is stored on the first source node, and the target value is less than or equal to M. The target value can be used to specify the maximum number of faulty nodes supported by the torus topology network.
[0183] In the present application, the number of at least one second data path may be less than or equal to the target value. Thus, by selecting an appropriate target value, it is not only helpful to avoid deadlock in the ring topology network during the transmission of the first message, but also to increase the probability of the first message being successfully transmitted to the destination node.
[0184] The following describes how to determine the number of at least one second path through methods a and b.
[0185] Mode a: The first source node determines K second data paths with the least number of hops from M data paths.
[0186] Wherein, K is a positive integer less than or equal to the target value.
[0187] In the present application, after determining that a fault exists on the first data path, the first source node obtains a target value and determines K second data paths with the least number of hops from the M data paths indicated by the target routing information.
[0188] In one example, K is equal to a target value. For example, if the target value is 2, then K is equal to 2. In this way, as many data paths as possible can be used to transmit the first message, thereby helping to increase the probability of the first message being transmitted to the destination node.
[0189] In another example, K is smaller than the target value. For example, if the target value is 3, K can be 1 or 2. This allows fewer data paths to be used for message transmission, thereby reducing the number of occupied nodes and helping to avoid deadlock.
[0190] For example, Figure 6 As shown in (b) or (c), the number of hops of data path 2 is 5, the number of hops of data path 3 is 5, the number of hops of data path 4 is 3, and the number of hops of data path 5 is 3. Based on method a, if K is 2, node 9 can determine that data path 4 and data path 5 are used to transmit the first message.
[0191] In this manner, K second data paths with the least number of hops are determined for transmitting messages, which helps reduce the number of nodes occupied by transmitting messages and further helps avoid deadlock in the ring topology network.
[0192] Mode b: The first source node randomly determines K second data paths from M data paths.
[0193] In the present application, after the first source node obtains the target value, it can randomly determine K second data paths from the M data paths indicated by the target routing information.
[0194] For example, Figure 6As shown in (b) or (c), based on method b, if K is equal to 2, node 9 can determine data path 2 and data path 3, or data path 4 and data path 5, or data path 2 and data path 5, or data path 3 and data path 4 as the second data path. For example, node 9 determines data path 2 and data path 3 as the second data path.
[0195] It should be noted that for other related instructions of method b, please refer to the instructions of method a and will not be repeated here.
[0196] In this manner, K second data paths are randomly determined for transmitting messages, which helps to improve the selection diversity of the second data paths.
[0197] In the present application, when the first source node does not store the target data value, at least one second data path may be determined in the above-mentioned manner a or manner b, wherein the number of the at least one second data path may be less than or equal to M.
[0198] The following is an exemplary introduction to the process of the first source node obtaining the target value.
[0199] Optionally, based on steps 501 to 506, the message transmission method may further include: the first source node receives the target value.
[0200] Combine Figure 3 A user can use an electronic device to send a target value to a management node to specify the maximum number of faulty nodes supported by a torus topology network. After receiving the target value, the management node can distribute the target value to the torus topology network. For example, the target value can be sent to each of the multiple nodes in the torus topology network. After receiving the target value, the first source node stores it in a storage medium.
[0201] Exemplarily, the management node may send the target value to the network card device on each node, and the network card device may store the target value in a local storage medium of the network card device, or may also store the target value in a local storage medium of the node.
[0202] In this embodiment, by setting the first source node to receive and store the target value, the user can specify the maximum number of fault nodes supported by the ring topology network. In this way, the user can specify appropriate target values for the ring topology network in different scenarios based on the business needs in different scenarios, thereby helping to improve the performance of the ring topology network and the user experience. For example, for scenarios with low business reliability requirements, the target value can be set to a relatively small value. For scenarios with high business reliability requirements, the target value can be set to a relatively large value to increase the probability of successful message transmission.
[0203] Optionally, each second data path in the at least one second data path includes a fourth node, the fourth node is a next-hop node of the first source node, and the fourth node is a non-faulty node.
[0204] For example, Figure 6 In (d), node 5 is a faulty node. Node 5 is a node on data path 2 and data path 4, and is the next-hop node of the first source node. Therefore, data path 2 and data path 4 are not used as the second data path. In other words, the source node of the first message can select the second data path from data path 3 and data path 5.
[0205] In this embodiment, by setting the next hop node of the first source node on the second data path to a non-faulty node, it helps to increase the data on the second data path that can successfully transmit the first message, thereby helping to increase the probability of successfully transmitting the first message to the destination node.
[0206] Optionally, in a case where the at least one second data path includes a plurality of second data paths, different second data paths in the plurality of second data paths do not have the same node.
[0207] For example, Figure 6 As shown in (b) or (c), data path 2 and data path 4 have the same node (i.e., node 5), and data path 3 and data path 5 have the same node (i.e., node 13). Based on this, the plurality of second data paths may include data path 2 and data path 3, or data path 4 and data path 5, or data path 2 and data path 5, or data path 3 and data path 4.
[0208] In this embodiment, by ensuring that different second data paths within the multiple second data paths do not have the same nodes, it is possible to prevent first messages transmitted by different second data paths from being transmitted to the same node. This, on the one hand, increases the number of faulty nodes that can be tolerated during message transmission. For example, if two second data paths share the same node, then a faulty node will simultaneously prevent both second data paths from properly transmitting messages. Furthermore, it prevents multiple first messages from simultaneously occupying multiple virtual channels of the same node, thereby helping to avoid deadlock during the transmission of the first message and, in turn, increasing the probability that the first message will be transmitted to its destination node.
[0209] Step 505: The first source node transmits the first message through at least one second data path.
[0210] In the present application, after determining at least one second data path, the first source node transmits the first message through each second data path of the at least one second data path.
[0211] For example, in combination with the above-mentioned method b, the first source node determines data path 2 and data path 5 as the second data path. Figure 6 In (c) or (d) above, node 9 may transmit the first message to node 5 to transmit the first message through data path 2. In addition, node 9 may also transmit the first message to node 13 to transmit the first message through data path 5.
[0212] In the above embodiment, when a fault occurs on the first data path, the first source node directly re-determines a data path (i.e., at least one second data path) that can reach the first destination node based on the multiple data paths indicated by the currently stored target routing information, and transmits the first message through the re-determined at least one second data path. In this way, there is no need for the faulty node to report the fault information to the management node to trigger the management node to re-issue routing information. Therefore, the interruption of the first message transmission service caused by the fault is avoided, thereby preventing the fault from affecting the transmission efficiency of the first message, and further ensuring that the transmission efficiency of the first message will not be seriously reduced when a fault occurs.
[0213] Optionally, if the result of the judgment in step 502 is yes, the message transmission method may further include: the first source node marking the first message as a fault message. The fact that a message is a fault message indicates that a fault has been detected on the data path transmitting the message. For example, the fact that the first message is a fault message may indicate that a fault has been detected on the data path transmitting the first message (e.g., the first data path).
[0214] In this way, when an intermediate node transmitting a message detects a fault, it can determine whether the currently detected fault is the first detected fault by judging whether the received message is a fault message. This allows the intermediate node to implement relevant policies based on the number of faults detected, such as discarding the message when the number of faults detected is greater than 1. An intermediate node refers to a node other than the source node and the destination node among multiple nodes transmitting a message. For example: Figure 6 Node 10, node 11, node 12 and node 8 shown in (a) are intermediate nodes of the first message.
[0215] There are multiple implementations for marking the first message as a fault message. Two implementations are exemplarily introduced below.
[0216] In a first implementation, the first message may be updated to mark the first message as a fault message. Based on this, the message transmission method may include the following S5.
[0217] S5: The first source node updates the first message to obtain a third message, wherein the third message includes a fault identifier, which is used to indicate the number of times a fault is detected during the transmission of the first message.
[0218] The first message may be referred to as an original message of the third message, and the third message may be referred to as an updated message of the first message.
[0219] In one example, the fault flag is used to indicate that a fault has been detected during the transmission of the first message, or in other words, that a fault has been detected once. In another example, the fault flag is used to indicate that a fault has been detected S times during the transmission of the first message, where S is a positive integer greater than 1. In other words, the fault flag can be used to indicate that a fault has been detected, or to indicate the specific number of times a fault has been detected.
[0220] In the present application, in conjunction with S5, step 505 may include: the first source node transmitting the third message via at least one second data path. In this way, the intermediate node transmitting the first message can directly determine the number of faults detected during the transmission of the original message (i.e., the first message) of the third message based on the received third message, thereby improving the efficiency of determining the number of faults.
[0221] In this implementation, the first message is marked as a fault message by updating the first message, which helps to improve marking efficiency.
[0222] In a second implementation, the first message may be marked as a fault message through a message. Based on this, the message transmission method may further include the following S6.
[0223] S6: The first source node generates a first message, wherein the first message is used to indicate the number of faults detected during the transmission of the first message.
[0224] The first message may be used to indicate that a fault has been detected, or to indicate the specific number of times a fault has been detected. For other descriptions of the first message, please refer to the description of the fault identifier, which will not be repeated here.
[0225] Based on this, and in addition to step S505, the message transmission method may further include: the source node of the first message transmits a first message to each node on at least one second data path. In this way, the intermediate node transmitting the first message can receive the first message and further determine the number of faults detected during the transmission of the first message based on the second message.
[0226] It should be noted that for other related instructions of S6, please refer to the above-mentioned instructions of S5, which will not be repeated here.
[0227] In this implementation, the first message is marked as a fault message through a message, which helps to ensure the accuracy of the first message.
[0228] In the present application, marking the first message as a fault message when the first source node detects a fault on the first data path is merely an optional method. That is, the first source node may not mark the first message as a fault message. This helps reduce the number of times a fault is detected during the transmission of the first message, thereby helping to delay the execution of the policy of discarding messages, and further helping to improve the success rate of message transmission.
[0229] It should be noted that, below, the present application is exemplarily explained by taking the example that the first source node does not mark the first message as a fault message.
[0230] The above content details the scheme for transmitting messages when the source node of the first message is the target node. Next, Figure 7 and Figure 8 , introducing a scheme for transmitting messages when a node on the first data path serves as a target node.
[0231] It should be noted that, below, only the differences between the two solutions are described in detail, and the similarities are not described in detail.
[0232] Figure 7 This is a flowchart of another message transmission method provided by the present application. Exemplarily, the method may include the following steps 701-705.
[0233] Exemplarily, the judgment result of step 502 is no, that is, after the first source node generates the first message, it transmits the first message through the first data path. Based on this, the nodes on the first data path can receive the first message.
[0234] In this application, the node on the first data path may be referred to as a first intermediate node. In the following, steps 701 to 705 are described as an example when the first intermediate node is used as the target node.
[0235] In the following embodiments, the operations performed by the first intermediate node, such as steps 701 to 705, can be performed by a network card device on the first intermediate node. For example, the network card device of the first intermediate node can execute steps 701 to 705 by running a target routing algorithm.
[0236] Optionally, based on steps 701 to 705, the message transmission method may further include the following S7. By configuring the message transmission method to include S7, a node receiving a message may execute the message transmission method provided by the present application when the node is not the destination node.
[0237] S7: The first intermediate node determines whether the current node is the destination node of the first message based on the received first message.
[0238] If the judgment result is yes, then the process ends. If the judgment result is no, then step 701 is executed.
[0239] In this application, a first source node generates a first message, determines that a first data path is used to transmit the first message, and then sends the first message to a node on the first data path (i.e., a first intermediate node). The first message indicates the source node (i.e., the first source node) and the destination node (i.e., the first destination node) of the first message. After receiving the first message, the first intermediate node determines whether the node (i.e., the first intermediate node) is the destination node (i.e., the first destination node) of the first message.
[0240] For example, the first intermediate node can parse the first message to determine the node identifier of the first destination node and determine whether the node identifier of the current node is the same as the node identifier of the first destination node. If they are the same, the current node is the first destination node; if they are not the same, the current node is not the first destination node.
[0241] For example, Figure 8 As shown in (a), after node 10 receives the first message sent by node 9, it determines whether node 10 is the destination node of the first message. Since the destination node of the first message is node 4, the determination result is no.
[0242] It should be noted that S7 is an optional step. That is, after receiving the first message, the first intermediate node may directly execute step 701 without executing S7.
[0243] In this embodiment, setting the message transmission method includes determining whether the node receiving the message is the destination node. In this way, the destination node can avoid executing the message transmission method provided in this application after receiving the message, thereby helping to avoid occupying the processing resources of the destination node.
[0244] Step 701: A first intermediate node determines a first data path from multiple data paths for transmitting a first message.
[0245] In the present application, after receiving the first message, the first intermediate node can determine the first data path for transmitting the first message from the data path between the source node and the destination node indicated by the first message (i.e., multiple fourth data paths), that is, the multiple fourth data paths between the first source node and the first destination node.
[0246] For example, Figure 8As shown in (a), after node 10 receives the first message sent by node 9, it can determine that data path 1 is used to transmit the first message from multiple fourth data paths between node 9 (i.e., the first source node) and node 4 (i.e., the first destination node) of the first message.
[0247] Since each node on the ring topology network determines the data path for transmitting messages based on the same routing algorithm, that is, the nodes on the first data path and the first source node determine the data path for transmitting the first message based on the same routing algorithm (such as the target routing algorithm), therefore, the first data path for transmitting the first message determined by the first intermediate node and the first source node is the same.
[0248] It should be noted that for other relevant descriptions of step 701, reference can be made to the description of step 501, which will not be repeated here.
[0249] Step 702: The first intermediate node determines whether there is a fault on the first data path.
[0250] If the judgment result is no, then execute step 703. If the judgment result is yes, then execute step 704.
[0251] It should be noted that for the relevant description of step 702, reference can be made to the description of step 502, which will not be repeated here.
[0252] Step 703: The first intermediate node transmits the first message through the first data path.
[0253] It should be noted that for the relevant description of step 703, reference can be made to step 503, which will not be repeated here.
[0254] Step 704: The first intermediate node determines at least one second data path from the multiple data paths.
[0255] In this application, step 704 includes multiple implementation methods, and the following is an exemplary description using Method A to Method B.
[0256] Mode A: Step 704 includes: the first intermediate node determines at least one second data path from a plurality of fourth data paths between the first source node and the first destination node.
[0257] In the present application, when the first intermediate node detects a fault on the first data path, it can determine at least one second data path based on multiple fourth data paths between the source node and the destination node indicated by the first message, and the fourth node on each second data path is the next hop node of the first intermediate node.
[0258] For example, Figure 8As shown in (b), after node 10 detects that node 11 is a faulty node, it can determine that at least one of data path 2 and data path 5 is the second data path from data path 2 to data path 5 between node 9 and node 4, wherein node 6 on data path 2 is the next hop node of node 10, and node 14 on data path 3 is the next hop node of node 10.
[0259] In this implementation, the second data path is screened from the plurality of fourth data paths between the first source node (i.e., the initial source node) and the first destination node. In this way, the logic for determining the second data path is the same as the logic for determining the first data path, which helps to reduce the complexity of the target routing algorithm.
[0260] Mode B: Step 704 includes: the first intermediate node determines at least one second data path from a plurality of target data paths between the first intermediate node and the first destination node.
[0261] In the present application, when a first intermediate node detects a fault on a first data path, the first intermediate node is used as a new source node, and at least one second data path from the new source node to the first destination node is determined.
[0262] For example, Figure 8 As shown in (b), node 10 is the first intermediate node. After detecting that node 11 is a faulty node, node 10 uses node 10 as the new source node for the first message and determines at least one second data path through which node 10 can reach node 4. For example, the data paths through which node 10 can reach node 4 include data path 6, data path 7, and data path 8. Data path 6 includes nodes 6, 7, and 8; data path 7 includes nodes 14, 15, and 16; and data path 8 includes nodes 9, 12, and 8. Based on this, node 10 can determine data paths 6, 7, and 8 as second data paths.
[0263] In this implementation, the first intermediate node is used as the new source node, and the second data path is selected from multiple data paths between the new source node and the first destination node. This not only helps to reduce the number of hops of the second data path, but also helps to increase the selection range of the second data path.
[0264] For example, Figure 8As shown in (b), if node 11 of data path 1, where node 10 is located, is a faulty node, if a second data path is selected from the multiple data paths a between node 9 (i.e., the initial source node) and node 4, then node 10's next-hop nodes can include nodes 6 and 14. In other words, node 10 can only select a second data path from the portion of data paths a that include nodes 6 and 14, so that node 10 can transmit the message to the selected data path a. If a second data path is selected from the multiple data paths b between node 10 and node 4, then node 10's next-hop nodes can include nodes 6, 9, and 14. This shows that using node 10 as the new source node and determining at least one second data path helps increase the number of data paths to be selected.
[0265] In this application, when the first intermediate node selects the second data path using method B, the first source node and the first intermediate node determine the second data path differently. The first source node determines the second data path using the first source node as the source node of the first message, while the first intermediate node determines the second data path using the first intermediate node as the source node of the first message. The following describes two exemplary implementations of the method in which the first intermediate node selects the second data path using itself as the source node.
[0266] First implementation: Based on steps 701 to 705, the message transmission method may include the following S8, so that the first source node and the first intermediate node can implement different solutions for screening the second data path according to different judgment results.
[0267] S8: The first intermediate node determines whether the current node is the source node of the first message.
[0268] If the judgment result is yes, the first solution is executed. If the judgment result is no, the second solution is executed.
[0269] The first solution includes determining at least one second data path based on the source node of the first message and the destination node of the first message, that is, determining at least one second data path from multiple fourth data paths, which is also the solution of Method A. The second solution includes determining at least one second data path based on the first intermediate node and the destination node of the first message, that is, determining at least one second data path from multiple target data paths between the first intermediate node and the destination node of the first message, which is also the solution of Method B.
[0270] In the present application, after receiving the first message, the first intermediate node can also determine whether the node is the source node of the received first message. Since the first intermediate node is a node on the first data path, the first intermediate node is not the source node of the first message, and the judgment result of S8 is no.
[0271] Optionally, the first intermediate node may determine whether the local node is the source node of the first message based on the node identifier of the local node and the node identifier of the source node indicated by the first message. If the node identifier of the local node is the same as the node identifier of the source node indicated by the first message, then the local node is the source node of the first message, and the judgment result of S8 is yes. Otherwise, the local node is not the source node of the first message, and the judgment result of S8 is no.
[0272] In this embodiment, when a target node (e.g., the first source node) detects a fault on the first data path, it continues to determine whether the current node is the source node of the first message, and executes different logic based on different judgment results. In this way, the target node can determine the second data path through different strategies based on the different types of the current node, thereby helping to improve the diversity of the message transmission method. Furthermore, by configuring different strategies for different types of nodes, the reliability and stability of the message transmission method can be improved.
[0273] Based on the first implementation method, in one example, the step of "determining whether this node is the source node of the first message" can be performed only by the node receiving the first message. This helps to simplify the steps of the message transmission method executed by the first source node, thereby improving the message transmission efficiency.
[0274] In another example, Figure 5 The message transmission method shown may also include the step of "determining whether the current node is the source node of the first message." In this way, the steps of the message transmission method performed by different nodes transmitting the message (i.e., the initial source node and the intermediate node) are the same, thereby helping to simplify the logic of the target routing algorithm and reduce the complexity of the target routing algorithm.
[0275] The intermediate node refers to any node other than the original source node and the destination node among the nodes transmitting the message, such as the first intermediate node on the first data path.
[0276] In a second implementation, when the first source node and the first intermediate node detect a fault on the first data path, a second data path is determined with the current node as the source node. That is, at least one second data path is determined from the data paths between the current node and the first destination node.
[0277] Based on this, when a fault is detected at the first source node, a second data path is determined with the first source node as the source node. When a fault is detected at the first intermediate node, a second data path is determined with the first intermediate node as the source node.
[0278] In this implementation, by setting the node as the source node to determine the second data path when the node is detected, it can not only ensure that the initial source node and the intermediate node determine the second data path through different strategies, but also simplify the steps of the message transmission method, thereby improving the message transmission efficiency.
[0279] Step 705: The first intermediate node transmits the first message through at least one second data path.
[0280] In the present application, step 705 may include multiple implementations. Two implementations are exemplified below.
[0281] In the first implementation, in combination with method A in step 704, when executing step 705, the first intermediate node can directly transmit the first message through at least one second data path. In other words, the first intermediate node does not update the source node of the first message. This helps improve transmission efficiency. For example, in combination with method A in step 704, Figure 8 As shown in (b), node 10 can transmit the first message to node 14 and node 6.
[0282] The second implementation method, combined with method B in step 704, is that during the execution of step 705, the first intermediate node can update the source node of the first message, and update the source node of the first message to the first intermediate node, so that the intermediate node that subsequently receives the first message can determine the data path based on the updated source node. This helps to improve the accuracy of other subsequent intermediate nodes in determining the data path.
[0283] There are multiple implementation methods for updating the source node of the first message. The following is an illustrative introduction using Method A to Method B.
[0284] Mode A: updating the source node of the first message by updating the first message. Based on this, the message transmission method may further include the following S9.
[0285] S9: The first intermediate node updates the first message to obtain a second message; the second message indicates that the source node is the first intermediate node, and the second message indicates that the destination node is the first destination node.
[0286] The first message may be referred to as an original message of the second message, and the second message may be referred to as an updated message of the first message.
[0287] In the present application, when the first intermediate node determines that a fault exists on the first data path, it can update the first message to obtain a second message. The second message indicates that the source node is the first intermediate node, that is, the first intermediate node is the new source node, and the second message indicates that the destination node is the destination node of the first message, that is, the destination node remains unchanged.
[0288] In this application, the difference between the first message and the second message includes that the source node indicated by the first message is different from the source node indicated by the second message. Figure 8 As shown in (b), the source node indicated by the first message is node 9, and the source node indicated by the second message is node 10. It can be considered that node 9 is the initial source node and node 10 is the new source node.
[0289] It should be noted that the data field in the first message is the same as the number field in the second message, that is, the data carried by the first message is the same as the data carried by the second message.
[0290] Optionally, the second message includes a target identifier, which is the node identifier of the first intermediate node. In this way, the second message can indicate the first intermediate node as the new source node through the target identifier, which helps ensure the accuracy of the new source node indicated by the second message.
[0291] Optionally, the target identifier may include an Internet Protocol (IP) address, an identity document (ID), or a Media Access Control (MAC) address. Indicating the first intermediate node by an IP address, ID, MAC address, etc. helps increase the diversity of ways in which the second message indicates the source node.
[0292] Optionally, the header of the second message includes a destination identifier. In this way, the destination identifier can be added to the header of the first message to generate the second message. This can avoid modifying the content of the data field of the first message when updating the first message, thereby helping to ensure that the data field of the second message is identical to the data field of the first message, and further helping to ensure the accuracy of the data content conveyed by the second message.
[0293] Exemplarily, the message header of the first message may include the fields shown in Table 1.
[0294] Table 1
[0295] Flag Class Number Option length Option Data
[0296] Among them, Flag, Class, and Number are used to represent the optional type (Type). "Flag" is used to represent the identifier, "Class" is used to represent the type, and "Number" is used to represent the value. "Option length" is used to represent the length of the optional type, and "Option Data" is used to represent the data of the optional type. For example, the "Flag" field is 1 bit (bits), the "Class" field is 2 bits, and the "Number" field is 5 bits. The "Option length" field is 8 bits. The length of the "Option Data" field is variable (variable). The value of the "Class" field can include 0 or 2, where 0 is used to represent control and 2 is used to represent debugging and measurement (For debugging and measurement).
[0297] Exemplarily, the first intermediate node may write the destination identifier of the first intermediate node into the "Option Data" field of the first message, thereby obtaining the second message. The second message indicates that the source node is the first intermediate node through the destination identifier stored in the "Option Data" field.
[0298] It should be noted that the present application does not limit the field in the second message indicating that the first intermediate node is the source node. The above is merely an exemplary description.
[0299] In the above embodiment, after the first intermediate node detects a fault on the first data path, it updates the first message, thereby obtaining a second message indicating that the first intermediate node is the new source node. In this way, after the intermediate node on the second data path receives the second message, it can determine the data path based on the new source node and destination node indicated by the second message, thereby helping to ensure that the data path determined by the node on the second data path is the same as the data path determined by the first intermediate node, and further helping to avoid message disorder when transmitting multiple messages.
[0300] In the present application, in combination with S9, step 705 may include: the first intermediate node transmits the second message through at least the first second data path.
[0301] For example, Figure 8 As shown in (b), node 10 can transmit the second message to node 6, node 9, and node 14, so as to transmit the second message through data path 6, data path 7, and data path 8.
[0302] It should be noted that for other relevant descriptions of step 705, reference can be made to the description of step 505, which will not be repeated here.
[0303] Combine Figure 5 、 Figure 7 In the illustrated embodiment, a first source node can be set as a target node. When the first source node detects a fault on the first data path, it can re-determine at least one second data path for message transmission. Furthermore, a node on the first data path can be set as a target node. During the transmission of the first message, if any node on the first data path detects a fault on the first data path, it can re-determine at least one second data path for message transmission. Thus, if a fault is detected on the data path for message transmission at any time, the solution of the present application can be used to re-determine the data path for message transmission. This not only helps expand the scope of nodes applicable to the solution of the present application, but also helps improve the success rate of message transmission.
[0304] Optionally, the second message may include a fault identifier, where the fault identifier is used to indicate the number of times a fault is detected during the transmission of the first message.
[0305] If the number of fault detections is greater than or equal to 1, the second message can be considered a fault message. A message being a fault message can indicate that a fault has been detected on the data path transmitting the message (or the original message transmitting the message). For example, a second message being a fault message can indicate that a fault has been detected on the data path transmitting the first message (i.e., the original message of the second message) (i.e., the first data path).
[0306] In the present application, the fault identifier may be recorded in the message header, which helps to ensure the accuracy of the data field of the first message, thereby helping to ensure the accuracy of the content conveyed by the first message.
[0307] It should be noted that, for the relevant description of the fault identifier, reference can be made to the description of the fault identifier in S5, which will not be repeated here. In addition, for the recording method of the fault identifier, reference can be made to the recording method of the target identifier, which will not be repeated here.
[0308] In the above embodiment, the second message can be set to carry a fault identifier. In this way, after other intermediate nodes receive the second message, they can directly determine the number of times the fault is detected based on the second message. This not only helps to improve the efficiency of determining the number of faults, but also helps to provide the accuracy of the determined number of faults. When other intermediate nodes detect a fault, they can accurately determine whether it is the first time the fault is detected, and then accurately determine whether to discard the second message to avoid performing erroneous operations.
[0309] Mode B can update the source node of the first message through a message. Based on this, the message transmission method may include the following S10.
[0310] S10: The first intermediate node generates a target message, where the target message is used to indicate that the source node of the first message is the first intermediate node.
[0311] Based on this, and in addition to step 705, the message transmission method may include: the first intermediate node transmitting a target message to each node on at least one second data path. In this way, the other intermediate nodes transmitting the first message can determine, based on the received target message, that the source node of the first message is the first intermediate node, and thus determine the data path for transmitting the first message based on the first intermediate node. In other words, the data path for transmitting the first message is determined from the data paths between the first intermediate node and the first destination node.
[0312] For example, the target message may further indicate that the destination node of the first message is the first destination node. In this way, other intermediate nodes can simultaneously determine the source node and destination node of the first message through the target message, which helps improve the convenience of determining the source node and destination node, thereby helping improve the convenience of determining the data path.
[0313] Optionally, the target message may also be used to indicate the number of faults detected during the transmission of the first message. This not only helps to increase the diversity of the target message indication content, but also helps to ensure the accuracy of the first message.
[0314] It should be noted that for other related instructions of S10, please refer to the instructions of S9, S6, etc., which will not be repeated here.
[0315] In this implementation, indicating the update source node of the first message through the target message helps to ensure the accuracy of the first message.
[0316] The above content details the scheme of using the source node of the first message (or the first intermediate node on the first data path) as the target node to transmit the message. Figure 9 , introduces a scheme for nodes on the second data path to transmit messages.
[0317] It should be noted that the following only details the differences between the two solutions, and the similarities are not repeated here.
[0318] Figure 9 This is a flowchart of another message transmission method provided by the present application. Exemplarily, the method may include steps 901 to 906.
[0319] In Example 1, the judgment result of step 502 is yes, that is, the first source node generates a first message and, after determining that the first data path is used to transmit the first message, detects a fault on the first data path. Based on this, the first source node transmits the first message via at least one second data path. In conjunction with step 505, the second intermediate node may receive the first message transmitted by the first source node, or receive a third message transmitted by the first source node, or receive first target information transmitted by the first source node, where the first target information includes the first message and the first message.
[0320] In Example 2, the judgment result of step 702 is yes, that is, after the first intermediate node receives the first message and determines that the first data path is used to transmit the first message, it detects a fault on the first data path. Based on this, the first intermediate node transmits the first message via at least one second data path. In conjunction with step 705, the second intermediate node may receive the first message transmitted by the first intermediate node, or receive the second message transmitted by the first intermediate node, or receive the second target information transmitted by the first intermediate node, where the second target information includes the first message and the target message.
[0321] It should be noted that, below, steps 901 to 906 are exemplarily described by taking the above-mentioned Example 2 and the second intermediate node receiving the second message as an example.
[0322] In the present application, the at least one second data path includes a target second data path, wherein the target second data path may be any second data path among the at least one second data path.
[0323] The node on the second data path may be referred to as a second intermediate node (or a first node). Below, steps 701 to 705 are exemplarily described by taking the second intermediate node on the target second data path as an example.
[0324] Step 901: A second intermediate node determines a target second data path from multiple data paths for transmitting a second message.
[0325] Step 902: The second intermediate node determines whether there is a fault on the target second data path.
[0326] If the judgment result is no, then execute step 903. If the judgment result is yes, then execute step 904 or step 905.
[0327] Step 903: The second intermediate node transmits the second message through the target second data path.
[0328] It should be noted that for the relevant descriptions of steps 901 to 903, reference can be made to the descriptions of steps 501 to 503 and steps 701 to 703, which will not be described in detail here.
[0329] Optionally, based on the judgment result of step 902 being no, the message transmission method may include the following S11, so that it can be determined whether the currently detected fault is a fault detected for the first time, thereby determining the target strategy to be executed, such as: discarding the message, redetermining the data path for transmitting the message, etc.
[0330] S11: The second intermediate node determines whether the fault detected on the second data path is the first fault detected.
[0331] If the judgment result is no, the third strategy is executed. If the judgment result is yes, the fourth strategy is executed.
[0332] The third strategy includes discarding the received message, such as the solution of step 904. The fourth strategy includes determining at least one data path for transmitting the received message, such as the solution of steps 905-906.
[0333] In this application, when the second intermediate node detects a fault on the target second data path, it can determine whether it is the first detected fault. If the judgment result is no, it means that at least one fault has been detected during the transmission of the message, and the fault on the target second data path is not the first detected fault.
[0334] Exemplarily, the second intermediate node parses the received message to determine whether the received message includes a fault identifier, or whether the received information includes the first message / target message, thereby determining whether the fault is detected for the first time. If the received message includes the fault identifier, or the received information includes the first message / target message, the determination result is negative. Conversely, if the received message does not include the fault identifier, or the received information does not include the first message / target message, the determination result is positive.
[0335] In combination with the above examples 1 and 2, if the first message is received, the judgment result is No. If the second message and the third message are received, the judgment result is Yes. If the first target information and the second target information are received, the judgment result is Yes.
[0336] In this embodiment, the message transmission method is set to include determining whether the fault detected by the current node is the first time the fault is detected, so that the message can be discarded when the fault is detected multiple times. This not only helps to ensure the success rate of message transmission, but also helps to avoid occupying too many data paths and causing deadlock.
[0337] In this application, Figure 7 The message transmission method shown may also include the step of "determining whether this is the first time a fault has been detected." In this way, different intermediate nodes transmitting messages (e.g., the first intermediate node and the second intermediate node) perform the same message transmission method steps, thereby simplifying the logic of the target routing algorithm and reducing its complexity.
[0338] It can be understood that the first source node transmits the first message to the first intermediate node on the first data path when no fault is detected on the first data path. Therefore, the fault detected by the first intermediate node on the first data path is the first detected fault, that is, the judgment result is yes. Based on this, the first intermediate node executes a plan to determine at least one data path for transmitting the received message, that is, executes the plan of steps 704-705.
[0339] Optionally, based on the judgment result of step 902 being negative, the message transmission method may include the following S12, so that the number of times the fault is detected can be judged, thereby determining the target strategy to be executed, such as: discarding the message, re-determining the data path for transmitting the message, etc.
[0340] S12: The second intermediate node determines whether the number of times the fault is detected is greater than or equal to a number threshold.
[0341] If the judgment result is yes, the third strategy is executed. If the judgment result is no, the fourth strategy is executed.
[0342] It should be noted that this application does not limit the specific value of the number threshold, which can be set by the user according to the actual scenario.
[0343] In this embodiment, whether to discard a message is determined by the number of detected failures. This not only helps to increase the diversity of message transmission methods, but also helps to increase the flexibility of message transmission methods, thereby helping to improve user experience.
[0344] It should be noted that for other related instructions of S11, please refer to the instructions of S10 and will not be repeated here.
[0345] Step 904: The second intermediate node discards the received second message.
[0346] In the present application, when the second intermediate node detects a fault on the target second data path, it discards the received second message.
[0347] In this embodiment, when a node on the second data path detects a fault again, it may directly discard the received second message. This helps reduce the number of data paths that continue to transmit the first message, thereby helping to avoid deadlock.
[0348] Step 905: The second intermediate node determines at least one third data path from the multiple data paths.
[0349] Step 906: The second intermediate node transmits the second message through at least one third data path.
[0350] It should be noted that for the relevant descriptions of step 905-step 906, reference can be made to step 504-step 505 and step 704-step 705, which will not be repeated here.
[0351] In this embodiment, when a node on the second data path detects a fault again, at least one third data path may be re-determined for transmitting the message based on the target routing information. This helps increase the number of data paths for transmitting the second message, thereby helping to improve the transmission success rate of the second message.
[0352] In the present application, the scheme for transmitting messages by the node on the third data path may refer to the scheme for transmitting messages by the second intermediate node on the second data path, and will not be repeated here.
[0353] The above mainly introduces the solution provided by the present application from the perspective of the method. In order to realize the above functions, the message transmission device includes a hardware structure and / or software module corresponding to the execution of each function. It should be easy for those skilled in the art to realize that, in combination with the units and algorithm steps of each example described in the embodiments disclosed herein, the present application can be implemented in the form of hardware or a combination of hardware and computer software. Whether a function is executed in the form of hardware or computer software driving hardware depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of this application.
[0354] The present application can exemplarily divide the message transmission device into functional modules according to the above method. For example, the message transmission device can include functional modules corresponding to the functional divisions, or two or more functions can be integrated into one processing module. The above integrated modules can be implemented in the form of hardware or software functional modules. It should be noted that the division of modules in this application is schematic and is only a logical functional division. In actual implementation, other division methods may be used.
[0355] For example, Figure 10A possible structural diagram of the message transmission device (denoted as message transmission device 1000) involved in the above embodiment is shown. The actions performed by the message transmission device 1000 are implemented by the nodes of the ring topology network or by the nodes executing corresponding software. The ring topology network stores target routing information, and the target routing information is used to indicate multiple data paths on the ring topology network. The message transmission device 1000 may include a routing module 1001 and a transmission module 1002. The routing module 1001 is used to determine, from multiple data paths, a first data path for transmitting a first message. For example, Figure 5 S501 shown, and Figure 7 The routing module is further configured to determine at least one second data path from a plurality of data paths if a fault occurs on the first data path. Figure 5 S502 and S504 shown, and Figure 7 The transmission module is configured to transmit the first message via at least one second data path. Figure 5 S505 shown, and Figure 7 S705 shown.
[0356] Optionally, the target node includes a source node of the first message, or the target node includes a node on the first data path.
[0357] Optionally, if the target node is a node on the first data path, the routing module 1001 is configured to: determine at least one second data path from a plurality of target data paths; the plurality of target data paths include a data path between the target node and a destination node of the first message.
[0358] Optionally, if the destination node is a node on the first data path, the routing module 1001 is configured to determine at least one second data path from a plurality of fourth data paths; the plurality of fourth data paths include data paths between a source node of the first message and a destination node of the first message.
[0359] Optionally, the target node transmits the first message through at least one second data path, including: updating the first message to obtain a second message; the source node indicated by the second message is the target node, and the destination node indicated by the second message is the destination node of the first message; transmitting the second message through at least one second data path.
[0360] Optionally, the second message includes a fault identifier, where the fault identifier is used to indicate the number of times a fault is detected during the transmission of the first message.
[0361] In another possible implementation, the routing module 1001 is configured to: transmit a target message to at least a first node of the second data path; the target message is used to indicate that the source node of the first message is updated to the target node.
[0362] In another possible implementation, the target message is further used to indicate the number of times a fault is detected during the transmission of the first message.
[0363] Optionally, the at least one second data path includes a target second data path, and the target second data path includes the first node; the routing module 1001 is further configured to discard the received message if a fault occurs in the target second data path.
[0364] Optionally, the failure of the first data path includes at least one of a failure of the first port on the target node, a failure of the second port on the destination node of the first message, a failure of the second node on the first data path, and a failure of the third node on the first data path; wherein, the first port is connected to the node on the first data path, the second port is connected to the node on the first data path, the second node is the next hop node of the target node, there is at least one node between the third node and the target node, and at least one node is located on the first data path.
[0365] Optionally, the at least one second data path includes multiple second data paths, and different second data paths in the multiple second data paths do not have the same node.
[0366] Optionally, each second data path in the at least one second data path includes a fourth node, the fourth node is a next-hop node of the target node, and the fourth node is a non-faulty node.
[0367] Optionally, the routing module 1001 is further configured to: receive a target value; and the target value is used to determine the quantity of at least one second data path.
[0368] Optionally, the target routing information indicates that there are M data paths between the target node and the destination node of the first message, where M is a positive integer greater than 1; the routing module 1001 is specifically used to: determine K second data paths with the least number of hops from the M data paths; K is a positive integer less than or equal to the target value.
[0369] Optionally, the target routing information indicates that there are M data paths between the target node and the destination node of the first message, where M is a positive integer greater than 1; the routing module 1001 is specifically configured to randomly determine K second data paths from the M data paths.
[0370] For the detailed description of the above optional methods, please refer to the above method embodiments, which will not be repeated here. In addition, the explanation and description of the beneficial effects of any of the above message transmission devices 1000 can refer to the above corresponding method embodiments, which will not be repeated here.
[0371] The present application also provides a computing device, which may include a processor, a memory, and a computer program / instructions stored in the memory; the processor executes the computer program to enable the computing device to implement the steps of the above-mentioned message transmission method.
[0372] The present application also provides a node in a ring topology network, the node including a processor, a memory and a computer program / instructions stored in the memory; the processor executes the computer program to enable the node to implement the steps of the above-mentioned message transmission method.
[0373] The present application also provides a ring topology network, which includes multiple nodes, each of which includes a processor, a memory, and a computer program / instructions stored in the memory; the processor executes the computer program to enable the ring topology network to implement the steps of the above-mentioned message transmission method.
[0374] The present application also provides a computer program product, including a computer program / instruction, which implements the steps of the above-mentioned message transmission method when the computer program / instruction is executed by a processor.
[0375] The present application also provides a computer-readable storage medium, on which a computer program / instruction is stored. When the computer program / instruction is executed by a processor, the steps of the above-mentioned message transmission method are implemented.
[0376] In the above embodiments, all or part of the embodiments may be implemented using software, hardware, firmware, or any combination thereof. When implemented using a software program, all or part of the embodiments may be implemented in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer instructions are loaded and executed on a computer, the processes or functions described in this application are generated in whole or in part.
[0377] The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions may be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions may be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via a wired (e.g., coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) method. The computer-readable storage medium may be any available medium that a computer can access or a data storage device such as a server or data center that includes one or more available media integrated therein. The available medium may be a magnetic medium (e.g., a floppy disk, a magnetic disk, a tape), an optical medium (e.g., a digital video disc (DVD)), or a semiconductor medium (e.g., a solid-state drive (SSD)).
[0378] Through the description of the above embodiments, those skilled in the art will clearly understand that for the sake of convenience and brevity, only the division of the above functional modules is used as an example. In actual applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above. The specific working processes of the above-described systems, devices, and units can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.
[0379] In the several embodiments provided in this application, it should be understood that the disclosed methods, devices, etc. can be implemented in other ways. For example, the device embodiments described above are merely schematic. For example, the division of the modules or units is only a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.
[0380] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.
[0381] In addition, the functional units in the various embodiments of the present application may be integrated into a single processing unit, or each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.
[0382] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application is essentially or the part that contributes to the prior art or all or part of the technical solution can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) or a processor to execute all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes: various media that can store program codes, such as flash memory, mobile hard disk, read-only memory, random access memory, magnetic disk or optical disk.
[0383] The above is only a specific embodiment of the present application, but the scope of protection of the present application is not limited thereto. Any changes or replacements within the technical scope disclosed in this application should be included in the scope of protection of the present application. Therefore, the scope of protection of the present application should be based on the scope of protection of the claims.
Claims
1. A message transmission method, characterized in that: The method is applied to a torus topology network, wherein the torus topology network includes a plurality of nodes, each of the plurality of nodes stores target routing information, and the target routing information is used to indicate a plurality of data paths on the torus topology network; the method includes: The destination node among the multiple nodes determines, from the multiple data paths, a first data path for transmitting the first message; If there is a fault on the first data path, the target node determines at least one second data path from the multiple data paths; The target node transmits the first message through the at least one second data path.
2. The method according to claim 1, characterized in that The target node includes a source node of the first message, or the target node includes a node on the first data path.
3. The method according to claim 2, characterized in that If the target node includes a node on the first data path, the target node determines at least one second data path from the plurality of data paths, including: The target node determines the at least one second data path from a plurality of target data paths; the plurality of target data paths include a data path between the target node and a destination node of the first message.
4. The method according to claim 3, characterized in that The target node transmitting the first message through the at least one second data path includes: The target node updates the first message to obtain a second message; the source node indicated by the second message is the target node, and the destination node indicated by the second message is the destination node of the first message; The destination node transmits the second message through the at least one second data path.
5. The method according to any one of claims 1 to 4, characterized in that The at least one second data path includes a target second data path, the target second data path includes the first node; the method further includes: If the target second data path fails, the first node discards the received message.
6. The method according to any one of claims 1 to 5, characterized in that The failure of the first data path includes at least one of a failure of a first port on the target node, a failure of a second port on the destination node of the first message, a failure of a second node on the first data path, and a failure of a third node on the first data path; wherein, the first port is connected to a node on the first data path, the second port is connected to a node on the first data path, the first node is the next hop node of the target node, there is at least one node between the third node and the target node, and the at least one node is located on the first data path.
7. The method according to any one of claims 1 to 6, characterized in that The at least one second data path includes a plurality of second data paths, and different second data paths in the plurality of second data paths do not have the same node.
8. The method according to any one of claims 1 to 7, characterized in that Each second data path in the at least one second data path includes a fourth node, the fourth node is a next-hop node of the target node, and the fourth node is a non-faulty node.
9. The method according to any one of claims 1 to 7, characterized in that The method further comprises: The target node receives a target value; the target value is used to determine the quantity of the at least one second data path.
10. The method according to claim 9, characterized in that The target routing information indicates that there are M data paths between the target node and the destination node of the first message, where M is a positive integer greater than 1; The target node determines at least one second data path from the multiple data paths, including: The target node determines K second data paths with the least number of hops from the M data paths; K is a positive integer less than or equal to the target value; or The target node randomly determines K second data paths from the M data paths.
11. A message transmission device, characterized in that: Applicable to a ring topology network; the device includes: a routing module, configured to determine, from a plurality of data paths, a first data path for transmitting the first message; The routing module is further configured to determine at least one second data path from the plurality of data paths if a fault occurs on the first data path; A transmission module is configured to transmit the first message through the at least one second data path.
12. A computing device, characterized in that The computing device comprises a processor, a memory, and a computer program / instruction stored in the memory; the processor executes the computer program to enable the computing device to implement the steps of the method according to any one of claims 1 to 10.
13. A node in a torus topology network, characterized in that: The node includes a processor, a memory, and a computer program / instruction stored in the memory; the processor executes the computer program to enable the node to implement the steps of the method according to any one of claims 1 to 10.
14. A ring topology network, characterized in that: The torus topology network includes multiple nodes, each of the multiple nodes includes a processor, a memory, and a computer program / instructions stored on the memory; the processor executes the computer program so that the torus topology network implements the steps of the method as described in any one of claims 1-10.
15. A computer program product, characterized in that The method comprises a computer program / instruction, which, when executed by a processor, implements the steps of the method according to any one of claims 1 to 10.
16. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program / instruction, which, when executed by a processor, implements the steps of the method according to any one of claims 1 to 10.
Citation Information
Patent Citations
Self-adaptive fault-tolerant routing method and device for high-dimensional Torus network
CN111953589A
Torus network fault tolerance method based on topology reconstruction and path planning
CN113347029A
Fault-tolerant routing scheme for a multi-path interconnection fabric in a storage network
US20030023893A1