Communication method, communication device and communication equipment
By introducing serial number information into the message, the forwarding device can quickly detect and perceive service flow failures, solving the problem that BFD cannot detect data flow failures, and achieving efficient fault detection and response.
Patent Information
- Application Number
- CN202410165492.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-02-02
- Publication Date
- 2025-08-05
AI Technical Summary
The existing two-way forwarding detection (BFD) protocol can only detect the integrity of the link, but cannot sense the failures caused by the data flow, and cannot quickly detect and monitor gray failures in the network, such as silent failures, misconfiguration, etc., resulting in a degradation of network performance.
By introducing serial number information into the packets of the target service flow, the forwarding device can obtain the serial number information of the packets and use the serial number difference to determine whether the service flow has failed, including the serial number and confirmation number in the TCP header or the packet serial number in the RDMA message, so as to realize orderly transmission and fault detection of the service flow.
It realizes rapid detection and perception of service flow failures, reduces detection complexity, improves the efficiency of fault detection, and can quickly respond to network failures in a low-latency environment.
Smart Images

Figure CN120434153A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of communication technologies, and in particular, to a communication method, a communication device, and a communication equipment. Background Art
[0002] Bidirectional Forwarding Detection (BFD) is a network protocol used to quickly detect and monitor the forwarding connectivity of links or Internet Protocol (IP) routes in a network. BFD can improve the performance of existing networks. By quickly detecting communication failures, it enables network devices to establish backup channels more quickly to resume communication.
[0003] Between two network devices at both ends, BFD packets are periodically transmitted on the connected link. If a certain network device does not receive a BFD packet within a preset time period, the network device can determine that the link connecting to the peer network device has failed.
[0004] BFD can only detect the integrity of the link (i.e., whether the link has failed), and cannot sense the failures that occur to data flows. In view of this, a data flow-oriented fault detection scheme is urgently needed. Summary of the Invention
[0005] This application provides a communication method, a communication device, and a communication equipment for detecting service flow failures.
[0006] In a first aspect, this application provides a communication method. After the packets in a target service flow are sent from a source device and forwarded by a forwarding device, they can reach a destination device. Each packet in the target service flow includes sequence number information corresponding to the packet. In the embodiments of this application, the sequence number information of the packet is used to implement important functions such as orderly transmission of packets, retransmission of lost packets, and error recovery, thereby ensuring the reliable transmission of the service flow in the network.
[0007] In the embodiments of this application, the forwarding device receives a first packet in the target service flow, and the first packet includes first sequence number information corresponding to the first packet. The forwarding device can obtain second sequence number information of a second packet, where both the first packet and the second packet are from the target service flow, that is, the first packet and the second packet are different packets in the same service flow.
[0008] In the embodiments of the present application, the forwarding device can determine that a target service flow has failed based on the first sequence number information and the second sequence number information. In practical applications, the target service flow includes multiple packets, and these packets continuously pass through the forwarding device. Then, the forwarding device can execute the communication method in the embodiments of the present application based on these packets to determine whether the target service flow has failed and promptly perceive the failure that occurs in the target service flow. On the other hand, the communication method in the embodiments of the present application does not require modification and extension of the original packet format, has a low implementation complexity, and a high efficiency in detecting failures.
[0009] Based on the first aspect, in an optional implementation, the sequence number information of a packet can be the sequence number (Sequence Number) in the TCP header, and the sequence number of the packet is used to indicate the position of the data content of the packet in the complete data content of the target service flow; or, it can also be the acknowledgment number (Acknowledgment Number) in the TCP header, and the acknowledgment number of the packet is used to indicate the position of the packet that has been successfully received by the destination device in the target service flow; or, it can also be the sequence number and acknowledgment number in the TCP header; or, it can also be the packet sequence number (Packet Sequence Number, PSN) in the remote direct memory access (Remote Direct Memory Access, RDMA) packet, and the PSN of the RDMA packet is used to indicate the order of the RDMA packets sent or received through the RDMA queue pair (Queue Pair, QP).
[0010] Based on the first aspect, in an optional implementation, in a Transmission Control Protocol (TCP) scenario, the source IP address of the first packet is the same as the destination IP address of the second packet, the destination IP address of the first packet is the same as the source IP address of the second packet, the source port of the first packet is the same as the destination port of the second packet, and the destination port of the first packet is the same as the source port of the second packet. The first sequence number information of the first packet at least includes the sequence number (Sequence Number) in the TCP header of the first packet, and the second sequence number information of the second packet at least includes the acknowledgment number (Acknowledgment Number) in the TCP header of the second packet. Optionally, in practical applications, the transport layer protocols of the first packet and the second packet are the same.
[0011] Optionally, the first message and the second message are respectively the latest messages obtained by the forwarding device in two opposite transmission directions. Suppose the target service flow includes messages in two opposite transmission directions. If the first message is the latest message obtained by the forwarding device in one transmission direction, the second message is the latest message obtained by the forwarding device (or another device that has established a peer link with this forwarding device) in the other transmission direction.
[0012] The forwarding device obtains the first sequence number information (i.e., the sequence number of the first message) and the second sequence number information (i.e., the acknowledgment number of the second message). If the sequence number in the TCP header of the first message is greater than the acknowledgment number in the TCP header of the second message, and the difference between the current time and the first arrival time of the second message is greater than a preset threshold, the forwarding device determines that the target service flow has failed. Here, the current time is the time when the forwarding device performs the failure determination of the target service flow. In practical applications, the frequency, period, and time of the forwarding device's failure determination, as well as the size of the preset threshold, can be adaptively configured according to service requirements and network environments to improve the flexibility of the solution. Specifically, for the two opposite transmission directions of the target service flow, the sequence number of the message in one direction corresponds to the acknowledgment number of the message in the other direction. If the difference between the current time and the first arrival time of the second message is greater than the preset threshold, it means that in the transmission direction of the second message of the forwarding device, no message of the target service flow has been received for a long time. And the sequence number in the TCP header of the first message is greater than the acknowledgment number in the TCP header of the second message, which means that in the transmission direction of the first message, the target service flow is continuously transmitting. That is, at this time, it can be excluded that the long-term non-receipt of the message of the target service flow in the transmission direction of the second message of the forwarding device is caused by the stop of the target service flow. Therefore, the forwarding device can determine that the target service flow has failed.
[0013] Based on the first aspect, in an optional implementation manner, in the TCP scenario, the source IP address of the first message is the same as the source IP address of the second message, the destination IP address of the first message is the same as the destination IP address of the second message, the source port of the first message is the same as the source port of the second message, and the destination port of the first message is the same as the destination port of the second message. The first sequence number information of the first message includes the sequence number and acknowledgment number in the TCP header of the first message, and the second sequence number information of the second message includes the sequence number and acknowledgment number in the TCP header of the second message. Optionally, in practical applications, the transport layer protocols of the first message and the second message are the same.
[0014] The first packet is the latest packet obtained by the forwarding device currently, and the second packet is the previous packet of the first packet in the same transmission direction. That is, the forwarding device first receives the second packet and then receives the first packet. After the first packet arrives at the forwarding device, the forwarding device takes the time when the first packet arrives at the forwarding device as the current time to determine whether the target service flow fails. If the sequence number in the TCP header of the first packet is less than or equal to the sequence number in the TCP header of the second packet, and the acknowledgment number in the TCP header of the first packet is less than or equal to the acknowledgment number in the TCP header of the second packet, and the difference between the current time and the first arrival time of the second packet is greater than a preset threshold, it indicates that the first packet received by the forwarding device is a retransmitted packet, and the forwarding device can determine that the target service flow has failed. Here, the current time is the time when the first packet arrives at the forwarding device.
[0015] Based on the first aspect, in an optional implementation manner, in the TCP scenario, the source IP address of the first packet is the same as the destination IP address of the second packet, the destination IP address of the first packet is the same as the source IP address of the second packet, the source port of the first packet is the same as the destination port of the second packet, and the destination port of the first packet is the same as the source port of the second packet. The first sequence number information of the first packet includes the sequence number (Sequence Number) and acknowledgment number (Acknowledgment Number) in the TCP header of the first packet, and the second sequence number information of the second packet at least includes the sequence number (Sequence Number) and acknowledgment number (Acknowledgment Number) in the TCP header of the second packet.
[0016] In a possible implementation, the first packet and the second packet are respectively the latest packets obtained by the forwarding device in two opposite transmission directions. Taking the first packet as the upstream packet and the second packet as the downstream packet as an example, the first packet is the latest upstream packet obtained by the forwarding device in the upstream direction, and the second packet is the latest downstream packet obtained by the forwarding device (or other devices that establish a peer link with this forwarding device) in the downstream direction.
[0017] The forwarding device obtains the first sequence number information (i.e., the sequence number and acknowledgment number of the first packet) and the second sequence number information (i.e., the sequence number and acknowledgment number of the second packet). If the acknowledgment number in the TCP header of the first packet is less than the sequence number in the TCP header of the second packet, and the difference between the current time and the second arrival time of the first packet is greater than a preset threshold, it indicates that the forwarding device can receive the target traffic flow in the direction of the second packet, but the forwarding device has not received the target traffic flow in the same direction as the first packet for a long time (exceeding the preset duration). Therefore, the acknowledgment number of the first packet stops increasing. So, the forwarding device determines that the target traffic flow has a fault, and the upstream path of the second packet is normal and has no fault. Here, the upstream path of the second packet refers to the path that the second packet passes through from the sending end to this forwarding device. Thus, while sensing the target traffic flow, the forwarding device can further determine the normal path that the target traffic flow passes through, which is convenient for subsequent fault location.
[0018] Based on the first aspect, in an optional implementation, the forwarding device can also determine that the upstream path of the second packet is normal and has no fault. At this time, it indicates that the cause of the fault has nothing to do with the devices on the upstream path of the second packet (i.e., the upstream devices of the second packet). Therefore, the forwarding device can send a fault notice to the upstream devices of the second packet, and this fault notice indicates that the upstream path of the second packet is normal, so that the upstream devices of the second packet can avoid performing ineffective path switching.
[0019] Based on the first aspect, in an optional implementation, the first message and the second message are messages in the RDMA protocol. The source IP address of the first message is the same as the destination IP address of the second message, and the destination IP address of the first message is the same as the source IP address of the second message. The queue pairs (Queue Pair, QP) of the first message and the second message are different. The first sequence number information includes the PSN of the first message, and the second sequence number information includes the PSN of the second message. Specifically, the first message and the second message are messages in different transmission directions in the same RDMA flow. In other words, the first message and the second message have the same queue pair context (Queue Pair Context, QPC) and queue pair key (Queue Pair Key, QKey), but the destination QP (Destination QP) of the first message and the destination QP of the second message are different. Optionally, the first message can be a request message in the RDMA service flow, and the second message is a response message in the RDMA service flow; or, the second message can be a response message in the RDMA service flow, and the second message is a request message in the RDMA service flow. If the PSN of the first message is greater than the PSN of the second message, and the difference between the current time and the first arrival time of the second message is greater than a preset threshold, it is determined that the target service flow has failed. Specifically, the RDMA protocol provides an acknowledgment mechanism, which is used to ensure that RMDA messages can be reliably received. In the RDMA scenario where the acknowledgment mechanism is enabled, after the destination device of the target service flow receives a request message from the source device, it determines the PSN of the response message of the request message according to the PSN of the request message. The PSN of the response message indicates the position of the data content of the response message in the complete data content of the target service flow, and also indicates the position of the request message that the destination device has successfully received in the target service flow. After the source device receives the response message, it continues to determine the PSN of the new request message according to the PSN in the response message, so as to continue to send the new request message. Therefore, in the RDMA scenario, for the two opposite transmission directions of the target service flow, the PSN of the message in one direction corresponds to the PSN of the message in the other direction. If the difference between the current time and the first arrival time of the second message is greater than the preset threshold, it means that in the transmission direction of the second message, no message of the target service flow has been received for a long time. And the PSN of the first message is greater than the PSN of the second message, which means that in the transmission direction of the first message, the target service flow is continuously transmitting. That is, at this time, it can be excluded that the target service flow stops flowing, resulting in no message of the target service flow being received in the transmission direction of the second message by the forwarding device. Therefore, the forwarding device can determine that the target service flow has failed.
[0020] Based on the first aspect, in an optional implementation, the first message and the second message are of the same type of message in the RDMA protocol. The source IP address of the first message is the same as the source IP address of the second message, the destination IP address of the first message is the same as the destination IP address of the second message, the queue pair (QP) of the first message is the same as that of the second message. The first sequence number information includes the PSN of the first message, and the second sequence number information includes the PSN of the second message.
[0021] The first message is the latest RDMA message currently obtained by the forwarding device, and the second message is the previous RDMA message of the first message in the same transmission direction. That is, the forwarding device first receives the second message and then receives the first message. After the first message arrives at the forwarding device, the forwarding device uses the time when the first message arrives at the forwarding device as the current time to determine whether the target traffic flow fails. In the RDMA protocol, the PSN of the RDMA message indicates the position of the data content of the RDMA message in the complete data content of the target traffic flow. Therefore, if the PSN of the first message is less than or equal to the PSN of the second message, and the difference between the current time and the first arrival time of the second message is greater than the preset threshold, it means that the first message received by the forwarding device is a retransmitted message, and the forwarding device can determine that the target traffic flow has failed. Here, the current time is the time when the first message arrives at the forwarding device.
[0022] Based on the first aspect, in an optional implementation, the source IP address of the first packet is the same as the destination IP address of the second packet, the destination IP address of the first packet is the same as the source IP address of the second packet, the QP of the first packet is different from the QP of the second packet, the first sequence number information includes the PSN of the first packet, and the second sequence number information includes the PSN of the second packet. Specifically, the first packet and the second packet are packets in different transmission directions in the same RDMA flow. In other words, the first packet and the second packet have the same queue pair context (QPC) and queue pair key (QKey), but the destination QP (Destination QP) of the first packet is different from the destination QP of the second packet. Optionally, the first packet may be a request packet in the RDMA service flow, and the second packet is a response packet in the RDMA service flow; or, the second packet may be a response packet in the RDMA service flow, and the second packet is a request packet in the RDMA service flow. As can be seen from the above, in the RDMA scenario where the acknowledgment mechanism is enabled, the PSN of the RDMA packet indicates the position of the data content of the RDMA packet in the complete data content of the target service flow, and also indicates the position of the packet that the destination device has successfully received in the target service flow. That is, for the two opposite transmission directions of the target service flow, the PSN of the packet in one direction corresponds to the PSN of the packet in the other direction. If the PSN of the first packet is less than the PSN of the second packet, and the difference between the current time and the second arrival time of the first packet is greater than the preset threshold, it means that the target service flow in the direction where the forwarding device can receive the second packet exists, but the forwarding device has not received the target service flow in the same direction as the first packet for a long time (exceeding the preset duration). Therefore, the PSN of the first packet stops growing. So, the forwarding device determines that the target service flow has failed, and the upstream path of the second packet is normal and has not failed. The upstream path of the second packet refers to the path passed through during the process of the second packet being transmitted to the forwarding device.
[0023] Based on the first aspect, in an optional implementation, the second packet in the target service flow is transmitted to the forwarding device, and the second packet includes the second sequence number information. After receiving the second packet, the forwarding device can obtain the second sequence number information of the second packet. On the other hand, the forwarding device records the time when the second packet arrives at the forwarding device, so as to obtain the first arrival time of the second packet. In other words, the first arrival time of the second packet is the time when the second packet arrives at the forwarding device.
[0024] Based on the first aspect, in an optional implementation, the forwarding device establishes peer links with several other forwarding devices. The second packet in the target traffic flow is not delivered to this forwarding device, but to other forwarding devices that have established peer links with this forwarding device. Then, this forwarding device can synchronously receive the second sequence number information of the second packet through the peer link. Thus, the forwarding device can determine the fault of the target traffic flow based on the packets in the target traffic flow received by other forwarding devices, making the communication method in the embodiments of this application still applicable to the scenario where the transmission paths of the first packet and the second packet are inconsistent, improving the flexibility of the solution and increasing the applicable scenarios of the solution. The first arrival time of the second packet can be the time when the second packet is delivered to the peer link of this forwarding device, or it can also be the time when this forwarding device receives the second sequence number information. Specifically, it is not limited here.
[0025] Based on the first aspect, in an optional implementation, after the forwarding device determines that the target traffic flow has a fault, it generates fault information for the target traffic flow, and the fault information indicates that the target traffic flow has a fault.
[0026] Optionally, the fault information includes but is not limited to the time when the target traffic flow is determined to have a fault, the first sequence number information of the first packet, the second sequence number information of the second packet, and the five-tuple (or triple) of the target traffic flow. Exemplarily, in the TCP scenario, the five-tuple can be the source IP address, destination IP address, source port, destination port, and transport layer protocol of the target traffic flow; in the RDMA scenario, the triple can be the source IP address, destination IP address, and QP of the target traffic flow.
[0027] Optionally, the fault information of the target traffic flow can be queried by the management personnel in real time, or the forwarding device can further send the fault information for the target traffic flow to the controller, and the fault information is used to indicate that the target traffic flow has a fault. After receiving the fault information for the target traffic flow, the controller can analyze the current fault of the target traffic flow.
[0028] In the second aspect, this application provides a communication device, which includes:
[0029] A transceiver unit, configured to receive a first packet, where the first packet includes first sequence number information;
[0030] The transceiver unit is further configured to obtain the second sequence number information of a second packet, where the first packet and the second packet correspond to a target traffic flow;
[0031] A processing unit, configured to determine that the target traffic flow has a fault according to the first sequence number information and the second sequence number information.
[0032] Based on the second aspect, in an optional implementation, the source Internet Protocol (IP) address of the first packet is the same as the destination IP address of the second packet, the destination IP address of the first packet is the same as the source IP address of the second packet, the source port of the first packet is the same as the destination port of the second packet, the destination port of the first packet is the same as the source port of the second packet, the first sequence number information includes the sequence number in the Transmission Control Protocol (TCP) header of the first packet, and the second sequence number information includes the acknowledgement number in the TCP header of the second packet;
[0033] The processing unit is specifically configured to: when the sequence number in the TCP header of the first packet is greater than the acknowledgement number in the TCP header of the second packet, and the difference between the current time and the first arrival time of the second packet is greater than a preset threshold, determine that the target service flow has a fault.
[0034] Based on the second aspect, in an optional implementation, the source IP address of the first packet is the same as the source IP address of the second packet, the destination IP address of the first packet is the same as the destination IP address of the second packet, the source port of the first packet is the same as the source port of the second packet, the destination port of the first packet is the same as the destination port of the second packet, the first sequence number information includes the sequence number and the acknowledgement number in the TCP header of the first packet, and the second sequence number information includes the sequence number and the acknowledgement number in the TCP header of the second packet;
[0035] The processing unit is specifically configured to: when the sequence number in the TCP header of the first packet is less than or equal to the sequence number in the TCP header of the second packet, and the acknowledgement number in the TCP header of the first packet is less than or equal to the acknowledgement number in the TCP header of the second packet, and the difference between the current time and the first arrival time of the second packet is greater than a preset threshold, determine that the target service flow has a fault.
[0036] Based on the second aspect, in an optional implementation, the source IP address of the first packet is the same as the destination IP address of the second packet, the destination IP address of the first packet is the same as the source IP address of the second packet, the source port of the first packet is the same as the destination port of the second packet, the destination port of the first packet is the same as the source port of the second packet, the first sequence number information includes the sequence number and the acknowledgement number in the TCP header of the first packet, and the second sequence number information includes the sequence number and the acknowledgement number in the TCP header of the second packet;
[0037] The processing unit is specifically configured to: when the acknowledgement number in the TCP header of the first packet is less than the sequence number in the TCP header of the second packet, and the difference between the current time and the second arrival time of the first packet is greater than a preset threshold, determine that the upstream path of the second packet is normal.
[0038] Based on the second aspect, in an optional implementation, the transceiver unit is further configured to send a fault notice to the upstream device of the second message, and the fault notice is used to indicate that the upstream path of the second message is normal.
[0039] Based on the second aspect, in an optional implementation, the first message and the second message are messages in the Remote Direct Memory Access (RDMA) protocol. The source IP address of the first message is the same as the source IP address of the second message, the destination IP address of the first message is the same as the destination IP address of the second message, the queue pair (QP) of the first message is the same as the QP of the second message, the first sequence number information includes the packet sequence number (PSN) of the first message, and the second sequence number information includes the PSN of the second message;
[0040] The processing unit is specifically configured to: when the PSN of the first message is less than or equal to the PSN of the second message, and the difference between the current time and the first arrival time of the second message is greater than a preset threshold, determine that the target traffic flow has a fault.
[0041] Based on the second aspect, in an optional implementation, the first message and the second message are messages in the Remote Direct Memory Access (RDMA) protocol. The source IP address of the first message is the same as the destination IP address of the second message, the destination IP address of the first message is the same as the source IP address of the second message, the QP of the first message is different from the QP of the second message, the first sequence number information includes the PSN of the first message, and the second sequence number information includes the PSN of the second message;
[0042] The processing unit is specifically configured to: when the PSN of the first message is greater than the PSN of the second message, and the difference between the current time and the first arrival time of the second message is greater than a preset threshold, determine that the target traffic flow has a fault.
[0043] Based on the second aspect, in an optional implementation, the source IP address of the first message is the same as the destination IP address of the second message, the destination IP address of the first message is the same as the source IP address of the second message, the QP of the first message is different from the QP of the second message, the first sequence number information includes the PSN of the first message, and the second sequence number information includes the PSN of the second message;
[0044] The processing unit is specifically configured to: when the PSN of the first message is less than the PSN of the second message, and the difference between the current time and the second arrival time of the first message is greater than a preset threshold, determine that the upstream path of the second message is normal.
[0045] Based on the second aspect, in an optional implementation, the transceiver unit is specifically configured to receive a second message, and the second message includes second sequence number information.
[0046] Based on the second aspect, in an optional implementation, the second serial number information is received through a peer link.
[0047] Based on the second aspect, in an optional implementation, the transceiver unit is further configured to send fault information for a target service flow to a controller, where the fault information is used to indicate that the target service flow has a fault.
[0048] The content such as the information interaction and execution process of the embodiments shown in this aspect is based on the same concept as the embodiments shown in the first aspect. Therefore, for the description of the beneficial effects shown in this aspect, please refer to the above-mentioned first aspect for details, and will not be elaborated here specifically.
[0049] In a third aspect, the present application provides a communication device, including: a processor, the processor is coupled to a memory, and the memory is used to store instructions. When the instructions are executed by the processor, the computing device implements the method in the above-mentioned first aspect or any possible implementation manner of the first aspect.
[0050] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, on which instructions are stored. When the instructions are executed, the computer executes the method in the above-mentioned first aspect or any possible implementation manner of the first aspect.
[0051] In a fifth aspect, an embodiment of the present application provides a computer program product, in which computer-readable instructions are stored. When the computer-readable instructions are executed by a processor, the method in the first aspect or any possible implementation manner of the first aspect is implemented.
[0052] In a sixth aspect, an embodiment of the present application provides a chip, including: a processor, the processor is coupled to a memory, and the memory is used to store instructions. When the instructions are executed by the processor, the chip implements the method in the above-mentioned first aspect or any possible implementation manner of the first aspect.
[0053] Among them, the technical effects brought by any implementation manner in the third aspect to the sixth aspect can be referred to the technical effects brought by the implementation manner in the above-mentioned first aspect, and will not be elaborated here. BRIEF DESCRIPTION OF THE DRAWINGS
[0054] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only the embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained according to the provided drawings without creative efforts.
[0055] Figure 1 It is a schematic diagram of a BFD scenario;
[0056] Figure 2 It is a schematic diagram of a possible and non - restrictive network architecture for the communication method in the embodiments of the present application;
[0057] Figure 3 It is a schematic flowchart of a communication method in the embodiments of the present application;
[0058] Figure 4 It is a possible scenario example diagram for a forwarding device to obtain second sequence number information in the embodiments of the present application;
[0059] Figure 5 It is a possible scenario schematic diagram of the communication method in the embodiments of the present application;
[0060] Figure 6 It is another possible scenario schematic diagram of the communication method in the embodiments of the present application;
[0061] Figure 7 It is another possible scenario schematic diagram of the communication method in the embodiments of the present application;
[0062] Figure 8 It is a schematic diagram of the communication method in the embodiments of the present application applied to the RDMA scenario;
[0063] Figure 9 It is a possible scenario schematic diagram of the communication method in the embodiments of the present application;
[0064] Figure 10 It is a schematic structural diagram of a communication device provided by the embodiments of the present application;
[0065] Figure 11 It is a schematic logical structure diagram of a communication device provided by the embodiments of the present application. Detailed implementation manners
[0066] The embodiments of the present application provide a communication method, a communication device, and a communication device for detecting service flow failures.
[0067] The following describes the embodiments of the present application in conjunction with the accompanying drawings in the embodiments of the present application. The terms used in the implementation part of the present application are only used to explain the specific embodiments of the present application, and are not intended to limit the embodiments of the present application. Those of ordinary skill in the art know that with the development of technology and the emergence of new scenarios, the technical solutions provided by the embodiments of the present application are equally applicable to similar technical problems.
[0068] In the embodiments of the present application, "at least one" means one or more, and "multiple" means two or more. "And / or" describes the association relationship of associated objects, indicating that three relationships can exist. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone, where A and B can be singular or plural. "At least one (item)" or its similar expression refers to any combination of these items, including any combination of single item(s) or plural item(s). For example, at least one (item) of a, b, or c can represent: a, b, c, a - b, a - c, b - c, or a - b - c, where a, b, and c can be single or multiple.
[0069] In the description and claims of the present application and the above-mentioned drawings, terms such as "first", "second", "third", "fourth", etc. (if any) are used to distinguish similar objects and do not necessarily describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances so that the embodiments of the present application described here, for example, can be implemented in an order other than those illustrated or described here. In addition, the terms "comprise" and "have" and any of their variants are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that comprises a series of steps or units does not necessarily limit to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products, or devices.
[0070] Some nouns or terms used in the embodiments of the present application are explained below, and this noun or term is also part of the invention content.
[0071] Transmission Control Protocol (TCP): TCP is an important transport layer protocol of the Internet. TCP provides connection-oriented, reliable, ordered, and byte-stream transmission services. Before an application uses TCP, it must first establish a TCP connection; TCP achieves reliable transmission through mechanisms such as checksum, sequence number, acknowledgment, retransmission control, connection management, and window control.
[0072] Remote Direct Memory Access (RDMA): The RDMA technology is a high-performance network communication technology with advantages such as high bandwidth, low latency, no Central Processing Unit (CPU) overhead, and zero copy. RDMA directly transfers data into the storage area of a computer through the network, quickly moving data from one system to the remote system memory without any impact on the operating system, so that not much computer processing power is required. RDMA eliminates the overhead of external memory copying and context switching, thus liberating memory bandwidth and CPU cycles for improving the performance of application systems.
[0073] The Sequence Number field in the TCP header is a 32-bit unsigned integer, and the Sequence Number field plays a very crucial role in the TCP data transmission process. The sequence number is used to identify the unique position of each byte in the data stream sent from the sender to the receiver. Each time data is sent, the sequence number of the TCP packet is incremented so that the receiver can reassemble the received data in the correct order. The sequence number in the TCP header is necessary because network transmission may disrupt the order of data. By using the sequence number, the receiver can determine the correct position of each packet in the data stream and reassemble them to restore the original data.
[0074] The Acknowledgment Number field in the TCP header is also a 32-bit unsigned integer. In TCP communication, the acknowledgment number is used to indicate the position of the packets that the receiver has successfully received in the data stream. After the receiver correctly receives the data, it sends a TCP packet with an acknowledgment number as an acknowledgment response for the received data, so that the sender knows which data has been correctly received by the receiver, and the sender can continue to send subsequent data.
[0075] Gray Failure: Gray failure refers to those failures that are not easily detected or located. These failures usually do not immediately cause the system to completely crash or stop service, but may have potential negative impacts on system performance, stability, and availability, and may lead to problems such as data inconsistency or service quality degradation, for example, problems such as performance degradation, random packet loss, memory jitter, and non-fatal exceptions.
[0076] Next, the possible application scenarios involved in the embodiments of this application will be introduced.
[0077] In a real network environment, there are often gray failures such as silent failures, misconfigurations, routing black holes, hardware failures, or physical port false deaths. Currently, there are multiple definitions for gray failures. One definition is: "When at least one application observes that the system is unhealthy while the system fault detection tool observes that the system is healthy, the system is defined as experiencing a gray failure." Another definition is: "A gray failure is a hardware failure that causes non-instantaneous packet loss of the traffic forwarded on a forwarding device (switch), and congestion does not belong to a gray failure." The commonality of the two definitions is that they both exclude link failures (such as scenarios like port failure or link disconnection) from gray failures.
[0078] In a real network environment, gray failures are a major type of network failure, and they have various manifestations, such as performance degradation, random packet loss, and memory jitter. Currently, there is no effective sensing solution for these gray failures.
[0079] Bidirectional Forwarding Detection (BFD) is a network protocol used to quickly detect and monitor the forwarding connectivity of links or Internet Protocol (IP) routes in a network. BFD can improve the performance of the existing network. By quickly detecting communication failures, it enables network devices to establish backup channels and resume communication more quickly.
[0080] Between two network devices at both ends, BFD packets are periodically transmitted on the connected link. If a certain network device does not receive a BFD packet within a preset time period, then this network device can determine that the link connecting to the peer network device has failed. Please refer to Figure 1 , Figure 1 for a schematic diagram of a BFD scenario. As Figure 1 shown, there are two reachable links (Link 1 and Link 2) between Network Device A and Network Device B. Then Network Device A periodically sends BFD packets to Network Device B through Link 1 and Link 2 respectively. If Network Device A does not receive a BFD packet from Network Device B through Link 1 within the preset time period, then Network Device A determines that Link 1 connecting to Network Device B has failed; if Network Device A receives a BFD packet from Network Device B through Link 2, then Network Device A determines that Link 2 connecting to Network Device B is normal.
[0081] BFD can only detect the integrity of the link (i.e., whether the link has failed). However, the occurrence of gray failures is mainly due to reasons such as misconfiguration of the routing table on the network device, resulting in the inconsistency between the actual forwarding path and the expected forwarding path. That is, the occurrence of gray failures has nothing to do with the integrity of the link. Therefore, BDF cannot sense the failures that occur to the data stream.
[0082] On the other hand, BFD fault detection requires network devices to periodically generate BFD packets, which requires a certain amount of resource overhead. Moreover, in practical applications, the BFD fault detection and convergence time generally takes about 200 milliseconds. The BFD solution cannot meet the latency requirements for applications with low latency requirements (such as latency within 10 milliseconds) such as games, autonomous driving, or news live broadcasts.
[0083] In view of this, the embodiments of the present application provide a communication method, a communication device, and a communication device for detecting service flow faults. For the convenience of understanding, first, a possible and non-limiting network architecture of the communication method in the embodiments of the present application is introduced. Please refer to Figure 2 , Figure 2 is a schematic diagram of a possible and non-limiting network architecture of the communication method in the embodiments of the present application. In the Figure 2 illustrated scenario, the network architecture includes Server 1, Server 2, Switch 1, Switch 2, Switch 3, Switch 4, a controller, and a network. Among them, between Server 1 and Server 2, communication is carried out through a plurality of forwarding devices (such as Figure 2 the illustrated Switch 1, Switch 2, Switch 3, and Switch 4). As shown in Figure 2 , the packets sent by Server 1 reach Server 2 after being relayed by Switch 1 and Switch 3; the packets sent by Server 2 reach Server 1 after being relayed by Switch 4 and Switch 2. The controller, as an analyzer, is used to receive the fault information of the service flows from each forwarding device in order to analyze the service flows with faults.
[0084] It should be understood that the above Figure 2 illustrated network architecture is only an exemplary description. In practical applications, faults may occur on any forwarding path in any network structure. The communication method of the embodiments of the present application is applicable to at least one forwarding device on the transmission path of a service flow (packet) to detect service flow faults. Among them, a forwarding device is a communication entity used to send signals, or receive signals, or send and receive signals. Optionally, the communication entity may be a router, a switch, a virtual switch, a virtual router, or a smart network card, etc., and specific details are not limited here. Next, the communication method in the embodiments of the present application is introduced. Please refer to Figure 3 , Figure 3 is a schematic flowchart of the communication method in the embodiments of the present application. In the communication method of the embodiments of the present application, the forwarding device is used as the execution subject to illustrate the method, but the present application does not limit the execution subject of this interaction schematic, nor does it limit the specific hardware form or software form of the forwarding device. For example, Figure 3The forwarding device in Figure 3 can be a chip, a chip system, or a processor used to support the forwarding device in implementing this method; alternatively, the forwarding device can also be a logical node, a logical module, or software used to implement all or part of the functions of the forwarding device; the forwarding device can also be a collective term for multiple logical nodes, multiple logical modules, or multiple software used to implement the communication method. As
[0085] 101. The forwarding device receives a first packet.
[0086] The forwarding device is a transit device on the path from the source device to the destination device for the target traffic flow. After the packets in the target traffic flow are sent from the source device and forwarded by the forwarding device, they can reach the destination device. Among them, each packet in the target traffic flow includes the sequence number information corresponding to the packet. In the embodiments of this application, the sequence number information of the packet is used to implement important functions such as the ordered transmission of packets, retransmission of lost packets, and error recovery, thereby ensuring the reliable transmission of the traffic flow in the network. In a possible implementation, the sequence number information of the packet can be the sequence number (Sequence Number) in the TCP header, and the sequence number of the packet is used to indicate the position of the data content of the packet in the complete data content of the target traffic flow; or, it can also be the acknowledgment number (Acknowledgment Number) in the TCP header, and the acknowledgment number of the packet is used to indicate the position of the packet that the destination device has successfully received in the target traffic flow; or, it can also be the sequence number and acknowledgment number in the TCP header; or, it can also be the packet sequence number (Packet Sequence Number, PSN) in the remote direct memory access (Remote Direct Memory Access, RDMA) packet, and the PSN of the RDMA packet is used to indicate the order of the RDMA packets sent or received through the RDMA queue pair (QueuePair, QP).
[0087] In the embodiments of this application, the first packet in the target traffic flow is transmitted to the forwarding device, and the first packet includes the first sequence number information corresponding to the first packet. After receiving the first packet, the forwarding device sends the first packet to the next-hop device.
[0088] 102. The forwarding device obtains the second sequence number information of the second packet.
[0089] The forwarding device obtains the second sequence number information of the second packet. Here, the first packet and the second packet both come from the target service flow, that is, the first packet and the second packet are different packets in the same service flow. Between the first packet and the second packet, they can be in the same transmission direction, that is, the source IP address of the first packet is the same as the source IP address of the second packet, and the destination IP address of the first packet is the same as the destination IP address of the second packet; or, between the first packet and the second packet, they can be in different transmission directions, that is, the source IP address of the first packet is the same as the destination IP address of the second packet, and the destination IP address of the first packet is the same as the source IP address of the second packet.
[0090] In the embodiments of the present application, the execution order of step 101 and step 102 is not limited. For example, step 101 can be executed first, and then step 102; or, step 102 can be executed first, and then step 101; or, step 101 and step 102 can be executed simultaneously, which is not specifically limited here.
[0091] In a possible implementation, the second packet in the target service flow is transmitted to the forwarding device, and the second packet includes the second sequence number information. After the forwarding device receives the second packet, it can obtain the second sequence number information of the second packet. On the other hand, the forwarding device records the time when the second packet arrives at this forwarding device, so as to obtain the first arrival time of the second packet. In other words, the first arrival time of the second packet is the time when the second packet arrives at this forwarding device.
[0092] In a possible implementation, the forwarding device establishes peer links with several other forwarding devices. The second packet in the target service flow is not transmitted to this forwarding device, but to other forwarding devices that have established peer links with this forwarding device. Then this forwarding device can synchronously receive the second sequence number information of the second packet through the peer link. Thus, the forwarding device can perform fault determination on the target service flow based on the packets in the target service flow received by other forwarding devices, making the communication method in the embodiments of the present application still applicable to the scenario where the transmission paths of the first packet and the second packet are inconsistent, improving the flexibility of the solution and increasing the applicable scenarios of the solution. And the first arrival time of the second packet can be the time when the second packet is transmitted to the peer link of this forwarding device, or, it can also be the time when this forwarding device receives the second sequence number information, which is not specifically limited here.
[0093] For ease of understanding, please refer to Figure 4 , Figure 4 which is a possible scenario example diagram for the forwarding device to obtain the second sequence number information in the embodiments of the present application. As Figure 4As shown in the figure, the forwarding device A establishes a peer link with the forwarding device B. The first packet is delivered to the forwarding device A, and the second packet is delivered to the forwarding device B. Then, the forwarding device B can synchronize the second sequence number information of the second packet to the forwarding device A through the peer link, so that the forwarding device A can obtain the second sequence number information of the second packet; alternatively, the forwarding device A can also synchronize the first sequence number information of the first packet to the forwarding device B through the peer link, so that the forwarding device B can obtain the first sequence number information of the first packet.
[0094] 103. The forwarding device determines that the target service flow has a fault based on the first sequence number information and the second sequence number information.
[0095] In the embodiment of this application, the forwarding device can determine that the target service flow has a fault based on the first sequence number information and the second sequence number information. In practical applications, the target service flow includes multiple packets, and these packets will continuously pass through the forwarding device. Then, the forwarding device can execute the communication method in the embodiment of this application based on these packets to determine whether the target service flow has a fault and sense the fault that occurs in the target service flow in a timely manner. On the other hand, the communication method in the embodiment of this application does not require modification and extension of the original packet format, has a low implementation complexity, and a high efficiency in detecting faults.
[0096] In a possible implementation, after the forwarding device determines that the target service flow has a fault, it generates fault information for the target service flow, and the fault information indicates that the target service flow has a fault. Specifically, the fault information may include, but is not limited to, the time when the target service flow is determined to have a fault, the first sequence number information of the first packet, the second sequence number information of the second packet, and the five-tuple (or three-tuple) of the target service flow. Exemplarily, the five-tuple may be the source IP address, destination IP address, source port, destination port, and transport layer protocol of the target service flow in the TCP scenario; the three-tuple may be the source IP address, destination IP address, and QP of the target service flow in the RDMA scenario. Optionally, the fault information of the target service flow can be queried by the management personnel in real time, or the forwarding device can further send the fault information for the target service flow to the controller, and the fault information is used to indicate that the target service flow has a fault. After receiving the fault information for the target service flow, the controller can analyze the current fault of the target service flow.
[0097] In the embodiment of this application, after the forwarding device obtains the first sequence number information and the second sequence number information, it can determine whether the target service flow has a fault through various determination logics. Next, the various determination logics in the embodiment of this application will be introduced separately.
[0098] The first determination logic: The forwarding device performs a fault determination of the target service flow based on two packets in different transmission directions in the target service flow.
[0099] In the TCP scenario, the source IP address of the first packet is the same as the destination IP address of the second packet, the destination IP address of the first packet is the same as the source IP address of the second packet, the source port of the first packet is the same as the destination port of the second packet, and the destination port of the first packet is the same as the source port of the second packet. The first sequence number information of the first packet at least includes the sequence number (Sequence Number) in the TCP header of the first packet, and the second sequence number information of the second packet at least includes the acknowledgment number (Acknowledgment Number) in the TCP header of the second packet. Optionally, in practical applications, the transport layer protocols of the first packet and the second packet are the same.
[0100] In a possible implementation, the first packet and the second packet are respectively the latest packets obtained by the forwarding device in two opposite transmission directions. Suppose the target service flow includes packets in two opposite transmission directions. If the first packet is the latest packet obtained by the forwarding device in one transmission direction, and the second packet is the latest packet obtained by the forwarding device (or other devices that establish a peer link with this forwarding device) in the other transmission direction.
[0101] The forwarding device obtains the first sequence number information (i.e., the sequence number of the first packet) and the second sequence number information (i.e., the acknowledgment number of the second packet). If the sequence number in the TCP header of the first packet is greater than the acknowledgment number in the TCP header of the second packet, and the difference between the current time and the first arrival time of the second packet is greater than the preset threshold, then the forwarding device determines that the target service flow has a fault. Here, the current time is the time when the forwarding device performs the fault determination of the target service flow. In practical applications, the frequency, period, and time of the forwarding device's fault determination, as well as the size of the preset threshold, can be adaptively configured according to service requirements and network environment to improve the flexibility of the solution. Specifically, for the two opposite transmission directions of the target service flow, the sequence number of the packets in one direction corresponds to the acknowledgment number of the packets in the other direction. If the difference between the current time and the first arrival time of the second packet is greater than the preset threshold, it means that in the transmission direction of the second packet of the forwarding device, no packets of the target service flow have been received for a long time. And the sequence number in the TCP header of the first packet is greater than the acknowledgment number in the TCP header of the second packet, which means that in the transmission direction of the first packet, the target service flow is continuously transmitting. That is, at this time, it can be excluded that the target service flow stops flowing, resulting in no packets of the target service flow being received in the transmission direction of the second packet of the forwarding device for a long time. Therefore, the forwarding device can determine that the target service flow has a fault.
[0102] In the RMDA scenario, the above first determination logic also applies. Specifically, the first message and the second message are messages in the RDMA protocol. The source IP address of the first message is the same as the destination IP address of the second message, the destination IP address of the first message is the same as the source IP address of the second message, and the queue pairs (Queue Pair, QP) of the first message and the second message are different. Specifically, the first message and the second message are messages in different transmission directions in the same RDMA flow. In other words, the first message and the second message have the same queue pair context (Queue Pair Context, QPC) and queue pair key (Queue Pair Key, QKey), but the destination QP (Destination QP) of the first message and the destination QP of the second message are different. Optionally, the first message can be a request message in the RDMA service flow, and the second message is a response message in the RDMA service flow; or, the second message can be a response message in the RDMA service flow, and the second message is a request message in the RDMA service flow. The first sequence number information includes the PSN of the first message, and the second sequence number information includes the PSN of the second message. If the PSN of the first message is greater than the PSN of the second message, and the difference between the current time and the first arrival time of the second message is greater than a preset threshold, it is determined that the target service flow has failed. Specifically, the RDMA protocol provides an acknowledgment mechanism, which is used to ensure that RMDA messages can be reliably received. In the RDMA scenario where the acknowledgment mechanism is enabled, after the destination device of the target service flow receives a request message from the source device, it determines the PSN of the response message of the request message according to the PSN of the request message. The PSN of the response message indicates the position of the data content of the response message in the complete data content of the target service flow, and also indicates the position of the request message that the destination device has successfully received in the target service flow. After the source device receives the response message, it continues to determine the PSN of the new request message according to the PSN in the response message, so as to continue to send the new request message. Therefore, in the RDMA scenario, for the two opposite transmission directions of the target service flow, the PSN of the message in one direction corresponds to the PSN of the message in the other direction. If the difference between the current time and the first arrival time of the second message is greater than the preset threshold, it means that in the transmission direction of the second message, no message of the target service flow has been received for a long time. And the PSN of the first message is greater than the PSN of the second message, which means that in the transmission direction of the first message, the target service flow is continuously transmitting. That is, at this time, it can be excluded that due to the stop of the target service flow, the forwarding device has not received the message of the target service flow in the transmission direction of the second message for a long time. Therefore, the forwarding device can determine that the target service flow has failed.
[0103] In practical applications, a forwarding device can establish a traffic flow table for the traffic flows arriving at the forwarding device. The traffic flow table records the sequence number information and arrival time of the packets from each traffic flow, so that the forwarding device can perform fault determination based on the first packet and the second packet of each traffic flow.
[0104] Please refer to Figure 5 , Figure 5 a possible schematic diagram of a scenario of the communication method in an embodiment of the present application. As Figure 5 shown, both forwarding device A and forwarding device B can establish a traffic flow table for the traffic flows arriving at the forwarding device. The traffic flow table is used to record the sequence number information and arrival time of the packets from each traffic flow. Assume that the upstream packet and the downstream packet are packets in two opposite directions of the target traffic flow. The upstream packet and the downstream packet are relative. In Figure 5 the example, the packet from forwarding device A to forwarding device B is used as the upstream packet, and the packet from forwarding device B to forwarding device A is used as the downstream packet. Among them, the KEY field in each entry of the traffic flow table is used to record the identifier of the traffic flow received by the forwarding device, such as the five-tuple or three-tuple of the traffic flow; the additional data (AdditionalData, AD) field in the entry corresponding to the traffic flow is used to record the sequence number of the upstream packet (l2r seq in the table), the acknowledgment number of the downstream packet (r2l ack seq in the table), and the arrival time of the downstream packet (r2l timestamp in the table), etc.
[0105] In Figure 5In the scenario shown, the upstream packets and downstream packets of the target service flow share the same entry in the service flow table. Among them, the forwarding device can create a table based on the five-tuple or three-tuple of the upstream packet, or it can also create a table based on the five-tuple or three-tuple of the downstream packet. Next, taking the case where both forwarding device A and forwarding device B create a table based on the five-tuple of the upstream packet as an example, an introduction is made. When the first packet (regardless of whether it is an upstream packet or a downstream packet) of the target service flow arrives at the forwarding device, the forwarding device can identify whether the packet is an upstream packet or a downstream packet, and then create an entry for the target service flow in the service flow table according to the five-tuple of the packet. Suppose the packet is an upstream packet, then record the sequence number of the upstream packet. When a subsequent forwarding device receives a new upstream packet in the target service flow, update the sequence number of the upstream packet in the entry of the target service flow table. If a subsequent forwarding device receives a downstream packet in the target service flow, then swap the source IP address, destination IP address, source port, and destination port of the downstream packet to obtain a five-tuple that matches the entry of the target service flow table, and then update the sequence number and arrival time of the downstream packet in the entry. Next, the forwarding device periodically queries the entries of each service flow in the service flow table, takes the time when a certain service flow is queried as the current time (curTS), and triggers the forwarding device to execute the first determination logic provided in the embodiment of the present application. Specifically, when both l2r seq>r2l ack seq and curTS–r2l timestamp> a preset threshold (Threshold) are satisfied, it is determined that the service flow corresponding to the currently queried entry has a fault. After the forwarding device determines that the service flow has a fault, it generates fault information for the service flow, and the fault information indicates that the service flow has a fault. Specifically, the fault information may include, but is not limited to, the current time (curTS), the sequence number of the first packet (l2r seq), the acknowledgment number of the second packet (r2l ackseq), and the five-tuple (or three-tuple) of the service flow. Optionally, the fault information of the target service flow can be queried in real time by the management personnel, or the forwarding device can further send the fault information for the service flow to the controller, and the fault information is used to indicate that the service flow has a fault. After receiving the fault information for the service flow, the controller can analyze the current fault of the service flow.
[0106] Please refer to Figure 6 , Figure 6 for another possible schematic diagram of the communication method in the embodiment of the present application. As Figure 6 shown, both forwarding device A and forwarding device B can create a service flow table for the service flow arriving at the forwarding device. Suppose the upstream packet and the downstream packet are packets in two opposite directions of the target service flow. The upstream packet and the downstream packet are relative. In Figure 6In the example, the packets in the direction from forwarding device A to forwarding device B are taken as the upstream packets, and the packets in the direction from forwarding device B to forwarding device A are taken as the downstream packets. Among them, the upstream and downstream packets of a traffic flow respectively occupy an entry in the traffic flow table. Each entry records the sequence number of the upstream or downstream packet (l2rseq in the table), the acknowledgement number (r2l ack seq in the table), and the arrival time at the forwarding device (l2r timestamp in the table).
[0107] When the first packet of the target traffic flow (regardless of whether it is an upstream or downstream packet) arrives at the forwarding device, the forwarding device can identify whether the packet is an upstream or downstream packet, and then creates an entry for the target traffic flow in the traffic flow table according to the five-tuple of the packet. Suppose the packet is an upstream packet, then record the sequence number of the upstream packet (l2r seq in the table), the acknowledgement number (r2l ack seq in the table), and the arrival time at the forwarding device (l2r timestamp in the table) in the entry corresponding to the upstream packet. When a subsequent forwarding device receives a new upstream packet in the target traffic flow, it updates the sequence number (l2r seq in the table), the acknowledgement number (r2l ackseq in the table), and the arrival time at the forwarding device (l2r timestamp in the table) in the entry of the upstream packet; if the packet is a downstream packet, then record the sequence number of the downstream packet (l2r seq in the table), the acknowledgement number (r2l ack seq in the table), and the arrival time at the forwarding device (l2r timestamp in the table) in the entry corresponding to the downstream packet. When a subsequent forwarding device receives a new downstream packet in the target traffic flow, it updates the sequence number (l2rseq in the table), the acknowledgement number (r2l ack seq in the table), and the arrival time at the forwarding device (l2r timestamp in the table) in the entry of the downstream packet.
[0108] Next, the forwarding device periodically queries the entries of each traffic flow in the traffic flow table. Take the time when a certain entry is queried as the current time (curTS), and swap the KEY of this entry (by swapping the source IP address and the destination IP address, and swapping the source port and the destination port) to obtain the entry of the traffic flow that belongs to the same traffic as this entry and has the opposite transmission direction. When both seq1 > seq4 and curTS - ts2 > the preset threshold (Threshold) are satisfied, it is determined that the traffic flow corresponding to the currently queried entry has a fault.
[0109] The second determination logic: The forwarding device performs fault determination on the target traffic flow based on two packets in the same transmission direction in the target traffic flow.
[0110] In the TCP scenario, the source IP address of the first packet is the same as that of the second packet, the destination IP address of the first packet is the same as that of the second packet, the source port of the first packet is the same as that of the second packet, and the destination port of the first packet is the same as that of the second packet. The first sequence number information of the first packet includes the sequence number and the acknowledgment number in the TCP header of the first packet, and the second sequence number information of the second packet includes the sequence number and the acknowledgment number in the TCP header of the second packet. Optionally, in practical applications, the transport layer protocols of the first packet and the second packet are the same.
[0111] The first packet is the latest packet obtained by the forwarding device currently, and the second packet is the previous packet of the first packet in the same transmission direction. That is, the forwarding device first receives the second packet and then receives the first packet. After the first packet arrives at the forwarding device, the forwarding device takes the time when the first packet arrives at the forwarding device as the current time to determine whether the target service flow fails. If the sequence number in the TCP header of the first packet is less than or equal to the sequence number in the TCP header of the second packet, and the acknowledgment number in the TCP header of the first packet is less than or equal to the acknowledgment number in the TCP header of the second packet, and the difference between the current time and the first arrival time of the second packet is greater than a preset threshold, it indicates that the first packet received by the forwarding device is a retransmitted packet, and the forwarding device can determine that the target service flow fails. Here, the current time is the time when the first packet arrives at the forwarding device.
[0112] In the RMDA scenario, the above second determination logic is also applicable. Specifically, the first packet and the second packet are the same type of packets in the RDMA protocol, the source IP address of the first packet is the same as that of the second packet, the destination IP address of the first packet is the same as that of the second packet, the queue pair (QueuePair, QP) of the first packet is the same as that of the second packet, the first sequence number information includes the PSN of the first packet, and the second sequence number information includes the PSN of the second packet.
[0113] The first packet is the latest RDMA packet obtained by the forwarding device currently, and the second packet is the previous RDMA packet of the first packet in the same transmission direction. That is, the forwarding device first receives the second packet and then receives the first packet. After the first packet arrives at the forwarding device, the forwarding device uses the time when the first packet arrives at the forwarding device as the current time to determine whether the target traffic flow fails. In the RDMA protocol, the PSN of the RDMA packet indicates the position of the data content of the RDMA packet in the complete data content of the target traffic flow. Therefore, if the PSN of the first packet is less than or equal to the PSN of the second packet, and the difference between the current time and the first arrival time of the second packet is greater than a preset threshold, it means that the first packet received by the forwarding device is a retransmitted packet, and the forwarding device can determine that the target traffic flow has failed. Here, the current time is the time when the first packet arrives at the forwarding device.
[0114] In practical applications, the forwarding device can establish a traffic flow table for the packets in a single transmission direction of the traffic flow. The traffic flow table records the sequence number information and arrival time of the packets in a single transmission direction from each traffic flow, so that the forwarding device can perform fault determination based on the first packet and the second packet of each traffic flow.
[0115] Please refer to Figure 7 , Figure 7 which is another possible scenario schematic diagram of the communication method in the embodiments of this application. As Figure 7 shown, both the forwarding device A and the forwarding device B can establish a traffic flow table for the traffic flow arriving at the forwarding device. The traffic flow table is used to record the sequence number information and arrival time of the packets in a single transmission direction from each traffic flow. In the Figure 7 scenario shown as an example, the forwarding device only establishes a traffic flow table for the upstream packets of the traffic flow. Taking the example that both the forwarding device A and the forwarding device B build tables according to the five-tuple of the upstream packets, the upstream packets and the downstream packets are relative. In Figure 7In the example, the packets in the direction from forwarding device A to forwarding device B are used as upstream packets, and the packets in the direction from forwarding device B to forwarding device A are used as downstream packets. Among them, the KEY field in each entry of the service flow table is used to record the identifier of the service flow received by the forwarding device, such as the five-tuple or three-tuple of the service flow; the Additional Data (AD) field in the entry corresponding to the service flow is used to record the sequence number of the upstream packet (l2r seq in the table), the acknowledgment number of the upstream packet (l2r ack seq in the table), and the arrival time of the upstream packet (l2r timestamp in the table), etc. In practical applications, the forwarding device can build a table according to the five-tuple or three-tuple of the upstream packet, and the forwarding device determines faults based on two packets in the upstream direction; alternatively, it can also build a table according to the five-tuple or three-tuple of the downstream packet, and the forwarding device determines faults based on two packets in the downstream direction.
[0116] When the first upstream packet of the target service flow arrives at the forwarding device, the forwarding device can identify that the packet belongs to an upstream packet, and then creates an entry for the target service flow in the service flow table according to the five-tuple of the upstream packet. This entry includes the sequence number of the upstream packet (l2r seq in the table), the acknowledgment number of the upstream packet (l2r ack seq in the table), and the arrival time of the upstream packet (l2r timestamp in the table). When the forwarding device subsequently receives a new upstream packet in the target service flow, it triggers the forwarding device to execute the second determination logic provided in the embodiments of this application. Among them, the forwarding device uses the time when the new upstream packet arrives at the forwarding device as the current time (curTS), uses the packet recorded in the service flow table as the second packet, and uses the newly received packet as the first packet. Then, the sequence number of the packet recorded in the service flow table is the second sequence number of the second packet (l2r seq in the table), the acknowledgment number of the packet recorded in the service flow table is the second acknowledgment number of the second packet (l2r ack seq in the table), the arrival time of the packet recorded in the service flow table is the first timestamp of the second packet (l2r timestamp in the table), the sequence number of the newly received packet is the first sequence number of the first packet (seq), and the acknowledgment number of the newly received packet is the first acknowledgment number of the first packet (ack seq).
[0117] Specifically, when seq <= l2r seq, and ack seq <= l2r ack seq, and curTS - l2rtimestamp > the preset threshold (Threshold), it is determined that a failure has occurred in the service flow corresponding to the currently queried table entry. If the above conditions are not met, it is determined that no failure has occurred in the current target service flow, and the newly received packet will be used as the second packet, and the sequence number, acknowledgment number, and arrival time of the newly received packet will be used to update the sequence number, acknowledgment number, and arrival time of the table entry for the target service flow in the service flow table.
[0118] After the forwarding device determines that a service flow has failed, it generates failure information for the service flow, which indicates that the service flow has failed. Specifically, the failure information may include, but is not limited to, the current time (curTS), the sequence number (seq) of the first packet, the acknowledgment number (ack seq) of the first packet, the sequence number (l2r seq) of the second packet, the acknowledgment number (l2r ack seq) of the second packet, and the five-tuple (or triple) of the service flow. Optionally, the failure information of the target service flow can be queried in real time by the management personnel, or the forwarding device can further send the failure information for the service flow to the controller, and the failure information is used to indicate that a failure has occurred in the service flow. After receiving the failure information for the service flow, the controller can analyze this failure of the service flow.
[0119] In the RMDA scenario, the above second determination logic is also applicable. Specifically, the first packet and the second packet are of the same type of packet in the RDMA protocol, the source IP address of the first packet is the same as the source IP address of the second packet, the destination IP address of the first packet is the same as the destination IP address of the second packet, the queue pair (QueuePair, QP) of the first packet is the same as that of the second packet, the first sequence number information includes the PSN of the first packet, and the second sequence number information includes the PSN of the second packet. If the PSN of the first packet is less than or equal to the PSN of the second packet, and the difference between the current time and the arrival time is greater than the preset threshold, it is determined that the target service flow has failed. Specifically, RDMA packets include, but are not limited to, various packet types such as SEND packets, RECEIVE packets, READ packets, and WRITE packets, and the PSNs of different types of RDMA packets are independent of each other. In the communication method of the embodiments of the present application, failure determination is performed based on two packets of the same type (the first packet and the second packet) in the RDMA protocol.
[0120] Please refer to Figure 8 , Figure 8 which is a schematic diagram of the application of the communication method in the RDMA scenario in the embodiments of the present application. As Figure 8As shown, both forwarding device A and forwarding device B can establish a traffic flow table for the traffic flows arriving at the forwarding device. The traffic flow table is used to record the sequence number information and arrival time of the packets in a single transmission direction from each traffic flow. In Figure 8 In the example scenario shown, the forwarding device only establishes a traffic flow table for the upstream packets of the traffic flow. The upstream packets and downstream packets are relative. In Figure 8 the example, the packets in the direction from forwarding device A to forwarding device B are used as upstream packets, and the packets in the direction from forwarding device B to forwarding device A are used as downstream packets. Taking the example that both forwarding device A and forwarding device B build tables according to the triple of a certain RDMA type of upstream packets, the introduction is as follows. Among them, the KEY field in each entry of the traffic flow table is used to record the identifier of the traffic flow received by the forwarding device, such as the triple (source IP address, destination IP address, and QP) of the RDMA flow; the Additional Data (AD) field in the entry corresponding to the RDMA flow is used to record the PSN of the upstream packet (l2r PSN shown in the table) and the arrival time of the upstream packet (l2r timestamp in the table), etc.
[0121] In practical applications, the forwarding device can build a table according to the triple of the upstream packet, and the forwarding device determines faults based on two packets in the upstream direction; or, it can also build a table according to the triple of the downstream packet, and the forwarding device determines faults based on two packets in the downstream direction.
[0122] When the first upstream packet of the target traffic flow arrives at the forwarding device, the forwarding device can identify that the packet belongs to an upstream packet, and then establishes an entry for the target traffic flow in the traffic flow table according to the quintuple of the upstream packet. The entry includes the PSN of the upstream packet (l2r PSN shown in the table) and the arrival time of the upstream packet (l2r timestamp in the table). When the forwarding device subsequently receives a new upstream packet in the target traffic flow, it triggers the forwarding device to execute the second determination logic provided in the embodiment of the present application. Among them, the forwarding device uses the time when the new upstream packet arrives at the forwarding device as the current time (curTS), uses the packet recorded in the traffic flow table as the second packet, and uses the newly received packet as the first packet. Then the PSN of the packet recorded in the traffic flow table is the second PSN of the second packet (l2r PSN shown in the table), the arrival time of the packet recorded in the traffic flow table is the first timestamp of the second packet (l2r timestamp in the table), and the PSN of the newly received packet is the first PSN (PSN) of the first packet.
[0123] Specifically, when PSN <= l2r PSN and curTS - l2r timestamp > the preset threshold (Threshold), it is determined that a fault has occurred in the service flow corresponding to the currently queried table entry. If the above conditions are not met, it is determined that no fault has occurred in the current target service flow. Then, the newly received packet is used as the second packet, and the sequence number, acknowledgment number, and arrival time of the newly received packet are used to update the sequence number, acknowledgment number, and arrival time of the table entry for the target service flow in the service flow table.
[0124] The third determination logic: The forwarding device performs fault determination on the target service flow based on two packets in different transmission directions in the target service flow.
[0125] In the TCP scenario, the source IP address of the first packet is the same as the destination IP address of the second packet, the destination IP address of the first packet is the same as the source IP address of the second packet, the source port of the first packet is the same as the destination port of the second packet, and the destination port of the first packet is the same as the source port of the second packet. The first sequence number information of the first packet includes the sequence number (Sequence Number) and acknowledgment number (Acknowledgment Number) in the TCP header of the first packet. The second sequence number information of the second packet at least includes the sequence number (Sequence Number) and acknowledgment number (Acknowledgment Number) in the TCP header of the second packet.
[0126] In a possible implementation, the first packet and the second packet are respectively the latest packets obtained by the forwarding device in two opposite transmission directions. Taking the first packet as the upstream packet and the second packet as the downstream packet as an example, the first packet is the latest upstream packet obtained by the forwarding device in the upstream direction, and the second packet is the latest downstream packet obtained by the forwarding device (or other devices that establish a peer link with this forwarding device) in the downstream direction.
[0127] The forwarding device obtains the first sequence number information (i.e., the sequence number and acknowledgment number of the first packet) and the second sequence number information (i.e., the sequence number and acknowledgment number of the second packet). If the acknowledgment number in the TCP header of the first packet is less than the sequence number in the TCP header of the second packet, and the difference between the current time and the second arrival time of the first packet is greater than the preset threshold, it indicates that the forwarding device can receive the target traffic flow in the direction of the second packet, but the forwarding device has not received the target traffic flow in the same direction as the first packet for a long time (exceeding the preset duration). Therefore, the acknowledgment number of the first packet stops increasing. So, the forwarding device determines that the target traffic flow has a fault, and the upstream path of the second packet is normal and has no fault. Here, the upstream path of the second packet refers to the path passed through during the transmission of the second packet to this forwarding device. Thus, while perceiving the target traffic flow, the forwarding device can further determine the normal path through which the target traffic flow passes, facilitating subsequent fault location.
[0128] In the RMDA scenario, the above-mentioned third determination logic is also applicable. Specifically, the source IP address of the first packet is the same as the destination IP address of the second packet, the destination IP address of the first packet is the same as the source IP address of the second packet, and the QP of the first packet is different from the QP of the second packet. Specifically, the first packet and the second packet are packets in different transmission directions in the same RDMA flow. In other words, the first packet and the second packet have the same Queue Pair Context (QPC) and Queue Pair Key (QKey), but the Destination QP of the first packet is different from the Destination QP of the second packet. Optionally, the first packet may be a request packet in the RDMA service flow, and the second packet is a response packet in the RDMA service flow; or, the second packet may be a response packet in the RDMA service flow, and the second packet is a request packet in the RDMA service flow. The first sequence number information includes the PSN of the first packet, and the second sequence number information includes the PSN of the second packet. As can be seen from the above, in the RDMA scenario where the acknowledgment mechanism is enabled, the PSN of the RDMA packet indicates the position of the data content of the RDMA packet in the complete data content of the target service flow, and also indicates the position of the packet that the destination device has successfully received in the target service flow. That is, for the two opposite transfer directions of the target service flow, the PSN of the packet in one direction corresponds to the PSN of the packet in the other direction. If the PSN of the first packet is less than the PSN of the second packet, and the difference between the current time and the second arrival time of the first packet is greater than the preset threshold, it means that the target service flow in the direction where the second packet can be received by the forwarding device, but the forwarding device has not received the target service flow in the same direction as the first packet for a long time (exceeding the preset duration). Therefore, the growth of the PSN of the first packet stops. So, the forwarding device determines that the target service flow has failed, and the upstream path of the second packet is normal and has not failed, where the upstream path of the second packet refers to the path passed through during the transfer of the second packet to the forwarding device.
[0129] Please refer to Figure 9 , Figure 9 a possible scenario diagram of the communication method in the embodiment of the present application. As Figure 9 shown, both the forwarding device A and the forwarding device B can establish a service flow table for the service flows arriving at the forwarding device. The service flow table is used to record the sequence number information and arrival time of the packets from each service flow. Assume that the upstream packet and the downstream packet are packets in two opposite directions of the target service flow. The upstream packet and the downstream packet are relative. In Figure 7In the example, the packets in the direction from forwarding device B to forwarding device A are used as upstream packets, and the packets in the direction from forwarding device A to forwarding device B are used as downstream packets. Among them, the KEY field in each entry of the service flow table is used to record the identifier of the service flow received by the forwarding device, such as the five-tuple or three-tuple of the service flow; the additional data (AdditionalData, AD) field in the entry corresponding to the service flow is used to record the acknowledgment number of the upstream packet (l2r ack seq in the table), the sequence number of the downstream packet (r2l seq in the table), and the arrival time of the downstream packet (r2l timestamp in the table), etc. Taking Figure 9 the forwarding device B shown in the figure as an example, if seq7 > seq6, and the current time (curTS) - the second arrival time (ts3) > the preset threshold (Threshold), then the forwarding device B determines that the target service flow has a fault, and the upstream path of the second packet is normal and no fault has occurred. The specific table building process is similar to the description of the embodiment shown in the foregoing Figures 5 to 8 and will not be elaborated here specifically.
[0130] In a possible implementation, the transmission path of the target service flow includes multiple forwarding devices. Among them, each forwarding device can apply the communication method provided by the embodiments of the present application to determine the faults that occur in the target service flow. Among the respective forwarding devices, one or more determination logics provided by the embodiments of the present application can be arbitrarily configured. If the first packet and the second packet received by the forwarding device meet any one of the determination logics of the embodiments of the present application, the forwarding device can determine that the target service flow has a fault.
[0131] Furthermore, as Figure 9 shown, in the third determination logic in the embodiments of the present application, the forwarding device can further determine that the upstream path of the second packet is normal and no fault has occurred. At this time, it means that the cause of the fault has nothing to do with the device on the upstream path of the second packet (i.e., the upstream device of the second packet). Therefore, the forwarding device can send a fault notice to the upstream device of the second packet, and the fault notice is used to indicate that the upstream path of the second packet is normal, so that the upstream device of the second packet can avoid performing an invalid path switch. Exemplarily, the fault notice includes but is not limited to the five-tuple or three-tuple of the target service flow.
[0132] Correspondingly, the embodiments of the present application further provide related devices for implementing the above solutions. Specifically, please refer to Figure 10 , Figure 10 which is a schematic structural diagram of the communication device provided by the embodiments of the present application. Figure 10The communication device in Figure 10 may be a chip, a chip system, or a processor for supporting a forwarding device to implement the method; alternatively, the communication device may also be a logical node, a logical module, or software for implementing all or part of the functions of the communication device; the communication device may also be a collective term for multiple logical nodes, multiple logical modules, or multiple software for implementing the communication method. As
[0133] shown, the communication device includes:
[0134] a transceiver unit 201, configured to receive a first packet, where the first packet includes first sequence number information;
[0135] The transceiver unit 201 is further configured to obtain second sequence number information of a second packet, where the first packet and the second packet correspond to a target service flow;
[0136] A processing unit 202 is configured to determine that a fault occurs in the target service flow according to the first sequence number information and the second sequence number information.
[0137] In a possible implementation, the source Internet Protocol (IP) address of the first packet is the same as the destination IP address of the second packet, the destination IP address of the first packet is the same as the source IP address of the second packet, the source port of the first packet is the same as the destination port of the second packet, the destination port of the first packet is the same as the source port of the second packet, the first sequence number information includes the sequence number in the Transmission Control Protocol (TCP) header of the first packet, and the second sequence number information includes the acknowledgement number in the TCP header of the second packet;
[0138] Specifically, the processing unit 202 is configured to: when the sequence number in the TCP header of the first packet is greater than the acknowledgement number in the TCP header of the second packet, and the difference between the current time and the first arrival time of the second packet is greater than a preset threshold, determine that a fault occurs in the target service flow.
[0139] In a possible implementation, the source IP address of the first packet is the same as the source IP address of the second packet, the destination IP address of the first packet is the same as the destination IP address of the second packet, the source port of the first packet is the same as the source port of the second packet, the destination port of the first packet is the same as the destination port of the second packet, the first sequence number information includes the sequence number and the acknowledgement number in the TCP header of the first packet, and the second sequence number information includes the sequence number and the acknowledgement number in the TCP header of the second packet;
[0140] In a possible implementation, the source IP address of the first packet is the same as the destination IP address of the second packet, the destination IP address of the first packet is the same as the source IP address of the second packet, the source port of the first packet is the same as the destination port of the second packet, the destination port of the first packet is the same as the source port of the second packet, the first sequence number information includes the sequence number and acknowledgment number in the TCP header of the first packet, and the second sequence number information includes the sequence number and acknowledgment number in the TCP header of the second packet;
[0141] The processing unit 202 is specifically configured to: when the acknowledgment number in the TCP header of the first packet is less than the sequence number in the TCP header of the second packet, and the difference between the current time and the second arrival time of the first packet is greater than a preset threshold, determine that the upstream path of the second packet is normal.
[0142] In a possible implementation, the transceiver unit 201 is further configured to send a fault notice to the upstream device of the second packet, and the fault notice is used to indicate that the upstream path of the second packet is normal.
[0143] In a possible implementation, the first packet and the second packet are packets in the Remote Direct Memory Access (RDMA) protocol. The source IP address of the first packet is the same as the source IP address of the second packet, the destination IP address of the first packet is the same as the destination IP address of the second packet, the queue pair (QP) of the first packet is the same as the QP of the second packet, the first sequence number information includes the Packet Sequence Number (PSN) of the first packet, and the second sequence number information includes the PSN of the second packet;
[0144] The processing unit 202 is specifically configured to: when the PSN of the first packet is less than or equal to the PSN of the second packet, and the difference between the current time and the first arrival time of the second packet is greater than a preset threshold, determine that the target traffic flow has a fault.
[0145] In a possible implementation, the first packet and the second packet are packets in the Remote Direct Memory Access (RDMA) protocol. The source IP address of the first packet is the same as the destination IP address of the second packet, the destination IP address of the first packet is the same as the source IP address of the second packet, the QP of the first packet is different from the QP of the second packet, the first sequence number information includes the PSN of the first packet, and the second sequence number information includes the PSN of the second packet.
[0146] The processing unit 202 is specifically configured to: when the PSN of the first packet is greater than the PSN of the second packet, and the difference between the current time and the first arrival time of the second packet is greater than a preset threshold, determine that the target traffic flow has a fault.
[0147] In a possible implementation, the source IP address of the first packet is the same as the destination IP address of the second packet, the destination IP address of the first packet is the same as the source IP address of the second packet, the QP of the first packet is different from the QP of the second packet, the first sequence number information includes the PSN of the first packet, and the second sequence number information includes the PSN of the second packet;
[0148] The processing unit 202 is specifically configured to: when the PSN of the first packet is less than the PSN of the second packet, and the difference between the current time and the second arrival time of the first packet is greater than a preset threshold, determine that the upstream path of the second packet is normal.
[0149] In a possible implementation, the transceiver unit 201 is specifically configured to receive the second packet, and the second packet includes the second sequence number information.
[0150] In a possible implementation, the second sequence number information is received through a peer link.
[0151] In a possible implementation, the transceiver unit 201 is further configured to send fault information for a target traffic flow to the controller, and the fault information is used to indicate that the target traffic flow has a fault.
[0152] It should be noted that the information interaction, execution process, etc. between the modules / units in the communication device are based on the same concept as the corresponding method embodiments in this application. For specific content, reference can be made to the descriptions in the method embodiments shown above in this application, and details are not repeated here. Figure 3 For details, please refer to
[0153] Please refer to Figure 11 , Figure 11 which is a schematic logical structure diagram of the communication device 30 provided in an embodiment of this application. Figure 11 The communication device 30 in Figure 10 can be a chip, a chip system, or a processor used to support a forwarding device to implement this method; alternatively, the communication device 30 can also be a logical node, a logical module, or software used to implement all or part of the functions of the communication device 30; the communication device 30 can also be a collective term for multiple logical nodes, multiple logical modules, or multiple software used to implement the communication method. The communication device 30 can be deployed with Figure 3 the communication device described in the corresponding embodiment, for implementing
[0154] The memory 301 can be a read only memory (ROM), a static storage device, a dynamic storage device, or a random access memory (RAM). The memory 301 can store a program. When the program stored in the memory 301 is executed by the processor 302, the processor 302 and the communication interface 303 are used to execute steps 101-103 of the above communication method embodiments.
[0155] The processor 302 can be a central processing unit (CPU), a microprocessor, an application specific integrated circuit (ASIC), a graphics processing unit (GPU), a digital signal processor (DSP), a field programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, or any combination thereof, and is used to execute relevant programs to implement one or more of steps 101-103 in the communication method embodiments of the present application. The steps of the data processing method disclosed in the embodiments of the present application can be executed by a compiler and an executor, where the compiler and the executor can be executed by a hardware decoding processor, or executed by a combination of hardware and software modules in the decoding processor. The software module can be located in a mature storage medium in the art such as a random access memory, a flash memory, a read only memory, a programmable read only memory, or an electrically erasable programmable memory, a register, etc. This storage medium is located in the memory 301, and the processor 302 reads the information in the memory 301 and combines its hardware to execute one or more of steps 101-103 in the communication method embodiments of the present application.
[0156] The communication interface 303 uses a transceiver device such as, but not limited to, a transceiver to implement communication between the communication device 30 and other devices or communication networks.
[0157] The bus 304 can implement a path for transmitting information between various components of the computer device 30 (for example, the memory 301, the processor 302, and the communication interface 303). The bus 304 can be a Peripheral Component Interconnect (PCI) bus, an Extended Industry Standard Architecture (EISA) bus, or the like. The bus can be divided into an address bus, a data bus, a control bus, etc. For the sake of representation, Figure 11 only a thick line is used to represent it in Figure 11 , but it does not mean that there is only one bus or one type of bus.
[0158] It should be noted that the information interaction, execution process, etc. between the modules / units in the communication device are based on the same concept as the corresponding method embodiments in this application. For specific content, refer to the description in the method embodiments shown above in this application, and details will not be repeated here. Figure 3 The corresponding method embodiments are based on the same concept, and for specific content, refer to the description in the method embodiments shown above in this application, and details will not be repeated here.
[0159] The embodiments of this application also provide a computer program product containing instructions. The computer program product can be software or a program product containing instructions that can run on a computing device or be stored in any available medium. When the computer program product runs on at least one computer device, it causes at least one computer device to execute the method described in the embodiments shown above Figure 3 as described.
[0160] The embodiments of this application also provide a computer-readable storage medium. The computer-readable storage medium can be any available medium that a computing device can store or a data storage device such as a data center containing one or more available media. The available medium can be a magnetic medium (for example, a floppy disk, a hard disk, a magnetic tape), an optical medium (for example, a DVD), or a semiconductor medium (for example, a solid-state drive), etc. The computer-readable storage medium includes instructions that instruct the computing device to execute the method described in the above Figure 3 as described in the embodiments shown.
[0161] The communication device provided in the embodiments of this application can specifically be a chip. The chip includes a processing unit and a communication unit. The processing unit can be, for example, a processor, and the communication unit can be, for example, an input / output interface, a pin, or a circuit, etc. The processing unit can execute the computer execution instructions stored in the storage unit to cause the chip to execute the above Figure 3The method described in the illustrated embodiment. Optionally, the storage unit is a storage unit within the chip, such as a register, cache, etc. The storage unit can also be a storage unit located outside the chip within the wireless access device, such as a read-only memory (ROM) or other types of static storage devices that can store static information and instructions, a random access memory (RAM), etc.
[0162] It should be further noted that the device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. In addition, in the drawings of the device embodiments provided in the embodiments of the present application, the connection relationships between the modules indicate that they have communication connections, which can be specifically implemented as one or more communication buses or signal lines.
[0163] Through the description of the above embodiments, those skilled in the art can clearly understand that the embodiments of the present application can be implemented by means of software plus necessary general hardware, and of course, can also be implemented by dedicated hardware including dedicated integrated circuits, dedicated CPUs, dedicated memories, dedicated components, etc. Generally, functions completed by computer programs can be easily implemented by corresponding hardware, and the specific hardware structures used to implement the same function can also be various, such as analog circuits, digital circuits or dedicated circuits. However, for the embodiments of the present application, in more cases, software program implementation is a better implementation method. Based on such an understanding, the technical solution of the embodiments of the present application, in essence, or the part that contributes to the prior art can be embodied in the form of a software product. This computer software product is stored in a readable storage medium, such as a floppy disk, USB flash drive, mobile hard disk, ROM, RAM, magnetic disk or optical disc of a computer, etc., and includes several instructions to enable a computer device (which can be a personal computer, training device, or network device, etc.) to execute the methods described in various embodiments of the present application.
[0164] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product.
[0165] The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions described in the embodiments of the present application are generated in whole or in part. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions may be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions may be transmitted from one website, computer, training device, or data center to another website, computer, training device, or data center via wired (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (such as infrared, wireless, microwave, etc.) means. The computer-readable storage medium may be any available medium that a computer can store or a data storage device such as a training device or a data center that includes one or more integrated available media. The available medium may be a magnetic medium (such as a floppy disk, a hard disk, a magnetic tape), an optical medium (such as a DVD), or a semiconductor medium (such as a solid state disk (SSD)).
Claims
1. A communication method, characterized in that: include: receiving a first message, wherein the first message includes first sequence number information; Obtaining second sequence number information of a second message, wherein the first message and the second message correspond to a target service flow; It is determined that a failure occurs in the target service flow based on the first sequence number information and the second sequence number information.
2. The method according to claim 1, characterized in that The source Internet Protocol (IP) address of the first message is the same as the destination IP address of the second message, the destination IP address of the first message is the same as the source IP address of the second message, the source port of the first message is the same as the destination port of the second message, the destination port of the first message is the same as the source port of the second message, the first sequence number information includes a sequence number in a Transmission Control Protocol (TCP) header of the first message, and the second sequence number information includes an acknowledgment number in a TCP header of the second message; Determining, according to the first sequence number information and the second sequence number information, that a failure occurs in the target service flow includes: If the sequence number in the TCP header of the first message is greater than the acknowledgment number in the TCP header of the second message, and the difference between the current time and the first arrival time of the second message is greater than a preset threshold, it is determined that the target service flow has failed.
3. The method according to claim 1, characterized in that The source IP address of the first message is the same as the source IP address of the second message, the destination IP address of the first message is the same as the destination IP address of the second message, the source port of the first message is the same as the source port of the second message, the destination port of the first message is the same as the destination port of the second message, the first sequence number information includes the sequence number and acknowledgment number in the TCP header of the first message, and the second sequence number information includes the sequence number and acknowledgment number in the TCP header of the second message; Determining, according to the first sequence number information and the second sequence number information, that a failure occurs in the target service flow includes: If the sequence number in the TCP header of the first message is less than or equal to the sequence number in the TCP header of the second message, and the confirmation number in the TCP header of the first message is less than or equal to the confirmation number in the TCP header of the second message, and the difference between the current time and the first arrival time of the second message is greater than a preset threshold, it is determined that a failure has occurred in the target service flow.
4. The method according to claim 1, wherein The source IP address of the first message is the same as the destination IP address of the second message, the destination IP address of the first message is the same as the source IP address of the second message, the source port of the first message is the same as the destination port of the second message, the destination port of the first message is the same as the source port of the second message, the first sequence number information includes the sequence number and acknowledgment number in the TCP header of the first message, and the second sequence number information includes the sequence number and acknowledgment number in the TCP header of the second message; Determining, according to the first sequence number information and the second sequence number information, that a failure occurs in the target service flow includes: If the confirmation number in the TCP header of the first message is smaller than the sequence number in the TCP header of the second message, and the difference between the current time and the second arrival time of the first message is greater than a preset threshold, it is determined that the upstream path of the second message is normal.
5. The method according to claim 4, characterized in that The method further comprises: A fault notification is sent to an upstream device of the second message, where the fault notification is used to indicate that an upstream path of the second message is normal.
6. The method according to claim 1, characterized in that The first message and the second message are messages in a Remote Direct Memory Access (RDMA) protocol, a source IP address of the first message is the same as a destination IP address of the second message, the destination IP address of the first message is the same as a source IP address of the second message, a QP of the first message is different from a QP of the second message, the first sequence number information includes a PSN of the first message, and the second sequence number information includes a PSN of the second message; Determining, according to the first sequence number information and the second sequence number information, that a failure occurs in the target service flow includes: If the PSN of the first message is greater than the PSN of the second message, and the difference between the current time and the first arrival time of the second message is greater than a preset threshold, it is determined that a failure occurs in the target service flow.
7. The method according to claim 1, characterized in that The first message and the second message are messages of the same type in a remote direct memory access (RDMA) protocol, a source IP address of the first message is the same as a source IP address of the second message, a destination IP address of the first message is the same as a destination IP address of the second message, a queue pair (QP) of the first message is the same as a QP of the second message, the first sequence number information includes a packet sequence number (PSN) of the first message, and the second sequence number information includes a PSN of the second message; Determining, according to the first sequence number information and the second sequence number information, that a failure occurs in the target service flow includes: If the PSN of the first message is less than or equal to the PSN of the second message, and the difference between the current time and the first arrival time of the second message is greater than a preset threshold, it is determined that a failure occurs in the target service flow.
8. The method according to claim 1, characterized in that The source IP address of the first message is the same as the destination IP address of the second message, the destination IP address of the first message is the same as the source IP address of the second message, the QP of the first message is different from the QP of the second message, the first sequence number information includes the PSN of the first message, and the second sequence number information includes the PSN of the second message; Determining, according to the first sequence number information and the second sequence number information, that a failure occurs in the target service flow includes: If the PSN of the first message is smaller than the PSN of the second message, and the difference between the current time and the second arrival time of the first message is greater than a preset threshold, it is determined that the upstream path of the second message is normal.
9. The method according to any one of claims 2 to 8, characterized in that The method further comprises: A second message is received, where the second message includes second sequence number information.
10. The method according to any one of claims 2 to 8, characterized in that The second sequence number information is received through a peer link.
11. The method according to any one of claims 1 to 10, characterized in that The method further comprises: Fault information for the target service flow is sent to a controller, where the fault information is used to indicate that a fault occurs in the target service flow.
12. A communication device, characterized in that: include: a transceiver unit, configured to receive a first message, wherein the first message includes first sequence number information; The transceiver unit is further configured to obtain second sequence number information of a second message, wherein the first message and the second message are from a target service flow; A processing unit is used to determine whether a failure occurs in the target service flow based on the first sequence number information and the second sequence number information.
13. A communication device, characterized in that: comprising a processor coupled to a memory; The memory is used to store instructions; The processor is configured to execute instructions in the memory, so that the communication device performs the method according to any one of claims 1 to 11.
14. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the method according to any one of claims 1 to 11 is implemented.
15. A computer program product, characterized in that The computer program product stores computer-readable instructions, and when the computer-readable instructions are executed by a processor, the method according to any one of claims 1 to 11 is implemented.
16. A chip, characterized in that: The chip includes a processor coupled to a memory; The memory is used to store instructions; The processor is configured to execute instructions in the memory, so that the chip executes the method according to any one of claims 1 to 11.
Citation Information
Cited By
Communication method, communication apparatus, and communication device
WO2025161854A1