Reducing retransmission latency in out-of-order data transmissions
By detecting network errors and sending notifications to nodes, the problem that the receiver finds it difficult to distinguish between delay and packet discarding in out-of-order data transmission is solved, and the effect of reducing retransmission delay and improving application performance is achieved.
Patent Information
- Application Number
- CN202411625505.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-11-14
- Filing Date
- 2024-11-14
- Publication Date
- 2025-05-16
AI Technical Summary
In out-of-order data transmission, it is difficult for the receiver to distinguish between the packets delayed by the network and the packets discarded, resulting in high tail delays and affecting application performance.
By detecting network errors and sending notifications to connected nodes, the receiver can more intelligently distinguish delayed and discarded packets, thereby reducing retransmission delay.
It effectively reduces the retransmission delay in out-of-order data transmission, reduces other problems caused by packet discarding, and improves the performance of the application.
Smart Images

Figure CN120017729A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure generally relates to systems, methods, and devices for error checking, and more particularly to systems, methods, and devices for reducing retransmission delays in out-of-order data transmission. Background Art
[0002] Data can be split into multiple packets that are routed to their destination via multiple paths by hopping from one router to another. One node can send multiple messages to another node, which needs to identify which packets belong together. In addition, packets may arrive out of order. This is especially likely to happen if two packets follow different paths to their destination. Packets may be corrupted, meaning that for some reason the received data no longer matches the data that was originally sent. Packets may also be lost due to problems with the physical layer or the router's forwarding table. If even one packet of a message is lost, it may not be possible to put the message back together in a reasonable way. Similarly, packets may be duplicated due to accidental retransmissions of the same packet. Transmission Control Protocol (TCP) and User Datagram Protocol (UDP) are data transmission protocols used for packet sequencing, retransmission, and data integrity. Summary of the invention
[0003] Transport layer protocols, such as Transmission Control Protocol (TCP), Remote Direct Memory Access (RDMA), and User Datagram Protocol (UDP), are data transmission protocols used for packet sequencing, retransmission, and data integrity. Transport layer protocols guarantee lossless data transmission from A to B by retransmitting data lost along the path. Some transport layer protocols support out-of-order data transmission, so that the message transmitted from node A to node B is broken into multiple smaller packets that can be delivered out of order between nodes A and B. In such protocols, the receiver receives out-of-order packets, reorders them, and then sends an acknowledgment.
[0004] When a receiver receives packets out of order, it may be difficult to distinguish between packets that were delayed by the network and will arrive soon (e.g., delayed packets); and packets that were dropped (or otherwise lost) by the network and will never arrive (e.g., lost packets). Because the receiver is uncertain whether a packet was delayed or lost, the receiver delays its retransmission requests to prevent the transmission of duplicate packets, which in turn causes high latency when packets are dropped. In other words, the uncertainty of whether a packet is delayed or lost causes high "tail latency", which affects application performance because the entire application running on thousands of nodes is affected by a single node delaying its retransmission message.
[0005] In a network, "physical layer errors" or physical coding sublayer (PCS) block errors (e.g., random errors on an analog medium due to thermal or cosmic noise or optical analog device failures) cause most packet drops (e.g., up to 99% of packet drops are of this type). When a packet is dropped due to a physical layer error or a physical coding sublayer (PCS) error, the network device can detect that a PCS error has occurred and can signal the sender, the receiver, or both that the packet may have been dropped. In an embodiment, the present disclosure may be able to determine the specific sender and / or receiver of the dropped packet. In other embodiments, the sender and / or receiver of the dropped packet may not be determined (e.g., if the packet is damaged), and all nodes connected to the device or a subset of all nodes are notified of the packet drop. For example, when a packet transmitted via a specific port on a switch is dropped, all nodes connected to that port may be notified. Nodes that receive an alert of a PCS error / packet drop reduce the time used to transmit requests by a predetermined amount of time. In other words, within a predetermined amount of time after an alert, a node may become more sensitive to packet drops and reduce the time it waits to see if a delayed packet is received before sending a retransmission request.
[0006] In an embodiment, a network device may notify a sender and / or a receiver of packets that may be discarded by sending a packet with an indication of at least one packet loss to the sender and / or the receiver. In other embodiments, a network device may mark all packets that flow through (e.g., on a specific port) with a dedicated flag within a predetermined time period. For example, a flag may be added to a packet header that may be turned on or off to indicate that a packet has been lost (possibly) recently. The present disclosure utilizes transport layer information to reduce retransmission delays in out-of-order data transmission. In other words, a receiver may more intelligently distinguish between delayed packets and discarded packets. In addition, or as an alternative, a sender may initiate a retransmission without waiting for a retransmission request.
[0007] According to one or more embodiments described herein, a network device (e.g., a switch) can enable various nodes (e.g., switches, servers, personal computers, and other computing devices) to communicate across a network. The ports of the network device can be used as communication endpoints, allowing the network device to manage multiple simultaneous network connections with one or more nodes.
[0008] Each port of a network device can be considered as a pathway and can be associated with an egress queue of data (e.g., in the form of packets) waiting to be sent via the port. In practice, each port can be used as an independent channel for data communication with the network device. Each port of a network device can be connected to one or more ports of one or more other devices. Ports allow for concurrent network communications, enabling a network device to perform multiple data exchanges with different network nodes at the same time.
[0009] Load balancing of network traffic between multiple paths is often a computationally difficult task. Consider a network switch that receives packets from one or more sources. Each packet that flows through the switch is associated with a specific destination. In simple topologies, the switch may have only one port from which the packet must be sent to reach the destination. However, in modern network topologies, such as a cluster of graphics processing units (GPUs) used for artificial intelligence (AI) related tasks, there may be many possible ports from which packets can be transmitted to reach the associated destination. Therefore, due to the presence of multiple paths in the network, a decision must be made as to which of the many possible ports each packet should be transmitted from. In many applications, the goal of the switch in this case is to route packets to the destination in a way that provides the maximum overall throughput and avoids congestion.
[0010] The present disclosure describes a system and method for reducing retransmission delays in out-of-order data transmission. Embodiments of the present disclosure are intended to address the above-mentioned shortcomings and other problems by implementing an improved method for detecting discarded packets. The system and method described herein reduces retransmission delays and other problems caused by packet discards.
[0011] The methods depicted and described herein may be applied to any suitable type of device known or yet to be developed. In an illustrative example, a method is disclosed that includes: detecting a network error; and in response to detecting the network error, sending a notification of the detection of the network error to connected nodes. The method also includes changing, by each node receiving the notification, at least one transport layer parameter, wherein after a predetermined amount of time, the at least one transport layer parameter is restored back to an original value.
[0012] In another example, a system is disclosed that includes one or more circuits for: detecting a network error; and in response to detecting the network error, sending a notification of the detected network error to at least one of a plurality of connected nodes, wherein the notification causes each node receiving the notification to change at least one transport layer parameter.
[0013] In another example, a device is disclosed, comprising: an interface for detecting a network error; and a processing circuit for sending a notification of the detection of the network error to connected nodes in response to detecting the network error, wherein the notification causes each node receiving the notification to reduce a latency threshold.
[0014] Any of the above-described example aspects include: wherein sending the notification includes sending a packet with an indication that a network error was detected.
[0015] Any of the above-described example aspects include: wherein the packet with the indication that a network error was detected includes a timestamp.
[0016] Any of the above example aspects include: wherein sending the notification includes: a dedicated flag in a subsequent packet is turned on.
[0017] Any of the above example aspects include: wherein the network error comprises an optical or electrical device error.
[0018] Any of the above example aspects include: wherein the network errors include thermal noise or cosmic noise.
[0019] Any of the above-described example aspects include: wherein detecting a network error comprises detecting an egress port associated with the detected network error.
[0020] Any of the above example aspects includes where sending the notification includes a dedicated flag being turned on in subsequent packets sent on the egress port associated with the detected network error.
[0021] Any of the above example aspects includes where sending a notification of the detected network error to a connected node includes sending a notification to a node connected to an egress port associated with the detected network error.
[0022] Any of the above-described example aspects include: wherein changing at least one transport layer parameter includes: reducing a timeout for requesting a retransmission within a predetermined amount of time.
[0023] Any of the above example aspects include: wherein the notification includes a packet with an indication that a network error was detected, and wherein the packet with the indication that a network error was detected includes a timestamp.
[0024] Any of the above example aspects include: wherein the latency threshold is reduced for a predetermined amount of time, and after the predetermined amount of time, the latency threshold is restored back to the original value.
[0025] Additional features and advantages are described herein, and will be apparent from the following description and accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0026] The present disclosure is described in conjunction with the accompanying drawings, which are not necessarily drawn to scale:
[0027] Figure 1 is a block diagram depicting an illustrative configuration of a computing system in accordance with at least some embodiments of the present disclosure;
[0028] Figure 2 A block diagram depicting out-of-order data transmission according to at least some embodiments of the present disclosure is shown;
[0029] Figure 3 shows an example packet header in accordance with at least some embodiments of the present disclosure; and
[0030] Figure 4 is a flow chart depicting a method in accordance with at least some embodiments of the present disclosure. DETAILED DESCRIPTION
[0031] Before explaining any embodiment of the present disclosure in detail, it should be understood that the application of the present disclosure is not limited to the details of the construction and component arrangement set forth in the following description or shown in the accompanying drawings. The present disclosure can have other embodiments and can be practiced or executed in various ways. In addition, it should be understood that the wording and terminology used herein are for descriptive purposes and should not be considered as limiting. "Including", "comprising" or "having" and its variants used herein are intended to cover the items listed thereafter and their equivalents and additional items. In addition, the present disclosure can use examples to illustrate one or more aspects thereof. Unless otherwise expressly stated, using or listing one or more examples (which may be represented by "for example", "by example", "for example", "such as" or similar language) is not intended to and does not limit the scope of the present disclosure.
[0032] The details of one or more aspects of the disclosure are set forth in the accompanying drawings and the description below. Other features, objects, and advantages of the techniques described in this disclosure will be apparent from the description and drawings, and from the claims.
[0033] The phrases "at least one", "one or more", and "and / or" are open expressions that can be used as both conjunctions and disjuncts. For example, the expressions "at least one of A, B, and C", "at least one of A, B, or C", "one or more of A, B, and C", "one or more of A, B, or C", and "A, B, and / or C" each represent A alone, B alone, C alone, A and B together, A and C together, B and C together, or A, B, and C together. When each of A, B, and C in the above expressions refers to an element (e.g., X, Y, and Z) or a class of elements (e.g., X1-Xn, Y1-Ym, and Z1-Zo), the phrase is intended to refer to a single element selected from X, Y, and Z, a combination of elements selected from the same class (e.g., X1 and X2), and a combination of elements selected from two or more classes (e.g., Y1 and Zo).
[0034] The term "a" or "an" entity refers to one or more of that entity. Therefore, the terms "a" (or "an"), "one or more" and "at least one" can be used interchangeably herein. It should also be noted that the terms "including", "comprising" and "having" can be used interchangeably.
[0035] The above is a brief overview of the present disclosure, which is intended to provide an understanding of certain aspects of the present disclosure. This overview is not a broad or exhaustive overview of the present disclosure and its various aspects, embodiments, and configurations. Its purpose is not to identify the key or important elements of the present disclosure, nor to define the scope of the present disclosure, but to present the selected concepts of the present disclosure in a simplified form as an introduction to the more detailed description provided below. It should be understood that other aspects, embodiments, and configurations of the present disclosure may use one or more features described above or in detail below, alone or in combination.
[0036] Numerous additional features and advantages are described herein and will become apparent to those skilled in the art upon consideration of the following detailed description and reference to the accompanying drawings.
[0037] The following description provides only embodiments and is not intended to limit the scope, applicability or configuration of the claims. Instead, the following description will provide a feasible description of implementing the embodiments for those skilled in the art. It should be understood that various changes may be made to the functions and arrangements of the elements without departing from the spirit and scope of the appended claims.
[0038] It will be appreciated from the following description that, for reasons of computational efficiency, the components of the system may be arranged at any suitable location within a distributed network of components without affecting the operation of the system.
[0039] In addition, it should be appreciated that the various links connecting the elements may be wired, trace or wireless links, or any suitable combination thereof, or any other suitable known or later developed element capable of providing and / or transmitting data to the connected elements. The transmission medium used as the link may be, for example, any suitable electrical signal carrier, including coaxial cable, copper wire and optical fiber, electrical traces on a printed circuit board (PCB), etc.
[0040] As used herein, the phrases "at least one", "one or more", "or", and "and / or" are open-ended expressions that can be used as both conjunctions and disjuncts. For example, each of the expressions "at least one of A, B, and C", "at least one of A, B, or C", "one or more of A, B, and C", "one or more of A, B, or C", "A, B and / or C", and "A, B, or C" means: A alone, B alone, C alone, A and B together, A and C together, B and C together, or A, B, and C together.
[0041] As used herein, the term "automatic" and variations thereof refer to any suitable process or operation that is accomplished without substantial human input in the performance of the process or operation. However, a process or operation may be automatic even if the performance of the process or operation uses substantial or insubstantial human input if the input is received prior to the performance of the process or operation. Human input is considered substantial if it affects the manner in which the process or operation is performed. Human input that consents to the performance of the process or operation is not considered "substantial."
[0042] As used herein, the terms "determine," "calculate," and "estimate," and variations thereof, are used interchangeably and include any suitable type of method, process, operation, or technique.
[0043] Various aspects of the disclosure are described herein with reference to the figures that are schematic illustrations of idealized configurations.
[0044] Any of the steps, functions, and operations discussed herein can be performed continuously and automatically.
[0045] The exemplary systems and methods of the present disclosure have been described with respect to switch networks; however, in order to avoid unnecessarily obscuring the present disclosure, the foregoing description omits many known structures and devices. Such omissions should not be construed as limiting the scope of the claimed disclosure. Specific details are set forth to provide an understanding of the present disclosure. However, it should be appreciated that the present disclosure may be implemented in a variety of ways beyond the specific details set forth herein.
[0046] The present disclosure may be subject to many changes and modifications. Some features of the present disclosure may be provided without providing other features.
[0047] References to "one embodiment", "an embodiment", "example embodiment", "some embodiments", etc. in the specification indicate that the embodiments may include specific features, structures or characteristics, but each embodiment does not necessarily include specific features, structures or characteristics. In addition, these phrases do not necessarily refer to the same embodiment. In addition, when a specific feature, structure or characteristic is described in conjunction with an embodiment, the description of such feature, structure or characteristic may apply to any other embodiment unless otherwise specified and / or readily apparent to a person skilled in the art from the description. In various embodiments, configurations and aspects, the present disclosure includes components, methods, processes, systems and / or devices substantially the same as those depicted and described herein, including various embodiments, sub-combinations and subsets thereof. A person skilled in the art will understand how to make and use the systems and methods disclosed herein after understanding the contents of the present disclosure. In various embodiments, configurations and aspects, the present disclosure includes providing devices and processes in the absence of items not depicted and / or described herein, or providing devices and processes in various embodiments, configurations or aspects thereof, including providing devices and processes in the absence of items that may have been used in previous devices or processes, for example, to improve performance, achieve simplicity and / or reduce implementation costs.
[0048] Unless otherwise defined, all terms (including technical and scientific terms) used herein have the same meaning as commonly understood by those skilled in the art to which the present disclosure belongs. It should also be understood that terms such as those defined in commonly used dictionaries should be interpreted as having a meaning consistent with their meaning in the context of the relevant field and the present disclosure.
[0049] As used herein, the singular forms "a", "an", and "the" are intended to include the plural forms as well, unless the context clearly indicates otherwise. It should also be understood that the terms "include", "comprise", and / or "comprising" when used in this specification specify the presence of the features, integers, steps, operations, elements, and / or components, but do not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof. The term "and / or" includes any and all combinations of one or more of the associated listed items.
[0050] Reference now Figures 1 to 4, various systems and methods for out-of-order data transmission and retransmission requests between communication nodes will be described. The concepts of packet routing depicted and described herein can be applied to routing information from one computing device to another computing device. The term packet used herein should be interpreted as any appropriate discrete amount of digitized information. The information being routed can be in the form of a single packet or multiple packets without departing from the scope of the present disclosure. In addition, certain embodiments will be described in conjunction with systems configured to make centralized routing decisions, while other embodiments will be described in conjunction with systems configured to make distributed and potentially uncoordinated routing decisions. It should be understood that the features and functionality of a centralized architecture can be applied or used in a distributed architecture, or vice versa.
[0051] According to one or more embodiments described herein, Figure 1 The computing system 103 shown can enable various systems (such as switches, servers, personal computers and other computing devices) to communicate across a network. Such computing system 103 described herein can be, for example, a switch or any computing device including multiple ports 106a-106d for connecting to nodes on a network.
[0052] The ports 106a-106d of the computing system 103 can be used as communication endpoints, allowing the computing system 103 to manage multiple simultaneous network connections with one or more nodes. Each port 106a-106d can be used to transmit data associated with one or more flows. Each port 106a-106d can be associated with a queue 121a-121d, thereby enabling the port 106a-106d to process incoming and outgoing data packets associated with the flow.
[0053] Each port 106a-106d of the computing system can be viewed as a path and is associated with a corresponding egress queue 121a-121d for data (e.g., in the form of packets) waiting to be sent via the port 106a-106d. In effect, each port 106 can be used as an independent channel for data communication with the computing system 103. The ports 106 allow for concurrent network communications, enabling the computing system 103 to conduct multiple data exchanges with different network nodes at the same time. When a packet or other form of data is ready to be sent from the computing system 103, the packet can be assigned to the port 106 and the packet will be sent from the port 106 by storing it in the queue 121 associated with the port 106.
[0054] The ports 106a-106d of the computing system 103 may be physical connection points that allow a network cable (e.g., an Ethernet cable) to connect the computing system 103 to one or more network nodes. Each port 106a-106d may be of a different type, including, for example, 100 Mbps, 1000 Mbps, or 10 Gigabit Ethernet ports, each providing a different level of bandwidth.
[0055] Because each port 106a-106d can be used to send a particular packet, when a packet is received, created, or otherwise processed by computing system 103 and is to be transmitted from computing system 103, one or more ports 106a-106d of computing system 103 may be selected to transmit the packet. Transmitting a packet from a port 106a-106d may include placing the data in a queue 121a-121d associated with another port 106a-106d.
[0056] The computing system's switching hardware 109 may include an internal structure or path within the computing system 103 through which data is transmitted between the two ports 106a-106d. In some embodiments, the switching hardware 109 may include one or more network interface cards (NICs). For example, in some embodiments, each port 106a-106d may be associated with a different NIC. The NIC or multiple NICs may include hardware and / or circuitry that can be used to transmit data between the ports 106a-106d.
[0057] The switching hardware 109 may also or alternatively include one or more application specific integrated circuits (ASICs) to perform tasks such as determining which port a received packet should be sent to. The switching hardware 109 may include various components, such as port controllers that manage the operation of each port, network interface cards that facilitate data transmission, and internal data paths that direct data flows within the computing system 103. The switching hardware 109 may also include memory elements for temporarily storing data and management software for controlling the operation of the hardware. This configuration enables the switching hardware 109 to accurately track port usage and provide data to the processor 115 upon request.
[0058] Packets received by computing system 103 may be placed in buffer 112 until placed in queue 121a-121d and then transmitted by corresponding port 106a-106d. Buffer 112 may effectively serve as an ingress queue in which received data packets may be temporarily stored. As described herein, the port 106a-106d via which a given packet is to be sent may be determined based on a variety of factors.
[0059] like Figure 1As shown, the computing system 103 may also include a processor 115, such as a CPU, a microprocessor, or any circuit or device capable of reading instructions from the memory 118 and performing actions. The processor 115 may execute software instructions to control the operation of the computing system 103.
[0060] Processor 115 may serve as the central processing unit of computing system 103 and perform the operating functions of the system. Processor 115 communicates with other components of computing system 103 to manage and execute computing operations, thereby ensuring optimal system functionality and performance.
[0061] In more detail, processor 115 can be designed to perform various computing tasks. Its functions may include executing program instructions, managing data within the system, and controlling the operation of other hardware components (e.g., switching hardware 109). Processor 115 can be a single-core or multi-core processor, and may include one or more processing units, depending on the specific design and requirements of computing system 103. The architectural design of processor 115 can allow efficient instruction execution, data processing, and overall system management, thereby enhancing the performance and practicality of computing system 103 in various applications. In addition, processor 115 can be programmed or adjusted to perform specific tasks and operations according to application requirements, thereby potentially enhancing the versatility and adaptability of computing system 103.
[0062] The system 103 may also include one or more memory 118 components. The memory 118 may be configured to communicate with the processor 115 of the computing system 103. The communication between the memory 118 and the processor 115 may enable various operations including, but not limited to, data exchange, command execution, and memory management.
[0063] Memory 118 may be comprised of a variety of physical components, depending on the specific type and design. At the core, memory 118 may include one or more memory cells capable of storing data in the form of binary information. These memory cells may be comprised of transistors, capacitors, or other suitable electronic components, depending on the memory type, such as DRAM, SRAM, or flash memory. To enable data transfer and communication with other parts of computing system 103, memory 118 may also include data lines or buses, address lines, and control lines. These physical components may collectively comprise memory 118, contributing to its ability to store and manage data.
[0064] The data in the memory 118 may contain information about various aspects of port, buffer, and system usage. Such information may include data about active connections, the amount of data in the queues 121a-121d, the amount of data in the buffer 112, the state of each port within the ports 106a-106d, and the like. The data may include, for example, buffer occupancy, the number of active ports 106a-106d, the number of total ports 106a-106d, and the queue depth or length of each port 106a-106d, as described in more detail herein. The data may be stored, accessed, and utilized by the processor 115 to manage port operations and network communications. For example, the processor 115 may utilize the data in the memory 118 to manage network traffic, determine priorities, or otherwise control the flow of data through the computing system 103, as described in more detail herein. Therefore, the potential combination of the memory 118 and the processor 115 may play a key role in optimizing the use and performance of the ports 106 of the computing system 103.
[0065] The data stored in the memory 118 may include various indicators, such as the amount of data or the number of packets in each queue 121a-121d, the amount of data or the number of packets in the buffer 112, and / or other information, such as data transmission rate, error rate, and the status of each port. After receiving the data, the processor 115 may perform further operations based on the information obtained, such as optimizing port usage, balancing network load, or solving problems, as described herein.
[0066] In one or more embodiments of the present disclosure, the computing system 103 (eg, a switch) may communicate with a plurality of network nodes 200, such as Figure 2 As shown. Each network node 200 may be a computing system capable of sending and receiving data. Each node 200 may be any of a variety of devices, including but not limited to a switch, a personal computer, a server, or any other device capable of sending and receiving data in packet form.
[0067] Computing system 103 may establish a communication channel with network node 200 via ports 106a-106f. Such a channel may support data transmission in the form of a stream of packets, following a predetermined protocol governing the format, size, transmission method, and other aspects of the packets.
[0068] Each network node 200 may interact with computing system 103 in various ways. Node 200 may send data packets to computing system 103 for processing, transmission, or other operations, or for forwarding to another node 200. Conversely, each node 200 may receive data from computing system 103, which may come from computing system 103 itself, or from other network nodes 200 via computing system 103. In this way, computing system 103 and nodes 200 may together form a network to facilitate data exchange, resource sharing, and a range of other collaborative operations.
[0069] Node 200 can be connected to multiple computing systems 103, thereby forming a network of nodes 200 and computing systems 103. For example, the systems and methods described herein may include multiple interconnected switches. Multiple computing systems 103 (e.g., switches) can be interconnected in various topologies, such as star, ring, or mesh, depending on the specific requirements and resilience required for the network. For example, in a star topology, multiple switches can be connected to a central switch, while in a ring topology, each switch can be connected to two other switches in a closed loop. In a mesh topology, each switch can be interconnected with each other switch in the network.
[0070] exist Figure 2 In the example shown, node 200 receives packets (e.g., packets 210-216). The packets may be received out of order as shown. At t=0, node 200 has received packets 210, 213, and 215. As shown by the bold line, packet 213 is the most recently received packet. At t=0, node 200 is still waiting for packets 211, 212, 214, and 216. In some cases, packets 211 and 212 are delayed. However, either of packets 211 and 212 may be discarded and therefore never received. In the event that a packet is discarded, node 200 may wait for a predetermined time before requesting a retransmission.
[0071] At t=1, packet 212 is received. At t=2, packet 214 is received. At t=3, packet 216 is received. At t=3, node 200 is still waiting for packet 211. In an embodiment, if packet 211 had been discarded at or before t=0, the receiver may have requested a retransmission of packet 211 without waiting until a later time (e.g., t=4).
[0072] Figure 3An example packet header 300 is shown. The packet header 300 may include a flag 301 that may be turned off / on to indicate to the node 200 (e.g., a receiver) that they should adjust one or more transport layer parameters (e.g., the time of a retransmission request). The size of the packets described herein may be measured in bits or bytes. The size of the packet may depend, for example, on the size of the payload of the packet. For example, packets generated by an application may include payloads of various sizes.
[0073] like Figure 4 As shown, and according to Figure 1 The computing system 103 shown, and as described herein, may perform the method 400 to reduce retransmission requests. Although the description of the method 400 provided herein describes the steps of the method 400 as being performed by the processor 115 of the computing system 103, the steps of the method 400 may be performed by one or more processors 115, switching hardware 109, one or more controllers, one or more circuits, a packet transmitter, a packet receiver, or some combination thereof in the computing system 103. It should be understood that the method 400 may be implemented in hardware or software.
[0074] At step 403, method 400 may begin with processor 115 of computing system 103 detecting a network error (e.g., a PCS error). When a packet is dropped due to a physical layer error or a PCS error, the network device may detect that a PCS error has occurred and may signal the sender, the receiver, or both that the packet may have been dropped. In other examples, computing system 103 may detect that a packet has been dropped.
[0075] At step 406, method 400 may include sending a notification to relevant nodes. In an embodiment, computing system 103 may be able to determine the specific sender and / or recipient of the discarded packet. In other embodiments, the sender and / or recipient of the discarded packet may not be determined (e.g., if the packet is damaged), and all nodes connected to the device or a subset of all nodes may receive notification that the packet was discarded. For example, when a packet sent via a specific port on a switch is discarded, all nodes connected to that port may be notified.
[0076] In step 409, method 400 includes sending an alarm of PCS error / packet drop to the node. For example, an alarm is sent to the affected node. In one example, if the recipient of the dropped packet can be determined, an alarm is sent to the recipient. In another example, an alarm is sent to a node connected via a port associated with the error / packet drop. In another example, if an error / packet drop occurs on a switch, an alarm can be sent to all nodes connected to the switch. In response to an alarm (e.g., a flag, a packet, etc.), the node can adjust its transport layer parameters within a predetermined amount of time. For example, the affected node can reduce the time of transmission requests for a period of time. In other words, within a predetermined amount of time after the alarm, the node may be more sensitive to packet drops and can reduce the time that the node waits to see if a delayed packet is received before sending a retransmission request.
[0077] Embodiments of the present disclosure include a system comprising one or more circuits for: receiving a packet; determining a size of the packet; determining one of a plurality of groups of packets based on the size of the packet; determining a port for the packet using a loop for the group of packets; and sending the packet via the port.
[0078] An embodiment also includes a system comprising one or more circuits configured to: receive multiple packet sizes from an application; initialize multiple packet arbitrator circuits, wherein each packet arbitrator circuit is associated with one of the multiple packet sizes; receive a first packet associated with the application; determine a size of the packet; based on the determined packet size, associate the packet with one of the packet arbitrator circuits; select one of a plurality of ports using the associated packet arbitrator circuit; and route the packet to the selected port of the plurality of ports.
[0079] An embodiment also includes a switch comprising one or more circuits configured to: receive a packet; match a size of the packet to a packet size class; determine a port for the packet based on the packet size class matching the packet size using a loop associated with the packet size class of the packet; and send the packet via the port.
[0080] Aspects of the above-described systems and switches include wherein the one or more circuits are further configured to: receive application data from an application; and create a plurality of groups based on the application data.
[0081] An embodiment of the present disclosure includes a method comprising: detecting a network error; and in response to detecting the network error, sending a notification of the detection of the network error to connected nodes. The method also includes: changing at least one transport layer parameter by each node receiving the notification, wherein after a predetermined amount of time, the at least one transport layer parameter is restored back to an original value.
[0082] An embodiment of the present disclosure also includes a system comprising one or more circuits for: detecting a network error; and in response to detecting a network error, sending a notification of the detection of the network error to at least one of a plurality of connected nodes, wherein the notification causes each node receiving the notification to change at least one transport layer parameter.
[0083] An embodiment of the present disclosure also includes a device comprising: an interface for detecting a network error; and a processing circuit, the processing circuit being used to send a notification of a detected network error to a connected node in response to detecting a network error, wherein the notification causes each node receiving the notification to reduce a delay threshold.
[0084] Aspects of the above methods, systems, and devices include: wherein sending the notification includes sending a packet with an indication that a network error was detected.
[0085] Aspects of the above methods, systems, and devices include: wherein the packet with the indication that a network error was detected includes a timestamp.
[0086] Aspects of the above-described methods, systems, and devices include: wherein sending a notification includes a dedicated flag in a subsequent packet being turned on.
[0087] Aspects of the above-described methods, systems, and devices include: wherein the network error comprises an optical or electrical device error.
[0088] Aspects of the above methods, systems and devices include: wherein the network errors include thermal noise or cosmic noise.
[0089] Aspects of the above methods, systems, and devices include: wherein detecting a network error includes detecting an egress port associated with the detected network error.
[0090] Aspects of the above methods, systems, and devices include where sending the notification includes a dedicated flag being turned on in subsequent packets sent on the egress port associated with the detected network error.
[0091] Aspects of the above methods, systems, and devices include: wherein sending a notification of a detected network error to a connected node includes sending a notification to a node connected to an egress port associated with the detected network error.
[0092] Aspects of the above methods, systems, and devices include: wherein changing at least one transport layer parameter includes: reducing a timeout for requesting a retransmission within a predetermined amount of time.
[0093] Aspects of the above methods, systems, and devices include where the notification includes a packet with an indication that a network error was detected, and where the packet with the indication that a network error was detected includes a timestamp.
[0094] Aspects of the above methods, systems, and devices include where the latency threshold is reduced for a predetermined amount of time, and after the predetermined amount of time, the latency threshold is restored back to the original value.
[0095] It should be understood that any feature described herein may be claimed in combination with any other feature described herein, whether or not such features are from the same described embodiment.
[0096] Specific details are given in the description to provide a comprehensive understanding of the embodiments. However, it will be appreciated by those skilled in the art that the embodiments may be implemented without these specific details. In other cases, well-known circuits, processes, algorithms, structures, and techniques may not be shown in unnecessary detail to avoid obscuring the embodiments.
[0097] Although illustrative embodiments of the present disclosure have been described in detail herein, it should be understood that the concepts of the present invention may be embodied and used in other various ways, and that the appended claims are intended to be interpreted to include such variations unless limited by prior art.
Claims
1. A method for reducing delay in out-of-order data transmission, the method comprising: Detect network errors; In response to detecting the network error, sending a notification of the detection of the network error to a connected node; as well as At least one transport layer parameter is changed by each node receiving the notification, wherein after a predetermined amount of time, the at least one transport layer parameter is restored back to an original value.
2. The method according to claim 1, wherein: Sending the notification includes sending a packet with an indication that the network error was detected.
3. The method of claim 2, wherein the packet with the indication that the network error was detected includes a timestamp.
4. The method of claim 1 , wherein sending the notification comprises: The dedicated flag in subsequent packets is turned on. The method of claim 1 , wherein the network error comprises an optical or electrical device error. The method of claim 1 , wherein the network errors include thermal noise or cosmic noise.
7. The method of claim 1 , wherein detecting the network error comprises: An egress port associated with the detected network error is detected.
8. The method of claim 7, wherein sending the notification comprises: A dedicated flag is turned on in subsequent packets sent on the egress port associated with the detected network error.
9. The method of claim 7, wherein sending the notification of a detected network error to a connected node comprises: The notification is sent to a node connected to the egress port associated with the detected network error.
10. The method of claim 1, wherein changing the at least one transport layer parameter comprises: A timeout for requesting a retransmission is reduced within a predetermined amount of time.
11. A system for reducing latency in out-of-order data transmission, the system comprising one or more circuits configured to: Detect network errors; and In response to detecting the network error, sending a notification of the detection of the network error to at least one of the plurality of connected nodes, wherein: The notification causes each node receiving the notification to change at least one transport layer parameter.
12. The system of claim 11, wherein the notification comprises a packet with an indication that the network error was detected, and wherein the packet with the indication that the network error was detected comprises a timestamp.
13. The system of claim 11, wherein the notification comprises a dedicated flag in a subsequent packet being turned on.
14. The system of claim 11, wherein the network error comprises at least one of: an optical device error, an electrical device error, and thermal noise or cosmic noise.
15. The system of claim 11, wherein the network error is associated with at least one egress port.
16. The system of claim 15, wherein the notification comprises: A dedicated flag is turned on in subsequent packets sent on at least one associated said egress port.
17. The system of claim 11, wherein changing the at least one transport layer parameter comprises: A timeout for requesting a retransmission is reduced within a predetermined amount of time.
18. A device for reducing delay in out-of-order data transmission, the device comprising: Interface for detecting network errors; as well as Processing circuitry is configured to, in response to detecting the network error, send a notification of the detection of the network error to connected nodes, wherein the notification causes each node receiving the notification to reduce a latency threshold.
19. The apparatus according to claim 18, wherein: The delay threshold is reduced for a predetermined amount of time, and after the predetermined amount of time, the delay threshold is restored back to the original value.
20. The apparatus of claim 18, wherein: Sending the notification includes sending a packet with an indication that the network error was detected or turning on a dedicated flag in a subsequent packet.