Reducing the impact of incast congestion in a network
The method addresses incast congestion in networks by enabling nodes to differentiate and manage incast and non-incast traffic flows, reducing congestion's impact on non-incast streams and enhancing network efficiency.
Patent Information
- Application Number
- PCT/EP2023/087525
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2023-12-22
- Publication Date
- 2025-06-26
AI Technical Summary
Incast congestion in networks leads to increased service completion time for non-incast traffic and generates additional control traffic, harming non-incast streams.
A method that involves a destination node determining whether a traffic flow is an incast flow, transmitting this information to the source node, and having the source node indicate whether packets belong to incast or non-incast traffic, allowing each node to select appropriate traffic queues based on this indication.
This approach effectively reduces the impact of incast congestion on non-incast streams by instantaneously identifying and managing incast traffic, thereby avoiding harm to non-incast streams and improving overall network performance.
Smart Images

Figure EP2023087525_26062025_PF_FP_ABST
Abstract
Description
[0001] REDUCING THE IMPACT OF INCAST CONGESTION IN A NETWORK
[0002] TECHNICAL FIELD
[0003] The present disclosure relates, in general, to reducing the impact of incast congestion in a network. Aspects of the disclosure relate to avoiding harm to non-incast streams by incast streams.
[0004] BACKGROUND
[0005] Incast refers to a networking phenomenon where multiple sources send data simultaneously to a single receiver, causing congestion in the network. This problem leads to increased overall service completion time for any sort of traffic, particularly non-incast traffic, competing at the same bottleneck resource. Additionally, congestion can lead to additional control traffic for congestion control, as is the case with, for example, telemetry.
[0006] SUMMARY
[0007] An objective of the present disclosure is to provide a mechanism for reducing the impact of incast congestion to non- incast streams in a network.
[0008] The foregoing and other objectives are achieved by the features of the independent claims.
[0009] Further implementation forms are apparent from the dependent claims, the description and the Figures.
[0010] A first aspect of the present disclosure provides a method of reducing the impact of incast congestion in a network on nonincast streams, the method comprising receiving, by a destination node in the network, a session request from a source node in the network, determining, by the destination node, whether a traffic flow associated with the session request comprises an incast flow, transmitting a result of the determination from the destination node to the source node, based on the received result of the determination, indicating, by the source node in the network, whether packets generated for the traffic flow comprise incast traffic, and selecting, at each node of the multiple nodes in the network as the traffic flow traverses the multiple nodes, a traffic queue for the traffic flow based on the indication.
[0011] Accordingly, the impact of incast congestion can be avoided regardless of the length / duration of the traffic flow. Instantaneously and short-lived incast can be dealt with in an effective manner, while avoiding harm to incast streams.
[0012] Determining, by the destination node, whether the traffic flow associated with the session request comprises the incast flow may comprise determining, by the source node in the network, that a traffic flow in relation to which the session request is to be sent comprises the incast flow, and sending the result of the determination to the destination node as part of the session request, whereby to enable the destination node to determine whether the traffic flow associated with the session request comprises the incast flow.
[0013] Receiving, by the destination node in the network, the session request from the source node in the network may comprise receiving, by the destination node in the network, a first session request from the source node, and the method may further comprise determining that the traffic flow associated with the first session request does not comprise the incast flow, receiving, by the destination node in the network, a second session request from a node of the multiple nodes in the network, updating the result of determination associated with the first session request to thereby indicate that the traffic flow associated with the first session comprises the incast flow, and transmitting the updated result of the determination from the destination node to the source node. The method may further comprise determining, by the destination node, that the traffic flow associated with the session request comprises the incast flow based on determining that the session request relates to the same memory address as an existing session, and / or determining that the session request relates to sending traffic to a single port associated with a specific service.
[0014] Indicating, by the source node in the network, whether the packets generated for the traffic flow comprise incast traffic may comprise inserting a marker into each packet of the packets generated for the traffic flow, respectively, before transmitting the packet from the source node.
[0015] The marker may comprise a differentiated services code point, a differentiated services code point group, and / or a service level.
[0016] Selecting, at each node of multiple nodes in the network as the traffic flow traverses the multiple nodes, the traffic queue for the traffic flow based on the indication may comprise, in response to the indication indicating that the packets generated for the traffic flow comprise the incast traffic, selecting a first traffic pathway for the traffic flow, the first traffic pathway comprising a set of first traffic queues for the traffic flow, and / or, in response to the indication indicating that the packets generated for the traffic flow comprise non-incast traffic, selecting a second traffic pathway for the traffic flow, the second traffic pathway comprising a set of second traffic queues for the traffic flow.
[0017] The method may further comprise, for non- incast traffic, allocating packets to a queue of the set of first traffic queues, and, for incast traffic, allocating packets to a queue of the set of second traffic queues based on a destination address of each packet.
[0018] The method may further comprise allocating packets from respective ones of the set of second traffic queues to an incast queue of the first traffic pathway according to a round robin procedure.
[0019] The method may further comprise allocating packets from respective queues of the set of first traffic queues and respective queues of the set of second traffic queues to an out queue according to a round robin procedure.
[0020] The method may further comprise allocating packets from respective queues of the set of first traffic queues and respective queues of the set of second traffic queues to an out queue according to a weighted round robin procedure.
[0021] A second aspect of the present disclosure provides a network apparatus arranged to perform the method described herein.
[0022] A third aspect of the present disclosure provides a computer readable storage medium comprising computer program code, accessible by an apparatus comprising a processor, to provide instructions and / or data to the apparatus, the computer program code configured to, with the processor, cause the apparatus to receive, by a destination node in the network, a session request from a source node in a network, determine, by the destination node, whether a traffic flow associated with the session request comprises an incast flow, transmit a result of the determination from the destination node to the source node, based on the received result of the determination, indicate, by the source node in the network, whether packets generated for the traffic flow comprise incast traffic, and select, at each node of multiple nodes in the network as the traffic flow traverses the multiple nodes, a traffic queue for the traffic flow based on the indication.
[0023] The computer program code configured to, with the processor, cause the apparatus to determine by the destination node, whether the traffic flow associated with the session request comprises the incast flow may further comprise computer program code further configured to, with the processor, cause the apparatus to determine, by the source node in the network, that a traffic flow in relation to which the session request is to be sent comprises the incast flow, and send the result of the determination to the destination node as part of the session request, whereby to enable the destination node to determine whether the traffic flow associated with the session request comprises the incast flow.
[0024] Receiving, by the destination node in the network, the session request from the source node in the network may comprise receiving, by the destination node in the network, a first session request from the source node, and the computer program code may be further configured to, with the processor, cause the apparatus to determine that the traffic flow associated with the first session request does not comprise the incast flow, receive, by the destination node in the network, a second session request from a node of the multiple nodes in the network, update the result of determination associated with the first session request to thereby indicate that the traffic flow associated with the first session comprises the incast flow, and transmit the updated result of the determination from the destination node to the source node.
[0025] These and other aspects of the invention will be apparent from the embodiment(s) described below.
[0026] BRIEF DESCRIPTION OF THE DRAWINGS
[0027] In order that the present invention may be more readily understood, embodiments of the invention will now be described, by way of example, with reference to the accompanying drawings, in which:
[0028] Figure 1 is a flow chart of a method of reducing the impact of incast congestion in a network according to an example;
[0029] Figure 2 is a schematic depiction of a network according to an example;
[0030] Figure 3 is a flow chart of a queue selection mechanism according to an example; and
[0031] Figure 4 is schematic depiction of a network apparatus according to an example.
[0032] DETAILED DESCRIPTION
[0033] Example embodiments are described below in sufficient detail to enable those of ordinary skill in the art to embody and implement the systems and processes herein described. It is important to understand that embodiments can be provided in many alternate forms and should not be construed as limited to the examples set forth herein.
[0034] Accordingly, while embodiments can be modified in various ways and take on various alternative forms, specific embodiments thereof are shown in the drawings and described in detail below as examples. There is no intent to limit to the particular forms disclosed. On the contrary, all modifications, equivalents, and alternatives falling within the scope of the appended claims should be included. Elements of the example embodiments are consistently denoted by the same reference numerals throughout the drawings and detailed description where appropriate.
[0035] The terminology used herein to describe embodiments is not intended to limit the scope. The articles “a,” “an,” and ‘The” are singular in that they have a single referent, however the use of the singular form in the present document should not preclude the presence of more than one referent. In other words, elements referred to in the singular can number one or more, unless the context clearly indicates otherwise. It will be further understood that the terms “comprises,” “comprising,” “includes,” and / or “including,” when used herein, specify the presence of stated features, items, steps, operations, elements, and / or components, but do not preclude the presence or addition of one or more other features, items, steps, operations, elements, components, and / or groups thereof.
[0036] Unless otherwise defined, all terms (including technical and scientific terms) used herein are to be interpreted as is customary in the art. It will be further understood that terms in common usage should also be interpreted as is customary in the relevant art and not in an idealized or overly formal sense unless expressly so defined herein. The incast problem is a well-known problem in networking. As such, many different methods have been used in order to address this issue. In computer networking, an elephant flow is an extremely large flow set up by a TCP (or any other protocol) flow measured over a network link. While the elephant flows are typically not numerous, due to their size, they can occupy a disproportionate share of the total bandwidth over a period of time. In particular, the elephant flows may starve bottleneck links of resources, causing delays to smaller (i.e., mice) flows.
[0037] To alleviate the impact of elephant flows, differential scheduling solutions have been utilised, deprioritising the elephant flow to enable the mice flows to pass loss-free at low latency. Unfortunately, elephant flows need detection. In other words, deprioritising the elephant flow is a reactive mechanism that takes time. As such, this approach may not be suitable for highspeed applications, such as RDMA. Furthermore, the differential scheduling can lead to inadequate buffer management, since deprioritising may starve the elephant flows of resources once the number of mice flows increases.
[0038] Another approach involves explicit congestion notification (ECN) to complement the priority -based flow control (PFC) at Ethernet level. In general terms, this approach utilises ECN marking in the bottleneck resource. Unfortunately, similar to differential scheduling described above, the ECN use is also a reactive mechanism having a latency associated with resolving the congestion, i.e., sending of ECN-marked packets and action in the endpoint.
[0039] Telemetry-based congestion control, e.g., done by high precision congestion control, has also been utilised in hopes of addressing the above-described problem. This approach involves estimating flow rates through a bottleneck, using in-band telemetry to estimate flow rates for balanced transfer through the bottleneck. However, once again, this approach is a reactive mechanism with a latency in finding the estimate and adjusting the sending rate as a result. This approach causes telemetry traffic.
[0040] According to an example, there is provided a mechanism to reduce the impact of incast congestion in a network. Advantageously, the approach described herein reduces the impact of incast congestion regardless of the length / duration of the flow, i.e., it is suitable for dealing with instantaneous and short-lived incast, while avoiding harm to non-incast streams by incast streams. The approach utilises application-awareness of communication (i.e., whether a particular data flow comprises incast or non-incast traffic) for proactive marking, thus addressing the incast problem instantaneously, without any latency associated therewith.
[0041] Examples in the present disclosure can be provided as methods, systems or machine-readable instructions, such as any combination of software, hardware, firmware or the like. Such machine-readable instructions may be included on a computer readable storage medium (including but not limited to disc storage, CD-ROM, optical storage, etc.) having computer readable program codes therein or thereon.
[0042] The present disclosure is described with reference to flow charts and / or block diagrams of the method, devices and systems according to examples of the present disclosure. Although the flow diagrams described above show a specific order of execution, the order of execution may differ from that which is depicted. Blocks described in relation to one flow chart may be combined with those of another flow chart. In some examples, some blocks of the flow diagrams may not be necessary and / or additional blocks may be added. It shall be understood that each flow and / or block in the flow charts and / or block diagrams, as well as combinations of the flows and / or diagrams in the flow charts and / or block diagrams can be realized by machine readable instructions.
[0043] The machine-readable instructions may, for example, be executed by a machine such as a general-purpose computer, user equipment such as a smart device, e.g., a smart phone, a special purpose computer, an embedded processor or processors of other programmable data processing devices to realize the functions described in the description and diagrams. In particular, a processor or processing apparatus may execute the machine-readable instructions. Thus, modules of apparatus (for example, a module implementing a comparator unit, or a firewall structure and so on) may be implemented by a processor executing machine readable instructions stored in a memory, or a processor operating in accordance with instructions embedded in logic circuitry. The term 'processor' is to be interpreted broadly to include a CPU, processing unit, ASIC, logic unit, or programmable gate set etc. The methods and modules may all be performed by a single processor or divided amongst several processors.
[0044] Such machine-readable instructions may also be stored in a computer readable storage that can guide the computer or other programmable data processing devices to operate in a specific mode. For example, the instructions may be provided on a non- transitory computer readable storage medium encoded with instructions, executable by a processor.
[0045] Figure 1 is a flow chart of a method of reducing the impact of incast congestion in a network according to an example. In block 101 , the method comprises receiving, by a destination node in the network, a session request from a source node in the network. The source of the communication may send a session request to the sink of the communication. The session request may be sent from the source to the sink via the source's ingress and egress switch.
[0046] In block 102, the method comprises determining, by the destination node, whether a traffic flow associated with the session request comprises an incast flow. The sink may determine whether the session may be part of a communication causing an incast traffic pattern, for example, whether the source is requesting the memory content of the same memory address / range as another outstanding session (i.e., remote direct memory access, RDMA), or whether the source is requesting to send traffic to the same port, indicating a specific service such as real-time video (i.e., Internet traffic example). In other words, an application may deduce their potential incast nature from its communication pattern, e.g., sending data to a centralised point with many sources, or high frequency of a source-sink pattern.
[0047] For RDMA-based sessions, queue pairs (QPs) may be used to identify possible incast situations. An application may implement a typical scatter / gather semantic, often leading to an incast traffic pattern. The application may wish to indicate the incast session through a suitable interface, for example, in the QP creation phase. This enables direct indication of incast when creating the QP. The skilled person would appreciate that the implementation of such indication may require a suitable verb application programming interface (API) extension. Certain kinds of traffic in RDMA, e.g., source node requests to specific memory areas (not subject to concurrent access), may be determined to comprise non-incast flow. Furthermore, for RDMA-based sessions, certain situations may be derived from existing verb API interaction to detect an incast situation rather than explicitly indicating that the traffic flow is associated with an incast flow. Examples of such situations comprise a sink creating several QPs with different source IDs, with those QPs being (or being set to be) in a ready to receive (RTR) state, i.e., the ibv modify _qp() is being called, and / or a sink issuing rdma _post_recv() or rdma _post_read() verbs to its API.
[0048] For transport protocol-based sessions, connections requests to the same port at the sink may be used as an indication that the traffic flow comprises an incast flow. Additionally, the destination node may employ an application-specific counter for the number of sessions that must be exceeded for incast to be considered. Furthermore, ingress marking may replace / complement application-based marking. The DCN operator may use policy-based ingress marking that marks traffic from certain (well- known) applications as incast traffic, instead of relying on application-based marking.
[0049] In block 103, the method comprises transmitting a result of the determination from the destination node to the source node. The destination node (i.e., the sink) may return the session request, together with the result of the previous decision, indicated by a unique marker. For example, a marker 1=0 may be used to indicate non-incast traffic, whereas a marker 1= 1 may be used to indicate incast traffic. The source node may maintain context information that associates the new session with the incast determination result (e.g., 1 or O).
[0050] In an embodiment, the source node may have previously requested a session at the destination node, which has been found to be non- incast, i.e., a marker 1=0 has been returned. However, due to new sessions arriving at the destination node, a decision may be made by the destination node that there may now exist incast traffic due to the set of source node / destination node relations. In such case, the destination node may return a revised decision to the source node by, for example, including a marker indicative of incast traffic, thus changing an ongoing session into an incast one. The source node may then revise the context information by changing the stored incast determination result accordingly.
[0051] In block 104, the method comprises, based on the received result of the determination, indicating, by the source node in the network, whether packets generated for the traffic flow comprise incast traffic. In other words, for generation of packets from the source node to the destination node, the source node may retrieve the incast determination result from its stored context information and insert the marker (e.g., 1=0 or 1=1) into the generated packet prior to sending it.
[0052] In block 105, the method comprises selecting, at each node of the multiple nodes in the network as the traffic flow traverses the multiple nodes, a traffic queue for the traffic flow based on the indication. Figure 2 is a schematic depiction of a network according to an example. The network 200 may comprise a source node 201, an ingress switch 202, an egress switch 203, and a destination node 204. The egress switch 203 may comprise a top of rack (ToR) switch - the switch immediately before the destination node 204. The judgment whether a particular switch is the egress switch to the destination node 204 may be performed based on an address of the destination node 204.
[0053] Either or both of the switches 202 and 203 may perform the de-incaster queuing, i.e., select the traffic queue based on the indication whether the packets generated by the source node 201 comprise incast traffic. Each node (switch) may maintain a plurality of different queues, or groups of queues. In the example described herein, each switch may maintain four different queues for the traffic. A first traffic pathway comprising a set of first traffic queues may be selected for the traffic flow comprising the incast traffic. A second traffic pathway comprising a set of second traffic queues may be selected for non-incast traffic.
[0054] To better illustrate the queue selection mechanism, reference is made to Figure 3, which is a flow chart of a queue selection mechanism according to an example. A non-incast queue 301 may be used for non-incast packets, i.e., packets associated with the appropriate marker (in this case, 1=0). For incast packets, the packets may be allocated to a queue of the sink (destination node) queues 302, sorted by destination address. The sink queues 302 may support N destination nodes, and a common queue for all destination nodes above N, i.e., unsupported destination nodes. An incast queue 303 may be used for all packets comprising the 1=1 marker. Packets may be allocated to the incast queue 303 from the sink queues 302 using a round-robin procedure. Finally, an output queue 304 may be used for outgoing packets comprising the 1=1 marker. Packets may be allocated to the output queue 304 from the non-incast queue 301 and the incast queue 303 using a round-robin procedure. The sink queues 302 and the incast queue 303 may form the second traffic pathway. The non-incast queue 301 may form the first traffic pathway. Once a packet has been received by a node (e.g., a switch), if the marker associated therewith comprises 1=0, the packet may be placed into the non-incast queue 301. Else, the packet may be placed into the respective queue of the sink queues 202 that has previously been used for the packet destination, or, if the number of destination nodes is greater than the number of destination nodes supported, into the common queue. Then, a round-robin procedure may be applied to all destination queues (including the common queue) to place the packet(s) into the incast queue 303. Finally, the packets from the non- incast queue 301 and the incast queue 303 may be allocated to the output queue 304 using the round-robin procedure. This procedure may be performed for each packet received at a switch.
[0055] Instead of using the round robin procedure to place the packets from the non- incast queue 301 and the incast queue 303 into the output queue 304, equally sharing the output queue 304 between the incast and non-incast traffic, a weighted round robin procedure may be used. The weighted round robin procedure may define a desired split, for example, 20% for non- incast traffic and 80% for incast traffic. This can enable reasonable flow time completion times for incast flows, while still allowing for reasonable completion time for non-incast flows. Weights may implement a specific network / tenant policy for sharing bandwidth between typical incast applications (such as scatter / gather semantics) and typical non-incast applications (such as status retrievals).
[0056] In addition, instead of using a static Weighted-RR to place packets from the non-incast queue 301 and the incast queue 303 into the output queue 304, the weights may be adjusted based on telemetry information, thus representing a dynamic share between the traffic. An exponential average may be used to control the dynamic adjustment of the weights over time.
[0057] Alternatively, instead of using the round-robin procedure, priority queuing could also be used, giving priority to either incast or non-incast flows before serving any other kind of flow(s). While such approach can be useful for traffic scenarios in which incast applications are background applications - for example, Al reasoning with scatter / gather semantics, which may tolerate degradation compared to non-incast, priority applications - strict priorities could lead to similar starvation problems as those described in the prior art.
[0058] Furthermore, any extended queuing could be signalled through different DSCPs, e.g., different CPs may indicate different weights or different key performance indicators (such as low latency for incast, drop rates, or similar).
[0059] Importantly, the method described herein in relation to blocks 101-105 may either be implemented across a single transport protocol (for example, RDMA), or across transport protocols. In order to implement the method in a single protocol only, integration into the transport protocol API may be required at endpoints in order to signal the I bit information (i.e., marker indicating whether the traffic has been determined to comprise incast traffic), as well as integration into an appropriate protocol header information. Similarly, in order to implement the method described herein across different transport protocols, the API and the protocol header information implementation may be performed across all of the transport protocols.
[0060] Problems may arise when traffic competes at a bottleneck switch with transport protocols that do not utilise the method described herein, particularly if that competing protocol causes incast traffic. To address this, network isolation may be used to separate traffic compliant with the method described herein from non-compliant applications.
[0061] In addition, the method described herein may also be implemented in a single datagram service between a transport layer and a network layer. In particular, the method may be implemented in the buffer management of a single shim layer between all transports (e.g., RDMA / RoCE, TCP, and so on) and the network layer. The shim layer may expose a single API for signalling of the I bit information, while using a single protocol header integration for signalling the information across the network, therefore simplifying the implementation of the method across different protocols. While a single datagram service (EQDS) may also implement incast detections according to blocks 101-104, application-specific usage of the I bit information may still require integration into the specific applications.
[0062] Figure 4 is schematic depiction of a network apparatus according to an example. The network apparatus may comprise a processor 403, and a memory 405 coupled to the processor 403 and configured to store instructions or program code 407, executable by the processor 403. The network apparatus 400 can be, e.g., a network device (physical or virtual), or part thereof. The network apparatus 400 may comprise the program code 407 arranged to cause the apparatus to perform the method of reducing the impact of incast congestion in a network as described herein.
[0063] According to an example, machine-readable instructions can be loaded onto a computer or other programmable data processing devices, so that the computer or other programmable data processing devices perform a series of operations to produce computer-implemented processing, thus the instructions executed on the computer or other programmable devices provide an operation for realizing functions specified by flow(s) in the flow charts and / or block(s) in the block diagrams.
[0064] Further, the teachings herein may be implemented in the form of a computer or software product, such as a non-transitory machine-readable storage medium, the computer software or product being stored in a storage medium and comprising a plurality of instructions, e.g., machine readable instructions, for making a computer device implement the methods recited in the examples of the present disclosure.
[0065] In some examples, some methods can be performed in a cloud-computing or network-based environment. Cloud-computing environments may provide various services and applications via the Internet. These cloud-based services (e.g., software as a service, platform as a service, infrastructure as a service, etc.) may be accessible through a web browser or other remote interface of the user equipment for example. Various functions described herein may be provided through a remote desktop environment or any other cloud-based computing environment.
[0066] While various embodiments have been described and / or illustrated herein in the context of fully functional computing systems, one or more of these exemplary embodiments may be distributed as a program product in a variety of forms, regardless of the particular type of computer-readable-storage media used to actually carry out the distribution. The embodiments disclosed herein may also be implemented using software modules that perform certain tasks. These software modules may include script, batch, or other executable files that may be stored on a computer-readable storage medium or in a computing system. In some embodiments, these software modules may configure a computing system to perform one or more of the exemplary embodiments disclosed herein. In addition, one or more of the modules described herein may transform data, physical devices, and / or representations of physical devices from one form to another.
[0067] The preceding description has been provided to enable others skilled in the art to best utilize various aspects of the exemplary embodiments disclosed herein. This exemplary description is not intended to be exhaustive or to be limited to any precise form disclosed. Many modifications and variations are possible without departing from the spirit and scope of the instant disclosure. The embodiments disclosed herein should be considered in all respects illustrative and not restrictive. Reference should be made to the appended claims and their equivalents in determining the scope of the instant disclosure.
Claims
CLAIMS1. A method of reducing incast congestion in a network, the method comprising: receiving, by a destination node in the network, a session request from a source node in the network (101); determining, by the destination node, whether a traffic flow associated with the session request comprises an incast flow (102); transmitting a result of the determination from the destination node to the source node (103); based on the received result of the determination, indicating, by the source node in the network, whether packets generated for the traffic flow comprise incast traffic (104); and selecting, at each node of the multiple nodes in the network as the traffic flow traverses the multiple nodes, a traffic queue for the traffic flow based on the indication (105).
2. The method of claim 1, wherein determining, by the destination node, whether the traffic flow associated with the session request comprises the incast flow (102) comprises: determining, by the source node in the network, that a traffic flow in relation to which the session request is to be sent comprises the incast flow; and sending the result of the determination to the destination node as part of the session request, whereby to enable the destination node to determine whether the traffic flow associated with the session request comprises the incast flow.
3. The method of claim 1 or 2, wherein receiving, by the destination node in the network, the session request from the source node in the network (101) comprises receiving, by the destination node in the network, a first session request from the source node, wherein the method further comprises: determining that the traffic flow associated with the first session request does not comprise the incast flow; receiving, by the destination node in the network, a second session request from a node of the multiple nodes in the network; updating the result of determination associated with the first session request to thereby indicate that the traffic flow associated with the first session comprises the incast flow; and transmitting the updated result of the determination from the destination node to the source node.
4. The method of claim 1 , 2 or 3, further comprising: determining, by the destination node, that the traffic flow associated with the session request comprises the incast flow based on determining that the session request relates to the same memory address as an existing session, and / or determining that the session request relates to sending traffic to a single port associated with a specific service.
5. The method of any one of claims 1 to 4, wherein indicating, by the source node in the network, whether the packets generated for the traffic flow comprise incast traffic (104) comprises: inserting a marker into each packet of the packets generated for the traffic flow, respectively, before transmitting the packet from the source node.
6. The method of claim 5, wherein the marker comprises a differentiated services code point, a differentiated services code point group, and / or a service level.
7. The method of any preceding claim, wherein selecting, at each node of multiple nodes in the network as the traffic flow traverses the multiple nodes, the traffic queue for the traffic flow based on the indication (105) comprises: in response to the indication indicating that the packets generated for the traffic flow comprise the incast traffic, selecting a first traffic pathway for the traffic flow, the first traffic pathway comprising a set of first traffic queues for the traffic flow; and / or in response to the indication indicating that the packets generated for the traffic flow comprise non-incast traffic, selecting a second traffic pathway for the traffic flow, the second traffic pathway comprising a set of second traffic queues for the traffic flow.
8. The method of claim 7, further comprising: for non-incast traffic, allocating packets to a queue of the set of first traffic queues; for incast traffic, allocating packets to a queue of the set of second traffic queues based on a destination address of each packet.
9. The method of claim 8, further comprising: allocating packets from respective ones of the set of second traffic queues to an incast queue of the first traffic pathway according to a round robin procedure.
10. The method of claim 9, further comprising: allocating packets from respective queues of the set of first traffic queues and respective queues of the set of second traffic queues to an out queue according to a round robin procedure.
11. The method of claim 10, further comprising: allocating packets from respective queues of the set of first traffic queues and respective queues of the set of second traffic queues to an out queue according to a weighted round robin procedure.
12. A network apparatus (400) arranged to perform the method of any one of claims 1 to 11.
13. A computer readable storage medium comprising computer program code, accessible by an apparatus comprising a processor, to provide instructions and / or data to the apparatus, the computer program code configured to, with the processor, cause the apparatus to: receive, by a destination node in the network, a session request from a source node in a network; determine, by the destination node, whether a traffic flow associated with the session request comprises an incast flow; transmit a result of the determination from the destination node to the source node; based on the received result of the determination, indicate, by the source node in the network, whether packets generated for the traffic flow comprise incast traffic; and select, at each node of multiple nodes in the network as the traffic flow traverses the multiple nodes, a traffic queue for the traffic flow based on the indication.
14. The computer readable storage medium of claim 13, wherein the computer program code configured to, with the processor, cause the apparatus to determine by the destination node, whether the traffic flow associated with the session request comprises the incast flow further comprises computer program code further configured to, with the processor, cause the apparatus to:determine, by the source node in the network, that a traffic flow in relation to which the session request is to be sent comprises the incast flow; and send the result of the determination to the destination node as part of the session request, whereby to enable the destination node to determine whether the traffic flow associated with the session request comprises the incast flow.
15. The computer readable storage medium of claim 13 or 14, wherein receiving, by the destination node in the network, the session request from the source node in the network comprises receiving, by the destination node in the network, a first session request from the source node, wherein the computer program code is further configured to, with the processor, cause the apparatus to: determine that the traffic flow associated with the first session request does not comprise the incast flow; receive, by the destination node in the network, a second session request from a node of the multiple nodes in the network; update the result of determination associated with the first session request to thereby indicate that the traffic flow associated with the first session comprises the incast flow; and transmit the updated result of the determination from the destination node to the source node.
Citation Information
Patent Citations
In-band signaling for latency guarantee service (LGS)
WO2021174236A2