Controlled network routing for machine learning workloads

The method optimizes network resources for ML workloads by using a host-controller system to determine and assign paths based on topology and traffic patterns, reducing latency and congestion for improved ML data flow management.

WO2026002374A1PCT designated stage Publication Date: 2026-01-02HUAWEI TECH CO LTD +2
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
PCT/EP2024/067825
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-06-25
Publication Date
2026-01-02

AI Technical Summary

Technical Problem

Conventional methods fail to optimize network resources for efficient handling of Machine Learning Workloads (MLWLs), leading to inefficiencies in path assignment and increased data congestion due to complex network topologies and varying data flow scales.

Method used

A method involving a host and a controller that determines ML data flows, retrieves network topology and traffic patterns, and uses a path assignment algorithm to optimize path selection, ensuring efficient and reliable transmission.

Benefits of technology

This approach minimizes latency and congestion, enhancing network performance and reliability for ML-intensive computations by ensuring dedicated resource allocation and customized routing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure EP2024067825_02012026_PF_FP_ABST
    Figure EP2024067825_02012026_PF_FP_ABST
Patent Text Reader

Abstract

A method for a network having a network topology, the network comprising a host and a controller comprising a processing unit of the host for determining that a data flow for an executing Machine Learning process is to be transmitted to a receiver. Furthermore, determining that the data flow is part of a Machine Learning Workload (MLWL), thereby being a Machine Learning (ML) data flow, and in response sending a MLWL path assignment request for the data flow to the controller. Furthermore, receiving the MLWL path assignment request retrieving the network topology and data traffic patterns, determining a path based on a path assignment algorithm based on the network topology, the data traffic patterns and active ML data flows, sending a path assignment response to the host indicating the assigned path to the host, receiving the path assignment response, and transmitting the ML data flow according to the assigned path.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] CONTROLLED NETWORK ROUTING FOR MAC HINE LEARNING WORKLOADS

[0002] TECHNICAL FIELD

[0003] The present disclosure relates generally to the field of network computing and more specifically, to a method for a network having network topology and to a network having network topology.

[0004] BACKGROUND

[0005] In modem network environments, the demand for efficient handling of Machine Learning Workloads (MLWLs) has significantly increased. Machine learning (ML) applications often involve the transmission of large volumes of data, necessitating optimized routing strategies to ensure timely and reliable delivery. Moreover, a path assignment in large-scale network having MLWL is a critical affair due to complex network topology. The complex network topology allows for multiple possible paths for data flow between a source and a destination pair. However, due to multiple paths selection of an optimal path for data flow becomes a challenging task due to an increased possibility of overall data transmission latency, data congestion, and the like. Additionally, traffic in large-scale network is often composed of two entirely different scales of data flow, such as an elephant flow and a mouse flow. The elephant flow refers to a data flow that requires high bandwidth and is transmitted for a long period of time while the mouse flow refers to a data flow that requires lower bandwidth and can be transmitted over a short period of time. Therefore, elephant flows are known to be more challenging than mouse flow in terms of path assignment, since if multiple elephant flows are incidentally routed through the same path, it can lead to deadlocks, data congestion, and potential data loss.

[0006] Conventionally, certain attempts have been made to provide an optimized path assignment process for data flows, such as by balancing data transmission load between multiple paths by computing a hash in the packet header and using the hash value for selecting the path for each data flow or utilizing a centralized node to allocate paths for each data flow according to a network-wide perspective or based on performance metrics. However, such attempts fail to optimize network resources leading to inefficiencies in path assignment and network utilization. Additionally, such attempts often struggle to manage data flows effectively and reliably, with a higher possibility of data congestion. Thus, there exists a technical problem of how to control path assignment of data flows in the MLWL thereby improving network utilization during data flow.

[0007] Therefore, in light of the foregoing discussion, there exists a need to overcome the aforementioned drawbacks associated with the conventional method and network for controlling path assignment of data flow having network topology.

[0008] SUMMARY

[0009] The present disclosure provides a method for a network having network topology. The present disclosure further provides the network having network topology. The present disclosure provides a solution to the existing problem of how to control path assignment of data flows in the MLWL thereby improving network utilization during data flow. An objective of the present disclosure is to provide a solution that overcomes at least partially the problems encountered in the prior art and provides the method for the network having network topology, such as by providing a controlled network routing for machine learning workloads.

[0010] One or more objectives of the present disclosure are achieved by the solutions provided in the enclosed independent claims. Advantageous implementations of the present disclosure are further defined in the dependent claims. In one aspect, the present disclosure provides the method for a network having a network topology, the network comprising a host and a controller. The method comprises a processing unit of the host determining that a data flow for an executing Machine Learning (ML) process is to be transmitted to a receiver, the data flow having a size. Furthermore, the method includes determining that the data flow is part of a Machine Learning Workload, (MLWL) thereby being a ML data flow and in response thereto sending a MLWL path assignment request for the data flow to the controller, wherein the method further comprises a processing unit of the controller thereby receiving the MLWL path assignment request retrieving the network topology and data traffic patterns, determining a path based on a path assignment algorithm based on the network topology, the data traffic patterns and active ML data flows, sending a path assignment response to the host indicating the assigned path to the host. The method further comprises the processing unit of the host thereby receiving the path assignment response and transmitting the ML data flow according to the assigned path.

[0011] Advantageously, the method is used to provide an optimized path assignment for handling MLWL in the network having network topology. The processing unit of the host is configured to identify data flows associated with executing ML processes based on the size and data traffic patterns, distinguishing them from other network data traffic to allow prioritization and management of the MLWL, enabling customized network routing and prevent the possibility of the data congestion. Furthermore, upon determining that the data flow is part of an ML workload, the host is configured to send a specialized path assignment request to the controller, ensuring dedicated handling and resource allocation for the corresponding ML data flows. The processing unit of the controller receives the path assignment request, retrieves the network topology and data traffic patterns, and determines an optimized path in order to assign the path to having minimum latency. The controller then sends the assigned path to the host, facilitating seamless communication and coordination for efficient, reliable ML data flow transmission according to the designated path. Moreover, such a coordinated approach reduces the overall data transmission delays and errors that are associated with manual assignment, leading to the efficient execution of large-scale MLWL. As a result, the method is used to ensure efficient resource allocation, minimize congestion, and improve performance and reliability for ML-intensive computations in the network.

[0012] In another aspect, the present disclosure provides a network having a network topology, the network comprising a host and a controller. Moreover, a processing unit of the host is configured to determine that a data flow for an executing ML process is to be transmitted to a receiver, the data flow having a size, determine that the data flow is part of a MLWL thereby being a ML, data flow, and in response thereto send a MLWL path assignment request for the data flow to the controller. A processing unit of the controller is configured to thereby receive the MLWL path assignment request retrieve the network topology and data traffic patterns, determine a path based on a path assignment algorithm based on the network topology, the data traffic patterns and active ML data flows, send a path assignment response to the host indicating the assigned path to the host, wherein the processing unit of the host is configured to thereby receive the path assignment response and transmit the ML data flow according to the assigned path.

[0013] The network achieves all the advantages and technical effects of the method of the present disclosure.

[0014] In yet another aspect, the present disclosure provides a method for a network having a network topology. Moreover, the network includes a host and a controller, wherein the method comprises a processing unit of the controller receiving a MLWL path assignment request from the host for a data flow for a ML process executing on the host, retrieving the network topology and data traffic patterns, determining a path based on a path assignment algorithm based on the network topology, the data traffic patterns and active ML data flows, sending a path assignment response to the host indicating the assigned path to the host, thereby causing the host transmitting the ML data flow according to the assigned path.

[0015] Advantageously, the method is used to determine the optimal path for transmitting ML data flows within the network in order to ensure that ML-related data traffic is routed efficiently, minimizing latency, and maximizing throughput. Furthermore, the utilization of a path assignment algorithm based on network topology and the data traffic patterns allows for the efficient allocation of network resources in order to optimize resource utilization, ensuring that sufficient bandwidth is allocated to support ML computations.

[0016] In another aspect, the present disclosure provides a controller configured to be used in a network having a network topology, the network comprising a host and the controller. A processing unit of the controller is configured to receive a MLWL path assignment request from the host for a data flow for a ML process executing on the host, retrieve the network topology and data traffic patterns, determine a path based on a path assignment algorithm based on the network topology, the data traffic patterns and active ML data flows, send a path assignment response to the host indicating the assigned path to the host, thereby causing the host transmit the ML data flow according to the assigned path.

[0017] The controller achieves all the advantages and technical effects of the method of the present disclosure.

[0018] In yet another aspect, the present disclosure provides a method for a network having a network topology, the network comprising a host and a controller, wherein a processing unit of the host is configured to determining that a data flow for an executing Machine Learning process is to be transmitted to a receiver, the data flow having a size, determining that the data flow is part of a MLWL, thereby being a ML, data flow, and in response thereto sending a Machine Learning Workload, MLWL, path assignment request for the data flow to the controller, and then receiving a path assignment response indicating path determined by a path assignment algorithm based on the network topology, the data traffic patterns and active ML data flows and transmitting the ML data flow according to the assigned path.

[0019] Advantageously, the method is used to optimize the ML data flow handling and routing within the network by leveraging dynamic path assignment based on the network topology, data traffic patterns, and active MLWL. By proactively determining and assigning paths for ML data flows, the method is used to ensure an efficient resource allocation with reduced network congestion, ultimately enhancing overall network performance and responsiveness to MLWL requirements. Additionally, the method is used to streamline the communication between the host and the controller, facilitating seamless coordination and transmission of ML data flows thereby improving network efficiency and reliability.

[0020] In another aspect, the present disclosure provides a host configured to be used in a network having a network topology, the network comprising the host and a controller. A processing unit of the host is configured to determine that a data flow for an executing Machine Learning process is to be transmitted to a receiver, the data flow having a size, determine that the data flow is part of a MLWL, thereby being a ML, data flow, and in response thereto send a MLWL path assignment request for the data flow to the controller, and then receive a path assignment response indicating path determined by a path assignment algorithm based on the network topology, the data traffic patterns and active ML data flows and transmit the ML data flow according to the assigned path.

[0021] The host achieves all the advantages and technical effects of the method of the present disclosure.

[0022] It is to be appreciated that all the aforementioned implementation forms can be combined.

[0023] It has to be noted that all devices, elements, circuitry, units and means described in the present application could be implemented in the software or hardware elements or any kind of combination thereof. All steps which are performed by the various entities described in the present application, as well as the functionalities described to be performed by the various entities are intended to mean that the respective entity is adapted to or configured to perform the respective steps and functionalities. Even if, in the following description of specific embodiments, a specific functionality or step to be performed by external entities is not reflected in the description of a specific detailed element of that entity which performs that specific step or functionality, it should be clear for a skilled person that these methods and functionalities can be implemented in respective software or hardware elements, or any kind of combination thereof. It will be appreciated that features of the present disclosure are susceptible to being combined in various combinations without departing from the scope of the present disclosure as defined by the appended claims.

[0024] Additional aspects, advantages, features, and objects of the present disclosure would be made apparent from the drawings and the detailed description of the illustrative implementations construed in conjunction with the appended claims that follow.

[0025] BRIEF DESCRIPTION OF THE DRAWINGS

[0026] The summary above, as well as the following detailed description of illustrative embodiments, is better understood when read in conjunction with the appended drawings. For the purpose of illustrating the present disclosure, exemplary constructions of the disclosure are shown in the drawings. However, the present disclosure is not limited to specific methods and instrumentalities disclosed herein. Moreover, those in the art will understand that the drawings are not to scale. Wherever possible, like elements have been indicated by identical numbers.

[0027] Embodiments of the present disclosure will now be described, by way of example only, with reference to the following diagrams wherein:

[0028] FIG. 1 is a flowchart of a method for a network having network topology, in accordance with an embodiment of the present disclosure;

[0029] FIG. 2 is a block diagram that illustrates a network having network topology, in accordance with an embodiment of the present disclosure;

[0030] FIG. 3 is a flowchart of a method for a processing unit of a controller of a network having a network topology, in accordance with another embodiment of the present disclosure;

[0031] FIG. 4 is a block diagram that illustrates a controller to be used in a network having network topology, in accordance with an embodiment of the present disclosure;

[0032] FIG. 5 is a flowchart of a method for a processing unit of a host of a network having a network topology, in accordance with another embodiment of the present disclosure;

[0033] FIG. 6 is a block diagram that illustrates a host to be used in a network having network topology, in accordance with an embodiment of the present disclosure;

[0034] FIG. 7 is a diagram that illustrates a path assignment for a network having network topology, in accordance with an embodiment of the present disclosure; and

[0035] FIG. 8 is a diagram that illustrates a prediction of MLWL assignment request for an ML data flow, in accordance with an embodiment of the present disclosure.

[0036] In the accompanying drawings, an underlined number is employed to represent an item over which the underlined number is positioned or an item to which the underlined number is adjacent. A non-underlined number relates to an item identified by a line linking the non-underlined number to the item. When a number is non-underlined and accompanied by an associated arrow, the non-underlined number is used to identify a general item at which the arrow is pointing. DETAILED DESCRIPTION OF EMBODIMENTS

[0037] The following detailed description illustrates embodiments of the present disclosure and ways in which they can be implemented. Although some modes of carrying out the present disclosure have been disclosed, those skilled in the art would recognize that other embodiments for carrying out or practicing the present disclosure are also possible.

[0038] FIG. 1 is a flowchart of a method for a network having network topology, in accordance with an embodiment of the present disclosure. With reference to FIG. 1, there is shown a flowchart of a method 100 for the network having the network topology. The method 100 includes steps 102 to 118.

[0039] There is provided the method 100 for the network having a network topology. Moreover, the network includes a host and a controller. In an implementation, the host and the controller of the network are configured to execute the method 100 in order to provide a controlled network for machine learning workloads (MEWL). The method 100 is used to provide a routing solution for the data transmission in the network to enhance the overall network performance.

[0040] At step 102, the method 100 includes a processing unit of the host for determining that a data flow for an executing Machine Learning (ML) process is to be transmitted to a receiver, the data flow having a size. The data flow refers to a sequence of data packets that are transmitted between two hosts, such as from a source (sender) to a destination (receiver), over the network. In an implementation, the ML data flow is an elephant data flow. Moreover, the processing unit of the host is configured to determine the data flow associated with the executing ML process that is required to be transmitted to the receiver in order to distinguish the MLWL data traffic from other types of network traffic by recognizing the data flows originating from the ML processes and the corresponding sizes. The processing unit monitors the execution of the ML processes and identifies the ML processes associated with the data flows. As a result, the method 100 is used to prioritize and handle ML workload traffic differently, optimize resource allocation based on the data flow sizes, and implement customized routing and congestion control mechanisms tailored for the MLWL. Additionally, the method 100 is used to facilitate proactive network management by identifying ML data flows in order to ensure efficient and reliable data transmission for the ML computations for large-scale data transfers, leading to improved network performance and reliability in ML-intensive environments.

[0041] At step 104, the method 100 includes determining that the data flow is part of a Machine Learning Workload (MLWL) thereby being the ML data flow. In an implementation, the processing unit of the host is configured to determine that the data flow is the part of the MLWL. Moreover, the determination of the data flow as the part of the MLWL that can be further utilized to handle resources that are associated with the data transmission for efficient, reliable, and optimized data transmission within the network. In accordance with an embodiment, the method 100 further includes the processing unit of the host for determining that the data flow is part of the MLWL based on determining that the data flow is for the executing ML process. In an implementation, the processing unit of the host is configured to monitor the execution of processes on the host and further identify the ML processes, such as by analysing process identifiers, program names, or other attributes associated with ML computations. Thereafter, the processing unit of the host is configured to associate the data flows originating from that ML process with the ML workloads in order to allow targeted optimization and handling. As a result, the determination of the data flow as a part of the MLWL based on the determination that the data flow is for executing the ML process allows the network to allocate resources efficiently and reliably in order to support the ML computations with the required bandwidth and overall data processing capacity effectively and accurately. Therefore, the determination of the data flows based on the ML processes allows for the implementation of customized network path routing, and congestion control mechanisms that are optimized for the MLWL thereby leading to improved network performance and reliability in ML-intensive environments.

[0042] In accordance with an embodiment, the method 100 further includes the processing unit of the host determining that the data flow is part of the MLWL based on the size of the data flow. Moreover, the data flow is determined to be an ML data flow if the size of the data flow exceeds a threshold level. The threshold value can be determined by the processing unit of the host or by the processing unit of the controller. However, the threshold value is based on the ML application. In an implementation, if the size of the data flow exceeds a predefined threshold level, then, in that case, the data flow is determined as the ML data flow that indicates the association of the ML data flow with the MLWL. For example, the method 100 is used to determine whether the data flow is part of the MLWL, such as by measuring the size of the data flow (i.e., in terms of bytes, packets, or other relevant metrics). After that, the method 100 is used to compare the size of the data flow against the predefined threshold level and if the size of the data flow exceeds the pre-defined threshold level, then, in that case, the data flow is classified as an ML data flow. As a result, the determination of the data flow as a part of the ML data flow based on the size of the data flow enabled the method 100 to categorize the data flow accurately, efficiently, and reliably, such as by offering flexibility through customizable threshold selection based on network needs.

[0043] At step 106, the method 100 includes sending the MLWL path assignment request for the data flow to the controller in response to the determination that the data flow is part of the MLWL thereby being the ML data flow. In an implementation, the MLWL path assignment request refers to a request that is sent by the processing unit of the host to the processing unit of the controller for the assignment of the dedicated network path for the ML data flows. Moreover, the MLWL path assignment request includes information, such as the origin of the data flow, destination, size, and the like that is required for the identification of the ML data flow. The processing unit of the controller receives the path assignment request and initiates the process of determining an optimal path for the ML data flow based on the network topology, traffic patterns, and the like. As a result, the method 100 ensures dedicated handling for the ML workloads by initiating specialized path assignment requests and optimizing resource allocation and routing for improved performance and reliability.

[0044] At step 108, the method 100 further includes the processing unit of the controller thereby receiving the MLWL path assignment request. In an implementation, upon receiving the MLWL path assignment request, the processing unit of the controller is configured to allocate the optimized path from the network having network topology in order to ensure an efficient and reliable ML data flow thereby ensuring reliable and consistent data flow with reduced overall data transmission time.

[0045] At step 110, the method 100 further includes retrieving the network topology and data traffic patterns. In an implementation, the controller is configured to retrieve the network topology that includes information, such as a layout of network nodes of the network, and interconnections of the corresponding network nodes. Additionally, the controller retrieves the data traffic patterns, which involve analysing the type, volume, and direction of the data flows within the network. Moreover, the retrieval of the network topology and the data traffic is used to identify the potential bottlenecks, congestion paths, and the like that are required for the path assignment. As a result, by retrieving the network topology and the data traffic patterns, the method 100 enables the processing unit of the controller to make informed decisions regarding the path assignment for ML workloads.

[0046] In accordance with an embodiment, the data traffic patterns are for the MLWL. Firstly, the processing unit of the controller is configured to distinguish the data traffic patterns that are associated with the ML workloads from other types of traffic within the network. Thereafter, the processing unit of the controller is configured to retrieve the information about the data traffic patterns for the MLWL, such as volume, frequency, and destinations of the ML data flows generated by the ML computations. In an implementation, the data traffic patterns are based on the executing ML process. As a result, the processing unit of the controller utilizes the data traffic patterns in order to take path assignment decisions, ensuring that the relevant paths are selected to optimize the performance of the MLWL thereby ensuring sufficient bandwidth and capacity to support the ML computations.

[0047] In an implementation, the method 100 further includes the processing unit of the controller retrieving the network topology from a local memory included in the controller. The local memory refers to a memory that is configured to store network topology data of the network. Examples of implementation of the local memory may include, but are not limited to, Electrically Erasable Programmable Read-Only Memory (EEPROM), Dynamic Random Access Memory (DRAM), Random Access Memory (RAM), Read-Only Memory (ROM), Hard Disk Drive (HDD), Flash memory, a Secure Digital (SD) card, Solid-State Drive (SSD), and / or CPU cache memory. In another implementation, the method 100 further includes the processing unit of the controller retrieving the data traffic patterns from a local memory comprised in the controller. As a result, the controller is configured to allow efficient access to the network topology data with reduced overall processing time without relying on any external source for the same. In addition, the method 100 includes updating the data traffic patterns based on the assigned path.

[0048] At step 112, the method 100 includes determining a path based on a path assignment algorithm based on the network topology, the data traffic patterns, and active ML data flows. In an implementation, the processing unit of the controller is configured to execute a path assignment algorithm, such as a greedy path algorithm is used to allocate a path to the ML data flow. However, other path assignment algorithms can also be used to determine the path, without affecting the scope of the present disclosure. By considering the network topology, traffic patterns, and active ML data flows, the algorithm can be used to assign paths having minimum latency, congestion, and the like in order to optimize data performance and efficiency for the MLWL, leading to an improved overall network performance.

[0049] In accordance with an embodiment, the method 100 further includes the processing unit of the controller for predicting future data ML flows based on previous ML data flows and the ML data flow of the received MLWL path assignment request. In an implementation, the processing unit of the controller is configured to utilize a prediction algorithm to predict the next upcoming flows and pre-compute the path assignment for the predicted flows. Moreover, such prediction is based on known properties of the data traffic pattern and known inter-dependencies between hosts. An implementation scenario of future data ML flow prediction is described in detail in FIG. 7. As a result, the prediction of the future data ML flows reduces the overall data flow assignment time (i.e., the time from the host to the controller) in order to improve the overall network performance with optimized path assignments and minimize the possibility of overall data latency and data congestion.

[0050] At step 114, the method 100 includes sending a path assignment response to the host indicating the assigned path to the host. In an implementation, the processing unit of the controller is configured to send the path assignment response to the processing unit of the host, which indicates the assigned path to the host to initiate the data transmission along the designated network route. In an implementation, the processing unit of the controller is configured to send the path assignment response to the host, such as through a communication protocol, for example, a Transmission Control Protocol / Intemet Protocol (TCP / IP), User Datagram Protocol (UDP), and the like. As a result, by sending the path assignment response to the host, the controller is configured to ensure effective coordination between the controller and the host for ML data flow transmission with minimized latency and maximum throughput. Additionally, the transmission of the path assignment response to the host indicating the assigned path to the host also reduces the overall delays and errors that are associated with manual path assignment thereby leading to an efficient, effective, and reliable data transmission.

[0051] Furthermore, at step 116, the method 100 includes the processing unit of the host for receiving the path assignment response and transmitting the ML data flow according to the assigned path, such as at step 118. In other words, the processing unit of the host is configured to send the path assignment request to the processing unit of the controller. Thereafter, the processing unit of the controller is configured to determine the path from the network that can be assigned to the ML data flow for the data transmission. After that, the processing unit of the controller is configured to send the path assignment response to the processing unit of the host. The processing unit the host transmits the ML data flow according to the assigned path based on the received path assignment response. Moreover, the path assignment response includes information about the assigned path for the ML data flow, such as network nodes, links, and the like. As a result, the method 100 facilitates seamless communication between the controller and the host, ensuring prompt and error-free delivery of the assigned path for ML data flows. Additionally, the method 100 is used to enhance the overall network performance and reliability of the network by facilitating the efficient execution of the MLWL. In accordance with an embodiment, the method 100 further includes the processing unit of the host including the size of the data flow in the MLWL path assignment request, and the processing unit of the controller for determining the path also based on the size of the data flow. The determination of the path based on the size of the data flow is to identify the volume of the data being transmitted between the hosts. In an implementation, upon receiving the MLWL path assignment request, the processing unit of the controller determines the size of the data flow. Moreover, the controller is configured to incorporate the data flow size as a parameter when determining the optimal path for the ML data flow in order to allow the selection of the network path that can accommodate the data volume efficiently with minimized data congestion. In another implementation, the method 100 further includes the processing unit of the controller to determine the path based on the ML process and hosts included in the ML process. As a result, the method 100 can optimize network resource allocation and path selection to accommodate varying data flow sizes effectively.

[0052] In accordance with an embodiment, the method 100 further includes the processing unit of the controller determining a maximum transmission rate including the maximum transmission rate in the path assignment response and the processing unit of the host transmits data flow at or below the maximum transmission rate. Firstly, the processing unit of the controller is configured to calculate the maximum transmission rate, such as by calculating network conditions, available bandwidth, and the like. Moreover, the maximum transmission rate represents the highest allowable transmission speed by which the data can be transferred between the hosts in order to prevent data congestion. Subsequently, the controller includes this calculated rate in the path assignment response sent to the host, instructing it on the maximum allowable transmission speed. Upon receipt of such a response, the processing unit of the host ensures that the data flow is transmitted at or below the specified maximum transmission rate, thus preventing network congestion and ensuring adherence to capacity constraints. In an implementation, the maximal transmission rate is a product of the path assignment algorithm. For example, after the path has been assigned for a new flow the maximal transmission rate can be determined by a max-min fairness allocation. Similarly, the controller is configured to use the determined data traffic patterns for the maximal rate computation. For example, an all-reduce operation is performed by the host for sending the data traffic to the network, in this case, the completion of the all-reduce depends on the slowest flow in the network. Thus, the controller is configured to assign the same maximal transmission rate to all the ML data flows in the all-reduce ring and allows more bandwidth for other ML data flows and ensure smooth data transmission within the network. Additionally, the method further comprises the processing unit of the controller determining the maximal transmission rate by a max-min fairness allocation. The max-min fairness is used to ensure an optimized resource utilization with equitable distribution of bandwidth and enhanced network performance along with reduced risk of data congestion.

[0053] In an implementation, the method 100 further includes the processing unit of the host transmitting the ML data flow according to the assigned path utilizing segment routing. Moreover, the segment routing refers to a network routing process that allows for the explicit designation of a series of segments (i.e., the network nodes) through which a data packet should traverse. The segment routing provides a flexible and efficient means to guide data packets along specific paths within the network, optimizing network performance and reliability. The method 100 is used to determine the optimal path for the ML data flow based on various network factors and then communicate the assigned path to the host through the path assignment response. Furthermore, upon receiving the assigned path, the processing unit of the host is configured to initiate the transmission of the ML data flow, with the data packets following the designated path thereby optimizing the overall network efficiency and the overall performance of the network for the data transmission. In another implementation, the method ' 100 further comprises the processing unit of the host thereby transmitting the ML data flow according to the assigned path utilizing specific fields in a packet header for data packets to be transmitted for the ML data flow, in order to affect load balancing decisions of switches and routers in the network. The transmission of the ML data flow from the assigned path by manipulating specific fields in a packet header enables the host to influence load-balancing decisions made by switches and routers in the network thereby ensuring that data packets follow the designated path accurately. The processing unit of the host is configured to identify the relevant fields within the packet header. Thereafter, the processing unit of the host is configured to modify the packet headers in order to reflect the assigned path. Subsequently, the processing unit of the host initiates the transmission of the ML data flow, with the modified packet header guiding the routing decisions of network devices. In an example, switches or routers use a hash function computed over specific header fields to compute the selected path from the network. Furthermore, the controller is configured to determine the selected path by setting a specific field value in the header and sending this field value to the host, such as in RoCEv2 traffic, a common to use the UDP source port value as a means of changing the path. Therefore, in this case, the controller can include an explicit UDP source port value in the path assignment reply message, and thus the host uses the UDP source port value without the need to be aware of the network topology. Hence, the method 100 is used to provide precise path control with reduced overhead, enhanced scalability, and improved fault tolerance, making the data transmission well-suited for large-scale networks like data centres and enterprise environments.

[0054] Advantageously, the method 100 is used to provide an optimized path assignment for handling MLWL in the network having network topology. The processing unit of the host is configured to identify data flows associated with executing ML processes based on the size and data traffic patterns, distinguishing them from other network data traffic to allow prioritization and management of the MLWL, enabling customized network routing and prevent the possibility of the data congestion. Furthermore, upon determining that the data flow is part of an ML workload, the host is configured to send a specialized path assignment request to the controller, ensuring dedicated handling and resource allocation for the corresponding ML data flows. The processing unit of the controller receives the path assignment request, retrieves the network topology and data traffic patterns, and determines an optimized path in order to assign the path to having minimum latency. The controller then sends the assigned path to the host, facilitating seamless communication and coordination for efficient, reliable ML data flow transmission according to the designated path. Moreover, such a coordinated approach reduces the overall data transmission delays and errors that are associated with manual assignment, leading to the efficient execution of large-scale MLWL. As a result, the method 100 is used to ensure efficient resource allocation, minimize congestion, and improve performance and reliability for ML-intensive computations in the network.

[0055] The steps 102 to 118 are only illustrative, and other alternatives can also be provided where one or more steps are added, one or more steps are removed, or one or more steps are provided in a different sequence without departing from the scope of the claims herein.

[0056] There is further provided a computer program product comprising program instructions for performing the method 100 when executed by one or more processors in the ML network. The computer program product is implemented as an algorithm, embedded in a software stored in a non-transitory computer-readable storage medium. The non-transitory computer-readable storage means may include but are not limited to, an electronic storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the foregoing. Examples of implementation of computer-readable storage medium, but are not limited to, Electrically Erasable Programmable Read-Only Memory (EEPROM), Random Access Memory (RAM), Read Only Memory (ROM), Hard Disk Drive (HDD), Flash memory, a Secure Digital (SD) card, Solid-State Drive (SSD), a computer-readable storage medium, and / or CPU cache memory.

[0057] FIG. 2 is a block diagram that illustrates a network having network topology, in accordance with an embodiment of the present disclosure. With reference to FIG. 2, there is shown a block diagram of a network 200 having the network topology that includes a host 202, a communication channel 204, and a controller 206.

[0058] The host 202 may include suitable logic, circuitry, interfaces, and / or code that are configured to determine if the ML data flow is the part of the MLWL and further send the ML path assignment request to the controller 206 in order to transmit the ML data flow with reduced latency and enhanced reliability. Examples of the host 202 may include but are not limited to a computer, a personal digital assistant, a portable computing device, or an electronic device. The communication channel 204 includes a medium (e.g., a communication channel) through which the host 202 potentially communicates with the controller 206. Examples of the communication channel 204 may include, but are not limited to, a cellular network (e.g., a 2G, a 3G, long-term evolution (LTE) 4G, a 5G, or 5G New Radio (NR) network, such as sub 6 GEIz, cmWave, or mm-Wave communication network or a communication network in the future), a wireless sensor network (WSN), a cloud network, a Local Area Network (LAN), a vehicle-to-network (V2N) network, a Metropolitan Area Network (MAN), and / or the Internet.

[0059] The controller 206 may include suitable logic, circuitry, interfaces, and / or code that are configured to perform memory allocation. Examples of implementation of the controller 206 may include but are not limited to a central data processing device, a microprocessor, a microcontroller, a complex instruction set computing (CISC) processor, an application-specific integrated circuit (ASIC) processor, a reduced instruction set (RISC) processor, a very long instruction word (VLIW) processor, a state machine, and other processors or control circuitry.

[0060] In operation, a processing unit (i.e. , a first processing unit 208) of the host 202 is configured to determine that the data flow for the executing ML process is to be transmitted to the receiver, the data flow having a size. Furthermore, the processing unit of the host 202 is configured to determine that the data flow is part of the MLWL, thereby being a Machine Learning (ML) data flow, and in response thereto send the MLWL path assignment request for the data flow to the controller 206. Moreover, the determination of the data flow is part of the MLWL that can be further utilized to handle resources that are associated with the data transmission for efficient, reliable, and optimized data transmission within the network. Thereafter, the processing unit (i.e., a second processing unit 210) of the controller 206 is configured to receive the MLWL path assignment request and retrieve the network topology and data traffic patterns. The processing unit of the controller 206 receives the path assignment request and initiates the process of determining an optimal path for the ML data flow based on the network topology, traffic patterns, and the like. As a result, by retrieving the network topology and the data traffic patterns, the processing unit of the controller 206 is configured to make informed decisions regarding the path assignment for MLWL. Furthermore, the processing unit (i.e., the second processing unit 210) of the controller 206 is configured to determine a path based on a path assignment algorithm based on the network topology, the data traffic patterns and active ML data flows and send the path assignment response to the host 202 indicating the assigned path to the host 202. By considering network topology, traffic patterns, and active ML data flows, the algorithm can be used to assign paths having minimum latency, congestion, and the like in order to optimize data performance and efficiency for MLWL, leading to an improved overall network performance. Furthermore, the processing unit (i.e., the first processing unit 208) of the host 202 is configured to receive the path assignment response and transmit the ML data flow according to the assigned path. In other words, the processing unit of the host 202 is configured to send the path assignment request to the processing unit of the controller 206. Thereafter, the processing unit of the controller 206 is configured to determine the path from the network that can be assigned to the ML data flow for the data transmission. After that, the processing unit of the controller 206 is configured to send the path assignment response to the processing unit of the host 202. The processing unit the host 202 transmits the ML data flow according to the assigned path based on the received path assignment response. Moreover, the path assignment response includes information about the assigned path for the ML data flow, such as network nodes, links, and the like. As a result, seamless communication can be established between the controller 206 and the host 202 in order to ensure prompt and error-free delivery of the assigned path for ML data flows.

[0061] Advantageously, the controller 206 and the host 202 of the network 200 having network topology provide an optimized path assignment for handling the MLWL in the network 200. The processing unit of the host 202 is configured to identify data flows associated with executing the ML processes based on the size and data traffic patterns, distinguishing them from other network data traffic to allow prioritization and management of the MLWL, enabling customized network routing and prevent the possibility of the data congestion. Furthermore, upon determining that the data flow is part of an ML workload, the host 202 is configured to send a specialized path assignment request to the controller 206, ensuring dedicated handling and resource allocation for the corresponding ML data flows. The processing unit of the controller 206 receives the path assignment request, retrieves the network topology and data traffic patterns, and determines an optimized path in order to assign the path to having minimum latency. The controller 206 then sends the assigned path to the host 202, facilitating seamless communication and coordination for efficient, reliable ML data flow transmission according to the designated path. Moreover, such a coordinated approach reduces the overall data transmission delays and errors that are associated with manual assignment, leading to the efficient execution of large-scale MLWL. As a result, the controller 206 is configured to ensure efficient resource allocation, minimize congestion, and improve performance and reliability for ML-intensive computations in the network 200.

[0062] FIG. 3 is a flowchart of a method for a processing unit of a controller of a network having a network topology, in accordance with another embodiment of the present disclosure. With reference to FIG. 3, there is shown a flowchart of method 300 for use in the processing unit of the controller 206. The method 300 includes steps 302 to 310.

[0063] In operation, the method 300 includes the processing unit of the controller 206 for receiving an MLWL path assignment request from the host 202 for a data flow for a Machine Learning process executing on the host 202, such as at step 302. At step 304, the method 300 includes retrieving the network topology and data traffic patterns. By retrieving the network topology and the data traffic patterns, the method 300 enables the processing unit of the controller 206 to make informed decisions regarding the path assignment for ML workloads. At step 306, the method 300 includes determining a path based on a path assignment algorithm based on the network topology, the data traffic patterns, and active ML data flows. In an implementation, the processing unit (i.e., the second processing unit 210) of the controller 206 is configured to execute a path assignment algorithm, such as a greedy path algorithm to allocate a path to the ML data flow. However, other path assignment algorithms can also be used to determine the path, without affecting the scope of the present disclosure. By considering network topology, traffic patterns, and active ML data flows, the algorithm can be used to assign paths having minimum latency, congestion, and the like in order to optimize data performance and efficiency for the MLWL, leading to an improved overall network performance. At step 308, the method 300 includes sending a path assignment response to the host 202 indicating the assigned path to the host 202, thereby causing the host 202 to transmit the ML data flow according to the assigned path, such as at step 310. In an implementation, by sending the path assignment response to the host 202, the controller 206 is configured to ensure effective coordination between the controller 206 and the host 202 for ML data flow transmission with minimized latency and maximum throughput. Additionally, the transmission of the path assignment response to the host 202 indicating the assigned path to the host 202 reduces the overall delays and errors that are associated with manual path assignment thereby leading to an efficient, effective, and reliable data transmission.

[0064] Advantageously, the method 300 is used to determine the optimal path for transmitting ML data flows within the network 200 in order to ensure that ML-related data traffic is routed efficiently, minimizing latency, and maximizing throughput. Furthermore, the utilization of a path assignment algorithm based on network topology and the data traffic patterns allows for the efficient allocation of network resources in order to optimize resource utilization, ensuring that sufficient bandwidth is allocated to support ML computations.

[0065] The steps 302 to 310 are only illustrative, and other alternatives can also be provided where one or more steps are added, one or more steps are removed, or one or more steps are provided in a different sequence without departing from the scope of the claims herein.

[0066] There is further provided a computer program product comprising program instructions for performing the method 300 when executed by one or more processors in the ML network. The computer program product is implemented as an algorithm, embedded in a software stored in a non-transitory computer-readable storage medium. The non-transitory computer-readable storage means may include but are not limited to, an electronic storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the foregoing. Examples of implementation of computer-readable storage medium, but are not limited to, Electrically Erasable Programmable Read-Only Memory (EEPROM), Random Access Memory (RAM), Read Only Memory (ROM), Hard Disk Drive (HDD), Flash memory, a Secure Digital (SD) card, Solid-State Drive (SSD), a computer-readable storage medium, and / or CPU cache memory.

[0067] FIG. 4 is a block diagram that illustrates a controller to be used in a network having network topology, in accordance with an embodiment of the present disclosure. With reference to FIG. 4, there is shown a block diagram 400 that illustrates the controller 206 to be used in the network 200 having network topology. The controller 206 includes a second local memory 402, a second network interface 404, and the first processing unit 208.

[0068] The second local memory 402 may include suitable logic, circuitry, and / or interfaces that are configured to store machine code and / or instructions executable by the controller 206. Examples of implementation of the second local memory 402 may include, but are not limited to, an Electrically Erasable Programmable Read-Only Memory (EEPROM), Random Access Memory (RAM), Read Only Memory (ROM), Hard Disk Drive (HDD), Flash memory, a Secure Digital (SD) card, Solid-State Drive (SSD), a computer-readable storage medium, and / or CPU cache memory.

[0069] The second network interface 404 may include hardware or software that is configured to establish communication between the second processing unit 210 and the second local memory 402. Examples of the second network interface 404 may include but are not limited to, a computer port, a network socket, a network interface controller (NIC), and any other network interface device.

[0070] In operation, the processing unit (i.e., the second processing unit 210) of the controller 206 is configured to receive an MEWL path assignment request from the host 202 for the data flow for a Machine Learning process executing on the host 202. Furthermore, the processing unit (i.e. , the second processing unit 210) of the controller 206 is configured to retrieve the network topology and the data traffic patterns. Thereafter, the processing unit (i.e., the second processing unit 210) of the controller 206 is configured to determine a path based on a path assignment algorithm based on the network topology, the data traffic patterns, and active ML data flows and send a path assignment response to the host 202 indicating the assigned path to the host, thereby causing the host 202 to transmit the ML data flow according to the assigned path. By considering network topology, traffic patterns, and active ML data flows, the processing unit (i.e., the second processing unit 210 of the controller 206) is configured to assign paths having minimum latency, congestion, and the like in order to optimize data performance and efficiency for MLWL, leading to an improved overall network performance.

[0071] FIG. 5 is a flowchart of a method for a processing unit of a host of a network having a network topology, in accordance with another embodiment of the present disclosure. With reference to FIG. 5, there is shown a flowchart of method 500 for use in the processing unit of the host 202. The method 500 includes steps 502 to 510.

[0072] In operation, the method 500 includes determining that a data flow for an executing Machine Learning (ML) process is to be transmitted to a receiver, the data flow having a size, for example, at step 502. By differentiating the ML processes based on the size of the ML data flow, the method 500 is used to facilitate an efficient handling and routing of the ML data flow within the network 200. The processing unit of the host 202 of the network 200 is configured to execute algorithms in order to analyse the ongoing processes and identify ML-related data flows. Moreover, such determination is used to prioritize the ML-related data flows and perform further required necessary actions, such as assigning the path to the ML data flow.

[0073] At step 504, the method 500 includes determining that the data flow is part of a Machine Learning Workload, thereby being a Machine Learning (ML) data flow, and in response thereto sending a Machine Learning Workload (MLWL) path assignment request for the data flow to the controller 206, such as at step 506. The determination of the data flow as the part of the MLWL that can be further utilized to handle resources that are associated with the data transmission for efficient, reliable, and optimized data transmission within the network 200. Moreover, the transmission of the MLWL path assignment request for the data flow to the controller 206 ensures dedicated handling for ML workloads by initiating specialized path assignment requests and optimizing resource allocation and routing for improved performance and reliability. Thereafter, at step 508, the method 500 includes receiving a path assignment response indicating the path determined by a path assignment algorithm based on the network topology, the data traffic patterns, and active ML data flows and at step 510, the method 500 includes transmitting the ML data flow according to the assigned path. As a result, the method 500 is used to facilitate seamless communication between the controller 206 and the host 202, ensuring prompt and error-free delivery of the assigned path for ML data flows with an improved overall network performance and reliability of the network 200 by facilitating the efficient execution of the MLWL.

[0074] Advantageously, the method 500 is used to optimize the ML data flow handling and routing within the network 200 by leveraging dynamic path assignment based on the network topology, data traffic patterns, and active MLWL. By proactively determining and assigning paths for ML data flows, the method 500 is used to ensure an efficient resource allocation with reduced network congestion, ultimately enhancing overall network performance and responsiveness to MLWL requirements. Additionally, the method 500 is used to streamline the communication between the host 202 and the controller 206, facilitating seamless coordination and transmission of ML data flows thereby improving network efficiency and reliability.

[0075] The steps 502 to 510 are only illustrative, and other alternatives can also be provided where one or more steps are added, one or more steps are removed, or one or more steps are provided in a different sequence without departing from the scope of the claims herein.

[0076] There is further provided a computer program product comprising program instructions for performing the method 500 when executed by one or more processors in the ML network. The computer program product is implemented as an algorithm, embedded in a software stored in a non-transitory computer-readable storage medium. The non-transitory computer-readable storage means may include but are not limited to, an electronic storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the foregoing. Examples of implementation of computer-readable storage medium, but are not limited to, Electrically Erasable Programmable Read-Only Memory (EEPROM), Random Access Memory (RAM), Read Only Memory (ROM), Hard Disk Drive (HDD), Flash memory, a Secure Digital (SD) card, Solid-State Drive (SSD), a computer-readable storage medium, and / or CPU cache memory.

[0077] FIG. 6 is a block diagram that illustrates a host to be used in a network having network topology, in accordance with an embodiment of the present disclosure. With reference to FIG. 6, there is shown a block diagram 600 of the network 200 having the network topology that includes the host 202 configured to be used in the network200 (of FIG.2) having a network topology. The host 202 includes a first local memory 602, a first network interface 604, and the first processing unit 208.

[0078] The first local memory 602 may include suitable logic, circuitry, and / or interfaces that are configured to store machine code and / or instructions executable by the host 202. Examples of implementation of the first local memory 602 may include, but are not limited to, an Electrically Erasable Programmable Read-Only Memory (EEPROM), Random Access Memory (RAM), Read Only Memory (ROM), Hard Disk Drive (HDD), Flash memory, a Secure Digital (SD) card, Solid-State Drive (SSD), a computer-readable storage medium, and / or CPU cache memory.

[0079] The first network interface 604 may include hardware or software that is configured to establish communication between the first processing unit 208 and the first local memory 602. Examples of the first network interface 604 may include but are not limited to, a computer port, a network socket, a network interface controller (NIC), and any other network interface device.

[0080] In operation, the processing unit (i.e., the first processing unit 208) of the host 202 is configured to determine that the data flow for an executing ML process is to be transmitted to a receiver, the data flow having a size. The processing unit of the host 202 of the network 200 is configured to execute algorithms in order to analyse the ongoing processes and identify ML-related data flows. Moreover, such determination is used to prioritize the ML-related data flows and perform further required necessary actions, such as assigning a path to the ML data flow. Furthermore, the processing unit of the host 202 is configured to determine that the data flow is part of the MLWL, thereby being the ML data flow, and in response thereto send the MLWL, path assignment request for the data flow to the controller 206. The determination of the data flow as the part of the MLWL that can be further utilized to handle resources that are associated with the data transmission for efficient, reliable, and optimized data transmission within the network 200. Moreover, the transmission of the MLWL path assignment request for the data flow to the controller 206 ensures dedicated handling for ML workloads by initiating specialized path assignment requests and optimizing resource allocation and routing for improved performance and reliability. Furthermore, the processing unit of the host 202 is configured to receive a path assignment response indicating the path determined by a path assignment algorithm based on the network topology, the data traffic patterns, and active ML data flows and transmit the ML data flow according to the assigned path. As a result, the host 202 is configured to facilitate seamless communication between the controller 206 and the host 202, ensuring prompt and error-free delivery of the assigned path for ML data flows with an improved overall network performance and reliability of the network 200 by facilitating the efficient execution of the MLWL.

[0081] Advantageously, the host 202 is configured to optimize the ML data flow handling and routing within the network 200 by leveraging dynamic path assignment based on the network topology, data traffic patterns, and active MLWL. By proactively determining and assigning paths for ML data flows, the host 202 is configured to ensure an efficient resource allocation with reduced network congestion, ultimately enhancing overall network performance and responsiveness to MLWL requirements. Additionally, the host 202 is configured to streamline the communication between the host 202 and the controller 206, facilitating seamless coordination and transmission of ML data flows thereby improving network efficiency and reliability.

[0082] FIG. 7 is a diagram that illustrates a path assignment for a network having network topology, in accordance with an embodiment of the present disclosure. FIG. 7 is described in conjunction with elements from FIG. 1 to 6. With reference to FIG. 7, there is shown a block diagram 700 of the path assignment for the network having network topology.

[0083] In an exemplary scenario, a centralized controller 702 is configured to assign a path to each machine learning (ML) data flow (e.g., a new elephant data flow). When a host is required to start transmitting the ML data flow, then, in that case, the host sends an MLWL path assignment request to the centralized controller 702. In an implementation, the host (i.e., a first host 706A and a second host 706B) is configured to decide if the ML data flow is an elephant based on a predetermined threshold. In another implementation, the host (i.e., the first host 706A and the second host 706B) is configured to determine if remote direct memory access (RDMA) is used or not, such as by determining the RDMA message length (i.e., RDMA buffer occupancy). For example, at operation 708, the first host 706A is configured to provide RDMA buffer occupancy information to the centralized controller 702 to a data center cluster 704. Thereafter, the RDMA buffer occupancy information is further transmitted to the centralized controller 702, such as at operation 710. Similarly, the second host 706B is configured to provide the RDMA buffer occupancy information (i.e., at operation 712 and at operation 710) to the centralized controller 702 through the data center cluster 704. Thereafter, based on the ML assignment path assignment request, the centralized controller 702 is configured to assign a path for the ML data flow transmission. In an implementation, the centralized controller 702 utilizes a path assignment algorithm to allocate a path to the new ML data flow. In another implementation, the centralized controller 702 utilizes a prediction algorithm to predict the next upcoming ML data flows and pre-computes the path assignment for the predicted flows. Furthermore, the centralized controller 702 is configured to send the path assignment decision to the host, including the maximal transmission rate of the flow, such as at operation 714. In an example, at operation 716, the centralized controller 702 is configured to provide a routing plan to the first host 706A through the data center cluster 704. Similarly, at operation 718, the centralized controller 702 is configured to provide the routing plan to the second host 706B through the data center cluster 704. As a result, by utilizing such algorithms in order to assign a path for the ML data flow, the centralized controller 702 is configured to reduce the ML data flow assignment procedure time in order to improve the overall network performance. FIG. 8 is a diagram that illustrates a prediction of MLWL assignment request for an ML data flow, in accordance with an embodiment of the present disclosure. FIG. 8 is described in conjunction with elements from FIGs. 1 to 7. With reference to FIG. 8, there is shown a diagram of the prediction of MLWL assignment request for the ML data flow.

[0084] In an exemplary scenario, the host (e.g., the first host 706A or the second host 706B) is configured to report the size of the ML data flow to the controller when the ML data flow starts in order to identify the data traffic patterns. Furthermore, based on the number of hosts available in the network, ML computations are deployed, such as by defining sets of hosts that take part in the computation (e.g., a ring of hosts that use AlLReduce messages). Based on the ML data flows, the controller is configured to predict the next upcoming ML data flow and pre-computes the path assignment for the corresponding data flow. For example, when the controller receives an MLWL path assignment request from a first process 802, then, in that case, the controller is configured to predict that new ML data flows can be sent by processes, such as a second process 804, a third process 806, and a fourth process 808 in the near future. As a result, the prediction of the future data ML flows reduces the overall data flow assignment time (i.e., the time from the host to the controller) in order to improve the overall network performance with optimized path assignments and minimize the possibility of overall data latency and data congestion.

[0085] Modifications to embodiments of the present disclosure described in the foregoing are possible without departing from the scope of the present disclosure as defined by the accompanying claims. Expressions such as "including", "comprising", "incorporating", "have", "is" used to describe and claim the present disclosure are intended to be construed in a non-exclusive manner, namely allowing for items, components or elements not explicitly described also to be present. Reference to the singular is also to be construed to relate to the plural. The word "exemplary" is used herein to mean "serving as an example, instance or illustration". Any embodiment described as “exemplary” is not necessarily to be construed as preferred or advantageous over other embodiments or to exclude the incorporation of features from other embodiments. The word "optionally" is used herein to mean "is provided in some embodiments and not provided in other embodiments". It is appreciated that certain features of the present disclosure, which are, for clarity, described in the context of separate embodiments, may also be provided in combination in a single embodiment. Conversely, various features of the disclosure, which are, for brevity, described in the context of a single embodiment, may also be provided separately or in any suitable combination or as suitable in any other described embodiment of the disclosure.

Claims

CLAIMS1. A method (100) for a network (200) having a network topology, the network (200) comprising a host (202) and a controller (206), wherein the method (100) comprises: a processing unit of the host (202) determining that a data flow for an executing Machine Learning process is to be transmitted to a receiver, the data flow having a size, determining that the data flow is part of a Machine Learning Workload, thereby being a Machine Learning, ML, data flow, and in response thereto sending a Machine Learning Workload, MLWL, path assignment request for the data flow to the controller (206), wherein the method (100) further comprises a processing unit of the controller (206) thereby receiving the MLWL path assignment request retrieving the network topology and data traffic patterns, determining a path based on a path assignment algorithm based on the network topology, the data traffic patterns and active ML data flows, sending a path assignment response to the host (202) indicating the assigned path to the host (202), wherein the method (100) further comprises the processing unit of the host (202) thereby receiving the path assignment response and transmitting the ML data flow according to the assigned path.

2. The method (100) according to claim 1, wherein the method (100) further comprises the processing unit of the host (202) includes the size of the data flow in the MLWL path assignment request and the processing unit of the controller (206) determining the path also based on the size of the data flow.

3. The method (100) according to any preceding claim, wherein the data traffic patterns are for Machine Learning Workloads.

4. The method (100) according to any preceding claim, wherein the method (100) further comprises the processing unit of the controller (206) predicting future data ML flows based on previous ML data flows and the ML data flow of the received MLWL path assignment request.

5. The method (100) according to any preceding claim, wherein the method (100) further comprises the processing unit of the controller (206) retrieving the network topology from a local memory comprised in the controller (206).

6. The method (100) according to any preceding claim, wherein the method (100) further comprises the processing unit of the controller (206) retrieving the data traffic patterns from a local memory comprised in the controller (206).

7. The method (100) according to any preceding claim, wherein the method (100) further comprises the processing unit of the controller (206) determining a maximum transmission rate and including the maximum transmission rate in the path assignment response wherein the processing unit of the host (202) transmits data flow at or below the maximum transmission rate.

8. The method (100) according to any preceding claim, wherein the method (100) further comprises the processing unit of the host (202) transmitting the ML data flow according to the assigned path utilizing segment routing.

9. The method (100) according to any preceding claim, wherein the method (100) further comprises the processing unit of the host (202) thereby transmitting the ML data flow according to the assigned path utilizing specific fields in a packetheader for data packets to be transmitted for the ML data flow, in order to affect load balancing decisions of switches and routers in the network (202).

10. The method (100) according to any preceding claim, wherein the method (100) further comprises the processing unit of the host (202) determining that the data flow is part of a Machine Learning Workload based on determining that the data flow is for the executing ML process.

11. The method (100) according to any preceding claim, wherein the method (100) further comprises the processing unit of the host (202) determining that the data flow is part of a Machine Learning Workload based on a size of the data flow, wherein the data flow is determined to be a ML data flow if the size of the data flow exceeds a threshold level.

12. A network (200) having a network topology, the network (200) comprising a host (202) and a controller (206), wherein a processing unit of the host (202) is configured to: determine that a data flow for an executing Machine Learning process is to be transmitted to a receiver, the data flow having a size, determine that the data flow is part of a Machine Learning Workload, thereby being a Machine Learning, ML, data flow, and in response thereto send a Machine Learning Workload, MLWL, path assignment request for the data flow to the controller (206), and wherein a processing unit of the controller (206) is configured to thereby: receive the MLWL path assignment request retrieve the network topology and data traffic patterns, determine a path based on a path assignment algorithm based on the network topology, the data traffic patterns and active ML data flows, send a path assignment response to the host indicating the assigned path to the host, wherein the processing unit of the host (202) is configured to thereby: receive the path assignment response and transmit the ML data flow according to the assigned path.

13. A method (300) for a network (200) having a network topology, the network (200) comprising a host (202) and a controller (206), wherein the method (300) comprises a processing unit of the controller (206): receiving a MLWL path assignment request from the host (202) for a data flow for a Machine Learning process executing on the host (202), retrieving the network topology and data traffic patterns, determining a path based on a path assignment algorithm based on the network topology, the data traffic patterns and active ML data flows, sending a path assignment response to the host indicating the assigned path to the host, thereby causing the host (202) transmitting the ML data flow according to the assigned path.

14. A controller (206) configured to be used in a network (200) having a network topology, the network (200) comprising a host (202) and the controller (206), wherein a processing unit of the controller (206) is configured to: receive a MLWL path assignment request from the host (202) for a data flow for a Machine Learning process executing on the host (202), retrieve the network topology and data traffic patterns, determine a path based on a path assignment algorithm based on the network topology, the data traffic patterns and active ML data flows,send a path assignment response to the host (202) indicating the assigned path to the host (202), thereby causing the host (202) transmit the ML data flow according to the assigned path.

15. A method (500) for a network (200) having a network topology, the network (200) comprising a host and a controller (206), wherein a processing unit of the host (202) is configured to: determining that a data flow for an executing Machine Learning process is to be transmitted to a receiver, the data flow having a size, determining that the data flow is part of a Machine Learning Workload, thereby being a Machine Learning, ML, data flow, and in response thereto sending a Machine Learning Workload, MLWL, path assignment request for the data flow to the controller (206), and then receiving a path assignment response indicating path determined by a path assignment algorithm based on the network topology, the data traffic patterns and active ML data flows and transmitting the ML data flow according to the assigned path.

16. A host (202) configured to be used in a network (200) having a network topology, the network (200) comprising the host (202) and a controller (206), wherein a processing unit of the host (202) is configured to: determine that a data flow for an executing Machine Learning process is to be transmitted to a receiver, the data flow having a size, determine that the data flow is part of a Machine Learning Workload, thereby being a Machine Learning, ML, data flow, and in response thereto send a Machine Learning Workload, MLWL, path assignment request for the data flow to the controller (206), and then receive a path assignment response indicating path determined by a path assignment algorithm based on the network topology, the data traffic patterns and active ML data flows and transmit the ML data flow according to the assigned path.

17. A computer program product comprising program instructions for performing the method (100, 300, 500) according to any of claims 1 to 11, 13, or 15, when executed by one or more processors in a ML network.

Citation Information

Patent Citations

  • Accelerating multi-node performance of machine learning workloads

    US20210092069A1