Multi-path architecture for hardware offload
By allocating multiple source ports in a computer network system and detecting path characteristics, and establishing multiple paths to segment and transmit data, the problem of difficult multipath architectures in the prior art is solved, and efficient network bandwidth utilization and reliability are achieved.
Patent Information
- Application Number
- CN202411451082.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-12-21
- Filing Date
- 2024-10-17
- Publication Date
- 2025-06-24
AI Technical Summary
In existing computer network systems, multipath architectures are difficult to achieve effective end-side unawareness in data centers and cloud computing environments, resulting in difficulty in network congestion control and load balancing, and data transmission may be disordered, resulting in long-tail delay and waste of resources.
Multiple paths are established by allocating and changing multiple source ports in the connection and fixing the source Internet Protocol (IP) address, destination IP address, and/or destination port. The congestion level and bandwidth of each path are detected, the data is divided into multiple blocks, and the size of each block is determined according to the communication characteristics of the path for transmission through multiple paths.
It realizes increasing network bandwidth, reducing long-tail latency and jitter, enhancing network reliability, mitigating the impact of congestion and link failures, and avoiding the disordered delivery problems caused by multipaths.
Smart Images

Figure CN120200960A_ABST
Abstract
Description
Technical Field
[0001] Embodiments described herein generally relate to multipath architectures in computer network systems. More specifically, embodiments described herein relate to methods and systems for multipath architectures to achieve end-side unawareness of multipaths in computer network systems for (multiple) data centers and / or clouds. Background Art
[0002] Computer network performance requirements pose continuous challenges to computer network architectures, especially in computer applications including (multiple) data centers and / or clouds. Existing solutions such as multipaths for each network connection have been proposed to improve computer network performance. However, the improvement of computer network performance using existing solutions is limited, especially in network hardware, due to the out-of-order delivery problem in multiple paths for the same network connection.
[0003] Equal-cost multipath routing (ECMP) or other techniques are used in some multipath solutions. In some existing solutions, multiple paths share the same network hardware resources, such as output queues, but the hardware resources may not have detailed information for each path, especially path quality and congestion level, which makes network congestion control and load balancing between different paths very difficult. Even with an adaptive routing switch that can select a preferred path and split data flows, network endpoints still know very little about the paths, which makes it difficult for network congestion control and load balancing to make decisions about the paths. In other existing solutions, network hardware can treat each path separately. However, data may arrive out of order from different paths, which makes it difficult for network hardware to reorder the data received from multiple paths. Summary of the Invention
[0004] It should be understood that some computer network hardware (e.g., network interface cards and / or their field programmable gate arrays, etc.) may have relatively small buffers for receiving data. When using traditional multipath mechanisms, due to, for example, load balancing using polling or other (multiple) protocols (to balance the data packet load of different paths), the transmission speed of each path may be different (e.g., some may be faster while others may be slower). When data packets (data packets of the data to be transmitted) arrive at the destination, the data packets may be out of order and need to be stored in the hardware buffer and reordered. Since the hardware buffer may be small, many data packets may be discarded (which causes retransmission and may slow down the transmission speed). Features in the embodiments disclosed herein can increase bandwidth and reduce or prevent long tail latency, for example, by providing different source ports to data packets.
[0005] Features in the embodiments disclosed herein can provide (multiple) protocols / (multiple) algorithms to establish multiple paths within a connection by, for example, fixing the source and destination Internet Protocol (IP) addresses and / or destination ports and changing the (multiple) source ports. In such an embodiment, the path between a client (e.g., a client application, a client computer, a client device, etc.) and a server (e.g., a server application, a server computer, a server device, etc.) can be defined by a quadruple including the source IP address, the destination IP address, the source port, and the destination port. Each time the value of the source port is changed (changed to a new value), a different or new path can be established / used / selected. It should be understood that a computer network multipath transmission protocol can adopt multiple network paths for a single connection or link. The multipath protocol can have the benefit of "aggregating" network bandwidth, which can achieve high throughput and high available time of the network.
[0006] It should also be understood that techniques such as Receive Side Scaling (RSS) on a network interface card (NIC) can be used to set the number of data queues (e.g., receive queues, etc.) and set the number of cores (e.g., processors, controllers, etc.) that can handle data traffic. RSS can select the receive queue for a data packet, for example, based on a hash result calculated on a 5-tuple in the packet header (including the source IP address, the destination IP address, the source port, the destination port, and the protocol) (based on an Equal-Cost Multipath (ECMP) protocol / algorithm, etc.). In this way, the data traffic can be evenly dispersed across the queues (and cores). It should also be understood that each network switch can use a routing protocol to provide, for example, ECMP forwarding, where the data stream can be split across multiple paths based on the hash of the 5-tuple in each packet. When a packet is lost or times out, the original path for the packet may be considered severely congested or the link is broken, and the packet needs to be resent on a different path.
[0007] Features in the embodiments disclosed herein can increase throughput bandwidth, reduce long-tail latency and jitter, and enhance reliability, mitigating the impact of congestion and link failures. Features in the embodiments disclosed herein can also implement multipaths for each connection for data transmission without degrading the performance of other services (e.g., Transmission Control Protocol (TCP) services, etc.). Features in the embodiments disclosed herein can avoid, prevent, or reduce excessive resource consumption, such as hardware resources (e.g., the hardware resources of a NIC, such as the Field Programmable Gate Array (FPGA) of a NIC, etc.). Features in the embodiments disclosed herein can also solve the out-of-order delivery problem caused by multipaths. It should be understood that in (multiple) multipath protocols or (multiple) algorithms, too few paths may affect throughput, while too many paths may increase packet loss, resulting in an increase in long-tail latency. Features in the embodiments disclosed herein can provide a balanced / high-quality multipath protocol and / or algorithm with, for example, dynamic adjustment to establish multipaths.
[0008] In one exemplary embodiment, a method for multi-path data transmission in a computer network is provided. The method includes establishing multiple paths in a connection by allocating / changing multiple source ports and fixing the source Internet Protocol (IP) address, destination IP address, and / or destination port. The method further includes detecting at least one of a first congestion level and a first bandwidth of each of the multiple paths, and splitting first data into multiple blocks. The number of the multiple blocks of the first data corresponds to the number of the multiple paths. The size of each of the multiple blocks of the first data is determined based on at least one of the first congestion level and the first bandwidth of each of the multiple paths detected. The method further includes transmitting the multiple blocks of the first data via the multiple paths respectively.
[0009] In another exemplary embodiment, a computer network system for multi-path data transmission is provided. The system includes a processor and a memory storing first data. The processor is configured to establish multiple paths in a connection by allocating / changing multiple source ports and fixing the source Internet Protocol (IP) address, destination IP address, and / or destination port. The processor is further configured to detect at least one of a first congestion level and a first bandwidth of each of the multiple paths, and split the first data into multiple blocks. The number of the multiple blocks of the first data corresponds to the number of the multiple paths. The size of each of the multiple blocks of the first data is determined based on at least one of the first congestion level and the first bandwidth of each of the multiple paths detected. The processor is further configured to transmit the multiple blocks of the first data via the multiple paths respectively.
[0010] In yet another exemplary embodiment, a non-transitory computer-readable medium storing computer-executable instructions is provided. The instructions, when executed, cause one or more processors to perform operations including establishing multiple paths in a connection by allocating / changing multiple source ports and fixing the source Internet Protocol (IP) address, destination IP address, and / or destination port. The operations further include detecting at least one of a first congestion level and a first bandwidth of each of the multiple paths, and splitting first data into multiple blocks. The number of the multiple blocks of the first data corresponds to the number of the multiple paths. The size of each of the multiple blocks of the first data is determined based on at least one of the first congestion level and the first bandwidth of each of the multiple paths detected. The operations further include transmitting the multiple blocks of the first data via the multiple paths respectively. BRIEF DESCRIPTION OF THE DRAWINGS
[0011] The accompanying drawings illustrate various embodiments of the systems, methods, and various other aspects of the present disclosure. Any ordinary person skilled in the art will understand that the illustrated element boundaries in the drawings (e.g., boxes, groups of boxes, or other shapes) represent an example of a boundary. In some examples, one element can be designed as multiple elements, or multiple elements can be designed as one element. In some examples, an element shown as an internal component of one element can be implemented as an external component of another element, and vice versa. A non-limiting and non-exhaustive description is provided with reference to the following drawings. The components in the figures are not necessarily drawn to scale, and the emphasis is on illustrating the principles. In the following detailed description, the embodiments are described only by way of illustration, as various changes and modifications will become apparent to those skilled in the art from the following detailed description.
[0012] Figure 1 is a schematic diagram of an example computer network system for multipath data transmission arranged according to at least some of the embodiments described herein.
[0013] Figure 2 is a schematic diagram of an example multipath architecture in a computer network system arranged according to at least some of the embodiments described herein.
[0014] Figure 3A is a schematic diagram of an example data block partitioning for a multipath architecture arranged according to at least some of the embodiments described herein.
[0015] Figure 3B illustrates an example data block partitioning for a multipath architecture arranged according to at least some of the embodiments described herein.
[0016] Figure 4 is a schematic diagram of an example of transmitting data blocks using a multipath architecture arranged according to at least some of the embodiments described herein.
[0017] Figure 5 is a flowchart of an example processing flow of multipath data transmission in a computer network system shown according to at least some of the embodiments described herein.
[0018] Figure 6 is a schematic structural diagram of an example computer system suitable for implementing a device arranged according to at least some of the embodiments described herein. Detailed Description
[0019] In the following detailed description, specific embodiments of the present disclosure are described with reference to the accompanying drawings, which form a part of the description. In this specification and the accompanying drawings, unless the context otherwise requires, the same reference numerals denote elements that can perform the same, similar, or equivalent functions. Additionally, unless otherwise stated, the description of each successive drawing may refer to features from one or more of the previous drawings to provide a clearer context and a more substantial explanation of the current exemplary embodiment. However, the exemplary embodiments described in the detailed description, the drawings, and the claims are not intended to be limiting. Other embodiments may be utilized and other changes may be made without departing from the spirit or scope of the subject matter presented herein. It will be readily understood that the aspects of the present disclosure generally described and illustrated in the drawings can be arranged, substituted, combined, separated, and designed in a variety of different configurations, all of which are explicitly contemplated herein.
[0020] It should be understood that the disclosed embodiments are merely examples of the present disclosure, which can be embodied in various forms. Well-known functions or constructions are not described in detail to avoid obscuring the present disclosure with unnecessary details. Thus, the specific structural and functional details disclosed herein should not be construed as limiting, but merely as a basis for the claims and as a representative basis for teaching one of ordinary skill in the art to employ the present disclosure in virtually any appropriately detailed structure.
[0021] Additionally, the present disclosure may be described herein in terms of functional block components and various processing steps. It should be understood that such functional blocks can be implemented by any number of hardware and / or software components configured to perform the specified functions.
[0022] The scope of the present disclosure should be determined by the appended claims and their legal equivalents, rather than by the examples given herein. For example, the steps recited in any method claim can be performed in any order and are not limited to the order presented in the claim. Additionally, no element is essential to the practice of the present disclosure unless specifically described herein as "critical" or "necessary".
[0023] As used herein, "network" or "computer network" is a technical term and can refer to interconnected computing devices that can exchange data with each other and share resources. It should be understood that networked devices can use a system of rules (e.g., communication protocols, etc.) to transmit information via wired or wireless technologies.
[0024] As referenced herein, "packet", "data packet", "network packet", or "network data packet" in a computer network system are technical terms and can refer to a formatted data unit carried by a computer network system. It should be understood that a packet can include control information and user data. User data can refer to "payload". Control information can provide information for delivering the payload (e.g., source network address and destination network address, error detection code, sequencing information, etc.). It should also be understood that control information can typically be stored in a packet header and / or a packet trailer. It should also be understood that control information can include metadata. Metadata in a computer network can refer to descriptive information about user data or "data about data". Metadata can carry various characteristics related to the structure of the communication protocol and the packet.
[0025] As referenced herein, the "header" of a packet can refer to the initial part of the packet. It should be understood that the header can contain control information such as addressing, routing, protocol version, etc. The format of the header information can depend on the communication protocol used. For example, in the Open Systems Interconnection (OSI) model from the International Organization for Standardization (ISO) (the model used in the embodiments disclosed herein), communication in a computer network system can be divided into different layers: Physical Layer (L1), Data Link Layer (L2), Network Layer (L3), Transport Layer (L4), Session Layer (L5), Presentation Layer (L6), and Application Layer (L7). In another example, the Transmission Control Protocol / Internet Protocol (TCP / IP) model can include a Physical Layer (L1), Data Link Layer (L2), Internet Layer (L3), Transport Layer (L4), and Application Layer (L5). Each layer can have its own communication protocol. For example, an L2 header can contain fields such as the destination Media Access Control (MAC) address, source MAC address, etc.
[0026] As referenced herein, "connection" or "link" in a computer network system are technical terms and can refer to a communication facility between, for example, two nodes, applications, computers, devices, etc. For example, when a client application sends data to a server application over the Internet, a transport layer protocol (e.g., TCP, User Datagram Protocol (UDP), etc.) can be used to establish a connection between the client and the server and transfer data packets between the client and the server. The transport layer can provide communication between application processes running on different hosts (e.g., servers, clients, etc.) within the layered architecture of the protocol and other network components. That is, the transport layer can collect message segments from the application and transmit them into the network (e.g., Layer 3, etc.).
[0027] As used herein, "congestion control" in a computer network system is a technical term and can refer to a mechanism (e.g., software / hardware module, algorithm, engine, etc.) that controls the entry of data packets into the network, thereby achieving better utilization of the shared network infrastructure and avoiding congestion collapse. For example, TCP can use a congestion control algorithm that includes aspects such as additive increase / multiplicative decrease scheme, slow start, and / or congestion window (CWND) to achieve congestion avoidance.
[0028] As used herein, "gateway" or "network gateway" in a computer network system is a technical term and can refer to a device that connects different networks by converting communication from one communication protocol to another. The gateway can provide a connection between two networks, and the networks do not need to use the same network communication protocol. It should be understood that the gateway can serve as an entry point and an exit point of the network because all data must pass through or communicate with the gateway before being routed.
[0029] As used herein, "traffic", "traffic flow", or "network flow" in a computer network system can refer to data (e.g., a sequence of packets, etc.) from a source computer or device to a destination computer or device (which can be another host, multicast group, or broadcast domain). It should be understood that traffic can include all packets in a specific transmission connection and / or media stream, and can include a set of packets passing through an observation point in the network during a specific time interval.
[0030] As used herein, "load balancing" in a computer network system is a technical term and can refer to the process of distributing a set of tasks across a set of resources (e.g., computing units, etc.) with the aim of making their overall processing more efficient. Load balancing can optimize response time and avoid some computing nodes being unevenly overloaded while other computing nodes are idle. A "load balancer" can refer to a computer device or program configured to perform the load balancing process.
[0031] As used herein, "NIC" or "network interface card" in a computer network system is a technical term and can refer to a computer hardware component that connects a computer to a computer network. It should be understood that some NICs can provide an integrated field-programmable gate array (FPGA) for user-programmable processing of network traffic before it reaches the host computer, thereby allowing a significant reduction in latency in time-sensitive workloads.
[0032] As used herein, "equal-cost multi-path" or "ECMP" in a computer network system is a technical term and can refer to a network routing algorithm or protocol that allows traffic of the same session or flow (e.g., traffic with the same source and destination) to be transmitted across multiple equal-cost paths.
[0033] Figure 1 It is a schematic diagram of an example computer network system 100 for multipath data transmission arranged according to at least some of the embodiments described herein.
[0034] System 100 may include devices 105, 110, 115, 120, 130, 140, 150, network 160, gateway or router 170, and / or network switches 180 and 190. It should be understood that Figure 1 only an illustrative number of devices (105, 110, 115, 120, 130, 140, 150, 170, 180, and / or 190) and / or networks are shown. The embodiments described herein are not limited to the number and location / connection of the devices and / or networks described. That is, the number and location / connection of the devices and / or networks described herein are for illustrative purposes only and are not intended to be limiting.
[0035] According to at least some example embodiments, devices 105, 110, 115, 120, 130, 140, and / or 150 may be various electronic devices. The various electronic devices may include, but are not limited to, mobile devices (such as smart phones), tablet computers, e - book readers, laptop computers, desktop computers, servers, and / or any other suitable electronic devices.
[0036] According to at least some example embodiments, network 160 may be a medium for providing communication links between devices 105, 110, 115, 120, 130, 140, 150, 170, 180, and / or 190. Network 160 may be the Internet, a local area network (LAN), a wide area network (WAN), a local interconnect network (LIN), the cloud, etc. Network 160 may be implemented by various types of connections, such as wired communication links, wireless communication links, fiber optic cables, etc.
[0037] According to at least some example embodiments, one or more of the devices 105, 110, 115, 120, 130, 140, and / or 150 may be servers for providing various services to users using one or more of the other devices. The server may be implemented by a distributed server cluster including multiple servers, or may be implemented by a single server.
[0038] Users can interact with each other via network 160 using one or more of devices 105, 110, 115, 120, 130, 140, 150, 170, 180, and / or 190. Various applications or their localized interfaces (such as social media applications, online shopping services, machine learning services, data center services, high-performance computing services, artificial intelligence services, cloud services, etc.) can be installed on devices 105, 110, 115, 120, 130, 140, and / or 150.
[0039] It should be understood that software applications or services according to the embodiments described herein and / or according to the services provided by a service provider can be executed by devices 105, 110, 115, 120, 130, 140, 150, 170, 180, and / or 190. Accordingly, the apparatus for the software application and / or service can be arranged in devices 105, 110, 115, 120, 130, 140, 150, 170, 180, and / or 190.
[0040] It should also be understood that when the service is not executed remotely, system 100 may not include network 160 and may only include devices 105, 110, 115, 120, 130, 140, 150, 170, 180, and / or 190.
[0041] It should further be understood that each of devices 105, 110, 115, 120, 130, 140, 150, 170, 180, and / or 190 may include one or more processors, a memory, and a storage device storing one or more programs. Each of devices 105, 110, 115, 120, 130, 140, 150, 170, 180, and / or 190 may also include an Ethernet connector, a Wi-Fi receiver, etc. When executed by one or more processors, the one or more programs may cause the one or more processors to perform the methods described in any of the embodiments described herein. It should also be understood that, according to the embodiments described herein, a computer-readable non-volatile medium can be provided. The computer-readable medium stores a computer program. When executed by a processor, the computer program is used to perform the method(s) described in any of the embodiments described herein.
[0042] Figure 2 is a schematic diagram of an example multi-path architecture 200 in a computer network system arranged according to at least some of the embodiments described herein.
[0043] As Figure 2As shown, data can be transmitted between a source and a destination. The source and / or the destination can be an application, a computer, a device, a network node, a client, a server, etc. A connection or a link can be created, generated, or established between the source and the destination. The connection or the link can have multiple paths (Path 1, Path 2, Path 3, Path 4, etc.). On the connection or the link, there can be multiple network nodes (e.g., routers, switches, gateways, relays, computers, devices, etc.) 210A, 210B, 210C, 210D, 210E, 210F, 210G, etc. It should be understood that the number of multiple paths described herein, the routing, and the number of network nodes are only for descriptive purposes and are not intended to be limiting.
[0044] It should be understood that during connection / link creation, multiple paths (e.g., multiple paths of a connection or a link) can be detected and established (e.g., through the multipath algorithms or protocols described herein), where the multiple paths simultaneously transmit data for a single connection / link. If one or more paths encounter serious problems (severe congestion, failures, disconnections, etc.), then the data transmission on such (multiple) paths can be switched to other available paths. If the problematic (multiple) paths cannot be restored within a predetermined time period, then that path can be deleted, and new available paths can be detected.
[0045] It should also be understood that the multipath algorithms or protocols described herein can solve or overcome network-side problems (the network between the source and the destination) without imposing a burden on the end-side (e.g., the source and / or the destination). That is, the implementation of the multipath algorithms or protocols described herein can make the end-side unaware (or minimize the impact on) of network-side problems, where the end-side includes applications and underlying hardware such as NIC / FPGA.
[0046] In an example embodiment, the multipath algorithm or protocol implementation can be achieved by adjusting, changing, or selecting the source port. It should be understood that, for example, for the data to be transmitted, the path between the source and the destination can be defined by a quadruple (source IP address, destination IP address, source port, and destination port). The IP address pair (source IP address and destination IP address) can define multiple connections / links. Each connection / link can correspond to a queue pair (QP, e.g., a point-to-point connection), and each connection / link can have multiple paths. It should be understood that the (multiple) QPs can include a send queue and a receive queue. The send queue can send outbound messages that request operations such as remote direct memory access (RDMA) operations, remote procedure call (RPC) operations, etc. The receive queue can receive incoming messages or immediate data.
[0047] In an example embodiment, for a given pair of IP addresses (e.g., a given / fixed source IP address and a destination IP address), the destination port (or listening port) can be fixed and / or unique. For example, the value of the destination port can be fixed to / at 8888. In such a configuration, a predetermined number of source ports (i.e., the number of paths) can be reserved for each connection (e.g., having a fixed destination port for a given pair of IP addresses). For example, 8 source ports can be reserved for each connection. That is, each connection can have 8 paths, and each source port can represent one path (one of the 8 paths).
[0048] In an example embodiment, to ensure that the end side is not aware (of network side issues), a predetermined source port prefix can be used. For example, the value of the source port (or destination port) for a connection of an application can be a 16-bit binary value (e.g., 8892 in decimal or 0010_0010_1011_1100 in binary). A predetermined source port prefix (e.g., the 12-bit binary value 0010_0010_1011 in the 16-bit binary value of the source port) can be used to determine the number of source ports (and thus the number of paths) of a connection / link. In this case, the last 4 bits (XXXX) of the source port can represent a total of 16 different source ports (e.g., 0010_0010_1011_XXXX in binary), which belong to the same source port prefix (e.g., 0010_0010_1011 in binary). That is, the source port (e.g., 8892 in decimal or 0010_0010_1011_1100 in binary) and 15 other source ports belong to the same source port prefix (e.g., the 12-bit binary value 0010_0010_1011). It should be understood that in this case, the number of source ports (16, which can correspond to the same number of paths) is the upper limit / boundary of the same source port prefix. For example, four consecutive source ports belonging to the same prefix can be used to form four paths allocated for each connection. It should also be understood that by using source port prefixes of different sizes (e.g., 14-bit, 10-bit, etc.), a larger or smaller number of source ports (belonging to the same prefix) and the corresponding number of multiple paths can be determined.
[0049] Back to Figure 2 , for data to be transmitted on a connection / link between a source and a destination (e.g., a given pair of IP addresses with a fixed destination port), 4 different source ports (with the same source port prefix) can be used (e.g., for data chunks or data packets), which can establish 4 different paths, for example, to simultaneously transmit data chunks or data packets of the data to be transmitted.
[0050] It should be understood that by adjusting, changing, or selecting different source ports (with the same source port prefix) to establish different paths, only one rule or policy may be required in the hardware (e.g., NIC, FPGA, etc.) to handle all the different established paths (for the same source port prefix). The features in the embodiments disclosed herein can simplify the policies of the hardware because there is no need for port-specific policies. Instead, the policies of the link / connection may work well. It should also be understood that if a path is severely congested or the communication on that path is interrupted, the corresponding source port (for such a path) can be switched, for example, to an unused source port belonging to the same source port prefix to switch to a different path. It should also be understood that network policies are a set of rules for managing the behavior of network devices. If the multi-path algorithm / protocol is implemented in software, a single rule can be issued in the NIC to handle different paths (corresponding to different source ports with the same source port prefix) in the same way. If the multi-path algorithm / protocol is implemented in hardware (e.g., NIC or FPGA in the NIC, etc.), the queue pair (QP) in the FPGA can also use a similar prefix representation for the source port of a single QP, which can ensure that when the total number of QPs remains constant, the multi-path does not reduce the total number of connections.
[0051] In an example embodiment, regarding the application-side source port selection, to support different connections / links between (multiple) pairs of the same IP addresses, different source port prefixes can be used. Since the source ports within the same prefix represent the same connection / link, one connection / link can use 0010_0010_1011_XXXX (binary), while another connection / link uses 0010_0010_1100_XXXX (binary). It should be understood that different pairs of IP addresses can use the same source port prefix to solve the problem of excessive source port consumption.
[0052] It should also be understood that in the multi-path algorithm / protocol, each connection / link corresponds to a QP (queue pair), and the hardware (e.g., NIC, FPGA in the NIC, etc.) may not be aware of the existence of multiple paths. When the sending path becomes unavailable (e.g., due to communication timeout, severe congestion, etc.), that path can be deactivated. Then the data is switched to use another path by modifying the digits of the last predetermined number of the source port (e.g., if the source port prefix is 12 bits in binary, then the last predetermined number of digits can be the last 4 bits in binary because the value of the source port is 16 bits in binary), and round-trip time (RTT) probes can be sent periodically (or continuously) to detect the communication characteristics of the path. When the path is restored, the path can be reactivated by switching the source port back to the source port corresponding to that path.
[0053] It should also be understood that in a multi-path algorithm / protocol, multi-path transmission can be controlled and implemented in software, and the hardware is unaware of the multi-path. Each path in the software can have its own sliding window (e.g., CWND, etc.) independent of other paths. The sliding window is calculated and updated by the software, and the software load balancer can determine which path to use for data packet transmission by setting the last predetermined digit of the source port. The sum of the sliding window sizes of all paths can be sent to the hardware (e.g., the corresponding QP of the NIC or its FPGA), which can allow the hardware to reflect the total available bandwidth of all paths of the corresponding connection / link. The hardware can support the full number of connections / links without reduction due to multiple paths.
[0054] It should be understood that in a multi-path algorithm / protocol, the internal order of messages (e.g., data to be transmitted) can be preserved, but the order between messages may not be required. Therefore, the load balancer scheduling can be message-based to ensure that messages are transmitted through only one path. It should also be understood that in a multi-path algorithm / protocol, each path can have its own congestion control, which is independent of the congestion control of other paths. The congestion control can be software-driven and / or hardware-assisted congestion control.
[0055] In an example embodiment, a primary-backup path algorithm / protocol can be implemented. That is, a connection / link can have two effective paths, namely a primary path and a backup path. Only the primary path can transmit control messages and data packets, while the backup path can transmit messages such as RTT, keep-alive messages, etc. When the primary path experiences severe congestion or packet loss, the primary-backup path algorithm / protocol can check the latest RTT result of the backup path or immediately send an RTT message. If the RTT result indicates that the backup path is more preferable than the primary path (e.g., has less RTT and the difference between the RTTs of the paths is greater than a threshold, etc.), then a path switch (from the primary path to the backup path) can occur (by changing the source port). After a predetermined time interval, if the newly switched path meets the predetermined requirements, the backup path can be marked as the (new) primary path. If the original (primary) path has not recovered, a new backup path can be probed to replace the original (primary) path. If the original (primary) path has returned to normal, the original (primary) path can be marked as the (new) backup path.
[0056] Figure 3A is a schematic diagram of an example data block partitioning 300 for a multi-path architecture arranged according to at least some embodiments described herein.
[0057] In an example embodiment, in a multi-path algorithm / protocol, the data sender (e.g., Figure 2The source (etc.) can divide the data (or a part of the data) to be transmitted into multiple segments (e.g., consecutive segments such as data blocks), and each segment is transmitted through a different path. It should be understood that the size of the block can be determined based on the communication characteristics of each path (e.g., speed, bandwidth, congestion window, etc.) so as to complete the transmission through multiple paths (substantially) simultaneously. Each block can be converted into a separate work request to ensure the internal sorting within the block. Sorting between blocks may not be necessary. The receiving end can forward the (multiple) data blocks to, for example, the application layer application programming interface (API) based on the block address, leaving the sorting problem to be solved by the application layer. The load balancer can determine the block size and select paths based on information such as the number of available paths, bandwidth delay, etc.
[0058] As Figure 3A shown, after multiple paths (330A, 330B, 330B, 330B, etc.) have been established for the connection / link (e.g., as described in the description of Figure 2 ), the sender (e.g., the source of Figure 2 ) can monitor, check, probe, and / or poll the communication characteristics (e.g., speed, bandwidth, congestion window, etc.) of each path by using various congestion control mechanisms (such as sending RTT messages and checking RTT results, etc.). It should be understood that each path can have its own congestion control module (320A, 320B, 320C, 320D, etc.) independent of the congestion control modules of other (multiple) paths.
[0059] In an example embodiment, based on the (predetermined) number of source ports (with the same predetermined source port prefix) to be used, the number of available multiple paths can be determined. The sender can first optionally monitor, check, probe, and / or poll the communication characteristics (e.g., speed, bandwidth, congestion window, etc.) of each path to determine whether the available multiple paths can be used. It should be understood that when a path is severely congested or fails (e.g., the RTT result is greater than a predetermined threshold, etc.), the path is considered unavailable, and different paths can be checked and / or used. When a path is not severely congested or fails (e.g., the RTT result is equal to or less than a predetermined threshold, etc.), the path is considered available. It should also be understood that if the sender does not first optionally monitor, check, probe, and / or poll the communication characteristics (e.g., speed, bandwidth, congestion window, etc.) of each path, then all available paths can be considered equally available (i.e., available and having the same communication characteristics).
[0060] In an example embodiment, based on the number of available paths and the communication characteristics of each available path (e.g., speed, bandwidth, congestion window, etc.), the sender may divide or split the first portion of the data to be transmitted (e.g., 10%, 30%, 50%, 100%, etc.) into multiple segments (e.g., data blocks 310A, 310B, 310C, 310D, etc.). The size of the data blocks may be determined based on, for example, the communication characteristics of each available path so as to (substantially) simultaneously complete the transmission of the blocks (the first portion of the data to be transmitted) over multiple available paths.
[0061] It should be understood that the data to be transmitted may include one, two, or more portions, depending on the size of the data to be transmitted. When the data blocks (310A, 310B, 310C, 310D, etc.) of the first portion of the data to be transmitted are transmitted via the available paths, the sender may periodically or continuously monitor, check, probe, and / or poll the communication characteristics of each path (e.g., speed, bandwidth, congestion window, etc.), e.g., by using the congestion control modules (320A, 320B, 320C, 320D, etc.) of each path. When the first portion of the data to be transmitted is completed (e.g., when the first data block is completed) or before the first portion of the data to be transmitted is completed (e.g., when the first data block is completed), the sender may divide or split the second portion of the data to be transmitted into multiple segments (e.g., data blocks). The size of the data blocks may be determined based on, for example, the communication characteristics of each available path (acquired during the transmission of the first portion of the data to be transmitted) (and / or the completion time of the transmission of the data blocks of the previous data portion (if any)) so as to (substantially) simultaneously complete the transmission of the blocks (the second portion of the data to be transmitted) over multiple available paths. The same process as that for the second portion may be repeated until all portions of the data to be transmitted are completed.
[0062] In an example embodiment, block 340 of FIG. 3 may be a queue (e.g., a send queue), and each of 330A, 330B, 330B, 330B, etc. may be a queue for the corresponding available path. The congestion control module (320A, 320B, 320C, 320D, etc.) of each available path may have its own controller or processor or engine and may independently monitor, check, probe, poll, determine, and / or control the communication characteristics of each path (e.g., to determine congestion, etc.) and send the results back to the sender (e.g., in accordance with the sender's request).
[0063] Figure 3B An example data block division 301 for a multi-path architecture arranged in accordance with at least some embodiments described herein is shown.
[0064] As Figure 3BAs shown on the left side of, the data to be transmitted can be divided or split into multiple data blocks (360A, 360B, 360C, 360D, etc.), for example, based on the number of available or applicable paths (e.g., established according to the description of Figure 2 . In an example embodiment, the size of the data blocks (360A, 360B, 360C, 360D, etc.) can be determined based on, for example, the communication characteristics of each available or applicable path so as to complete the transmission of the blocks through multiple available paths (substantially) simultaneously. During the transmission of the data blocks (360A, 360B, 360C, 360D, etc.), the communication characteristics of the (multiple) paths can be changed, and due to, for example, the different bandwidths of the (multiple) paths (i.e., the width of each block) or congestion conditions, the completion time (i.e., the height of each block) may be different. In another example embodiment, if the communication characteristics of each available or applicable path are not available (e.g., due to non-execution of an optional check of the communication characteristics, etc.), the size of the data blocks (360A, 360B, 360C, 360D, etc.) can be predetermined (e.g., the same size for each block). During the transmission of the data blocks (360A, 360B, 360C, 360D, etc.), the communication characteristics of the (multiple) paths can be discovered / obtained / available, and due to, for example, the different bandwidths of the (multiple) paths (i.e., the width of each block) or congestion conditions, the completion time (i.e., the height of each block) may be different.
[0065] As Figure 3B shown on the right side of (an improved division or split of the data blocks compared to the left side of Figure 3B ), the first part of the data to be transmitted can be divided or split into multiple data blocks (370A, 370B, 370C, 370D, etc.), for example, based on the number of available or applicable paths (e.g., established according to the description of Figure 2 . In an example embodiment, the size of the data blocks (370A, 370B, 370C, 370D, etc.) can be determined based on, for example, the communication characteristics of each available or applicable path so as to complete the transmission of the blocks through multiple available paths (substantially) simultaneously. In another example embodiment, if the communication characteristics of each available or applicable path are not available (e.g., due to non-execution of an optional check of the communication characteristics, etc.), the size of the data blocks (370A, 370B, 370C, 370D, etc.) can be predetermined (e.g., the same size for each block).
[0066] During the transmission of the data blocks (370A, 370B, 370C, 370D, etc.), the communication characteristics of the (multiple) paths can be changed (or obtained if not available previously), and due to, for example, the different bandwidths of the (multiple) paths (i.e., the width of each block) or congestion conditions, the completion time (i.e., the height of each block) may be different.
[0067] A second portion of the data to be transmitted may be divided or split into a plurality of data chunks (380A, 380B, 380C, 380D, etc.), e.g., based on the number of available or applicable paths (e.g., established according to Figure 2 as described). In an example embodiment, the size of the data chunks (380A, 380B, 380C, 380D, etc.) may be determined based on, e.g., the communication characteristics of each available or applicable path (and / or the completion time of each chunk of the first portion of the data) such that the chunks are transmitted via multiple available paths (substantially) simultaneously. As shown on the right side of Figure 3B , the transmission completion times of the data chunks (380A, 380B, 380C, 380D, etc.) may be the same or substantially the same. It should be understood that those processes for the second portion of the data may be repeated until all portions of the data to be transmitted are completely transmitted.
[0068] It should be understood that during the transmission of the data chunks, it may not be allowed to modify the chunks (e.g., modify the size and / or address, etc. of the chunks). The data to be transmitted may be divided into one, two, or more portions, and each portion may be divided or split into a plurality of chunks corresponding to the number of available / applicable paths. When the data chunks of the first data portion are being transmitted or after the data chunks of the first data portion are transmitted, based on the communication characteristics of each available / applicable path (and / or the completion time of each chunk of the previous data portion (if any)), including the remaining untransmitted packets (if any) for each path, the size of the data chunks of the second / next data portion may be determined or re - determined. The features in the embodiments disclosed herein may help reduce the long - tail latency of the available / applicable paths such that all data chunks of each data portion (and / or all portions of the data) may be transmitted at the same or substantially the same time. It should also be understood that the data chunk splitting / division process may be done by software or hardware (e.g., NIC, FPGA of the NIC, etc.).
[0069] Figure 4 is a schematic diagram of an example 400 of transmitting data chunks using a multi - path architecture arranged according to at least some embodiments described herein.
[0070] It should be understood that in some applications, data may be read, written, or otherwise accessed between a sender (e.g., the source of Figure 2 ) and a receiver (e.g., the destination of Figure 2 ) by, e.g., defining the memory address of the data on the sender side and the memory address of the data on the receiver side. Such applications include, for example, remote procedure call (RPC), remote direct memory access (RDMA), etc.
[0071] It should be understood that RPC can refer to a software communication protocol by which a program can request services from a program in another computer located on a network without having to understand the details of the network. RPC can be used to call other processes on a remote system as if they were on a local system. It should also be understood that RDMA can refer to direct memory access from the memory of one computer to the memory of another computer without involving the operating system of either computer.
[0072] In an example embodiment, data at a starting memory address (SA0) from a sender can be written (or read or otherwise accessed between the sender and the receiver) to a starting memory address (RA0) of the receiver. The data can be divided into one or more parts (e.g., data parts with consecutive memory addresses where the starting memory address of the next data part immediately follows the ending memory address of the previous data part). Each part can be divided or segmented into a plurality of data blocks corresponding to multiple available / applicable paths (e.g., based on the number of available source ports having the same source port prefix). The size of the data blocks (410A, 410B, 410C, 410D, etc.) can be determined based on, for example, the communication characteristics of each available or applicable path (and / or the completion time of each block of the previous data part (if any)) so as to (substantially) simultaneously complete the transmission of the data blocks (410A, 410B, 410C, 410D, etc.) over the multiple available paths.
[0073] As Figure 4 shown, the starting memory address SA0 of the first data block 410A on the sender side can be assigned to be the same as the starting memory address SA0 of the data to be transmitted on the sender side. The starting memory address RA0 of the first data block 410A on the receiver side can be assigned to be the same as the starting memory address RA0 of the data to be transmitted on the receiver side.
[0074] The starting memory address SA1 of the second data block 410B on the sender side can be assigned to be the same as the starting memory address SA0 of the data to be transmitted on the sender side plus the size of the first data block 410A. The starting memory address RA1 of the second data block 410B on the receiver side can be assigned to be the same as the starting memory address RA0 of the data to be transmitted on the receiver side plus the size of the first data block 410A.
[0075] The starting memory address SA2 of the third data block 410C in the sender side can be assigned to be the same as the starting memory address SA0 of the data to be transmitted in the sender side plus the sizes of the first data block 410A and the second data block 410B. The starting memory address RA2 of the third data block 410C in the receiver side can be assigned to be the same as the starting memory address RA0 of the data to be transmitted in the receiver side plus the sizes of the first data block 410A and the second data block 410B.
[0076] The starting memory address SA3 of the fourth data block 410D in the sender side can be assigned to be the same as the starting memory address SA0 of the data to be transmitted in the sender side plus the sizes of the first data block 410A, the second data block 410B, and the third data block 410C. The starting memory address RA3 of the fourth data block 410D in the receiver side can be assigned to be the same as the starting memory address RA0 of the data to be transmitted in the receiver side plus the sizes of the first data block 410A, the second data block 410B, and the third data block 410C.
[0077] By repeating the same or a similar process, the starting address of each data block (410A, 410B, 410C, 410D, etc.) can be determined, as well as the determined size of each data block (and the determined starting address of the data portion (if any)).
[0078] It should be understood that the operations (OP0, OP1, OP2, OP3, etc.) can be read, write, or other data access operations, where the starting memory address in the sender side and the starting memory address in the receiver side can be given together with the size of the data block (e.g., as parameters, or transmitted between the sender and the receiver, etc.).
[0079] It should also be understood that for each path (and for all paths), the features in the embodiments disclosed herein can overcome (or eliminate) the out-of-order delivery problem. It should also be understood that after all data blocks (e.g., the data blocks of the last part of the data to be transmitted) are transmitted, the multipath algorithm or protocol can notify the upper layer (e.g., the application layer in a computer network system) of the completion of the data transmission, e.g., by sending a work completion message to the upper layer. That is, for the upper layer, there is only one memory transfer task, but the multipath algorithm or protocol described herein can divide a single task into multiple separate and independent data / block transfer tasks (e.g., via the established multiple paths). That is, the operation / transmission of each data block can be considered independent. It should be understood that the sender and / or receiver can communicate by following the same data transmission protocol (e.g., the multipath algorithm or protocol described herein) or by using the same data transmission protocol (e.g., the multipath algorithm or protocol described herein).
[0080] It should also be understood that each data block can be transmitted via the paths of the multipath, so that the data of the data block is ordered and can be written into the memory address on the receiver side. The hardware may not need to buffer the received data blocks because when the receiver receives the data packets of the data block, the data block can be written into the memory on the receiver side. When all data blocks (all data parts) are received or written on the receiver side, the hardware can notify the upper layer (e.g., the application layer) (e.g., send a message to it) indicating the completion of the reception. The multipath establishment / selection and data block segmentation can be transparent or invisible to the application layer.
[0081] Figure 5 is a flowchart of an example processing flow 500 for multipath data transmission in an illustrated computer network system according to at least some embodiments described herein.
[0082] It should be understood that unless otherwise specified, the processing flow 500 disclosed herein can be performed by one or more processors, which include, for example Figure 1 the processors of one or more of the devices 105, 110, 115, 120, 130, 140, 150, 170, 180, and / or 190 of Figure 6 the CPU 605 of
[0083] It should also be understood that the processing flow 500 may include one or more operations, actions, or functions as shown by one or more of the blocks such as blocks 510, 520, 530, 540, and 550. These various operations, functions, or actions may correspond, for example, to software, program code, or program instructions executable by a processor that causes the functions to be executed. Although illustrated as discrete blocks, obvious modifications can be made, for example, two or more of the blocks can be reordered; more blocks can be added; and depending on the desired implementation, the various blocks can be divided into additional blocks, combined into fewer blocks, or eliminated. It should be understood that operations including initialization, etc. can be performed before the processing flow 500. For example, system parameters and / or application parameters can be initialized. It should be understood that Figure 2 , 3A , the processes, operations, or actions described in 3B and 4 can be implemented or executed by a processor. The processing flow 500 can start at block 510.
[0084] At block 510 (establishing multiple paths), the processor can establish multiple paths for a connection / link based on, for example, a fixed source Internet Protocol (IP) address and a destination IP address and / or destination port, and change the (multiple) source port or assign different (multiple) source ports to the data to be transmitted (e.g., blocks, packets, etc.). The processor can change, assign, or select (multiple) source ports having the same predetermined source port prefix for the connection / link. The processing can proceed from block 510 to block 520.
[0085] At block 520 (detecting congestion), the processor can (optionally) monitor, check, probe, and / or poll the communication characteristics (e.g., speed, bandwidth, congestion window, etc.) of each path by using various congestion control mechanisms (such as sending RTT messages and checking RTT results, etc.). It should be understood that the processor can determine the applicable path among the available paths based on the determined communication characteristics. For example, if a path fails (e.g., loses (multiple) packets within a predetermined time period), the processor can determine that the path may not be applicable and the processor can determine to switch to or select another path. It should also be understood that the processor can perform the steps of block 520 periodically or continuously during the transmission of data (see the description of block 540). The processing can proceed from block 520 to block 530.
[0086] At block 530 (split data), the processor may (optionally) divide the data to be transmitted into one or more parts, e.g., depending on the size of the data, the communication characteristics of the path (or connection / link), etc. The processor may divide or split each data part into multiple data chunks based on, e.g., the number of available paths determined in block 510 or the number of applicable paths determined in block 520. The processor may also determine the size of the data chunks based on, e.g., the communication characteristics of each available or available path (and / or the completion time of each chunk of the previous data part, if any) so as to (substantially) simultaneously complete the transmission of the data chunks over multiple available or used paths. The processor may also determine the starting memory address for each data chunk (for the sender side and / or for the receiver side). Processing may proceed from block 530 to block 540.
[0087] At block 540 (transmit chunks), the processor may perform the data transmission of each data chunk by, e.g., sending each data chunk to the transmit queue for each corresponding path by utilizing, e.g., the memory address information. It should be understood that the processor may periodically or continuously perform the steps of block 520 during the transmission of the data. If there are remaining (multiple) data parts, the processing may proceed from block 540 to block 520 (or to block 530 since the process of block 520 is performed periodically or continuously). If there are remaining data parts and all data parts have been transmitted, the processing may proceed from block 540 to block 550.
[0088] At block 550 (notify upper layer), the processor may notify the upper layer (e.g., the application layer in a computer network system) of the completion of the data transmission, e.g., by sending an operation completion message to the upper layer.
[0089] The features (e.g., multi-path algorithms or protocols) in the embodiments disclosed herein may be implemented in software. In such a scenario, the software may perform block splitting (e.g., of a size of 4KB, etc.), determine the memory address for each data chunk (e.g., in the sender side or in the receiver side), and distribute the data chunks to multiple queue pairs in the hardware (e.g., NIC or FPGA of the NIC) using, e.g., a polling load balancing mechanism or based on a load balancing configuration. In such a scenario, the hardware may not be aware of the multi-path and may treat each path as a connection / link.
[0090] Features in the embodiments disclosed herein (e.g., multi-path algorithms or protocols) can be implemented in hardware (e.g., an NIC or an FPGA of an NIC). In such a scenario, each connection can have a fixed number of paths (e.g., 4 paths, etc.) corresponding to, for example, the number of available queue pairs. Software can initialize the configuration of the queue pairs within a connection. Hardware can split data blocks based on, for example, the communication characteristics of the paths in the connection. Hardware can treat the first data block (e.g., except for the last data block) as a multiple of the maximum transmission unit (MTU) and determine the starting memory address of each block. It should be understood that the corresponding queue pair for each path can transmit its corresponding data block. Each queue pair can split the data block based on, for example, the MTU and then send data packets. Each path has its own congestion control engine, and the data transmission on each path can be similar to the data transmission of a single-path transmission.
[0091] Features in the embodiments disclosed herein can eliminate disordered processing between multiple paths to eliminate disorder problems caused, for example, by the limited buffer size of the hardware for reordering or combining disordered data. When a problem (e.g., a fault, severe congestion, etc.) occurs on a path, software can sense or detect the problem. By modifying the source port in the context of the queue pair, a new path can be established, configured, or generated. It should be understood that the operation completion (message) can be submitted to the upper layer only after all data blocks (of all data parts) have been received on the receiving side.
[0092] Features in the embodiments disclosed herein can configure whether the hardware needs to perform multi-path data distribution. If multi-path data distribution is not required, the hardware can revert to single-path and / or primary backup path transmission. The hardware may need to determine the size of the data block and needs to determine load balancing during data block splitting such that all data blocks can (basically) complete transmission through multiple available or available paths simultaneously. Features in the embodiments disclosed herein can dynamically adjust the splitting size (i.e., the size of the data block) of each path based on, for example, the detected communication characteristics of each path (and / or the completion time of the data blocks of the previous data part (if any)).
[0093] Figure 6 is a schematic structural diagram of an example computer system 600 applicable to implementing an electronic device (e.g., Figure 1 one of the servers, terminal devices shown in Figure 6 and / or (multiple) switches and / or (multiple) routers) according to at least some of the embodiments described herein. It should be understood that
[0094] As shown in the figure, the computer system 600 may include a central processing unit (CPU) 605. The CPU 605 may perform various operations and processes based on a program stored in a read-only memory (ROM) 610 or a program loaded from a storage device 640 into a random access memory (RAM) 615. The RAM 615 may also store various data and programs required for the operation of the system 600. The CPU 605, ROM 610, and RAM 615 may be connected to each other via a bus 620. An input / output (I / O) interface 625 may also be connected to the bus 620.
[0095] Components connected to the I / O interface 625 may further include: an input device 630, which includes a keyboard, a mouse, a digital pen, a graphics tablet, etc.; an output device 635, which includes a display such as a liquid crystal display (LCD), a speaker, etc.; a storage device 640, which includes a hard disk, etc.; and a communication device 645, which includes a network interface card such as a LAN card, a modem, etc. The communication device 645 may perform communication processing via a network such as the Internet, a WAN, a LAN, a LIN, the cloud, etc. In an embodiment, a drive 650 may also be connected to the I / O interface 625. A removable medium 655 such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc. may be mounted on the drive 650 as needed, so that a computer program read from the removable medium 655 may be installed in the storage device 640.
[0096] It should be understood that the processes described in the flowchart with reference to Figure 5 and / or the processes described in other figures may be implemented as a computer software program or implemented in hardware. A computer program product may include a computer program stored in a computer-readable non-volatile medium. The computer program includes program code for performing the methods shown in the flowchart and / or the GUI. In this embodiment, the computer program may be downloaded and installed from a network via the communication device 645, and / or may be installed from the removable medium 655. When the computer program is executed by the central processing unit (CPU) 605, the above functions specified in the methods in the embodiments disclosed herein may be implemented.
[0097] It should be understood that the disclosed content and other solutions, examples, embodiments, modules, and functional operations described in this document can be implemented in digital electronic circuits, or in computer software, firmware, or hardware, including the structures disclosed in this document and their structural equivalents, or a combination of one or more of them. The disclosed content and other embodiments can be implemented as one or more computer program products, that is, one or more computer program instruction modules encoded on a computer-readable medium for execution by, or to control the operation of, a data processing apparatus. The computer-readable medium can be a machine-readable storage device, a machine-readable storage substrate, a storage device, a substance composition affecting a machine-readable propagated signal, or a combination of one or more of them. The term "data processing apparatus" encompasses all apparatuses, devices, and machines for processing data, including, for example, programmable processors, computers, or multiple processors or computers. In addition to hardware, the apparatus may also include code that creates an execution environment for the computer programs under discussion, for example, code constituting processor firmware, a protocol stack, a database management system, an operating system, or a combination of one or more of them.
[0098] A computer program (also referred to as a program, software, software application, script, or code) can be written in any form of programming language, including compiled or interpreted languages, and can be deployed in any form, including as a stand-alone program or as a module, component, subroutine, or other unit suitable for use in a computing environment. A computer program does not necessarily correspond to a file in a file system. The program can be stored in a part of a file that holds other programs or data (e.g., one or more scripts in a markup language document), stored in a single file dedicated to the program under discussion, or stored in multiple coordinated files (e.g., files that store one or more modules, subroutines, or portions of code). A computer program can be deployed to be executed on one or more computers located at one site or across multiple sites and interconnected by a communication network.
[0099] The processes and logical flows described in this document can be performed by one or more programmable processors executing one or more computer programs to perform functions by operating on input data and generating output. The processes and logical flows can also be performed by dedicated logic circuits (such as field programmable gate arrays, application specific integrated circuits, etc.), and the apparatus can also be implemented as dedicated logic circuits (such as field programmable gate arrays, application specific integrated circuits, etc.).
[0100] Processors suitable for executing computer programs include, for example, general and special-purpose microprocessors, as well as any one or more processors of any type of digital computer. Generally, a processor will receive instructions and data from a read-only memory or a random access memory or both. The basic elements of a computer are a processor for executing instructions and one or more memory devices for storing instructions and data. Generally, a computer will also include one or more mass storage devices for storing data, such as magnetic disks, magneto-optical disks, or optical disks, or be operatively coupled to receive data therefrom or transfer data thereto, or both. However, a computer need not have such devices. Computer-readable media suitable for storing computer program instructions and data include all forms of non-volatile memory, media, and memory devices, including, for example, semiconductor memory devices such as erasable programmable read-only memory, electrically erasable programmable read-only memory, and flash memory devices; magnetic disks, such as internal hard disks or removable disks; magneto-optical disks; and compact disc read-only memory and digital video disc read-only memory disks. The processor and memory may be supplemented by, or incorporated in, special-purpose logic circuitry.
[0101] It should be understood that different features, variations, and numerous different embodiments have been shown and described in various details. What is sometimes described in this application in terms of specific embodiments is for illustrative purposes only and is not intended to limit or teach that the contemplated subject matter is only a particular embodiment or specific embodiment. It should be understood that the present disclosure is not limited to any single particular embodiment or recited variation. Many modifications, variations, and other embodiments will be apparent to those skilled in the art, and these modifications, variations, and other embodiments are intended to be encompassed by the present disclosure. Indeed, the scope of the present disclosure should be determined by appropriate legal interpretation and construction of the present disclosure, including equivalents that those skilled in the art would understand upon a complete disclosure at the time of filing.
[0102] Aspects:
[0103] It should be understood that any of the aspects may be combined with each other.
[0104] Aspect 1. A method for multi-path data transmission in a computer network, the method comprising: establishing multiple paths in a connection by allocating multiple source ports and fixing a source Internet Protocol (IP) address, a destination IP address, and / or a destination port; detecting at least one of a first congestion level and a first bandwidth of each of the multiple paths; splitting first data into multiple blocks, the number of the multiple blocks of the first data corresponding to the number of the multiple paths, and the size of each of the multiple blocks of the first data being determined based on at least one of the detected first congestion level and the first bandwidth of each of the multiple paths; and transmitting the multiple blocks of the first data via the multiple paths respectively.
[0105] Aspect 2. The method according to Aspect 1, wherein the numbers of each of the plurality of source ports for the connection have the same prefix and the remaining numbers are different from each other.
[0106] Aspect 3. The method according to Aspect 1 or Aspect 2, further comprising: determining the source addresses and destination addresses of the plurality of blocks of the first data based on the source address and destination address of the first data and the determined sizes of each of the plurality of blocks of the first data.
[0107] Aspect 4. The method according to any one of Aspects 1-3, further comprising: detecting at least one of a second congestion level and a second bandwidth of each of the plurality of paths; splitting second data into a plurality of blocks, the number of the plurality of blocks of the second data corresponding to the number of the plurality of paths, and the size of each of the plurality of blocks of the second data being determined based on at least one of the detected second congestion level and the second bandwidth of each of the plurality of paths; and transmitting the plurality of blocks of the second data via the plurality of paths respectively.
[0108] Aspect 5. The method according to Aspect 4, further comprising: a receiving end receiving the plurality of blocks of the second data at a destination port substantially simultaneously; and the receiving end notifying an application layer of the computer network when all of the plurality of blocks of the first data and the plurality of blocks of the second data are received from a sending end.
[0109] Aspect 6. The method according to any one of Aspects 1-5, wherein establishing the plurality of paths and splitting the first data into a plurality of blocks are performed by running an algorithm of the computer network.
[0110] Aspect 7. The method according to any one of Aspects 1-6, wherein establishing the plurality of paths and splitting the first data into a plurality of blocks are performed via a network interface card or a field programmable gate array of the computer network.
[0111] Aspect 8. The method according to any one of Aspects 1-7, further comprising: when a failure occurs in a first path among the plurality of paths, switching from the first path to a second path among the plurality of paths, wherein the first path corresponds to a first source port among the plurality of source ports, and the second path corresponds to a second source port among the plurality of source ports.
[0112] Aspect 9. A computer network system for multi-path data transmission, the system comprising: a memory for storing first data; a processor for: establishing multiple paths in a connection by allocating multiple source ports and fixing a source Internet Protocol (IP) address, a destination IP address, and / or a destination port; detecting at least one of a first congestion level and a first bandwidth of each of the multiple paths; splitting the first data into multiple blocks, the number of the multiple blocks of the first data corresponding to the number of the multiple paths, and the size of each of the multiple blocks of the first data being determined based on at least one of the detected first congestion level and the first bandwidth of each of the multiple paths; and transmitting the multiple blocks of the first data via the multiple paths respectively.
[0113] Aspect 10. The system according to aspect 9, wherein the numbers of each of the multiple source ports for the connection have the same prefix and the remaining numbers are different from each other.
[0114] Aspect 11. The system according to aspect 9 or aspect 10, wherein the processor is further configured to: determine the source addresses and destination addresses of the multiple blocks of the first data based on the source address and destination address of the first data and the determined size of each of the multiple blocks of the first data.
[0115] Aspect 12. The system according to any one of aspects 9-11, wherein the processor is further configured to: detect at least one of a second congestion level and a second bandwidth of each of the multiple paths; split second data into multiple blocks, the number of the multiple blocks of the second data corresponding to the number of the multiple paths, and the size of each of the multiple blocks of the second data being determined based on at least one of the detected second congestion level and the second bandwidth of each of the multiple paths; and transmit the multiple blocks of the second data via the multiple paths respectively.
[0116] Aspect 13. The system according to aspect 12, wherein the processor is further configured to: receive the multiple blocks of the second data at the destination port substantially simultaneously; and notify the application layer of the computer network when all of the multiple blocks of the first data and the multiple blocks of the second data are received.
[0117] Aspect 14. The system according to any one of aspects 9-13, wherein the processor is further configured to: switch from a first path to a second path among the multiple paths when a failure occurs in the first path among the multiple paths, wherein the first path corresponds to a first source port among the multiple source ports, and the second path corresponds to a second source port among the multiple source ports.
[0118] Aspect 15. A non-transitory computer-readable medium storing computer-executable instructions that, when executed, cause one or more processors to perform operations including: establishing multiple paths in a connection by allocating multiple source ports and fixing a source Internet Protocol (IP) address, a destination IP address, and / or a destination port; detecting at least one of a first congestion level and a first bandwidth for each of the multiple paths; splitting first data into multiple blocks, the number of the multiple blocks of the first data corresponding to the number of the multiple paths, and the size of each of the multiple blocks of the first data being determined based on at least one of the detected first congestion level and the first bandwidth for each of the multiple paths; and transmitting the multiple blocks of the first data via the multiple paths respectively.
[0119] Aspect 16. The computer-readable medium according to aspect 15, wherein the numbers of each of the multiple source ports for the connection have the same prefix and the remaining numbers are different from each other.
[0120] Aspect 17. The computer-readable medium according to aspect 15 or aspect 16, the operations further including: determining source addresses and destination addresses of the multiple blocks of the first data based on a source address and a destination address of the first data and the determined sizes of each of the multiple blocks of the first data.
[0121] Aspect 18. The computer-readable medium according to any one of aspects 15 - 17, the operations further including: detecting at least one of a second congestion level and a second bandwidth for each of the multiple paths; splitting second data into multiple blocks, the number of the multiple blocks of the second data corresponding to the number of the multiple paths, and the size of each of the multiple blocks of the second data being determined based on at least one of the detected second congestion level and the second bandwidth for each of the multiple paths; and transmitting the multiple blocks of the second data via the multiple paths respectively.
[0122] Aspect 19. The computer-readable medium according to aspect 18, the operations further including: receiving the multiple blocks of the second data at a destination port substantially simultaneously, and notifying an application layer of the computer network when all of the multiple blocks of the first data and the multiple blocks of the second data are received.
[0123] Aspect 20. The computer-readable medium according to any one of aspects 15-19, wherein the operation further comprises: when a failure occurs in a first path among the plurality of paths, switching from the first path to a second path among the plurality of paths, wherein the first path corresponds to a first source port among the plurality of source ports, and the second path corresponds to a second source port among the plurality of source ports.
[0124] The terminology used in this specification is intended to describe particular embodiments and is not intended to be limiting. Unless otherwise expressly stated, the terms "a," "an," and "the" include plural forms. When used in this specification, the terms "comprising" and / or "including" specify the presence of the stated features, integers, steps, operations, elements, and / or components, but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, and / or components.
[0125] In regard to the foregoing description, it should be understood that changes may be made in detail, particularly in the materials of construction used and in the shape, size, and arrangement of the parts, without departing from the scope of the disclosure. This specification and the described embodiments are merely exemplary, and the true scope and spirit of the disclosure are indicated by the appended claims.
Claims
1. A method for multipath data transmission in a computer network, the method comprising: Establishing multiple paths in a connection by assigning multiple source ports and fixing the source Internet Protocol IP address, destination IP address, and destination port; detecting at least one of a first congestion level and a first bandwidth of each of the plurality of paths; splitting the first data into a plurality of blocks, the number of the plurality of blocks of the first data corresponding to the number of the plurality of paths, the size of each of the plurality of blocks of the first data being determined based on at least one of the first congestion level and the first bandwidth detected for each of the plurality of paths; as well as The plurality of blocks of the first data are transmitted respectively via the plurality of paths. 2 . The method according to claim 1 , wherein the numbers of each of the plurality of source ports for the connection have a same prefix and the remaining numbers are different from each other.
3. The method according to claim 1, further comprising: Source addresses and destination addresses of the plurality of blocks of the first data are determined based on the source addresses and destination addresses of the first data and the determined size of each of the plurality of blocks of the first data.
4. The method according to claim 1, further comprising: detecting at least one of a second congestion level and a second bandwidth of each of the plurality of paths; dividing the second data into a plurality of blocks, the number of the plurality of blocks of the second data corresponding to the number of the plurality of paths, the size of each of the plurality of blocks of the second data being determined based on at least one of the second congestion level and the second bandwidth detected for each of the plurality of paths; as well as The plurality of blocks of the second data are transmitted respectively via the plurality of paths.
5. The method according to claim 4, further comprising: The receiving end receives the plurality of blocks of the second data substantially simultaneously at the destination port; as well as The receiving end notifies the application layer of the computer network when receiving all the blocks of the first data and the blocks of the second data from the transmitting end.
6. The method of claim 1, wherein establishing the plurality of paths and segmenting the first data into a plurality of blocks is performed by running an algorithm of the computer network.
7. The method of claim 1, wherein establishing the plurality of paths and segmenting the first data into a plurality of blocks are performed via a network interface card or a field programmable gate array of the computer network.
8. The method according to claim 1, further comprising: When a failure occurs in a first path among the plurality of paths, switching from the first path to a second path among the plurality of paths, The first path corresponds to a first source port among the plurality of source ports, and the second path corresponds to a second source port among the plurality of source ports.
9. A computer network system for multipath data transmission, the system comprising: A memory for storing program instructions; A processor is configured to execute the program instructions, and when the program instructions are executed, the processor performs the method according to any one of claims 1-8.
10. A non-transitory computer-readable medium having computer-executable instructions stored thereon, which, when executed, cause one or more processors to perform the method according to any one of claims 1-8.