Multipath transmission apparatus and method
The multi-path transmission device addresses inefficiencies in RDMA by integrating NCCL and RDMA API, dynamically adjusting paths, and optimizing data transmission for high-performance computing environments.
Patent Information
- Application Number
- PCT/KR2024/010902
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-12-26
- Filing Date
- 2024-07-26
- Publication Date
- 2025-07-03
AI Technical Summary
Existing technologies face challenges in efficiently applying multi-path transmission to Remote Direct Memory Access (RDMA) due to its hardware-embedded communication processes, making it difficult to adapt existing hardware and inefficient in GPU-based RDMA communication.
A multi-path transmission device integrating an interface module, automatic path collection module, and dynamic data classification module, utilizing the NVIDIA Collective Communications Library (NCCL) and RDMA API, to optimize data transmission paths and dynamically adjust based on network conditions.
Enhances data transmission efficiency and reliability by automatically detecting and optimizing paths, minimizing latency and congestion, and supporting high-performance computing environments.
Smart Images

Figure KR2024010902_03072025_PF_FP_ABST
Abstract
Description
Multipath transmission device and method
[0001] The following embodiments relate to a multi-path transmission device and method thereof.
[0002] As data-intensive applications such as big data analytics, machine learning, and scientific simulations run in data centers or high-performance computing (HPC) environments,1 the importance of efficient network operations is growing.
[0003] In data center networks (DCNs), Remote Direct Memory Access (RDMA) is a promising networking technology for data-intensive applications requiring high bandwidth and ultra-low latency. RDMA supports zero-copy read / write operations by implementing transfer logic in hardware network interface cards (NICs). This logic can directly transfer data from the memory of one computing node to the memory of another.
[0004] A multi-path transmission device according to one embodiment may include an interface module that integrates different libraries, an automatic path collection module that automatically detects and collects data transmission paths available within a platform and selects optimal data transmission paths among the available data transmission paths to create a path set, a data transmission module that transmits data through the optimal data transmission paths, and a dynamic data classification module that monitors the data transmission module in real time to dynamically readjust the optimal data transmission path.
[0005] The above different libraries may include an NVIDIA Collective Communications Library (NCCL) and a Remote Direct Memory Access (RDMA) Application Programming Interface (API) library, and the interface module may transmit the number of GPUs to be used and IP information determined within the NCCL to the RDMA API library, thereby optimizing multi-GPU data transmission.
[0006] The above interface module can divide the data and transmit it through different paths only when the size of the data is greater than a threshold value.
[0007] The above automatic route collection module can detect network topology changes in real time and perform route search functions.
[0008] The above automatic route collection module can generate multiple virtual IP addresses within the server, transmit a search packet, and then compare the route information to which the search packet was transmitted with existing route information to select the optimal data transmission route.
[0009] The above automatic route collection module may include the new route in the route set if the route through which the search packet was transmitted is a new route.
[0010] The above dynamic data classification module can detect the network bandwidth and transmission delay of the data transmission module and dynamically readjust the optimal data transmission path.
[0011] The above dynamic data classification module can monitor performance indicators for the optimal data transmission paths, adjust the data transmission amount for each path, and optimize the transmission performance of the optimal data transmission paths.
[0012] An electronic device according to one embodiment may include a memory storing instructions and a processor, wherein the instructions, when executed by the processor, cause the electronic device to integrate different libraries, automatically detect and collect data transmission paths available within a platform, select optimal data transmission paths among the available data transmission paths to create a path set, transmit data through the optimal data transmission paths, and monitor the data transmission in real time to dynamically readjust the optimal data transmission path.
[0013] A multi-path transmission method according to one embodiment may include a step of integrating different libraries, a step of automatically detecting and collecting data transmission paths available within a platform, a step of selecting optimal data transmission paths among the available data transmission paths to generate a path set, a step of transmitting data through the optimal data transmission paths, and a step of monitoring the data transmission in real time to dynamically readjust the optimal data transmission path.
[0014] FIG. 1 is a schematic diagram illustrating RDMA according to one embodiment.
[0015] FIG. 2 is a schematic diagram illustrating a multi-path transmission device according to one embodiment.
[0016] FIG. 3 is a schematic flowchart illustrating a multi-path transmission method according to one embodiment.
[0017] FIG. 4 is a block diagram illustrating an electronic device according to one embodiment.
[0018] Specific structural or functional descriptions of the embodiments are disclosed for illustrative purposes only and may be modified and implemented in various forms. Therefore, the actual implementation is not limited to the specific embodiments disclosed, and the scope of this specification includes modifications, equivalents, or alternatives within the technical concepts described in the embodiments.
[0019] Although terms such as "first" or "second" may be used to describe various components, these terms should be interpreted solely to distinguish one component from another. For example, a first component may be referred to as a second component, and similarly, a second component may also be referred to as a first component.
[0020] When it is said that a component is "connected" to another component, it should be understood that it may be directly connected or connected to that other component, but there may also be other components in between.
[0021] Singular expressions include plural expressions unless the context clearly dictates otherwise. In this specification, the terms "comprises" or "has" should be understood to indicate the presence of a described feature, number, step, operation, component, part, or combination thereof, but not to exclude the possibility of the presence or addition of one or more other features, numbers, steps, operations, components, parts, or combinations thereof.
[0022] Unless otherwise defined, all terms used herein, including technical or scientific terms, have the same meaning as commonly understood by a person of ordinary skill in the art. Terms defined in commonly used dictionaries should be interpreted to have a meaning consistent with their meaning in the context of the relevant technology, and will not be interpreted in an idealized or overly formal sense unless explicitly defined herein.
[0023] Hereinafter, embodiments will be described in detail with reference to the attached drawings. In the description with reference to the attached drawings, identical components are assigned the same reference numerals regardless of the drawing numbers, and redundant descriptions thereof will be omitted.
[0024]
[0025] FIG. 1 is a schematic diagram illustrating RDMA according to one embodiment.
[0026] One or more blocks or combinations of blocks of FIG. 1 may be implemented by a special purpose hardware-based computer performing a specific function, or by a combination of special purpose hardware and computer instructions.
[0027] Remote Direct Memory Access (RDMA), according to one embodiment, may be a method for enabling direct memory access from the memory of one computer to the memory of another computer without using the computer's operating system, central processing unit, or cache. RDMA enables high-throughput, low-latency networking, which may be particularly advantageous in data centers and high-performance computing environments.
[0028] RDMA enables networking without memory copying. RDMA allows data (or messages) to be transferred directly from one system's memory to another, reducing the need for data copying between buffers.
[0029] RDMA can bypass the operating system kernel, reducing context switching, interrupts, and central processing unit overhead.
[0030] RDMA can be hardware-based. RDMA can be managed and executed by a network adapter (RDMA NIC, or RNIC), offloading the workload from the central processing unit.
[0031] RDMA can support a variety of network fabrics. RDMA can operate on various network fabrics, such as InfiniBand, Ethernet (RDMA over RoCE), and iWARP (Internet Wide Area RDMA Protocol).
[0032] RDMA can achieve low latency and high efficiency. By reducing the number of data copies and context switches, RDMA can significantly reduce latency and increase data transfer efficiency.
[0033] RDMA can be used in applications requiring high data throughput and low latency, such as large-scale database transactions, high-performance computing applications, and storage area networks. The ability to offload tasks from the central processing unit and minimize latency can be crucial in data-intensive computing environments.
[0034] Figure 1 illustrates the detailed architecture and operational flow of RDMA in a network environment. Key components of RDMA according to one embodiment may include Queue Pairs (QPs) and Memory Regions (MRs). A QP, consisting of a transmit and receive queue, may be essential for initiating and managing RDMA communication between two hosts. A memory region may be a specific area of memory prepared for direct access through RDMA operations. QP configuration may include establishing a communication channel between hosts with transmit and receive queues. In an MR, a host may identify a specific memory region that can be directly accessed for RDMA operations. In Work Queue Elements (WQEs), data transfer commands are queued in the QP, and these commands can specify the characteristics of the operation, such as transmitting or receiving data. The Network Interface Card (NIC) can process the command by performing a memory transfer directly according to the instructions in the WQE. After completing the operation, the NIC can report back through Completion Queue Elements (CQEs) to indicate the status of the RDMA operation.
[0035] More specifically, Figure 1 illustrates an overview of RDMA. It can illustrate the process of establishing and executing an RDMA task between two hosts. Key elements include establishing a queue pair (QP) between the hosts, each QP comprising a transmit and receive queue. Memory regions (MRs) can be defined for direct access by a network interface card (NIC). A task can be initiated by posting a work queue element (WQE) to the QP. The NIC can then process this WQE and transmit data according to the provided instructions. The completion of this task can be indicated by a completion queue element (CQE). This setup can demonstrate the direct, high-speed data transfer capabilities of RDMA, which bypass the operating system and CPU, reducing latency and increasing the efficiency of data transfer between networked computers.
[0036] Below, multipath transmission using RDMA is described.
[0037] Multipathing may be essential to fully utilize the various network links existing between servers within a data center. Multipathing can serve two purposes: it provides a bypass route in the event of a failure in a specific network link, and it can increase data transmission throughput by utilizing multiple paths simultaneously.
[0038] However, since all of the existing technologies were developed based on the classic TCP / IP (Transmission Control Protocol / Internet Protocol)-based network, it may be impossible to apply them to RDMA, a high-performance data transmission technology.
[0039] Unlike TCP / IP, RDMA has all of its communication processes embedded within the hardware, making it impossible to modify the RDMA transmission logic within the system kernel (OS). Therefore, applying multipath transmission technology to RDMA may require either 1) developing new hardware (RNIC, RDMA-specific NIC card) or 2) incorporating the technology into a user-level library that implements the RDMA API. In case of 1), it may be difficult to apply it to existing hardware, and method 2) may require the development of multipath transmission technology. However, in the case of conventional technologies, they mainly consider RDMA that performs data transmission targeting CPU memory, and when applied to GPU-based RDMA communication (GPUDirect RDMA) recently used in artificial intelligence, they may operate inefficiently or be difficult to utilize effectively.
[0040] The following description may be about a multi-path transmission technology that can be efficiently utilized for various platforms that operate based on high-performance computing utilizing GPUs.
[0041]
[0042] FIG. 2 is a schematic diagram illustrating a multi-path transmission device according to one embodiment.
[0043] The description referring to Fig. 1 can be equally applied to Fig. 2, and overlapping content may be omitted. One or more blocks and combinations of blocks of Fig. 2 may be implemented by a special-purpose hardware-based computer performing a specific function, or a combination of special-purpose hardware and computer instructions.
[0044] A multi-path transmission device according to one embodiment may include an interface module, an automatic path collection module, a data transmission module, and a dynamic data classification module. The multi-path transmission device may transmit data through multiple paths based on RDMA on a GPU. The multi-path transmission device may transmit data by efficiently linking different libraries, automatically collecting available paths, and dynamically classifying data.
[0045] The term "module" may mean, for example, a unit comprising one or a combination of two or more of hardware, software, or firmware. "Module" may be used interchangeably with terms such as unit, logic, logical block, component, or circuit. A "module" may be the smallest unit of an integrally formed component or a portion thereof. A "module" may also be the smallest unit or a portion thereof that performs one or more functions. A "module" may be implemented mechanically or electronically. For example, a "module" may include at least one of an application-specific integrated circuit (ASIC) chip, field-programmable gate array (FPGA), or programmable-logic device, known or to be developed in the future, that performs certain operations.
[0046] The interface module can integrate different libraries, including the NVIDIA Collective Communications Library (NCCL) and the Remote Direct Memory Access (RDMA) Application Programming Interface (API) library.
[0047] NCCL is a library provided by NVIDIA that supports high-performance, scalable collective communication in multi-GPU environments. NCCL is optimized for multi-GPUs, supports collective communication operations, large-scale clusters, and is compatible with various network interfaces.
[0048] The RDMA API can provide a programming interface for Remote Direct Memory Access. By using NCCL and the RDMA API together, data transfer and collective computation between multiple GPUs can be efficiently handled in high-performance computing environments.
[0049] The interface module can optimize multi-GPU data transfer by passing the number of GPUs to be used and IP information determined within the NCCL to the RDMA API library. The interface module can split the data and transmit it along separate paths only if the data size exceeds a threshold. For example, the interface module can perform multi-path transfer by splitting the data into multiple smaller pieces and transmitting them along separate paths only if the data size exceeds a threshold (e.g., 1 MB).
[0050] The interface module can multiplex the data transmission path by assigning different source IP addresses to data using ECMP characteristics and virtual IPs. ECMP (Equal-Cost Multi-Path) is an equal-cost multi-path routing method. ECMP can distribute traffic by using multiple paths with the same cost (e.g., delay, hop count) among multiple paths to the same destination. ECMP can reduce network congestion and improve overall network efficiency by transmitting traffic through multiple paths. The source IP (source IP address) can refer to the source address of a data packet on the network, i.e., the "origin IP address." The interface module can utilize ECMP characteristics and virtual IPs to diversify the source IP addresses of data, allowing data packets to be transmitted through multiple different paths. This can improve network load distribution, increase transmission efficiency, and provide higher data transmission reliability.
[0051] The interface module can choose to transmit data via a single path, as multipathing can be inefficient when the data is sufficiently small. Since data using a single path can be interrupted by data using multiple paths, data using a single path can be given higher priority to prevent performance degradation.
[0052] The interface module can convert and transmit information provided by NCCL into a form that the RDMA API can understand and utilize. The interface module can first collect information related to GPU usage from NCCL. This information may include the total number of GPUs to be used and the IP address assigned to each GPU.
[0053] The interface module can convert the collected information into a format understandable by the RDMA API. The interface module can then pass this information to the RDMA API, allowing RDMA to establish an efficient data transfer path between each GPU. Based on the information received from the interface module, the RDMA API can establish direct memory access to each GPU over the network. This minimizes network latency and optimizes data transfer speed.
[0054] The interface module can monitor whether the NCCL and RDMA APIs are working smoothly. It can also support optimized data transfer and communication between the NCCL and RDMA APIs.
[0055] The automatic route collection module automatically detects and collects available data transmission paths within the platform, and can select optimal data transmission paths from among them to create a route set. Examples of platforms include high-performance computing environments, cloud computing infrastructure, large-scale data centers, or enterprise environments with complex network architectures. The automatic route collection module can determine the optimal path based on the characteristics of each platform, taking into account various factors such as network bandwidth, latency, error rate, and congestion levels.
[0056] The automatic route collection module can detect network topology changes in real time and perform route discovery functions. Network topology can refer to the physical or logical structure of a network. For example, in high-performance computing environments, network topology can change over time. These changes can occur for a variety of reasons, including the addition of new devices (e.g., adding a switch), the removal of existing devices (e.g., removing a server), or changes in network connectivity. The automatic route collection module monitors the network in real time, continuously monitoring the current state of the network and detecting and identifying topology changes. The automatic route collection module then recalculates the path and optimizes the path to select the optimal data transmission path. This can be updated in the data transmission module to enable data transmission along the new path.
[0057] The automatic route collection module can create multiple virtual IP addresses within the server, send probe packets, and then compare the route information to which the probe packets were sent with the existing route information to select the optimal data transmission route.
[0058] The automatic route collection module can include a new route in the route set if the route along which the probe packet was sent is a new route.
[0059] When selecting a source IP address to multiplex data paths, since the selection method varies depending on the IP address depending on the switch equipment, it can be difficult to confirm which path each IP address goes to before it is actually sent. Therefore, the automatic route collection module can automatically find the number of available paths within the platform and the set of source IP addresses that should be used for each path. Here, the automatic route collection module creates multiple virtual IPs within the server, and sends a discovery packet with the virtual IP address through the path tracking function. Then, it compares the path information to which the packet was sent with the existing path information to determine whether it is a new path that did not exist before. If it is a new path, the automatic route collection module can add it to the existing path set and use the source IP address for multi-path transmission.
[0060] More specifically, the automatic route collection module can generate virtual IP addresses. These virtual IP addresses can be used to experimentally explore various data transmission paths on a network. These virtual IP addresses can be implemented in software rather than assigned to actual network devices. The automatic route collection module can transmit probe packets. These probe packets can use the generated virtual IP addresses to determine which paths will be used in actual network situations. The automatic route collection module can compare and analyze path information. The automatic route collection module can compare the paths through which probe packets are transmitted with existing paths to discover new paths and verify their validity. The automatic route collection module can add newly validated paths to the existing path set. The automatic route collection module can select a source IP address for multi-path transmission. The automatic route collection module can select an effective source IP address for each path by considering path selection that may vary depending on the network switch device. The automatic route collection module can perform the path discovery function by periodically performing the aforementioned process.
[0061] The data transmission module can transmit data through optimal data transmission paths selected by the automatic path collection module.
[0062] The dynamic data classification module monitors the data transmission module in real time, allowing it to dynamically adjust the optimal transmission path. The dynamic data classification module can monitor network traffic patterns, bandwidth usage, and latency in real time to understand the current state of the network.
[0063] The dynamic data classification module can detect network bandwidth and transmission delays of the data transmission module and dynamically adjust the optimal data transmission path. When detecting network congestion or changes in bandwidth, the dynamic data classification module can automatically adjust the optimal data transmission path. It can also work with the automatic path collection module to adjust the optimal data transmission path.
[0064] The dynamic data classification module can monitor performance indicators for optimal data transmission paths, adjust the data transmission volume for each path, and optimize the transmission performance of optimal data transmission paths.
[0065] The dynamic data classification module can multiplex paths based on data size when available bandwidth is limited due to poor network conditions. For example, even data of the same size can be sent through a single path under good network conditions. However, under poor network conditions, the paths may need to be multiplexed to divide the data into multiple smaller data pieces for transmission. The dynamic data classification module can determine whether to multiplex paths using a method as shown in mathematical equation 1 below. The method for determining whether to multiplex paths in the dynamic data classification module is not limited to the mathematical equation described, and various mathematical methods can be applied.
[0066] [Mathematical Formula 1]
[0067] size_th = max(1byte, 1MB * (current_throughput / max_throughput)
[0068] The dynamic data classification module can dynamically set the data segmentation threshold (size_th) based on the ratio of the network's current throughput to its maximum throughput. When network performance degrades (e.g., when the current throughput is lower than the maximum throughput), the data segmentation threshold can be reduced to more actively utilize path multiplication. For example, when network congestion occurs, the threshold (e.g., from 1 MB to 0.5 MB) can be reduced to allow smaller data to be transmitted via multiple paths, alleviating network congestion. In other words, while the multipath transmission device previously segmented data larger than 1 MB for multipath transmission, if a network condition abnormality occurs, the dynamic data classification device can reduce the threshold to 0.5 MB, allowing data larger than 0.5 MB to be segmented for multipath transmission.
[0069]
[0070] FIG. 3 is a schematic flowchart illustrating a multi-path transmission method according to one embodiment.
[0071] The description with reference to FIG. 1 and FIG. 2 can be equally applied to FIG. 3, and overlapping content can be omitted.
[0072] The operations of FIG. 3 may be performed in the order and manner illustrated, but the order of some operations may be changed or some operations may be omitted without departing from the spirit and scope of the illustrated embodiment. Multiple operations illustrated in FIG. 3 may be performed in parallel or simultaneously.
[0073] For convenience of explanation, steps (310 to 340) are described as being performed using the multi-path transmission device (200) illustrated in FIG. 2. However, these steps (310 to 340) may be utilized via any other suitable electronic device and within any suitable system.
[0074] In step (310), the multi-path transmission device can integrate different libraries. For example, the multi-path transmission device can optimize multi-GPU data transmission by transmitting the number of GPUs to be used and IP information determined within the NCCL to the RDMA API library. The multi-path transmission device can split the data and transmit it through different paths only when the data size exceeds a threshold.
[0075] In step (320), the multi-path transmission device can automatically detect and collect data transmission paths available within the platform, and select optimal data transmission paths among the available data transmission paths to create a path set. For example, the multi-path transmission device can detect network topology changes in real time and perform a path discovery function. The multi-path transmission device can create multiple virtual IP addresses within the server, transmit a discovery packet, and then compare the path through which the discovery packet was transmitted with existing path information to select an optimal data transmission path. If the path through which the discovery packet was transmitted is a new path, the multi-path transmission device can include the new path in the path set.
[0076] In step (330), the multi-path transmission device can transmit data through optimal data transmission paths.
[0077] In step (340), the multi-path transmission device can monitor data transmission in real time and dynamically readjust the optimal data transmission path. For example, the multi-path transmission device can detect the network bandwidth and transmission delay of the data transmission module and dynamically readjust the optimal data transmission path. The multi-path transmission device can monitor performance indicators for the optimal data transmission paths, adjust the data transmission amount for each path, and optimize the transmission performance of the optimal data transmission paths.
[0078]
[0079] FIG. 4 is a block diagram illustrating an electronic device according to one embodiment.
[0080] One or more blocks and combinations of blocks of FIG. 4 may be implemented by a special-purpose hardware-based computer performing a specific function, or by a combination of special-purpose hardware and computer instructions. The descriptions made with reference to FIGS. 1 through 3 may be equally applicable to FIG. 4 . For example, an electronic device (400) according to one embodiment may include a multi-path transmission device (200).
[0081] As shown in FIG. 4, the electronic device (400) may include a memory (410) and a processor (420). The electronic device (400) may further include a communication module, and the communication module may include a transmitter and a receiver.
[0082] An electronic device (400) according to one embodiment may include a memory (410) and a processor (420) connected to the memory (410) via a system bus or other suitable circuitry.
[0083] The electronic device (400) may store program code in memory (410). In one embodiment, the memory (410) may include one or more physical memory devices, such as local memory or one or more bulk storage devices. In this case, the local memory may include random access memory (RAM) or other volatile memory devices commonly used while actually executing the program code. The bulk storage device may be implemented as a hard disk drive (HDD), a solid state drive (SSD), or other non-volatile memory device.
[0084] As the executable program code stored in the memory (410) is executed by the electronic device (400), the processor (420) may perform various operations described in the present disclosure. For example, the memory (410) may store program code for causing the processor (420) to perform one or more operations described in FIGS. 1 to 3.
[0085] Depending on the specific type of device being implemented, the electronic device (400) may include fewer components than those illustrated or additional components not illustrated in FIG. 4. Additionally, one or more of the components may be incorporated into, or otherwise form part of, another component.
[0086] A processor (420) according to one embodiment is a hardware configuration that performs overall control functions for controlling the operations of an electronic device (400). For example, the processor (420) may control the electronic device (400) overall by executing programs stored in a memory (410) within the electronic device (400). The processor (420) may be implemented as a central processing unit (CPU), a graphics processing unit (GPU), an application processor (AP), a neural processing unit (NPU), or the like, provided within the electronic device (400), but is not limited thereto.
[0087] The processor (420) can integrate different libraries, automatically detect and collect data transmission paths available within the platform, select optimal data transmission paths among the available data transmission paths to create a path set, transmit data through the optimal data transmission paths, monitor data transmission in real time, and dynamically readjust the optimal data transmission path.
[0088]
[0089] The embodiments described above may be implemented using hardware components, software components, and / or a combination of hardware components and software components. For example, the devices, methods, and components described in the embodiments may be implemented using a general-purpose computer or a special-purpose computer, such as, for example, a processor, a controller, an arithmetic logic unit (ALU), a digital signal processor, a microcomputer, a field programmable gate array (FPGA), a programmable logic unit (PLU), a microprocessor, or any other device capable of executing instructions and responding to them. The processing device may execute an operating system (OS) and software applications running on the operating system. Furthermore, the processing device may access, store, manipulate, process, and generate data in response to the execution of the software. For ease of understanding, the processing device is sometimes described as being used alone; however, one of ordinary skill in the art will recognize that the processing device may include multiple processing elements and / or multiple types of processing elements. For example, a processing unit may include multiple processors, or a processor and a controller. Other processing configurations, such as parallel processors, are also possible.
[0090] Software may include a computer program, code, instructions, or a combination of one or more of these, which may configure a processing device to perform a desired operation or may, independently or collectively, command the processing device. The software and / or data may be permanently or temporarily embodied in any type of machine, component, physical device, virtual equipment, computer storage medium or device, or transmitted signal wave, for interpretation by the processing device or for providing instructions or data to the processing device. The software may also be distributed over networked computer systems and stored or executed in a distributed manner. The software and data may be stored on a computer-readable recording medium.
[0091] The method according to the embodiment may be implemented in the form of program commands that can be executed through various computer means and recorded on a computer-readable medium. The computer-readable medium may include program commands, data files, data structures, etc., alone or in combination, and the program commands recorded on the medium may be those specially designed and configured for the embodiment or may be known and available to those skilled in the art of computer software. Examples of the computer-readable recording medium include magnetic media such as hard disks, floppy disks, and magnetic tapes, optical media such as CD-ROMs and DVDs, magneto-optical media such as floptical disks, and hardware devices specially configured to store and execute program commands, such as ROMs, RAMs, and flash memories. Examples of program commands include not only machine language codes such as those generated by a compiler, but also high-level language codes that can be executed by a computer using an interpreter, etc.
[0092] The hardware device described above may be configured to operate as one or more software modules to perform the operations of the embodiment, and vice versa.
[0093] Although the embodiments described above have been described with limited drawings, those skilled in the art will appreciate that various technical modifications and variations can be applied based on the described embodiments. For example, appropriate results can still be achieved even if the described techniques are performed in a different order than described, and / or components of the described systems, structures, devices, circuits, etc. are combined or combined in a different manner than described, or are replaced or substituted with other components or equivalents.
[0094] Therefore, other implementations, other embodiments, and equivalents to the claims also fall within the scope of the claims described below.
Claims
1. In a multi-path transmission device, Interface module that integrates different libraries; An automatic path collection module that automatically detects and collects data transmission paths available within the platform and selects optimal data transmission paths among the available data transmission paths to create a path set; A data transmission module for transmitting data through the above optimal data transmission paths; and A dynamic data classification module that monitors the above data transmission module in real time and dynamically readjusts the optimal data transmission path. A multi-path transmission device comprising:
2. In paragraph 1, The above different libraries are NVIDIA Collective Communications Library (NCCL) and Remote Direct Memory Access (RDMA) Application Programming Interface (API) libraries Including, The above interface module A multi-path transmission device that optimizes multi-GPU data transmission by transmitting the number of GPUs to be used and IP information determined within the above NCCL to the above RDMA API library.
3. In paragraph 2, The above interface module A multi-path transmission device that divides the data and transmits it through different paths only when the size of the data is greater than a threshold value.
4. In paragraph 1, The above automatic route collection module A multi-path transmission device that detects network topology changes in real time and performs a path search function.
5. In paragraph 1, The above automatic route collection module A multi-path transmission device that creates multiple virtual IP addresses within a server, transmits a search packet, and then compares the route information along which the search packet was transmitted with existing route information to select the optimal data transmission route.
6. In paragraph 5, The above automatic route collection module A multi-path transmission device that includes the new path in the path set if the path along which the above search packet is transmitted is a new path.
7. In paragraph 1, The above dynamic data classification module A multi-path transmission device that detects the network bandwidth and transmission delay of the above data transmission module and dynamically readjusts the optimal data transmission path.
8. In paragraph 1, The above dynamic data classification module A multi-path transmission device that monitors performance indicators for the above optimal data transmission paths, adjusts the data transmission amount for each path, and optimizes the transmission performance of the above optimal data transmission paths.
9. In electronic devices, Memory for storing instructions; and Processor Including The above instructions, when executed by the processor, cause the electronic device to: Integrate different libraries, Automatically detects and collects data transmission paths available within the platform, selects optimal data transmission paths among the available data transmission paths, and creates a set of paths. Transmitting data through the above optimal data transmission paths, An electronic device that monitors the above data transmission in real time and dynamically readjusts the optimal data transmission path.
10. In a multi-path transmission method, Steps to integrate different libraries; A step of automatically detecting and collecting data transmission paths available within a platform, and selecting optimal data transmission paths among the available data transmission paths to generate a path set; A step of transmitting data through the above optimal data transmission paths; and A step of monitoring the above data transmission in real time and dynamically readjusting the optimal data transmission path. A multi-path transmission method comprising:
11. A computer program stored on a computer-readable recording medium to execute the method of clause 10 in combination with hardware.
Citation Information
Patent Citations
Communication Paths From An InfiniBand Host
US20080155107A1
Systems and methods for improved fault tolerance in solicited information handling systems
US20150200802A1
Data path selection for network transfer using high speed RDMA or non-RDMA data paths
US20160028819A1
Method for verifying data center network performance
US20230139774A1
Method and system for generating data packets on a heterogeneous network
US6212190B1