VRB dynamic scheduling method and device for low-delay computing power transmission, equipment and medium
By generating dynamic network topology maps and VRB protocol transmission data, the problems of high network transmission delay and packet loss rate in the prior art are solved, and efficient and reliable network transmission is achieved.
Patent Information
- Application Number
- CN202510659332.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-21
- Publication Date
- 2025-08-15
AI Technical Summary
The prior art may circulate through unnecessary nodes or links when scheduling tasks, resulting in increased network transmission delay and packet loss rate, reducing network transmission efficiency.
By obtaining the transmission parameters of multiple GPU servers and switches, a dynamic network topology diagram is generated, the optimal transmission path is selected based on the diagram, and data is transmitted through the VRB protocol, and the number of GPU servers is dynamically adjusted to adapt to traffic changes.
It realizes the selection of the optimal transmission path according to the real-time network status, avoids circulating through unnecessary nodes or links, improves network transmission efficiency and reliability, and optimizes resource utilization.
Smart Images

Figure CN120499083A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer technology, and in particular to a VRB-based dynamic network topology awareness scheduling method, device, equipment and medium. Background Art
[0002] With the development of science and technology, the scale of networks is constantly expanding and the network structure is becoming increasingly complex. People's requirements for efficiency when scheduling data involved in tasks are constantly increasing. In existing technologies, when scheduling tasks, unnecessary nodes or links may generally be bypassed, which increases transmission delay and packet loss rate and reduces network transmission efficiency. Summary of the Invention
[0003] In view of the above problems, embodiments of the present invention are proposed to provide a VRB-based dynamic network topology-aware scheduling method, apparatus, device, and medium that overcome the above problems or at least partially solve the above problems.
[0004] To solve the above problems, an embodiment of the present invention discloses a dynamic network topology-aware scheduling method based on VRB, involving multiple GPU servers and multiple switches. The method includes:
[0005] Obtaining transmission parameters of the multiple GPU servers and the multiple switches;
[0006] generating a dynamic network topology map according to the transmission parameters of the plurality of GPU servers and the plurality of switches;
[0007] The data involved in the task is transmitted based on the dynamic network topology graph.
[0008] Optionally, the data involved in the transmission task based on the dynamic network topology graph includes:
[0009] Based on the dynamic network topology diagram, data involved in the task is transmitted through the VRB protocol.
[0010] Optionally, the data involved in the transmission task based on the dynamic network topology graph includes:
[0011] Determining path costs of multiple paths for transmitting data involved in the task in the dynamic network topology graph;
[0012] determining a target transmission path according to the path costs of the multiple paths;
[0013] The data related to the task is transmitted through the target transmission path.
[0014] Optionally, determining path costs of multiple paths for transmitting data involved in the task in the dynamic network topology graph includes:
[0015] Determining load information of nodes involved in multiple paths in the dynamic network topology graph and link bandwidths of the multiple paths;
[0016] Path costs of the multiple paths are determined according to load information of nodes involved in the multiple paths and link bandwidths of the multiple paths.
[0017] Optionally, obtaining data flow of data involved in historical tasks in a preset time period;
[0018] Determining the data flow rate of data involved in the task to be transmitted in the next time period based on the data flow rate of the historical task involved in the preset time period;
[0019] The number of GPU servers in the VRB transmission network is adjusted according to the data flow of the data involved in the task to be transmitted in the next time period.
[0020] Optionally, adjusting the number of GPU servers in the VRB transmission network according to the data flow of the data involved in the task to be transmitted in the next time period includes:
[0021] If it is detected that the data flow of the data involved in the task to be transmitted in the next time period is greater than a first flow threshold, increasing the number of the GPU servers in the VRB transmission network;
[0022] If it is detected that the data flow of the data involved in the task to be transmitted in the next time period is less than a second flow threshold, the number of the GPU servers in the VRB transmission network is reduced, and the first flow threshold is greater than the second flow threshold.
[0023] Optionally, the transmission parameters include at least one of computing power information, load information, network bandwidth, and transmission delay.
[0024] The present invention also discloses a VRB-based dynamic network topology awareness scheduling device, involving multiple GPU servers and multiple switches, the device comprising:
[0025] An acquisition module, configured to acquire transmission parameters of the multiple GPU servers and the multiple switches;
[0026] A generating module, configured to generate a dynamic network topology map according to the transmission parameters of the plurality of GPU servers and the plurality of switches;
[0027] A transmission module is used to transmit data involved in the task based on the dynamic network topology diagram.
[0028] Optionally, the transmission module includes:
[0029] The first transmission submodule is configured to transmit the data involved in the task based on the dynamic network topology graph and through a VRB protocol.
[0030] Optionally, the transmission module includes:
[0031] A first determining submodule, configured to determine path costs of multiple paths for transmitting data involved in the task in the dynamic network topology graph;
[0032] A second determining submodule, configured to determine a target transmission path according to the path costs of the multiple paths;
[0033] The second transmission submodule is configured to transmit the data involved in the task through the target transmission path.
[0034] Optionally, the first determining submodule includes:
[0035] a first determining unit, configured to determine load information of nodes involved in a plurality of paths in the dynamic network topology graph, and link bandwidths of the plurality of paths;
[0036] The second determining unit is configured to determine the path costs of the multiple paths according to the load information of the nodes involved in the multiple paths and the link bandwidths of the multiple paths.
[0037] Optionally, it also includes:
[0038] A traffic acquisition module is used to obtain the data traffic of the historical tasks involved in a preset time period;
[0039] A determination module, configured to determine the data flow rate of the data involved in the task to be transmitted in the next time period based on the data flow rate of the data involved in the historical task in the preset time period;
[0040] The adjustment module is configured to adjust the number of GPU servers in the VRB transmission network according to the data flow of the data involved in the task to be transmitted in the next time period.
[0041] Optionally, the adjustment module includes:
[0042] an adding submodule, configured to increase the number of the GPU servers in the VRB transmission network if it is detected that the data flow rate of the data involved in the task to be transmitted in the next time period is greater than a first flow rate threshold;
[0043] The reducing submodule is configured to reduce the number of the GPU servers in the VRB transmission network if it is detected that the data flow of the data involved in the task to be transmitted in the next time period is less than a second flow threshold, and the first flow threshold is greater than the second flow threshold.
[0044] Optionally, the transmission parameters include at least one of computing power information, load information, network bandwidth, and transmission delay.
[0045] The present invention also discloses an electronic device, comprising: a processor, a memory, and a computer program stored in the memory and capable of running on the processor. When the computer program is executed by the processor, the steps of the above-mentioned VRB-based dynamic network topology-aware scheduling method are implemented.
[0046] The present invention also discloses a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the above-mentioned VRB-based dynamic network topology-aware scheduling method are implemented.
[0047] The embodiments of the present invention include the following advantages:
[0048] The present invention discloses a VRB-based dynamic network topology-aware scheduling method, apparatus, device, and storage medium. The method can generate a dynamic network topology diagram based on acquired transmission parameters, and can intuitively and accurately reflect the current network structure and the connection relationship between nodes, as well as their real-time status. When the device status in the network changes, such as when a GPU server is overloaded or a link is congested, the dynamic network topology diagram can be updated in a timely manner. The method can select the optimal transmission path for the data involved in the task based on the real-time network status, avoiding the problem of detouring through unnecessary nodes or links when scheduling the data involved in the task in the prior art, thereby improving network transmission efficiency. BRIEF DESCRIPTION OF THE DRAWINGS
[0049] Figure 1 This is a flowchart of a VRB-based dynamic network topology awareness scheduling method provided by an embodiment of the present invention;
[0050] Figure 2 is a schematic diagram of a dynamic network topology provided by an embodiment of the present invention;
[0051] Figure 3 This is a structural block diagram of another dynamic network topology diagram provided by an embodiment of the present invention;
[0052] Figure 4 This is a structural block diagram of a VRB-based dynamic network topology awareness scheduling device provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0053] In order to make the above-mentioned objects, features and advantages of the present invention more obvious and easy to understand, the present invention is further described in detail below with reference to the accompanying drawings and specific embodiments.
[0054] One of the core concepts of the embodiments of the present invention is that the present invention can generate a dynamic network topology map based on the acquired transmission parameters, which can intuitively and accurately reflect the structure of the current network and the connection relationship between each node, as well as their real-time status. When the status of the equipment in the network changes, such as when a GPU server is overloaded or a link is congested, the dynamic network topology map can be updated in time; the optimal transmission path can be selected for the data involved in the task based on the real-time network status, avoiding the problem of detouring through unnecessary nodes or links when scheduling the data involved in the task in the prior art, thereby improving network transmission efficiency.
[0055] Reference Figure 1 , shows a flowchart of the steps of a VRB-based dynamic network topology-aware scheduling method provided by an embodiment of the present invention, the method involves multiple GPU servers and multiple switches, and the method includes:
[0056] Step 101: Acquire transmission parameters of multiple GPU servers and multiple switches.
[0057] In an embodiment of the present invention, the method involves multiple GPU servers and multiple switches. The GPU servers have powerful graphics processing and parallel computing capabilities, and are often used to process large-scale data, perform deep learning and other computing-intensive tasks. The switches are responsible for forwarding data frames in the network, and can realize communication connections between different devices.
[0058] Transmission parameters: These refer to various parameters related to data transmission between GPU servers and switches, such as bandwidth utilization, transmission delay, packet loss rate, and device load. These parameters can reflect the real-time operating status of network devices and the quality of network links.
[0059] Step 102: Generate a dynamic network topology diagram based on the transmission parameters of the multiple GPU servers and the multiple switches.
[0060] In this embodiment of the present invention, the dynamic network topology diagram is a visual representation of the VRB transmission network structure. The diagram not only includes the location and connection relationship of each GPU server and switch, but also displays the real-time status of each node and link based on real-time transmission parameters. The diagram will be dynamically updated as the network status changes.
[0061] After obtaining the transmission parameters, these data can be processed and analyzed. First, the transmission parameters of each device are mapped to the corresponding nodes and links. Then, the network topology is constructed based on the physical connection relationship between the devices. Next, the topology is dynamically adjusted based on the transmission parameters. For example, different colors or line thicknesses are used to represent different bandwidth utilization, delay and other status information, thereby generating a dynamic network topology map that can reflect the real-time status of the network.
[0062] In the dynamic topology network diagram, each GPU server and switch is represented as a node in the diagram. The edges in the diagram represent the connection relationship between each GPU server and the switch. The weight of the edge can be used to represent the transmission parameters of multiple GPU servers and multiple switches.
[0063] like Figure 2 , shows a schematic diagram of a dynamic network topology diagram provided by an embodiment of the present invention, which may include 8 GPU servers, namely node1, node2, node3, node4, node5, node6, node7, and node0; the switches are respectively s0, s1, s2, s3, s4, s5, and s6. The edges in the diagram are used to represent the connection relationship between each GPU server and switch, and the weight of the edge can represent the transmission parameters of multiple GPU servers and multiple switches.
[0064] Step 103: transmitting data involved in the task based on the dynamic network topology graph.
[0065] In an embodiment of the present invention, when a task needs to be transmitted in a VRB transmission network, the optimal transmission path can be selected by referring to the dynamic network topology map. The dynamic network topology map comprehensively considers the status of each node and link, and preferentially selects a path with sufficient bandwidth, low latency, and low packet loss rate to transmit the data involved in the task. This can prevent the data involved in the task from detouring through unnecessary nodes or links, thereby improving the efficiency and reliability of the data transmission involved in the task.
[0066] The present invention can generate a dynamic network topology map based on the acquired transmission parameters, which can intuitively and accurately reflect the current network structure and the connection relationship between each node, as well as their real-time status. When the device status in the network changes, such as when a GPU server is overloaded or a link is congested, the dynamic network topology map can be updated in time; the optimal transmission path can be selected for the data involved in the task based on the real-time network status, avoiding the problem of detouring through unnecessary nodes or links when scheduling the data involved in the task in the prior art, thereby improving network transmission efficiency.
[0067] In an embodiment of the present invention, transmitting data involved in a task based on a dynamic network topology graph includes: transmitting data involved in the task based on the dynamic network topology graph and using a VRB protocol.
[0068] In an embodiment of the present invention, multiple GPU servers and multiple switches can form a VRB network. The VRB transmission network is a specific transmission network. The VRB protocol refers to V2V RDMA BAND. V2V refers to the Visual Networking Protocol, which is a network protocol for video communication. RDMA stands for Remote Direct Memory Access, which allows computers to directly access the memory of other computers without the intervention of the operating system. BAND stands for bandwidth, which refers to the on-chip remote direct memory access bandwidth technology based on the Visual Networking Protocol.
[0069] Each GPU server can be equipped with a VRB network card. When transmitting data involved in a task, the optimal path can be determined in combination with the dynamic network topology diagram, and the data involved in the task can be transmitted to the target GPU server through the VRB protocol. For example, there are four transmission paths between GPU server 1 and GPU server 3, namely transmission path 1, transmission path 2, transmission path 3, and transmission path 4. After comprehensive judgment, it can be determined that transmission path 3 is the optimal path. GPU server 1 can transmit the data involved in the task to GPU server 3 through transmission path 3 based on the VRB protocol.
[0070] The present invention directly transmits data between server memories through the VRB protocol without going through the TCP / IP protocol stack and the operating system kernel, thereby reducing CPU overhead and delay.
[0071] In one embodiment of the present invention, data involved in a transmission task based on a dynamic network topology graph includes: determining the path costs of multiple paths for transmitting data involved in the transmission task in the dynamic network topology graph; determining a target transmission path based on the path costs of the multiple paths; and transmitting the data involved in the transmission task through the target transmission path.
[0072] In an embodiment of the present invention, path cost: in a dynamic network topology diagram, each path has a corresponding path cost, which is a quantitative value obtained after comprehensive consideration of multiple transmission-related factors. These factors include bandwidth occupancy, transmission delay time, packet loss rate, device load level, etc. The path cost can be used to measure the resources consumed and the risks faced by the data involved in the transmission task on the path. The smaller the value, the more efficient and reliable the path is in transmitting the data involved in the task.
[0073] The target transmission path refers to the path that is most suitable for transmitting the data involved in the current task, selected from the many available paths in the dynamic network topology diagram based on the path cost. The target transmission path is usually the path with the lowest path cost, which can ensure that the data involved in the task is transmitted in an efficient and stable manner.
[0074] When the data involved in a task needs to be transmitted, all possible transmission paths can be found in the dynamic network topology graph based on the source node (the GPU server that initiates the data involved in the task) and the target node (the GPU server that receives the data involved in the task). This can be achieved with the help of graph search algorithms, such as breadth-first search (BFS), depth-first search (DFS), or Dijkstra's algorithm.
[0075] For each searched path, its path cost can be calculated by comprehensively considering multiple transmission parameters.
[0076] After calculating the path costs of all possible paths, the system will compare these path costs. Normally, the path with the smallest path cost will be selected as the target transmission path. This is because the path with the smallest path cost has advantages in resource utilization and transmission efficiency, and can minimize the time and resource consumption required to transmit the data involved in the task. Once the target transmission path is determined, the data involved in the task can be transmitted along this path. Specifically, the data involved in the task will be encapsulated at the source node, and then forwarded to the target node in sequence along the target transmission path through network devices such as switches. During the transmission process, the status of the target transmission path can be continuously monitored. If abnormal conditions such as path failure and congestion occur, the path cost will be recalculated in time and a new target transmission path will be selected to ensure that the data involved in the task can be transmitted smoothly.
[0077] In one application scenario, in a cloud computing data center, there is a VRB transmission network consisting of multiple GPU servers and switches. When a user initiates a data processing task involving data, the system first finds all possible paths from the user's server to the target server that processes the data involved in the task. Then, the path cost of each path is calculated based on factors such as bandwidth usage on each path, load on the switch and server, transmission delay, and packet loss rate. Assume that there are three paths, path A has sufficient bandwidth but a switch it passes through has a high load, path B has low bandwidth but low delay, and path C has medium bandwidth and delay. After calculation, the path cost of path B is the smallest. Then, the system will determine path B as the target transmission path and transmit the data involved in the task from the source server to the target server along path B to ensure that the data involved in the task is processed efficiently and stably.
[0078] This method can determine the path costs of multiple paths involved in a data transmission task within a dynamic network topology and select the target transmission path accordingly, comprehensively considering various network factors such as bandwidth, latency, packet loss rate, and device load. Compared with traditional transmission methods, this method no longer selects paths randomly or based on fixed rules. Instead, it accurately identifies the path that best suits the data transmission required by the current task, avoiding the problem of low transmission efficiency caused by unreasonable path selection.
[0079] In one embodiment of the present invention, determining the path costs of multiple paths of data involved in a transmission task in a dynamic network topology graph includes: determining the load information of the nodes involved in the multiple paths in the dynamic network topology graph, and the link bandwidth of the multiple paths; and determining the path costs of the multiple paths based on the load information of the nodes involved in the multiple paths and the link bandwidth of the multiple paths.
[0080] In an embodiment of the present invention, for each GPU server, switch, or other node in the VRB transmission network, load information can be collected through pre-installed monitoring software or a network management protocol (such as SNMP). The load information may include CPU usage, memory usage, processing queue length, etc. For example, for a GPU server, high CPU usage may mean that the server is processing data related to a large number of computing tasks, and its ability to process data related to new tasks may be limited. For a switch, a long processing queue length may indicate that a large amount of data is currently waiting to be forwarded, which may affect the transmission speed of new data.
[0081] The load information of these nodes can be updated regularly or in real time to ensure that the latest load conditions of the nodes in the dynamic network topology are obtained.
[0082] Link bandwidth refers to the maximum data rate at which a communication link between two nodes can transmit data. The bandwidth of each link can be obtained using network measurement tools or from the configuration information of network devices. For example, the bandwidth of a fiber optic link may be as high as 10 Gbps or higher, while the bandwidth of some wireless links may be relatively low.
[0083] Similarly, link bandwidth will change dynamically as network usage changes. This information can be continuously monitored and updated to ensure that the real-time bandwidth status of the link is reflected in the dynamic network topology diagram.
[0084] In order to incorporate the node load information into the calculation of the path cost, the load information needs to be quantified. Specifically, corresponding weights can be set according to the different load indicators of the node, and then a comprehensive load score can be calculated. For example, assuming that the weight of CPU utilization is 0.6, the weight of memory utilization is 0.3, and the weight of processing queue length is 0.1, then the node load score = 0.6 × CPU utilization + 0.3 × memory utilization + 0.1 × processing queue length (the utilization and queue length here need to be normalized so that their values range from 0 to 1).
[0085] The effect of link bandwidth on path cost is usually inverse, that is, the larger the bandwidth, the smaller the path cost. An inverse proportional function or other appropriate mathematical model can be used to convert link bandwidth into a cost factor. For example, path cost factor (link bandwidth) = constant / link bandwidth, where the constant can be adjusted according to the actual network situation.
[0086] Furthermore, for each path, the load scores of all nodes involved in the path and the cost factors of the links are accumulated or weighted to obtain the path cost of the path. For example, suppose a path passes through 3 nodes and 2 links, the node load scores are L1, L2, and L3, and the link cost factors are B1 and B2, respectively. The weight of the node load is α, and the weight of the link bandwidth is β (α+β=1). Then the path cost of the path = α×(L1+L2+L3)+β×(B1+B2).
[0087] The path cost of each path can be calculated using the above method.
[0088] The present invention can determine the path cost by comprehensively considering node load information and link bandwidth, and can more comprehensively evaluate the advantages and disadvantages of each path. Compared with considering only a single factor, this method can avoid selecting paths that have sufficient link bandwidth but too high node load to process data in a timely manner, thereby selecting a truly efficient transmission path.
[0089] In one embodiment of the present invention, the method further includes: obtaining data flow of data involved in historical tasks in a preset time period; determining data flow of data involved in tasks to be transmitted in a next time period based on the data flow of data involved in historical tasks in the preset time period; and adjusting the number of GPU servers in the VRB transmission network based on the data flow of data involved in tasks to be transmitted in the next time period.
[0090] In an embodiment of the present invention, a data traffic monitoring module can be deployed at each node in the VRB transmission network to collect data traffic information of data involved in historical tasks within a preset time period. The monitoring module can record detailed information such as the amount of data passing through the node at each time point and the data transmission direction. The preset time period can be set according to actual needs, for example, in units of hours or days, to facilitate subsequent data analysis.
[0091] Data analysis algorithms can be used to analyze the acquired historical data traffic. Common analysis methods include time series analysis and machine learning prediction algorithms. By analyzing the changing trends and periodic patterns of historical data traffic and combining them with business changes, the data traffic involved in the transmission tasks in the next time period can be predicted. For example, if historical data shows that data traffic increases significantly from 10 a.m. to 12 a.m. every day, and according to business arrangements, there are new business activities in this period the next day, the system will predict that the data traffic in this period the next day will increase significantly.
[0092] Furthermore, the predicted data flow of the data involved in the task to be transmitted in the next time period can be compared with the carrying capacity of the current VRB transmission network. If the predicted data flow exceeds the capacity that can be carried by the current number of GPU servers, the operation of adding GPU servers can be automatically started, such as enabling new GPU servers from the spare server pool, or dynamically renting additional GPU servers through cloud services to meet the needs of data processing and transmission; conversely, if the predicted data flow is low, in order to avoid resource waste, some GPU servers with low load can be appropriately shut down, and the data involved in the task can be concentrated on a few servers for processing, thereby improving resource utilization efficiency.
[0093] In one example, in an Internet video live streaming platform, the VRB transmission network is responsible for processing and transmitting a large amount of live video data. Before the live broadcast event begins, the platform will predict the data traffic involved in the data to be transmitted in the next time period (such as within 2 hours of the live broadcast) based on the historical data traffic of similar live broadcast events in the past, combined with factors such as the expected number of viewers of this live broadcast and the intensity of publicity and promotion. If it is predicted that the data traffic will increase significantly, the platform will increase the number of GPU servers in the VRB transmission network in advance to ensure that the live video stream can be processed and transmitted smoothly, avoiding situations such as lag and delay that affect the user's viewing experience; when the live broadcast ends, the data traffic decreases, and the platform will reduce the number of GPU servers based on the new predicted data to reduce operating costs.
[0094] The present invention can dynamically adjust the number of GPU servers according to data traffic, avoiding over-configuration or under-configuration of resources, improving the utilization efficiency of GPU server resources, and reducing operating costs.
[0095] In one embodiment of the present invention, the number of GPU servers in a VRB transmission network is adjusted according to the data flow of data involved in tasks to be transmitted in the next time period, including: if it is detected that the data flow of data involved in tasks to be transmitted in the next time period is greater than a first flow threshold, the number of GPU servers in the VRB transmission network is increased; if it is detected that the data flow of data involved in tasks to be transmitted in the next time period is less than a second flow threshold, the number of GPU servers in the VRB transmission network is reduced, and the first flow threshold is greater than the second flow threshold.
[0096] In this embodiment of the present invention, two traffic thresholds can be pre-set: a first traffic threshold and a second traffic threshold. The first traffic threshold is greater than the second traffic threshold. These two thresholds are determined based on factors such as the performance of GPU servers in the VRB transmission network and network bandwidth, combined with past service data and experience. They are used to determine the level of data traffic and serve as a basis for adjusting the number of GPU servers.
[0097] The data flow prediction value of the data involved in the task to be transmitted in the next time period can be continuously monitored and obtained, and the prediction value can be compared with the first flow threshold and the second flow threshold respectively.
[0098] If it is monitored that the data flow involved in the task to be transmitted in the next time period is greater than the first flow threshold, it indicates that the number of GPU servers in the current network may not be able to meet the data processing and transmission requirements, and there are risks such as backlog of data involved in the task and increased transmission delay. At this time, the operation of increasing the number of GPU servers will be triggered, such as enabling new GPU servers from the spare server pool, or leasing additional GPU server resources through the cloud computing platform to enhance the network's data processing and transmission capabilities.
[0099] If the data flow rate for the task to be transmitted in the next time period is detected to be less than the second flow threshold, it indicates that the GPU server resources in the current network are in excess, and some servers are underloaded, resulting in resource waste. The system will initiate a process to reduce the number of GPU servers, shutting down or deactivating some of the less-loaded GPU servers, and concentrating the data processing for the task on the remaining servers to improve resource utilization efficiency.
[0100] In one example, during a major promotion event on an e-commerce platform, such as during shopping festivals like "Double Eleven" and "618", a large number of users flock to the platform to browse products, place orders, and perform other operations, which generates huge data traffic. Before the start of the major promotion event, the e-commerce platform's system will predict the data traffic involved in the data to be transmitted in the next time period (such as the peak period on the day of the event) based on historical promotion data, current marketing promotion efforts, and other factors. If the predicted data traffic exceeds the first traffic threshold, the number of GPU servers can be increased in the VRB transmission network in advance to ensure that user requests and order data can be processed quickly, ensuring the smooth progress of the shopping process and avoiding problems such as slow page loading and order failures. When the major promotion event ends, the number of user visits and data traffic drops significantly. If it is monitored that the data traffic in the next time period is less than the second traffic threshold, the number of GPU servers can be reduced to reduce operating costs, while ensuring that the remaining servers are in a reasonable load state to maintain the normal operation of the platform.
[0101] By setting dual traffic thresholds, the present invention can accurately and dynamically adjust the number of GPU servers according to changes in data traffic, so that network resources are closely matched with actual business needs, avoiding resource waste or shortage, and improving resource utilization efficiency.
[0102] In one embodiment of the present invention, the transmission parameters include at least one of computing power information, load information, network bandwidth, and transmission delay.
[0103] In an embodiment of the present invention, computing power information represents the server's ability to execute data involved in computing tasks for a GPU server in a VRB transmission network. For example, the more cores a GPU has and the higher its main frequency, the more data it can process and the more data involved in computing tasks it can execute per unit time, and the stronger its computing power. In the data involved in deep learning tasks, a GPU server with high computing power can complete the data involved in model training and inference tasks more quickly. In a video transcoding scenario, a server with strong computing power can complete video encoding and conversion in a shorter time. Computing power information is an important indicator for evaluating the processing power of a GPU server and plays a key role in reasonably allocating data involved in computing tasks and ensuring that the data involved in the tasks can be completed within the specified time. After the system obtains the computing power information of each GPU server, it can preferentially allocate data involved in complex computing tasks to servers with stronger computing power, thereby improving overall computing efficiency.
[0104] Load information reflects the workload of the GPU server or switch in its current state. For GPU servers, load information can include CPU usage, memory usage, and the length of the data queue involved in the task. For example, when the CPU usage reaches above 80%, it means that the server is highly loaded and its ability to process data involved in new tasks is relatively weak. For switches, load information can be reflected in the port's data forwarding volume, queue waiting time, etc. If the data forwarding volume of a switch port is close to its bandwidth limit, then the load on that port is relatively large. Load information helps the system understand the current working status of network devices, avoid allocating too much data involved in tasks to devices that are already in a high-load state, achieve balanced distribution of data involved in tasks, and prevent performance degradation caused by device overload.
[0105] Network bandwidth refers to the amount of data that a network can transmit per unit time, usually measured in Mbps (megabits per second) or Gbps (gigabits per second). In a VRB transmission network, different links may have different bandwidths. For example, the link bandwidth between a GPU server and a switch may be 10Gbps, while the backbone link bandwidth between switches may be higher. Network bandwidth determines the speed at which data can be transmitted in the network. The higher the bandwidth, the faster the data transmission. After the system obtains the network bandwidth information of each link, it can select the link with the larger bandwidth for data transmission to reduce transmission time. At the same time, by rationally planning bandwidth resources, link congestion can be avoided and data can flow smoothly in the network.
[0106] Transmission delay refers to the time it takes for data to travel from a source node to a destination node. In VRB transmission networks, transmission delay is affected by a variety of factors, such as network topology, device processing speed, and link quality. For example, when there are too many intermediate nodes in the network or the link quality is poor, transmission delay increases. Transmission delay is crucial for applications with high real-time requirements (such as online gaming and video conferencing). After the system obtains transmission delay information, it can prioritize paths with lower delays for data transmission to ensure that data reaches the destination node in a timely manner, improving the real-time performance of the application and user experience.
[0107] Transmission parameters such as computing power, load, network bandwidth, and transmission delay can reflect the status of devices and links in the VRB transmission network from different perspectives. By obtaining these parameters, the system can gain a more comprehensive understanding of the network situation, thereby making more reasonable scheduling decisions and optimizing network performance.
[0108] like Figure 3, shows a structural block diagram of another dynamic network topology provided by an embodiment of the present invention. The dynamic network topology may include GPU1, GPU2, switch S0, switch S1, switch S2, and switch S3. There are three paths for GPU1 to transmit data involved in the task to GPU2. The first transmission path is: link A1+link A2; the second transmission path is: link A1+link B1+link B2+link B3; the third transmission path is: link A1+link B1+link C1+link C2+link B3; the path costs of the three paths can be calculated respectively, and then the path with the smallest path cost can be used as the target path, and the data involved in the task can be transmitted through the target path.
[0109] The present invention can generate a dynamic network topology map based on the acquired transmission parameters, which can intuitively and accurately reflect the current network structure and the connection relationship between each node, as well as their real-time status. When the device status in the network changes, such as when a GPU server is overloaded or a link is congested, the dynamic network topology map can be updated in time; the optimal transmission path can be selected for the data involved in the task based on the real-time network status, avoiding the problem of detouring through unnecessary nodes or links when scheduling the data involved in the task in the prior art, thereby improving network transmission efficiency.
[0110] It should be noted that for the sake of simplicity, the method embodiments are described as a series of actions. However, those skilled in the art should be aware that the embodiments of the present invention are not limited by the order of the actions described, because according to the embodiments of the present invention, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in this specification are all preferred embodiments, and the actions involved are not necessarily required by the embodiments of the present invention.
[0111] Reference Figure 4 , shows a VRB-based dynamic network topology awareness scheduling device provided by an embodiment of the present invention, involving multiple GPU servers and multiple switches, and the device includes:
[0112] An acquisition module 201 is configured to acquire transmission parameters of multiple GPU servers and multiple switches;
[0113] A generating module 202 is configured to generate a dynamic network topology map based on transmission parameters of multiple GPU servers and multiple switches;
[0114] The transmission module 203 is used to transmit data involved in the task based on the dynamic network topology graph.
[0115] The present invention discloses a dynamic network topology perception scheduling device based on VRB, which can generate a dynamic network topology map based on acquired transmission parameters, and can intuitively and accurately reflect the current network structure and the connection relationship between each node, as well as their real-time status. When the device status in the network changes, such as when a GPU server is overloaded or a link is congested, the dynamic network topology map can be updated in time; the optimal transmission path can be selected for the data involved in the task based on the real-time network status, avoiding the problem of detouring through unnecessary nodes or links when scheduling the data involved in the task in the prior art, thereby improving network transmission efficiency.
[0116] In one embodiment of the present invention, the transmission module includes:
[0117] The first transmission submodule is used to transmit data involved in the task based on the dynamic network topology diagram and through the VRB protocol.
[0118] In one embodiment of the present invention, the transmission module includes:
[0119] A first determination submodule is used to determine the path costs of multiple paths of data involved in the transmission task in the dynamic network topology graph;
[0120] A second determination submodule is configured to determine a target transmission path based on path costs of the multiple paths;
[0121] The second transmission submodule is used to transmit data involved in the task through a target transmission path.
[0122] In one embodiment of the present invention, the first determining submodule includes:
[0123] a first determining unit, configured to determine load information of nodes involved in multiple paths in a dynamic network topology graph and link bandwidths of the multiple paths;
[0124] The second determining unit is configured to determine the path costs of the multiple paths based on the load information of the nodes involved in the multiple paths and the link bandwidths of the multiple paths.
[0125] In one embodiment of the present invention, the present invention further includes:
[0126] A traffic acquisition module is used to obtain the data traffic of the historical tasks involved in a preset time period;
[0127] A determination module, configured to determine the data flow rate of data involved in tasks to be transmitted in a next time period based on the data flow rate of data involved in historical tasks in a preset time period;
[0128] The adjustment module is used to adjust the number of GPU servers in the VRB transmission network according to the data flow of the data involved in the task to be transmitted in the next time period.
[0129] In one embodiment of the present invention, the adjustment module includes:
[0130] an adding submodule, configured to increase the number of GPU servers in the VRB transmission network if it is detected that the data flow rate of the data involved in the task to be transmitted in the next time period is greater than a first flow threshold;
[0131] The reducing submodule is configured to reduce the number of GPU servers in the VRB transmission network if it is detected that the data flow of the data involved in the task to be transmitted in the next time period is less than the second flow threshold, and the first flow threshold is greater than the second flow threshold.
[0132] In one embodiment of the present invention, the transmission parameters include at least one of computing power information, load information, network bandwidth, and transmission delay.
[0133] The present invention discloses a dynamic network topology perception scheduling device based on VRB, which can generate a dynamic network topology diagram based on acquired transmission parameters, and can intuitively and accurately reflect the current network structure and the connection relationship between each node, as well as their real-time status. When the device status in the network changes, such as when a GPU server is overloaded or a link is congested, the dynamic network topology diagram can be updated in time; the optimal transmission path can be selected for the data involved in the task based on the real-time network status, avoiding the problem of detouring through unnecessary nodes or links when scheduling the data involved in the task in the prior art, thereby improving network transmission efficiency.
[0134] As for the device embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment.
[0135] An embodiment of the present invention further provides an electronic device, including:
[0136] The present invention includes a processor, a memory, and a computer program stored in the memory and capable of running on the processor. When the computer program is executed by the processor, each process of the embodiment of the above-mentioned dynamic network topology-aware scheduling method based on VRB is implemented, and the same technical effect can be achieved. To avoid repetition, it is not repeated here.
[0137] An embodiment of the present invention further provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the various processes of the above-mentioned embodiment of the dynamic network topology-aware scheduling method based on VRB are implemented, and the same technical effects can be achieved. To avoid repetition, they are not described here.
[0138] The various embodiments in this specification are described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same or similar parts between the various embodiments can be referenced to each other.
[0139] Those skilled in the art will appreciate that embodiments of the present invention may be provided as methods, apparatus, or computer program products. Thus, embodiments of the present invention may take the form of entirely hardware embodiments, entirely software embodiments, or embodiments combining software and hardware aspects. Furthermore, embodiments of the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0140] The embodiments of the present invention are described with reference to the flowcharts and / or block diagrams of the methods, terminal devices (systems), and computer program products according to the embodiments of the present invention. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as the combination of the processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing terminal device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing terminal device generate instructions for implementing the process in the flowchart and / or block diagram. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0141] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing terminal device to operate in a specific manner, so that the instructions stored in the computer readable memory produce a manufactured product including an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.
[0142] These computer program instructions can also be loaded onto a computer or other programmable data processing terminal device so that a series of operating steps are executed on the computer or other programmable terminal device to produce a computer-implemented process, thereby providing instructions for executing on the computer or other programmable terminal device to implement the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.
[0143] Although the preferred embodiments of the present invention have been described, those skilled in the art may make additional changes and modifications to these embodiments once they become aware of the basic creative concepts. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments and all changes and modifications that fall within the scope of the embodiments of the present invention.
[0144] Finally, it should be noted that, in this document, relational terms such as first and second, etc., are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprises," "comprising," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or terminal device that includes a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or terminal device. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of additional identical elements in the process, method, article, or terminal device that includes the element.
[0145] The above is a detailed introduction to the VRB-based dynamic network topology-aware scheduling method, device, equipment and storage medium provided by the present invention. Specific examples are used herein to illustrate the principles and implementation methods of the present invention. The description of the above embodiments is only used to help understand the method and core ideas of the present invention. At the same time, for those skilled in the art, according to the ideas of the present invention, there may be changes in the specific implementation methods and application scopes. In summary, the contents of this specification should not be understood as limiting the present invention.
Claims
1. A dynamic network topology-aware scheduling method based on VRB, characterized in that: Involving multiple GPU servers and multiple switches, the method includes: Obtaining transmission parameters of the multiple GPU servers and the multiple switches; generating a dynamic network topology map according to the transmission parameters of the plurality of GPU servers and the plurality of switches; The data involved in the task is transmitted based on the dynamic network topology graph.
2. The VRB-based dynamic network topology-aware scheduling method according to claim 1, characterized in that: The data involved in the task of transmitting the dynamic network topology graph includes: Based on the dynamic network topology diagram, data involved in the task is transmitted through the VRB protocol.
3. The VRB-based dynamic network topology-aware scheduling method according to claim 1, characterized in that: The data involved in the task of transmitting the dynamic network topology graph includes: Determining path costs of multiple paths for transmitting data involved in the task in the dynamic network topology graph; determining a target transmission path according to the path costs of the multiple paths; The data related to the task is transmitted through the target transmission path.
4. The VRB-based dynamic network topology-aware scheduling method according to claim 3, characterized in that: Determining the path costs of multiple paths for transmitting data involved in the task in the dynamic network topology graph includes: Determining load information of nodes involved in multiple paths in the dynamic network topology graph and link bandwidths of the multiple paths; Path costs of the multiple paths are determined according to load information of nodes involved in the multiple paths and link bandwidths of the multiple paths.
5. The VRB-based dynamic network topology-aware scheduling method according to claim 1, characterized in that: Also includes: Get the data flow of the historical tasks involved in the preset time period; Determining the data flow rate of data involved in the task to be transmitted in the next time period based on the data flow rate of the historical task involved in the preset time period; The number of GPU servers in the VRB transmission network is adjusted according to the data flow of the data involved in the task to be transmitted in the next time period.
6. The VRB-based dynamic network topology-aware scheduling method according to claim 5, characterized in that: The adjusting the number of GPU servers in the VRB transmission network according to the data flow of the data involved in the task to be transmitted in the next time period includes: If it is detected that the data flow of the data involved in the task to be transmitted in the next time period is greater than a first flow threshold, increasing the number of the GPU servers in the VRB transmission network; If it is detected that the data flow of the data involved in the task to be transmitted in the next time period is less than a second flow threshold, the number of the GPU servers in the VRB transmission network is reduced, and the first flow threshold is greater than the second flow threshold.
7. The VRB-based dynamic network topology-aware scheduling method according to claim 1, characterized in that: The transmission parameters include at least one of computing power information, load information, network bandwidth, and transmission delay.
8. A dynamic network topology awareness scheduling device based on VRB, characterized in that: Involving multiple GPU servers and multiple switches, the device includes: An acquisition module, configured to acquire transmission parameters of the multiple GPU servers and the multiple switches; A generating module, configured to generate a dynamic network topology map according to the transmission parameters of the plurality of GPU servers and the plurality of switches; A transmission module is used to transmit data involved in the task based on the dynamic network topology diagram.
9. An electronic device, characterized in that: include: A processor, a memory, and a computer program stored in the memory and capable of running on the processor, wherein when the computer program is executed by the processor, the steps of the VRB-based dynamic network topology-aware scheduling method according to any one of claims 1 to 7 are implemented.
10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the VRB-based dynamic network topology-aware scheduling method according to any one of claims 1 to 7 are implemented.
Citation Information
Cited By
Cluster node scheduling method, GPU server scheduling method, and device
WO2026026837A1