Resource allocation result determination method, electronic equipment and storage medium
By generating a topology diagram that describes the interconnection relationship of the server hardware structure, optimizing GPU resource allocation, the problem of unbalanced resource allocation in multi-GPU collaborative tasks is solved, and the accuracy and computing performance of resource allocation are improved.
Patent Information
- Application Number
- CN202510397673.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-31
- Publication Date
- 2025-07-18
AI Technical Summary
In the prior art, GPU resource allocation has an unbalanced problem in multi-GPU collaborative tasks, resulting in low accuracy of resource allocation results.
By generating a first topology map and a second topology map describing the interconnection relationship of the server hardware structure, a recommended resource list is generated based on these topology maps, GPU resource allocation is optimized, and communication performance is ensured to achieve optimal communication performance.
It improves the accuracy of GPU resource allocation and overall computing performance, reduces communication delay and resource waste, and improves the execution efficiency of AI tasks.
Smart Images

Figure CN120335998A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of data processing, and in particular, to a method for determining a resource allocation result, an electronic device, and a storage medium. Background Art
[0002] With the wide popularization of AI (Artificial Intelligence) applications, the Kubernetes cluster, as the mainstream resource management framework, is responsible for the resource allocation tasks required in the process of AI tasks.
[0003] In related technologies, the allocation of GPU (Graphics Processing Unit) resources is usually achieved by random allocation of resource management plugins running on servers. However, in multi-GPU collaborative tasks, there are usually large differences in the communication efficiency between GPU boards at different positions. Therefore, the method of randomly allocating GPU resources may result in uneven resource allocation, thus leading to the technical problem of low accuracy of GPU resource allocation results. Summary of the Invention
[0004] This application provides a method for determining a resource allocation result, an electronic device, and a storage medium, so as to at least solve the problem of low accuracy of resource allocation results that occurs in the process of resource allocation in related technologies.
[0005] According to one aspect of the embodiments of this application, a method for determining a resource allocation result is provided, including: in response to a target resource request, obtaining a default resource list composed of combined GPU resources required for executing a target task; obtaining a first topology map for describing the interconnection relationship of the hardware structure of a server, where the first topology map includes a second topology map for describing the interconnection relationship between at least one central processing unit and at least one GPU in the hardware structure; generating a recommended resource list composed of recommended combined GPU resources required for executing the target task based on the first topology map and the second topology map; and in the case that the recommended combined GPU resources in the recommended resource list include the combined GPU resources in the default resource list, determining to execute the target task according to the resource allocation result indicated by the combined GPU resources.
[0006] According to another aspect of the embodiments of the present application, there is also provided an apparatus for determining a resource allocation result, including: a first operation unit, configured to obtain a default resource list composed of combined resources of graphics processors required for executing a target task in response to a target resource request; a first processing unit, configured to obtain a first topology map for describing the interconnection relationship of the hardware structure of a server, where the first topology map includes a second topology map for describing the interconnection relationship between at least one central processing unit and at least one graphics processor in the hardware structure; a second processing unit, configured to generate a recommended resource list composed of recommended combined resources of graphics processors required for executing the target task based on the first topology map and the second topology map; a third processing unit, configured to determine to execute the target task according to the resource allocation result indicated by the combined resources of graphics processors when the recommended combined resources of graphics processors in the recommended resource list include the combined resources of graphics processors in the default resource list.
[0007] According to yet another aspect of the embodiments of the present application, there is also provided an electronic device, including a memory and a processor. A computer program is stored in the memory, and the processor is configured to execute the steps of any one of the above methods for determining a resource allocation result through the computer program.
[0008] According to yet another aspect of the embodiments of the present application, there is also provided a computer-readable storage medium storing a computer program, where the computer program is configured to execute the steps of any one of the above methods for determining a resource allocation result when running.
[0009] According to yet another aspect of the embodiments of the present application, there is provided a computer program product or a computer program. The computer program product or the computer program includes computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the steps of any one of the above methods for determining a resource allocation result.
[0010] Through the above embodiments provided by the present application, based on a deep analysis of the interconnection relationship of the hardware structure of the server, a first topology diagram and a second topology diagram representing the deep connection relationship among the central processing unit, the graphics processing unit, and the switch under the server are generated; based on these two topology diagrams, a recommended resource list with optimal communication performance is determined; by comparing the recommended resource list with the default resource list composed of the graphics processing unit combined resources randomly allocated based on the resource management plug-in, it is determined whether the resource allocation result of the random allocation meets the expected requirements. In other words, by deeply analyzing the complex interconnection relationship of the server hardware structure, a GPU resource allocation result closer to the actual performance requirements is given, and when the recommended GPU resource allocation result includes the random resource allocation result, it is determined that the random GPU resource allocation result is the optimal result. In this way, the technical problem of low accuracy of the resource allocation result caused by the inability to detect the rationality of the random resource allocation result in the related art can be solved. BRIEF DESCRIPTION OF THE DRAWINGS
[0011] To more clearly illustrate the embodiments of the present application, the drawings required for use in the embodiments will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0012] Figure 1 FIG. is a schematic diagram of an application scenario of a method for determining a resource allocation result according to an embodiment of the present application.
[0013] Figure 2 FIG. is a schematic flowchart of an optional method for determining a resource allocation result according to an embodiment of the present application.
[0014] Figure 3 FIG. is a schematic diagram of an optional method for determining a resource allocation result according to an embodiment of the present application.
[0015] Figure 4 FIG. is a schematic diagram of an optional default resource list and recommended resource list according to an embodiment of the present application.
[0016] Figure 5 FIG. is a schematic overall flowchart of an optional method for determining a resource allocation result according to an embodiment of the present application.
[0017] Figure 6 FIG. is a structural block diagram of an optional device for determining a resource allocation result according to an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0018] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments of the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.
[0019] It should be noted that in the description of the present application, the terms "include", "comprise" or any other variation thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article or device. The terms "first", "second", etc. in the present application are used to distinguish similar objects and not to describe a particular order or sequence.
[0020] In order to enable those skilled in the art of the present technology to better understand the solution of the present application, the present application will be further described in detail below in conjunction with the accompanying drawings and specific embodiments.
[0021] According to one aspect of the embodiments of the present application, a method for determining a resource allocation result is provided. Optionally, in this embodiment, the above method for determining a resource allocation result can be but is not limited to being applied to a hardware scenario as Figure 1 shown, where the server device may include one or more ( Figure 1 only one is shown in the figure) processors 102 (the processors 102 may include but are not limited to processing devices such as a microprocessor MCU or a programmable logic device FPGA) and a memory 104 for storing data. Among them, the above server device may further include a transmission device 106 for communication functions and an input / output device 108. Those of ordinary skill in the art can understand that Figure 1 the structure shown is only schematic and does not limit the structure of the above server device. For example, the server device may further include more or fewer components than Figure 1 shown in the figure, or have a different configuration from Figure 1 shown in the figure.
[0022] The memory 104 can be used to store computer programs, such as software programs and modules of application software, like the computer program corresponding to the control method of the server fan in the embodiments of the present application. The processor 102 executes various functional applications and data processing by running the computer program stored in the memory 104, that is, implements the above method. The memory 104 may include high-speed random access memory, and may also include non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memories. In some instances, the memory 104 may further include memories remotely disposed relative to the processor 102, and these remote memories can be connected to the server device through a network. Examples of the above network include but are not limited to the Internet, enterprise intranet, local area network, mobile communication network, and their combinations.
[0023] The transmission device 106 is used to receive or send data via a network. Specific examples of the above network may include a wireless network provided by the communication provider of the server device. In one instance, the transmission device 106 includes a network adapter (Network Interface Controller, abbreviated as NIC), which can be connected to other network devices through a base station and thus communicate with the Internet. In one instance, the transmission device 106 can be a radio frequency (RF) module, which is used to communicate with the Internet wirelessly.
[0024] The embodiments of the present application can be, but are not limited to, applicable to scenarios with GPU communication link selection requirements. For example, in an AI language large model, due to insufficient memory on 1 GPU board, multiple GPU boards are required to jointly deploy. At this time, multiple boards need to cooperate, and the communication between GPU boards is an important factor affecting performance. On different servers, the communication efficiency between GPU boards is often inconsistent. Then it is necessary to determine whether the communication efficiency of the currently selected GPU boards is the current optimal. And the object detection tool in the embodiments of the present application recommends the current optimal communication link according to the actual server configuration, so as to further verify whether the currently selected GPU board or the combination of GPU boards is reasonable, thereby assisting in optimizing the resource allocation of the GPUDevice-plugin plug-in, improving the communication performance between GPU boards, and further improving the overall performance.
[0025] The method for determining the resource allocation result in the embodiments of the present application can be executed by the server device, or can be executed by the server device in combination with at least one of the terminal devices (which can also be understood as the input / output device 108). Among them, the method for determining the resource allocation result in the embodiments of the present application executed by the terminal device can also be executed by the client installed thereon.
[0026] Taking the determination method of the resource allocation result in this embodiment executed by the server as an example, Figure 2 It is a schematic flowchart of an optional method for determining a resource allocation result according to an embodiment of the present application. As Figure 2 shown, the process of this method may include steps S202 to S208.
[0027] Step S202, in response to a target resource request, obtain a default resource list composed of the graphics processor combined resources required to execute the target task.
[0028] Step S204, obtain a first topology map for describing the interconnection relationship of the hardware structure of the server, where the first topology map includes a second topology map for describing the interconnection relationship between at least one central processor and at least one graphics processor in the hardware structure.
[0029] Step S206, based on the first topology map and the second topology map, generate a recommended resource list composed of the recommended combined resources of the graphics processors required to execute the target task.
[0030] Step S208, in the case where the recommended combined resources of the graphics processors in the recommended resource list include the combined resources of the graphics processors in the default resource list, determine to execute the target task according to the resource allocation result indicated by the combined resources of the graphics processors.
[0031] The method for determining the resource allocation result in this embodiment can be applied to task scenarios that require multiple GPU boards to cooperate in execution. For example, it can be applied to the execution scenario of AI tasks based on AI large language models. In actual application scenarios, as the structure of AI large language models becomes more and more complex and the number of model parameters also increases, multiple GPU boards need to cooperate with each other to efficiently complete AI tasks.
[0032] It should be noted that the number of CPUs in the server is different, and the situation of the PCIe switch (hereinafter referred to as the switch) is also relatively complex. For example, some GPUs come with a PCIe switch, there are independent PCIe switches, and the number of PCIe switches also varies.
[0033] However, during the collaborative work of the above multi-GPU boards, the communication problem between GPU boards is a key factor affecting the task execution efficiency, and there are obvious differences in the communication efficiency between GPU boards under different CPUs and different PCIe switches. For example, the communication efficiency between two GPU boards under the same CPU's NUMA (Non-Uniform Memory Access) architecture is different from that between two GPU boards under different CPUs. Similarly, the communication efficiency between two GPU boards under the same PCIe switch is also different from that between two GPU boards under different PCIe switches.
[0034] In addition, multiple GPU boards under the same PCIe switch usually also form groups. Generally, the communication efficiency between two GPU boards in the same group is better than that between two GPU boards under the same PCIe switch but in different groups.
[0035] It can be seen that during the GPU resource allocation process, the connection relationship and position relationship of multi-GPU boards in the hardware structure will affect the task execution efficiency of AI tasks executed through GPU collaboration.
[0036] In related technologies, usually, according to the GPU Device-plugin plugin (which can also be understood as a GPU device plugin or a resource management plugin), a GPU combination is randomly allocated from idle GPU boards, and the GPU resources corresponding to the GPU combination are allocated to the current task, such as the current AI task.
[0037] However, the method of randomly allocating GPU combinations by running the above Device-plugin plugin does not consider the connection relationship between idle GPU boards and the PCIe switch, nor the connection relationship between idle GPU boards and the CPU, and even less whether the combination of idle GPU boards meets the expected requirements of the optimal GPU resource combination. Therefore, the randomly allocated GPU boards may be multiple GPU boards under different PCIe switches (assuming there are multiple GPU boards under the same PCIe switch), or multiple GPU boards connected by different PCIe switches under different CPUs, etc., resulting in a low communication efficiency between GPU boards and causing an unreasonable problem with the GPU resource allocation result.
[0038] Based on the above problems, a method for determining resource allocation results is proposed in an embodiment of the present application, which aims to assist in detecting whether the default allocation of GPU board resources under the server by the GPU Device-plugin plug-in is the optimal solution. Specifically, the GPU Device-plugin plug-in allocates more reasonable GPU resources through overall resource integration optimization, thereby improving the overall performance of the server.
[0039] Specifically, by deploying a resource management plug-in and a target detection tool on the server, and upon receiving a target resource request, the resource management plug-in is run to allocate a random number of GPU resources, for example, to GPU0 and GPU1, and then the default resource list includes a resource combination of GPU0 and GPU1.
[0040] In this embodiment, the resource management plug-in on the server can be run in response to the target resource request, and the graphics processor combination resource in the above default resource list can be obtained, and the graphics processor combination resource is used to execute the target task. At the same time, the target detection tool can be run on the server, and a first topology map for describing the interconnection relationship of the hardware structure of the server can be generated.
[0041] In order to detect whether the GPU resources randomly allocated by the resource management plug-in are reasonable, the target detection tool on the server is run. During this process, all PCIe devices on the server are detected, some PCIe devices with PCIe switch are screened out, and a server PCIe tree (which can also be understood as the first topology diagram) is generated. The PCIe tree is used to represent the interconnection relationship between the hardware structures of the server.
[0042] Among them, PCIe (Peripheral Component Interconnect Express) is a high-speed serial computer expansion bus standard used to connect various peripherals in the computer, such as graphics cards, solid-state drives, network interfaces, etc. It uses a point-to-point connection method, and each device has its own channel bandwidth, supporting high-bandwidth, low-latency data transmission.
[0043] Based on the above understanding of PCIe, PCIe devices can have but are not limited to the following functions:
[0044] (1) High bandwidth and low latency: Because PCIe devices support multiple channel widths, bandwidth continues to increase with version upgrades, which enables it to meet the needs of high-performance computing and data-intensive applications;
[0045] (2) Multi-device expansion: Multiple devices can be connected through a PCIe switch to support complex system structures;
[0046] (3) Hot Plug and Power Management: PCIe supports hot plug functionality, allowing devices to be replaced while the system is running, and also has an efficient power management mechanism.
[0047] The first topology diagram formed by the interconnection relationship of the hardware structure under the server can be referred to Figure 3 as shown, including at least one CPU, and there may be at least one PCIe switch (which can also be understood as a switch) deployed under each CPU, and there may be multiple groups of GPU boards connected under each PCIe switch. For example, there are at least 2 groups of GPU boards connected under PCIe switch1.
[0048] Starting from each GPU board and searching upward along the Figure 3 shown tree-like topology diagram until reaching the RC (Root Complex) of PCIe, finally stitching together the PCIe tree of all GPUs in the server (which can also be understood as the second topology diagram). In other words, one GPU board corresponds to one second topology diagram, and this second topology diagram is used to describe the interconnection relationship between this GPU board and at least one central processing unit (CPU).
[0049] Based on the first topology diagram and the second topology diagram, generate a recommended combination resource of graphics processors (abbreviated as recommended GPU resources), and generate a recommended resource list based on the recommended GPU resources. For example, the recommended resource list may include various combinations of GPU boards, such as GPU0 and GPU1, GPU1 and GPU7, GPU2 and GPU3, etc.
[0050] Among them, the first topology diagram is a diagram for comprehensively describing the interconnection relationship of the server hardware structure, including but not limited to the topological structure of the connections between all hardware devices such as CPUs, GPUs, and PCIe switches; the second topology diagram, as a sub-diagram of the first topology diagram, focuses on the interconnection relationship between the CPU and the GPU, and describes the communication path and possible performance influencing factors between the GPU card and the CPU.
[0051] The central processing unit (CPU) is the main computing unit of the server, which can but is not limited to be responsible for executing instructions and processing data, and works in coordination with other hardware resources such as GPUs to support complex computing tasks; the switch (here referring to the PCIe switch) is a device for expanding the internal PCIe bus of the server, which can improve the scalability and communication efficiency of GPU resources, and especially plays an important role when multiple GPUs work together.
[0052] A Graphics Processing Unit (GPU) can be, but is not limited to, a dedicated processor for graphics processing and large-scale parallel computing. In the field of AI, especially in deep learning tasks, GPUs have become key acceleration components due to their powerful parallel computing capabilities.
[0053] After comparing the recommended resource list and the default resource list, when it is determined that the GPU combination in the recommended resource list includes the GPU combination in the default resource list, it is determined that the GPU resources randomly allocated by the traditional resource management plugin are the optimal GPU resource combination. For example, in the above embodiment, the recommended resource list also includes the combination of GPU0 and GPU1 in the default resource list.
[0054] In a specific embodiment, the above object detection tool can be used, but is not limited to, through the following steps to detect whether the GPU resources randomly allocated by the resource management plugin are optimal:
[0055] S11, Assume a server equipped with a GPU board and deploy the GPU-related software stack. On the GPU server, input the required number N of GPU resources randomly through the GPU Device-plugin plugin, and record the actually allocated GPU resource list A;
[0056] Specifically, reference can be made to Figure 4 the GPU resource list A shown.
[0057] S12, Run the object detection tool, input the above-mentioned number N of GPU resources and the occupied GPU resources, and obtain the recommended allocated GPU resource list B;
[0058] Specifically, reference can be made to Figure 4 the GPU resource list B shown, which includes multiple groups of GPU resource combinations, such as GPU0 and GPU1, GPU2 and GPU3, GPU0 and GPU7, etc.
[0059] S13, Determine whether list A is a subset of list B. If so, the GPU resources randomly allocated by the GPU Device-plugin plugin this time meet the expectations; otherwise, it means that there may be an optimization control in the GPU Device-plugin plugin and further verification is required.
[0060] Figure 4 The GPU resource list B shown includes the same group of GPU resource combinations as in the GPU resource list A, that is, both resource lists include the combination of GPU0 and GPU1. Therefore, the GPU resource combination randomly generated by the resource management plugin is the optimal GPU resource combination.
[0061] Through the embodiments provided in this application, based on a deep analysis of the interconnection relationship of the hardware structure of the server, a first topology diagram and a second topology diagram representing the deep connection relationship between the central processing unit, the graphics processing unit, and the switch under the server are generated; based on these two topology diagrams, a recommended resource list with optimal communication performance is determined; by comparing the recommended resource list with the default resource list composed of the graphics processing unit combined resources randomly allocated based on the resource management plug-in, it is determined whether the randomly allocated resource allocation result meets the expected requirements. In other words, by deeply analyzing the complex interconnection relationship of the server hardware structure, a GPU resource allocation result closer to the actual performance requirements is given, and when the recommended GPU resource allocation result includes the randomly allocated resource allocation result, it is determined that the randomly allocated GPU resource allocation result is the optimal result. In this way, the technical problem of low accuracy of the resource allocation result caused by the inability to detect the rationality of the randomly allocated resource allocation result in the related art can be solved.
[0062] In an exemplary embodiment, generating the recommended resource list composed of the recommended combined resources of the graphics processing unit required to execute the target task based on the first topology diagram and the second topology diagram includes: by searching the first topology diagram, determining at least one switch connected to at least one central processing unit in the hardware structure; based on a group of graphics processing units connected to at least one switch, determining a group of graphics processing units in the hardware structure; by searching the second topology diagram, determining the first part of the graphics processing units that have been occupied in the group of graphics processing units; based on the group of graphics processing units and the first part of the graphics processing units, determining the second part of the graphics processing units in an idle state; based on the second part of the graphics processing units, determining the recommended combined resources of the graphics processing units in the recommended resource list.
[0063] Combined with the description in the above embodiments, by using the first topology diagram scanned by the target detection tool, it is possible to find the PCIe switch connected to at least one CPU on the server. The PCIe switch is an important hub for GPU communication, and its location and quantity are crucial for optimizing GPU resource allocation.
[0064] After determining the PCIe switch connected to the CPU, the target detection tool further analyzes the group of GPUs connected to these PCIe switches. This process allows the detection tool to find out which GPUs have higher communication efficiency under the same PCIe topology path, thus providing key information for optimizing resource allocation.
[0065] The target detection tool utilizes the second topological graph to accurately identify which GPU resources are occupied by the currently running tasks. This precise information is essential for avoiding resource conflicts and optimizing resource utilization. After determining which GPU resources are occupied, the target detection tool can efficiently screen out which GPU boards are still idle and available for new task allocation. This process ensures that the GPU resource combination recommended by the target detection tool does not conflict with existing tasks, while enhancing the dynamic resource allocation ability.
[0066] Based on the identified idle GPU boards, as well as their topological positions, communication efficiency, and task requirements, a recommended GPU resource combination is generated, along with a recommended resource list. This combination aims to optimize GPU - to - GPU communication, improve overall computing performance, and avoid resource waste.
[0067] By leveraging the target detection tool, in - depth optimization of GPU resource allocation is achieved. It not only considers the physical locations of GPU resources but also fully evaluates the resource usage status and communication efficiency, ensuring the rationality and efficiency of the resource allocation scheme. Especially in multi - GPU collaborative tasks, this in - depth resource optimization strategy can significantly enhance the task execution speed, reduce communication latency, improve overall computing performance, while reducing resource waste and energy consumption.
[0068] In an exemplary embodiment, determining the recommended combined resources of the graphics processors in the recommended resource list based on the second part of the graphics processors includes:
[0069] Based on the second part of the graphics processors, the first topological graph, and the second topological graph, determine at least one graphics processor communication link, where one of the at least one graphics processor communication links is used for data transmission between graphics processors and between graphics processors and other expansion devices, and the other expansion devices include at least one switch and at least one central processor;
[0070] Obtain the first bandwidth between the target switch and the target central processor that are in a connected state among at least one switch and at least one central processor;
[0071] Obtain the second bandwidth between any two switches that are in a connected state among at least one switch;
[0072] Based on the first bandwidth and the second bandwidth, determine a set of path weights for each graphics processor communication link in the at least one graphics processor communication link;
[0073] Based on the set of path weights, determine the recommended combined resources of the graphics processors in the recommended resource list.
[0074] After generating the first topology map, the second topology map, and determining the GPU boards in the idle state in the manner of the above embodiments, it is possible but not limited to determine the possible GPU communication links through the following steps:
[0075] S21, Detect all CPUs and NUMA of the server to generate a CPU NUMA tree;
[0076] S22, Detect the NUMA numbers corresponding to each GPU board to generate a GPU NUMA tree;
[0077] S23, Analyze the number and groups of PCIe switches of the PCIe slots on the server where the GPU boards are located;
[0078] S24, When allocating GPU resources using the target detection tool, allocate preferentially according to the resource quantity required by the front end;
[0079] Among them, the front end includes but is not limited to application programs or user interfaces that submit GPU resource requests, or users, etc.
[0080] Preferential allocation includes but is not limited to that when allocating an even number of GPU resources, preferentially allocate GPUs of the same PCIe switch, and when allocating a single one, allocate it from the farthest isolated PCIe switch.
[0081] In addition, when allocating GPU resources, give priority to considering the NUMA of the same CPU because the communication efficiency of the CPU memory under the same NUMA structure is relatively high.
[0082] S25, Form a server PCIe bandwidth tree according to the PCIe tree of the GPU;
[0083] S26, Compare and analyze the PCIe server tree and the GPU tree;
[0084] If there are idle slots under the PCIe switch, detect whether the GPU boards are evenly installed; if not, prompt to install the GPU boards under the same PCIe switch as much as possible or evenly install the GPU boards under the PCIe switch.
[0085] S27, According to the input number of GPUs and the input occupied GPUs, give all optional idle GPU boards in combination with the above configurations;
[0086] S28, Generate available GPU Channels (which can also be understood as GPU communication links, hereinafter referred to as communication links) according to the above-known relevant PCIe trees.
[0087] After determining the available communication links according to the above steps S21 to S28, based on the bi-directional bandwidth of the connection between the CPU and the PCIe switch and the bi-directional bandwidth between the PCIe switches, determine the path weight of each available communication link (which can also be understood as a graphics processor communication link). Finally, according to the path weights, generate the graphics processor recommended combined resources in the recommended resource list, that is, the recommended GPU combined resources.
[0088] In this embodiment, first, based on the idle GPU boards, the server hardware interconnection relationship, and the GPU communication path, determine the communication path (which can also be understood as a GPU communication link) between the GPU board and other key hardware devices. Then, use the target detection tool to determine the first bandwidth and the second bandwidth respectively to quantify the performance of the communication link. Among them, the acquisition of the bandwidth data provides key parameters for the subsequent calculation of the path weights.
[0089] By synthesizing the first bandwidth and the second bandwidth data, and considering other factors that may affect the communication efficiency, such as the physical distance of the communication path, potential communication delays, etc., the target detection tool can assign a path weight value to each communication link. The calculation of the path weights is essentially a comprehensive evaluation system based on the hardware layout and communication performance, and its purpose is to identify the GPU resource combination with the highest communication efficiency.
[0090] Based on the above calculation of the path weights, the target detection tool generates an optimized GPU resource combination, that is, the graphics processor recommended combined resources in the recommended resource list. This resource combination aims to maximize the data transfer speed between GPUs, enhance the collaborative ability of GPU resources, and thus significantly improve the execution efficiency and overall performance of the computing tasks. Through this series of in-depth analysis and intelligent recommendations, not only the communication bottleneck problem that may occur in the traditional GPU resource allocation is solved, but also the efficient utilization and performance improvement of the server hardware resources in AI applications are promoted.
[0091] In an exemplary embodiment, the determining, based on the first bandwidth and the second bandwidth, a set of path weights for each graphics processor communication link in at least one graphics processor communication link includes:
[0092] Successively obtain each graphics processor communication link from at least one graphics processor communication link as the current graphics processor communication link;
[0093] Determine the current central processor and the current switch passed by the current graphics processor communication link from at least one switch and at least one central processor;
[0094] Based on the first bandwidth and the second bandwidth, determine the first sub-path weight between the current central processing units, the second sub-path weight between the current central processing unit and the current switch, and the third sub-path weight between the current switches;
[0095] Sum up the first sub-path weight, the second sub-path weight, and the third sub-path weight to obtain the path weight of the current graphics processing unit communication link.
[0096] After determining the GPU communication link through the method of the above embodiment, accumulate the bidirectional bandwidth of the connection between the central processing unit (CPU) and the PCIe switch and the bidirectional bandwidth between the PCIe switches as the weight sum, and calculate the bandwidth or weight sum passing through the CPU and the PCIe switch, so as to obtain the path weight of the communication link.
[0097] Specifically, select each communication link from all the pre-determined GPU communication links one by one for analysis, and regard each link as the object of the current analysis. For example, determine the CPU and the PCIe switch involved in the current communication link, that is, determine the CPU and / or PCIe switch passed by the current communication link. This process provides key information for the subsequent calculation of the path weight based on the hardware performance, and helps the tool accurately evaluate the actual communication efficiency of the communication link in the complex server hardware environment.
[0098] According to the determined first bandwidth between the CPU and the PCIe switch and the second bandwidth between the PCIe switches, calculate the sub-path weights involved between the current CPUs, between the current CPU and the PCIe switch, and between the current PCIe switches respectively. The determination of the sub-path weight takes into account the bandwidth difference and communication efficiency between the hardware devices in the communication link, and reflects the contribution degree of different components of the communication link to the overall performance through quantitative indicators.
[0099] Obtain the path weight of the current communication link by summing or weighted summing the first sub-path weight, the second sub-path weight, and the third sub-path weight. The calculation of the path weight comprehensively evaluates the communication efficiency of all key hardware devices in the communication link, and provides data support for determining the optimal GPU resource allocation strategy. This process ensures that the GPU combination resources recommended by the tool can maximize the performance of the GPU intercommunication link, improve the execution efficiency of the computing task, reduce the communication delay and energy consumption at the same time, and realize the optimization of resource utilization.
[0100] Through quantitative analysis of the performance of the GPU communication link, the deep optimization of the GPU resource allocation strategy is achieved. By systematically analyzing the hardware composition and bandwidth characteristics of each communication link, the accuracy of resource allocation is ensured, the communication performance between GPUs is improved, and the efficient execution of computing tasks is realized.
[0101] In addition, the technical solution of this application also fully considers the complexity of server hardware, provides an effective solution for the rational utilization and performance improvement of GPU resources, enhances the operation efficiency and economy of the AI cluster, and promotes the wide application of AI technology in various industries.
[0102] In an exemplary embodiment, determining the graphics processor recommended combined resources in the recommended resource list based on a set of path weights includes:
[0103] Based on a set of path weights, score each graphics processor communication link in at least one graphics processor communication link to obtain a set of scores;
[0104] Based on a set of scores, sort at least one graphics processor communication link to obtain the sorted graphics processor communication link;
[0105] Based on at least one switch and at least one non-uniform memory access architecture in the hardware structure, group the sorted graphics processor communication links to obtain N groups of graphics processor communication links, where N is a preset positive integer;
[0106] According to the preset priority judgment condition, determine the group of graphics processor communication links with the highest score from the N groups of graphics processor communication links;
[0107] Based on the combination method between the target graphics processors passed by a set of graphics processor communication links, determine the graphics processor recommended combined resources in the recommended resource list.
[0108] After calculating the path weights of each communication link according to the above embodiment, score each communication link according to the GPU Chnnel (GPU communication link) in combination with the path weights, and sort them from high to low.
[0109] Then, group the scored GPU communication links according to NUMA and PCIe switch to obtain N groups of graphics processor communication links (i.e., GPU communication links), and determine the group of GPU communication links with the highest score from the N groups of GPU communication links according to the preset priority.
[0110] Based on the GPUs passed by the group of GPU communication links with the highest score, determine the graphics processor recommended combined resources. For example, refer to Figure 3As shown in the figure, assume that Mn in the figure represents the nth CPU, and each PCIe switch is connected to multiple groups of GPU boards. Each group of GPU boards includes at least one GPU board, namely GPU1 to GPUn.
[0111] Assume that the group of graphics processor communication links with the highest scores are the GPU1 to GPUn connected to N1.1 under PCIe switch N1 connected to CPU M1, and the GPU1 to GPUn connected to N1.1 under PCIe switch N1 connected to CPU Mn. Then, according to the above calculation method of the path weight of the communication link, calculate the first sub-path weight corresponding to the sub-path passing through CPU M1 and CPU Mn, the second sub-path weight between CPU M1 and CPU Mn and the connected PCIe switch respectively, and the third sub-path weight between PCIe switch N1 and PCIe switch N2.
[0112] By multiplying the above first sub-path weight, second sub-path weight, and third sub-path weight, the score of the current GPU communication link is obtained. And in the same way, calculate the scores of other GPU communication paths to obtain a set of scores.
[0113] After sorting a set of scores in descending order, the sorted graphics processor communication links are obtained. Then, group the sorted graphics processor communication links according to NUMA and PCIe switch. Finally, according to the preset priority, determine a group of GPU communication links with the highest scores from the grouped GPU communication links.
[0114] According to the group of GPU communication links with the highest scores, determine at least one group of GPU resource combinations in the recommended resource list. For example, Figure 4 the 3 groups of GPU combinations in the GPU resource list as shown.
[0115] It should be noted that the above preset priority can be but is not limited to being set based on the performance, stability, and availability of the communication link. Its purpose is to ensure that the finally recommended GPU resource combination reaches the optimal state in multiple aspects.
[0116] It can be seen that the above scoring of each group of GPU communication links based on the path weight realizes the quantitative evaluation of all communication links. This scoring process ensures that the performance advantages and disadvantages of each communication link can be clearly identified and sorted.
[0117] Considering the impact of the PCIe switch and NUMA architecture in the server hardware structure on GPU communication, the sorted communication links are grouped according to these hardware characteristics to form N groups of communication links. The principle of grouping is to keep the communication links under the same PCIe switch or within the same NUMA architecture as much as possible to improve the communication efficiency and overall performance among GPU resources.
[0118] After the grouping is completed, according to the preset priority judgment conditions, the group with the highest score is selected from the N groups of communication links as the optimal link group. This selection process fully considers the performance, stability, and resource availability of the communication links to ensure that the finally recommended GPU resource combination not only achieves the optimal in communication efficiency but also meets other key indicators of business requirements.
[0119] In an exemplary embodiment, the above-mentioned grouping of the sorted graphics processor communication links based on at least one switch and at least one non-uniform memory access architecture in the hardware structure results in N groups of graphics processor communication links, including:
[0120] The first part of the sorted graphics processor communication links that are connected to the same switch and assigned to the same non-uniform memory access architecture is divided into the first group of graphics processor communication links. Among them, multiple graphics processors in the first group of graphics processor communication links assigned to the same non-uniform memory access architecture share local memory with the central processor connected to the same non-uniform memory access architecture;
[0121] The second part of the sorted graphics processor communication links that are connected to the same switch and assigned to different non-uniform memory access architectures is divided into the second group of graphics processor communication links;
[0122] The third part of the sorted graphics processor communication links that are connected to different switches and assigned to the same non-uniform memory access architecture is divided into the third group of graphics processor communication links;
[0123] The remaining graphics processor communication links in the sorted graphics processor communication links except the first group, the second group, and the third group of graphics processor communication links are determined as the fourth group of graphics processor communication links.
[0124] In this embodiment, the scored GPU Channels are grouped according to NUMA and PCIe switch. Those with the same PCIe switch and NUMA are determined as the first group, those with the same PCIe switch but different NUMA are determined as the second group, those with the same NUMA but different PCIe switches are determined as the third group, and the rest are determined as the fourth group.
[0125] By default, the priority of the first group is higher than that of the second group, the priority of the second group is higher than that of the third group, and the priority of the third group is higher than that of the fourth group. According to this priority rule, the group of GPU communication links with the highest score is determined.
[0126] Among them, when grouping the sorted communication links, first, the communication links connected to the same PCIe switch and assigned to the same NUMA architecture are divided into the first group. This means that the GPUs within the group not only share the same PCIe switch resources but also share local memory with the corresponding CPUs under the same NUMA architecture, thus achieving high efficiency and consistency at both the communication link level and the memory access level, providing the best hardware environment for GPU - to - GPU communication.
[0127] The communication links that belong to the same PCIe switch but are assigned to different NUMA architectures will be divided into the second group. This is because although these links share switch resources, due to the differences in NUMA architectures, they share memory with different CPUs, and may be slightly inferior to the first - group links in terms of communication efficiency and data access speed, but they are still a communication link combination with relatively good performance.
[0128] For the communication links connected to different PCIe switches but assigned to the same NUMA architecture, they are divided into the third group. This is because although the communication between GPUs within the group may need to cross switches, since they share the same memory access architecture, in some scenarios, especially in large - scale communication tasks, the third - group communication links can still provide relatively high communication performance and memory access efficiency.
[0129] In addition to the above three groups of links, the remaining communication links will be classified as the fourth group. This may include links connected to different switches and assigned to different NUMA architectures, which means that they have differences in both hardware resources and memory access, so they may show varying degrees of inferiority in terms of communication efficiency and data access speed. However, in certain specific situations, such as when resources are extremely limited, the fourth - group communication links can still be used as alternative solutions.
[0130] By using a target detection tool to deeply analyze the interconnection relationships in the server hardware structure, perform performance evaluation on each GPU link, and intelligently group and hierarchically allocate resource priorities, the overall optimization of the GPU resource allocation strategy is achieved. By quantitatively analyzing the performance of communication links and performing intelligent grouping in combination with hardware and memory access architecture characteristics, it is ensured that the recommended GPU resource combinations are optimized in terms of communication efficiency, data transfer speed, and resource utilization efficiency. The technical solution of this application is not only innovative and practical technically, but also fully considers the performance optimization requirements in complex hardware environments, provides an effective solution for the rational utilization and performance improvement of GPU resources, enhances the operation efficiency and economy of the AI cluster, and further promotes the wide application of AI technology in various industries.
[0131] In an exemplary embodiment, the above method further includes:
[0132] When the graphics processor recommended combined resources in the recommended resource list do not include the graphics processor combined resources in the default resource list and the graphics processor recommended combined resources correspond to at least one graphics processor communication link, determine the third bandwidth of the default graphics processor communication link corresponding to the graphics processor combined resources;
[0133] Determine the fourth bandwidth of the target graphics processor communication link with the highest score among at least one graphics processor communication link, where each graphics processor communication link in at least one graphics processor communication link is scored based on its path weight;
[0134] Compare the third bandwidth and the fourth bandwidth, and when the fourth bandwidth is greater than the third bandwidth, execute the target task according to the target graphics processor combined resources indicated by the target graphics processor communication link.
[0135] After determining the group of GPU communication links with the highest score in the above-described manner of the embodiment, perform a point-to-point bandwidth test on it. If the bandwidth of this group of GPU communication links is higher than the point-to-point bandwidth of the default GPU communication link randomly allocated by using the traditional GPU Device-plugin plug-in, then print the GPU communication link with the highest score; otherwise, prompt that no better GPU communication link is found and synchronously print the recommended GPU communication link with the highest score.
[0136] For example, if there is more than one PCIe switch on the GPU slot of the server, assuming the number of PCIe switches is N, the GPU resources can be divided into N groups, with each PCIe switch having a group under it; confirm the number of CPUs on the server. If there is only one CPU, the impact of the CPU on the resource allocation of the GPU Device-plugin plugin in this column can be ignored; if there is more than one CPU on the server, for example, there are M CPUs, the GPU resources can be divided into M groups. The specific division can refer to NUMA, and those under the same NUMA are in one group. When the number N of PCIe switches and the number M of CPUs are obtained, the allocation principle of GPU resources should preferably be the intersection of N and M, rather than random allocation. Priority should be given to M, followed by N, and the closer N is, the better. When distributing resource allocation requirements from the front-end usage requirements, try to distribute the larger requirements first, and allocate GPU resources under one M first, and try to ensure that the GPU resources under a subsequent M are close so that a large amount can be allocated, and try to avoid allocation under different CPUs.
[0137] Based on the above analysis, it can be seen that this embodiment is built on the basis of in-depth analysis and intelligent recommendation of the GPU communication link, further quantifying the difference in communication performance between the recommended resources and the default resources, providing data support for users to select the optimal GPU resource combination. Specifically, the target detection tool first checks whether there are differences between the recommended resource list and the resources allocated by the system by default, and whether the recommended resource combination is associated with the actual communication link. If this condition is met, it will deeply analyze the communication link under the system default allocation to determine its third bandwidth, that is, the data transmission rate and communication capacity of the default link, providing a benchmark for subsequent performance comparison.
[0138] Then select the target link with the highest score from the recommended communication links, and analyze and determine its fourth bandwidth, that is, the best data transmission ability and communication efficiency in the recommended link. This process scores based on the path weight of the link, ensuring that the scoring result can truly reflect the performance advantages and disadvantages of the communication link, providing a strong basis for optimizing resource allocation.
[0139] Finally, by comparing the third bandwidth and the fourth bandwidth, the difference in communication performance between the recommended scheme and the default scheme can be quantitatively evaluated. If the communication link of the recommended scheme shows obvious advantages in bandwidth, it will guide users to preferentially adopt the recommended GPU combined resources to execute specific tasks. This can not only improve the efficiency of GPU - to - GPU communication, reduce communication latency, but also optimize resource utilization, improve the overall computing performance, save costs for users, and accelerate task execution.
[0140] To more clearly understand the method for determining the above resource allocation results, the following further describes it in conjunction with Figure 5 the overall flowchart shown below.
[0141] S502, Screen out the PCIe devices with switches from the PCIe devices on the server;
[0142] For the description of the PCIe devices, reference can be made to the description in the above embodiments, and details will not be elaborated here.
[0143] S504, Generate a first topology diagram and a second topology diagram;
[0144] Among them, the first topology diagram is a diagram for comprehensively describing the interconnection relationship of the server hardware structure, including but not limited to the topology structure of the connections between all hardware devices such as CPUs, GPUs, PCIe switches, etc.; the second topology diagram, as a sub-diagram of the first topology diagram, focuses on the interconnection relationship between the CPU and the GPU, and describes the communication path between the GPU card and the CPU and possible performance influencing factors.
[0145] S506, Determine whether there are any idle slots under each switch by looking up the first topology diagram;
[0146] S508, Determine whether there are any idle slots under the same switch to which the GPU board is connected by looking up the second topology diagram;
[0147] If it is found that there are idle slots under the switch and there are also idle slots under the same switch to which the GPU board is connected, then execute step S510.
[0148] S510, Generate a GPU communication link based on the determined idle GPUs;
[0149] S512, Determine the path weights between CPUs, the path weights between the CPU and the switch, and the path weights between switches;
[0150] Specifically, reference can be made to the determination processes of the first sub-path weight, the second sub-path weight, and the third sub-path weight in the above embodiments, and details will not be elaborated here.
[0151] S514, Score each GPU communication link based on the respective path weights;
[0152] S516, Determine the optimal GPU resource combination recommended by the target detection tool based on the scored communication links;
[0153] Specifically, reference can be made to the above embodiments for grouping multiple GPU communication links according to NUMA and PCIe switch and determining the GPU communication link with the highest score according to the preset priority.
[0154] S518, perform a bandwidth test on the recommended optimal GPU resource combination and the GPU resource combination randomly allocated by using the resource management plug-in;
[0155] S520, determine whether the GPU resource combination defaultly allocated by the plug-in meets the expected requirements based on the bandwidth test results.
[0156] For example, if the GPU resource combination in the recommended resource list in the above embodiment includes the random GPU resource combination in the default resource list, then determine that the randomly allocated GPU resource combination is the optimal GPU resource combination.
[0157] In an exemplary embodiment, the above method further includes:
[0158] Obtain the hardware structure information of each server in multiple servers and the network topology information between the servers, where the hardware structure information includes the configuration information and connection relationships of the central processing unit, the graphics processing unit, and the switch, and the multiple servers are the servers in the server cluster;
[0159] In response to a target resource request, determine the resource requirements and communication requirements of the graphics processing unit resources required to execute the target task;
[0160] Based on the resource requirements and communication requirements, the internal communication performance of the multiple servers, and the network performance between the servers, execute a multi-objective optimization algorithm to obtain an optimal graphics processing unit resource allocation result, where the multi-objective optimization algorithm is used to determine a resource allocation strategy when simultaneously meeting multiple optimization objectives such as computing performance, communication efficiency, and resource utilization rate;
[0161] Obtain the recommended graphics processing unit combined resources in the recommended resource list;
[0162] In the case that the graphics processing unit recommended combined resources in the recommended resource list include the optimal graphics processing unit combined resources indicated by the optimal graphics processing unit resource allocation result, determine to execute the target task according to the optimal graphics processing unit combined resources.
[0163] The above embodiments are directed to the GPU resource allocation method under the same server. In addition, the target detection tool in the embodiments of the present application can also be applied to the GPU resource allocation in the cross-server resource optimization scenario. The specific process is as follows:
[0164] S31, obtain the hardware structure information of each server in the server cluster;
[0165] Deploy the target detection tool on each server to obtain detailed hardware structure information, including CPU, GPU model, PCIe switch configuration, PCIe link topology, and NUMA topology, etc.
[0166] In addition, network information is also obtained, specifically including the network topology information between servers, including network latency, bandwidth, routing, etc., and the network connection situation within the cluster.
[0167] S32, Server internal communication performance evaluation;
[0168] Use the target detection tool to test and evaluate the communication performance between GPU boards in each server, and establish a communication performance model based on PCIe bandwidth, latency, and NUMA topology.
[0169] At the same time, evaluate the communication performance between servers. For example, evaluate the communication performance between different servers, considering the physical distance between servers, the performance of network devices, and the communication performance within the servers, and establish a comprehensive communication performance evaluation system.
[0170] S33, Task requirement analysis and resource requirement analysis;
[0171] Task requirement analysis: Analyze the resource requirements and communication requirements of AI tasks, including the number of GPUs required, the frequency and data volume of communication between GPUs, etc.
[0172] Cluster resource availability analysis: Real-time monitor the GPU resource usage status of each server in the cluster, including information such as allocated resources, idle resources, and the physical location of resources.
[0173] S34, Implementation of cross-server resource allocation strategy.
[0174] Specifically, it includes two parts: multi-objective optimization algorithm and resource reservation and dynamic adjustment. Among them, the multi-objective optimization algorithm includes but is not limited to the following objectives:
[0175] (1) Maximize communication efficiency: Ensure that the communication latency between the allocated GPU resources is minimized and the communication bandwidth is maximized. Usually, it involves the physical location of GPUs within the server (such as within the same CPU NUMA node, under the same PCIe switch), and the network connection quality between servers (such as low-latency, high-bandwidth network links);
[0176] (2) Maximize computing performance: Ensure that the allocated GPU resources can provide sufficient computing power to meet the needs of AI tasks. This may require considering the model, performance, load status of GPUs, and the performance and load status of other resources such as CPUs and memory;
[0177] (3) Maximize resource utilization: Optimize resource allocation to minimize the idle time of GPU resources and avoid resource waste. At the same time, consider the usage of CPU, memory and other resources to achieve balanced utilization of overall resources;
[0178] (4) Minimizing energy consumption: While meeting performance requirements, reduce energy consumption during GPU resource allocation and use as much as possible.
[0179] Resource reservation and dynamic adjustment can refer to, but is not limited to, reserving the most suitable GPU resources based on predicted resource requirements and communication requirements before the task starts. During the task execution, the allocation of GPU resources is dynamically adjusted based on actual performance feedback to optimize overall performance.
[0180] Through the technical solution in this embodiment, accurate and optimized allocation of GPU resources in the server cluster is achieved. At the same time, the hardware configuration and network topology data of the server cluster are collected and analyzed from a global perspective, combined with the resource requirements and communication requirements of the computing tasks, to ensure that the resource allocation solution not only meets the task requirements, but also achieves the optimization of the overall performance and resource utilization efficiency of the server cluster.
[0181] Through the description of the above implementation methods, those skilled in the art can clearly understand that the method according to the above embodiment can be implemented by means of software plus a necessary general hardware platform, and of course by hardware, but in many cases the former is a better implementation method.
[0182] According to another aspect of the embodiments of the present application, a device for determining a resource allocation result is also provided, and the device for determining a resource allocation result can be used to implement the method for determining a resource allocation result provided in the above-mentioned embodiment, and the description thereof will not be repeated. As used below, the term "module" can implement a combination of software and / or hardware of a predetermined function. Although the device described in the following embodiments is preferably implemented in software, the implementation of hardware, or a combination of software and hardware, is also possible and conceived.
[0183] Figure 6 is a structural block diagram of an optional device for determining resource allocation results according to an embodiment of the present application, such as Figure 6 As shown in , the device for determining the resource allocation result includes a first operating unit 602 , a first processing unit 604 , a second processing unit 606 and a third processing unit 608 .
[0184] A first operating unit 602 is used to obtain a default resource list consisting of graphics processor combination resources required to execute a target task in response to a target resource request;
[0185] The first processing unit 604 is configured to obtain a first topology map for describing the interconnection relationship of the hardware structure of the server, where the first topology map includes a second topology map for describing the interconnection relationship between at least one central processing unit and at least one graphics processing unit in the hardware structure;
[0186] The second processing unit 606 is configured to generate a recommended resource list composed of recommended combined resources of graphics processing units required to execute a target task based on the first topology map and the second topology map;
[0187] The third processing unit 608 is configured to, when the recommended combined resources of graphics processing units in the recommended resource list include the combined resources of graphics processing units in the default resource list, determine to execute the target task according to the resource allocation result indicated by the combined resources of graphics processing units.
[0188] It should be noted that the first operation unit 602 in this embodiment may be used to execute the above step S202, the first processing unit 604 in this embodiment may be used to execute the above step S204, the second processing unit 606 in this embodiment may be used to execute the above step S206, and the third processing unit in this embodiment may be used to execute the above step S208.
[0189] Through the embodiments provided by the present application, based on the in-depth analysis of the interconnection relationship of the hardware structure of the server, a first topology map and a second topology map representing the deep connection relationship between the central processing unit, the graphics processing unit, and the switch under the server are generated; based on these two topology maps, a recommended resource list with the optimal communication performance is determined; by comparing the recommended resource list with the default resource list composed of randomly allocated combined resources of graphics processing units based on the resource management plug-in, it is determined whether the randomly allocated resource allocation result meets the expected requirements, so as to determine whether it is necessary to optimize the GPU resource allocation result and improve the accuracy of the GPU resource allocation result.
[0190] In an exemplary embodiment, the above second processing unit 606 includes: a first search module for determining at least one switch connected to at least one central processing unit in the hardware structure by searching the first topology map; a first processing module for determining a set of graphics processing units in the hardware structure based on a set of graphics processing units connected to at least one switch; a second search module for determining a first part of the graphics processing units that have been occupied in the set of graphics processing units by searching the second topology map; a second processing module for determining a second part of the graphics processing units in the idle state based on the set of graphics processing units and the first part of the graphics processing units; a third processing module for determining the recommended combined resources of graphics processing units in the recommended resource list based on the second part of the graphics processing units.
[0191] In an exemplary embodiment, the above-mentioned third processing module includes: a first processing sub-module, configured to determine at least one graphics processor communication link based on a second part of the graphics processor, a first topology map, and a second topology map, where one of the at least one graphics processor communication links is used for data transmission between graphics processors and between a graphics processor and other expansion devices, and the other expansion devices include at least one switch and at least one central processing unit; a first acquisition sub-module, configured to acquire a first bandwidth between a target switch and a target central processing unit that are in a connected state among the at least one switch and the at least one central processing unit; a second acquisition sub-module, configured to acquire a second bandwidth between any two switches that are in a connected state among the at least one switch; a second processing sub-module, configured to determine a set of path weights for each graphics processor communication link in the at least one graphics processor communication link based on the first bandwidth and the second bandwidth; a third processing sub-module, configured to determine a graphics processor recommended combination resource in a recommended resource list based on the set of path weights.
[0192] In an exemplary embodiment, the above-mentioned third processing module includes: a third acquisition sub-module, configured to sequentially acquire each graphics processor communication link in the at least one graphics processor communication link as the current graphics processor communication link; a fourth processing sub-module, configured to determine a current central processing unit and a current switch that the current graphics processor communication link passes through from the at least one switch and the at least one central processing unit; a fifth processing sub-module, configured to determine a first sub-path weight between the current central processing units, a second sub-path weight between the current central processing unit and the current switch, and a third sub-path weight between the current switches based on the first bandwidth and the second bandwidth; a sixth processing sub-module, configured to sum the first sub-path weight, the second sub-path weight, and the third sub-path weight to obtain the path weight of the current graphics processor communication link.
[0193] In an exemplary embodiment, the above-mentioned third processing module includes: a scoring sub-module, configured to score each graphics processor communication link in at least one graphics processor communication link based on a set of path weights to obtain a set of scores; a sorting sub-module, configured to sort at least one graphics processor communication link based on the set of scores to obtain the sorted graphics processor communication links; a grouping sub-module, configured to group the sorted graphics processor communication links based on at least one switch and at least one non-uniform memory access architecture in the hardware structure to obtain N groups of graphics processor communication links, where N is a preset positive integer; a seventh processing sub-module, configured to determine, according to a preset priority judgment condition, a group of graphics processor communication links with the highest score from the N groups of graphics processor communication links; an eighth processing sub-module, configured to determine a recommended combined resource of graphics processors in the recommended resource list based on the combination mode between the target graphics processors passed by a set of graphics processor communication links.
[0194] In an exemplary embodiment, the above-mentioned device further includes: a first partitioning sub-module, configured to partition a first part of the graphics processor communication links that are connected to the same switch and are assigned to the same non-uniform memory access architecture in the sorted graphics processor communication links into a first group of graphics processor communication links, where multiple graphics processors in the first group of graphics processor communication links assigned to the same non-uniform memory access architecture share local memory with a central processor connected to the same non-uniform memory access architecture; a second partitioning sub-module, configured to partition a second part of the graphics processor communication links that are connected to the same switch and are assigned to different non-uniform memory access architectures in the sorted graphics processor communication links into a second group of graphics processor communication links; a third partitioning sub-module, configured to partition a third part of the graphics processor communication links that are connected to different switches and are assigned to the same non-uniform memory access architecture in the sorted graphics processor communication links into a third group of graphics processor communication links; a ninth processing sub-module, configured to determine the remaining graphics processor communication links in the sorted graphics processor communication links except the first group of graphics processor communication links, the second group of graphics processor communication links, and the third group of graphics processor communication links as a fourth group of graphics processor communication links.
[0195] In an exemplary embodiment, the above device further includes: a fourth processing unit, configured to determine a third bandwidth of a default graphics processor communication link corresponding to a graphics processor combined resource when the graphics processor recommended combined resource in the recommended resource list does not include the graphics processor combined resource in the default resource list and the graphics processor recommended combined resource corresponds to at least one graphics processor communication link; a fifth processing unit, configured to determine a fourth bandwidth of a target graphics processor communication link with the highest score among at least one graphics processor communication link, where the at least one graphics processor communication link is scored based on the path weight of each graphics processor communication link therein; a comparison unit, configured to compare the third bandwidth and the fourth bandwidth, and when the fourth bandwidth is greater than the third bandwidth, execute a target task according to the target graphics processor combined resource indicated by the target graphics processor communication link.
[0196] In an exemplary embodiment, the above device further includes: an acquisition unit, configured to acquire the hardware structure information of each server among a plurality of servers and the network topology information among the servers, where the hardware structure information includes the configuration information and connection relationships of a central processing unit, a graphics processor, and a switch, and the plurality of servers are the servers in a server cluster; a sixth processing unit, configured to, in response to a target resource request, determine the resource requirements and communication requirements of the graphics processor resources required to execute a target task; a seventh processing unit, configured to execute a multi-objective optimization algorithm based on the resource requirements and communication requirements, the internal communication performance of the plurality of servers, and the network performance among the servers to obtain an optimal graphics processor resource allocation result, where the multi-objective optimization algorithm is used to determine a resource allocation strategy when multiple optimization objectives such as computing performance, communication efficiency, and resource utilization rate are simultaneously satisfied; an eighth processing unit, configured to acquire the recommended graphics processor combined resource in the recommended resource list; a ninth processing unit, configured to, when the graphics processor recommended combined resource in the recommended resource list includes the optimal graphics processor combined resource indicated by the optimal graphics processor resource allocation result, determine to execute the target task according to the optimal graphics processor combined resource.
[0197] It should be noted that the above-mentioned modules can be implemented by software or hardware. For the latter, it can be implemented in the following ways, but not limited thereto: the above-mentioned modules are all located in the same processor; or, the above-mentioned modules are respectively located in different processors in any combination form.
[0198] According to another aspect of the embodiments of the present application, an electronic device is further provided, including a memory and a processor, where a computer program is stored in the memory, and the processor is configured to run the computer program to execute the steps in any of the above method embodiments for determining a resource allocation result.
[0199] According to another aspect of the embodiments of the present application, there is also provided a computer-readable storage medium storing a computer program, wherein the computer program is configured to execute the steps in any of the above-described method embodiments for determining a resource allocation result when running.
[0200] In an exemplary embodiment, the above computer-readable storage medium may include, but is not limited to: various media such as USB flash drives, read-only memories (ROM), random access memories (RAM), mobile hard disks, magnetic disks, or optical discs that can store computer programs.
[0201] According to another aspect of the embodiments of the present application, there is also provided a computer program product, where the computer program product includes a computer program, and when the computer program is executed by a processor, it implements the steps in any of the above-described method embodiments for determining a resource allocation result.
[0202] The embodiments of the present application also provide another computer program product, including a non-volatile computer-readable storage medium storing a computer program, and when the computer program is executed by a processor, it implements the steps in any of the above-described method embodiments for determining a resource allocation result.
[0203] Those skilled in the art can further realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of the two. To clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described according to functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Skilled professionals can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present application.
[0204] The above has introduced in detail a method for determining a resource allocation result provided by the present application. Specific examples are used herein to elaborate on the principle and implementation manner of the present application. The description of the above embodiments is only used to help understand the method and its core idea of the present application. It should be noted that for those of ordinary skill in the art in this technical field, without departing from the principle of the present application, several improvements and modifications can be made to the present application, and these improvements and modifications also fall within the protection scope of the claims of the present application.
Claims
1. A method for determining a resource allocation result, characterized in that, Including: In response to a target resource request, obtain a default resource list composed of the combined resources of the graphics processors required to execute the target task; Obtain a first topology map for describing the interconnection relationship of the hardware structure of the server, wherein the first topology map includes a second topology map for describing the interconnection relationship between at least one central processing unit and at least one graphics processing unit in the hardware structure; Based on the first topology map and the second topology map, generate a recommended resource list composed of the recommended combined resources of the graphics processors required to execute the target task; In the case that the recommended combined resources of the graphics processors in the recommended resource list include the combined resources of the graphics processors in the default resource list, determine to execute the target task according to the resource allocation result indicated by the combined resources of the graphics processors.
2. The method according to claim 1, characterized in that The generating, based on the first topology map and the second topology map, a recommended resource list composed of the recommended combined resources of the graphics processors required to execute the target task includes: By searching the first topology map, determine at least one switch connected to the at least one central processing unit in the hardware structure; Based on a group of graphics processors connected to the at least one switch, determine a group of graphics processors in the hardware structure; By searching the second topology map, determine a first part of the graphics processors that have been occupied in the group of graphics processors; Based on the group of graphics processors and the first part of the graphics processors, determine a second part of the graphics processors in an idle state; Based on the second part of the graphics processors, determine the recommended combined resources of the graphics processors in the recommended resource list.
3. The method according to claim 2, wherein The determining, based on the second part of the graphics processors, the recommended combined resources of the graphics processors in the recommended resource list includes: Based on the second part of the graphics processors, the first topology map and the second topology map, determine at least one graphics processor communication link, wherein one communication link in the at least one graphics processor communication link is used for data transmission between graphics processors and between graphics processors and other expansion devices, and the other expansion devices include the at least one switch and the at least one central processing unit; Obtain a first bandwidth between a target switch and a target central processing unit in a connected state among the at least one switch and the at least one central processing unit; Obtain a second bandwidth between any two switches in a connected state among the at least one switch; Based on the first bandwidth and the second bandwidth, determine a set of path weights for each graphics processor communication link in the at least one graphics processor communication link; Based on the set of path weights, determine the recommended combined resources of the graphics processors in the recommended resource list.
4. The method according to claim 3, wherein The determining, based on the first bandwidth and the second bandwidth, a set of path weights for each graphics processor communication link in the at least one graphics processor communication link includes: Successively obtain each graphics processor communication link in the at least one graphics processor communication link as the current graphics processor communication link; Determine the current central processor and the current switch through which the current graphics processor communication link passes from the at least one switch and the at least one central processor; Based on the first bandwidth and the second bandwidth, determine a first sub-path weight between the current central processors, a second sub-path weight between the current central processor and the current switch, and a third sub-path weight between the current switches; Sum the first sub-path weight, the second sub-path weight, and the third sub-path weight to obtain the path weight of the current graphics processor communication link.
5. The method according to claim 3, wherein The determining the graphics processor recommended combined resources in the recommended resource list based on the set of path weights includes: Based on the set of path weights, score each graphics processor communication link in the at least one graphics processor communication link to obtain a set of scores; Based on the set of scores, sort the at least one graphics processor communication link to obtain the sorted graphics processor communication link; Based on the at least one switch and at least one non-uniform memory access architecture in the hardware structure, group the sorted graphics processor communication links to obtain N groups of graphics processor communication links, where N is a preset positive integer; According to a preset priority judgment condition, determine the group of graphics processor communication links with the highest score from the N groups of graphics processor communication links; Based on the combination mode between the target graphics processors passed by the set of graphics processor communication links, determine the graphics processor recommended combined resources in the recommended resource list.
6. The method according to claim 5, wherein The grouping the sorted graphics processor communication links based on the at least one switch and at least one non-uniform memory access architecture in the hardware structure to obtain N groups of graphics processor communication links includes: Divide the first part of the graphics processor communication links in the sorted graphics processor communication links that are connected to the same switch and assigned to the same non-uniform memory access architecture into the first group of graphics processor communication links, where multiple graphics processors in the first group of graphics processor communication links assigned to the same non-uniform memory access architecture share local memory with the central processor connected to the same non-uniform memory access architecture; Divide the second part of the graphics processor communication links in the sorted graphics processor communication links that are connected to the same switch and assigned to different non-uniform memory access architectures into the second group of graphics processor communication links; Divide the third part of the graphics processor communication links in the sorted graphics processor communication links that are connected to different switches and assigned to the same non-uniform memory access architecture into the third group of graphics processor communication links; Determine the remaining graphics processor communication links in the sorted graphics processor communication links except the first group of graphics processor communication links, the second group of graphics processor communication links, and the third group of graphics processor communication links as the fourth group of graphics processor communication links.
7. The method according to any one of claims 1 to 6, characterized in that The method further includes: When the graphics processor recommended combined resource in the recommended resource list does not include the graphics processor combined resource in the default resource list and the graphics processor recommended combined resource corresponds to at least one graphics processor communication link, determine the third bandwidth of the default graphics processor communication link corresponding to the graphics processor combined resource; Determine the fourth bandwidth of the target graphics processor communication link with the highest score among the at least one graphics processor communication link, where the score is based on the path weight of each graphics processor communication link in the at least one graphics processor communication link; Compare the third bandwidth and the fourth bandwidth, and when the fourth bandwidth is greater than the third bandwidth, execute the target task according to the target graphics processor combined resource indicated by the target graphics processor communication link.
8. The method according to any one of claims 1 to 6, characterized in that, The method further includes: Obtain the hardware structure information of each server among multiple servers and the network topology information between the servers, where the hardware structure information includes the configuration information and connection relationship of the central processing unit, graphics processor, and switch, and the multiple servers are the servers in the server cluster; In response to the target resource request, determine the resource requirements and communication requirements of the graphics processor resources required to execute the target task; Based on the resource requirements and communication requirements, the internal communication performance of the multiple servers, and the network performance between the servers, execute a multi-objective optimization algorithm to obtain an optimal graphics processor resource allocation result, where the multi-objective optimization algorithm is used to determine the resource allocation strategy when multiple optimization objectives such as computing performance, communication efficiency, and resource utilization rate are simultaneously satisfied; Obtain the recommended graphics processor combined resource in the recommended resource list; When the graphics processor recommended combined resource in the recommended resource list includes the optimal graphics processor combined resource indicated by the optimal graphics processor resource allocation result, determine to execute the target task according to the optimal graphics processor combined resource.
9. An electronic device, characterized in that, Includes: A memory for storing a computer program; A processor for implementing the steps of the method for determining the resource allocation result according to any one of claims 1 to 8 when executing the computer program.
10. A computer-readable storage medium, characterized in that, A computer program is stored in the computer-readable storage medium, where the computer program implements the steps of the method for determining the resource allocation result according to any one of claims 1 to 8 when executed by a processor.
Citation Information
Cited By
Bandwidth allocation method and device, electronic equipment and storage medium
CN120547067A