Resource scheduling method and device, computing equipment and storage medium
By obtaining the communication performance indicators of the processor set, determining the scheduling weights, and selecting the processor set with excellent internal and external communication performance, the unbalanced problem caused by the differences in communication performance in resource scheduling is solved, and the optimal resource allocation and the reservation effect of subsequent scheduling are achieved.
Patent Information
- Application Number
- CN202510221396.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-26
- Publication Date
- 2025-07-18
AI Technical Summary
The differences in communication performance between existing resource scheduling methods have not been fully considered, resulting in uneven resource allocation and affecting the effect of subsequent resource scheduling.
By acquiring the first communication performance indicators and the second communication performance indicators of the processor set, determining the scheduling weights, and selecting the processor set with excellent internal communication performance and good external communication performance as the goal, avoiding selecting the processor set with poor communication performance.
This achieves the optimal performance of this resource scheduling, and at the same time, processor resources with better communication performance are reserved for subsequent resource scheduling, improving the efficiency and accuracy of overall resource allocation.
Smart Images

Figure CN120335943A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer technologies, and in particular, to a resource scheduling method, apparatus, computing device, and storage medium. Background Art
[0002] When a resource requester uses a processor (such as a graphics processing unit) for high-performance computing, data exchange is usually required between processors. Since the communication performance between different processors may be different, when scheduling resources, it is necessary to consider the communication performance between the processors allocated to the resource requester.
[0003] Currently, when scheduling resources, usually according to the type of physical link between different processors in the processor topology and the number of processors required by the resource requester, a set of processors with the optimal communication performance between processors is determined from the processor resource pool and allocated to the resource requester.
[0004] However, in subsequent resource scheduling, it is only possible to schedule from the remaining idle processors, and it is very likely to select a set of processors with poor communication performance. Therefore, the result of this resource scheduling will affect the next resource scheduling, which is not conducive to the overall resource allocation effect. Summary of the Invention
[0005] Embodiments of this application provide a resource scheduling method, apparatus, computing device, and storage medium, which avoid the result of this resource scheduling from affecting the next resource scheduling and improve the overall resource allocation effect.
[0006] In a first aspect, an embodiment of this application provides a resource scheduling method, which includes: obtaining a resource application request, where the resource application request includes the number of processors; determining a set of processors according to the resource application request; traversing the set of processors, and determining a scheduling weight for each set of processors according to a first communication performance metric and a second communication performance metric corresponding to each set of processors, where the first communication performance metric represents the communication performance between different processors within the set of processors, and the second communication performance metric represents the communication performance between each processor within the set of processors and each processor outside the set of processors; and determining a target set of processors according to the scheduling weight of each set of processors, where the target set of processors is selected from the sets of processors in which the first communication performance metric is greater than or equal to a preset metric threshold among all sets of processors.
[0007] In this way, not only the communication performance between different processors within the processor set is fully considered, but also the communication performance between each processor within the processor set and each processor outside the processor set is taken into account. The scheduling weight can be understood as the probability that the processor set is scheduled. Based on this, taking the case where the scheduling weight is positively correlated with the first communication performance metric and negatively correlated with the second communication performance metric as an example, this indicates that according to the scheduling weight, the target processor set with a relatively high probability of being scheduled is the processor set with a relatively large first communication performance metric and a relatively small second communication performance metric. In this way, it is avoided that the processor set with a prominent second communication performance metric is selected as the target in this resource scheduling, thereby leaving available processor resources with better communication performance for subsequent resource scheduling. This not only achieves the optimal performance of this resource scheduling but also avoids having an adverse impact on subsequent processor resource scheduling. For example, a processor set with the optimal communication performance is reserved for the next scheduling.
[0008] In a possible implementation, before traversing the processor set and determining the scheduling weight of each processor set according to the first communication performance metric and the second communication performance metric corresponding to each processor set, it further includes: obtaining the communication performance between different processors in the processor resource pool; the processor resource pool includes processors within and outside the processor set. According to the communication performance between different processors in the processor resource pool, determine the first communication performance metric and the second communication performance metric corresponding to each processor set.
[0009] In this way, based on the communication performance between different processors in the processor resource pool, the first communication performance metric and the second communication performance metric are obtained, rather than fixed values set according to the type of physical link, which can better conform to the actual situation of the physical environment where the processors are located. Even in the case of a large number of processors and link types, the target processor set with the best actual communication performance can be determined.
[0010] In a possible implementation, obtaining the communication performance between different processors in the processor resource pool includes: obtaining the processor topology, where the processor topology is used to represent the physical links between different processors in the processor resource pool. Test the communication performance of each physical link in the processor topology to obtain the communication performance between different processors in the processor resource pool.
[0011] Based on the processor topology, by testing the communication performance of each physical link in the processor topology, the actual communication performance between different processors in the processor resource pool can be obtained in real time and accurately, which better conforms to the actual situation of the physical environment where the processors are located. In this way, the first communication performance metric and the second communication performance metric determined based on the test results further improve the accuracy and rationality of resource scheduling and better meet the resource scheduling requirements.
[0012] In a possible implementation, the communication performance includes bandwidth. Through the bandwidth, the data transfer rate between processors can be intuitively quantified, thus accurately reflecting the communication performance between processors.
[0013] In another possible implementation, the communication performance includes latency. Through the latency, the time cost of data exchange between processors can be accurately reflected, thus accurately reflecting the communication performance between processors.
[0014] In a possible implementation, the first communication performance metric is positively correlated with the sum of the bandwidths of the first type of physical links, where the first type of physical links are the physical links between different processors within each processor set. And / or, the first communication performance metric is negatively correlated with the sum of the latencies of the first type of physical links. Based on the first communication performance metric set accordingly, the communication performance situation within each processor set can be fed back, which is beneficial for efficiently screening out the processor sets with larger first communication performance metrics.
[0015] In a possible implementation, the second communication performance metric is positively correlated with the sum of the bandwidths of the second type of physical links, where the second type of physical links are the physical links between each processor within each processor set and each idle processor outside the processor set. And / or, the second communication performance metric is negatively correlated with the sum of the latencies of the second type of physical links. Based on the magnitude of the second communication performance metric set accordingly, the communication performance situation between each processor set and external processors can be fed back, which is beneficial for efficiently screening out the processor sets with smaller second communication performance metrics.
[0016] In a possible implementation, determining the target processor set according to the scheduling weight of each processor set includes: determining the processor set with the largest scheduling weight from the processor sets in which the first communication performance metric is greater than or equal to the preset metric threshold among all processor sets. The processor set with the largest scheduling weight is determined as the target processor set.
[0017] In this way, from the processor sets in which the first communication performance metric is greater than or equal to the preset metric threshold among all processor sets, the processor set with the largest scheduling weight is determined as the target processor set, avoiding selecting the processor set with better communication performance with external processors as the target. This not only achieves the optimal performance of this resource scheduling but also avoids having an adverse impact on subsequent processor resource scheduling. For example, reserving the processor set with the optimal communication performance for the next scheduling.
[0018] In a possible implementation, the scheduling weight is the difference between the first communication performance metric and the second communication performance metric.
[0019] In this way, by calculating the difference between the first communication performance metric and the second communication performance metric, a scheduling weight value is obtained to accurately obtain a set of processors with a relatively large first communication performance metric and a relatively small second communication performance metric, thereby more accurately selecting the target processor set.
[0020] In a possible implementation manner, the processor includes a graphics processing unit (GPU). In this way, for the application scenario of GPU scheduling, it is possible to avoid selecting the set of GPUs with prominent second communication performance metrics as the target for the current resource scheduling, thereby reserving available GPU resources with better communication performance for subsequent resource scheduling. This not only achieves the optimal communication performance among the GPUs for the current resource scheduling but also reserves a set of GPUs with the optimal communication performance for the next scheduling, improving the performance of subsequent GPU resource scheduling.
[0021] In a second aspect, an embodiment of the present application provides a resource scheduling device, which includes: an acquisition module for acquiring a resource application request, where the resource application request includes the number of processors; a first determination module for determining a set of processors according to the resource application request; a second determination module for traversing the set of processors and determining the scheduling weight value of each set of processors according to the first communication performance metric and the second communication performance metric corresponding to each set of processors, where the first communication performance metric represents the communication performance among different processors within the set of processors, and the second communication performance metric represents the communication performance between each processor within the set of processors and each processor outside the set of processors; and a third determination module for determining a target processor set according to the scheduling weight value of each set of processors, where the target processor set is selected from the set of processors in which the first communication performance metric is greater than or equal to a preset metric threshold among all the sets of processors.
[0022] In a third aspect, an embodiment of the present application provides a computing device, which includes: a processor and a memory for storing processor-executable instructions. When the processor is configured to execute the instructions, the computing device implements the method as described above.
[0023] In a fourth aspect, an embodiment of the present application provides a storage medium, on which computer program instructions are stored. When the computer program instructions are executed by a computing device, the computing device implements the method as described above.
[0024] Among them, the technical effects brought by any implementation manner in the second aspect to the fourth aspect can refer to the technical effects brought by different implementation manners in the first aspect, which will not be elaborated here.
[0025] Based on the implementation manners provided in the above aspects, the present application can be further combined to provide more implementation manners. BRIEF DESCRIPTION OF THE DRAWINGS
[0026] Figure 1 A schematic diagram of a processor topology provided by an embodiment of the present application;
[0027] Figure 2 A schematic diagram of a resource scheduling system provided by an embodiment of the present application;
[0028] Figure 3 A schematic diagram of the structure of a computing device provided by an embodiment of the present application;
[0029] Figure 4 A schematic flowchart of a resource scheduling method provided by an embodiment of the present application;
[0030] Figure 5 A schematic flowchart of another resource scheduling method provided by an embodiment of the present application;
[0031] Figure 6 A schematic diagram of the structure of a resource scheduling device provided by an embodiment of the present application;
[0032] Figure 7 A schematic diagram of the structure of another computing device provided by an embodiment of the present application. Detailed implementation manners
[0033] Next, the technical solutions in the embodiments of the present application will be described with reference to the accompanying drawings in the embodiments of the present application.
[0034] Among them, in the description of the present application, unless otherwise specified, " / " indicates that the objects associated before and after are in an "or" relationship. For example, A / B may represent A or B; "and / or" in the present application is only a description of the association relationship of the associated objects, indicating that there can be three relationships. For example, A and / or B may represent: A exists alone, A and B exist simultaneously, and B exists alone. These three situations, where A and B can be singular or plural.
[0035] Moreover, in the description of the present application, unless otherwise specified, "a plurality of" means two or more than two. "At least one (item)" or its similar expression below refers to any combination of these items, including any combination of a single item or plural items. For example, at least one (item) of a, b, or c may represent: a, b, c, a - b, a - c, b - c, or a - b - c, where a, b, c can be single or multiple.
[0036] In addition, in order to facilitate the clear description of the technical solutions of the embodiments of the present application, in the embodiments of the present application, the words "first", "second" and the like are used to distinguish the same items or similar items with substantially the same functions and effects. Those skilled in the art will understand that the words "first", "second" and the like do not limit the quantity and execution order, and the words "first", "second" and the like do not necessarily limit the differences. At the same time, in the embodiments of the present application, the words "exemplary" or "for example" are used to indicate examples, illustrations or explanations. Any embodiment or design described as "exemplary" or "for example" in the embodiments of the present application should not be interpreted as being more preferred or more advantageous than other embodiments or design solutions. Specifically, the use of words such as "exemplary" or "for example" is intended to present related concepts in a concrete manner for ease of understanding.
[0037] The following is an exemplary introduction to the application scenarios of the embodiments of the present application.
[0038] When resource demanders (such as computing devices, applications, processes, etc.) use processors (such as graphics processors) for high-performance computing, data exchange is usually required between processors. Since the communication performance between different processors may be different, the communication performance between processors allocated to resource demanders needs to be considered during resource scheduling.
[0039] In some embodiments, during resource scheduling, the processor resources allocated to the resource demander are determined based on the communication performance of the physical links between different processors in the processor topology and the resource application request. The processor topology is used to represent the connection relationship between different processors in the processor resource pool, that is, the physical links. Figure 1 As shown, an embodiment of the present application provides a schematic diagram of a processor topology, Figure 1 In the example, the processor resource pool includes 12 idle processors, namely processors 1 to 12. It should be noted that there may be a physical link (referred to as link) between any two processors in the processor resource pool. Figure 1 Only some links are shown, such as links 1-20. For example, the communication performance between different physical links may be different. For example, if weights are used to represent the communication performance, the weights of links 3, 11, and 13 may be 100, the weights of links 9 and 15 may be 50, the weights of links 1 and 5 may be 10, and so on.
[0040] In view of this, an embodiment of the present application provides a resource scheduling method. In this resource scheduling method, for each processor set, the probability that it is scheduled as the optimal processor set (i.e., the following scheduling weight) is determined. When determining this probability, not only the communication performance between different processors within the processor set is considered, but also the influence of the communication performance between the processors within the processor set and the processors outside the processor set is introduced. Thus, it is possible to avoid, as much as possible, the processor set selected during this resource scheduling from having an adverse effect on subsequent processor resource scheduling.
[0041] Specifically, first, a resource application request is obtained. The resource application request includes the number of processors. Second, a processor set is determined according to the resource application request. Then, the processor set is traversed, and according to the first communication performance index and the second communication performance index corresponding to each processor set, the scheduling weight of each processor set is determined. Among them, the first communication performance index represents the communication performance between different processors within the processor set, and the second communication performance index represents the communication performance between each processor within the processor set and each processor outside the processor set. Finally, according to the scheduling weight of each processor set, a target processor set is determined. The target processor set is selected from the processor sets in which the first communication performance index is greater than or equal to a preset index threshold among all processor sets.
[0042] In the embodiment of the present application, not only the communication performance between different processors within the processor set is fully considered, but also the communication performance between each processor within the processor set and each processor outside the processor set is considered. The scheduling weight can be understood as the probability that the processor set is scheduled. Based on this, taking the scheduling weight as being positively correlated with the first communication performance index and negatively correlated with the second communication performance index as an example, this indicates that according to the scheduling weight, the target processor set with a relatively large probability of being scheduled is the processor set with a relatively large first communication performance index and a relatively small second communication performance index. In this way, it is avoided that the processor set with a prominent second communication performance index is selected as the target during this resource scheduling, so as to reserve processor resources with better communication performance for subsequent resource scheduling. This not only achieves the optimal performance of this resource scheduling, but also avoids having an adverse effect on subsequent processor resource scheduling. For example, a processor set with the optimal communication performance is reserved for the next scheduling.
[0043] Next, an exemplary introduction to the system architecture of the embodiment of the present application is given.
[0044] As Figure 2 shown, an embodiment of the present application provides a resource scheduling system. The system includes a resource requester, a processor resource pool, and a resource scheduling platform.
[0045] Among them, the resource requester is any object that needs to apply for processor resources, such as: devices (computing devices, terminal devices, etc.), application programs, or processes, etc.
[0046] The processor resource pool includes multiple processors, and different processors are communicatively connected through a certain type of physical link. The processor can be a graphics processing unit (GPU), a central processing unit (CPU), a data processing unit (DPU), and so on. In the embodiments of the present application, the processor can specifically be a GPU. Among them, the processor resource pool includes processors inside and outside the processor set.
[0047] The resource scheduling platform can be used to obtain the resource application request of the resource requester, and the resource application request includes the number of processors. It can also be used to determine the processor set according to the resource application request. It can also be used to traverse the processor set, and determine the scheduling weight value of each processor set according to the first communication performance index and the second communication performance index corresponding to each processor set; the first communication performance index represents the communication performance between different processors within the processor set, and the second communication performance index represents the communication performance between each processor within the processor set and each processor outside the processor set. In addition, the resource scheduling platform can also be used to determine the target processor set according to the scheduling weight value of each processor set. The target processor set is selected from the processor sets in which the first communication performance index is greater than or equal to the preset index threshold among all processor sets.
[0048] Exemplarily, the resource scheduling platform can be a container scheduling platform (such as the application software kubernetes, etc.), and the container scheduling platform can schedule the required processor (such as GPU) resources for the upper-layer container application program.
[0049] Exemplarily, the resource scheduling platform can also be a virtual machine scheduling platform (such as the virtual machine scheduling platform vmware and the virtual machine scheduling platform kvm, etc.), and the virtual machine scheduling platform can schedule the required processor (such as GPU) resources for the upper-layer virtual machine application program.
[0050] In some embodiments, the resource scheduling platform is configured with a preset test tool, and the preset test tool is used to test the communication performance of each physical link in the processor topology. It can be understood that the communication performance of each physical link in the processor topology represents the communication performance between the corresponding every two processors in the processor resource pool.
[0051] In one implementation, the resource scheduling platform is further configured to determine a first communication performance metric and a second communication performance metric corresponding to each processor set according to the communication performance between different processors in the processor resource pool. Exemplarily, the communication performance is characterized by bandwidth or latency, and the preset speed measurement tool may be a CUDA peer-to-peer bandwidth latency (p2pbandwidthlatencytest) test tool.
[0052] Optionally, the resource requester and the processor resource pool may be deployed on the same computing device or the same computing device cluster.
[0053] Optionally, the resource scheduling platform may be deployed on the same computing device or computing device cluster as the resource requester and the processor resource pool, or may be deployed on different computing devices or different computing device clusters.
[0054] In the embodiments of the present application, the computing device may specifically be a network device, such as a server. Among them, the server may be a physical server, or may be two or more physical servers sharing different responsibilities and cooperating with each other to implement the various functions of the server. When the computing device is multiple servers, the scheduling system is a server cluster with high availability capabilities.
[0055] Exemplarily, the server may be a blade server, a high-density server, a rack server, or a tower server, etc.
[0056] Among them, the hardware part of the computing device includes a processor, a basic input / output system (BIOS) chip, an out-of-band controller, and a memory, and the software part mainly includes BIOS, an out-of-band management module, and an operating system (OS), as Figure 3 shown.
[0057] The processor may include a central processing unit CPU, and the CPU includes one or more CPU cores. All operations of the CPU for processing data are executed by the CPU cores. The more CPU cores included in the CPU, the faster the data processing speed. The processor may further include a graphics processing unit GPU. In the embodiments of the present application, the central processing unit in the computing device may be used as a scheduler to execute the above workflow scheduling method. Exemplarily, the graphics processing unit GPU may be used as a resource to be scheduled.
[0058] The BIOS chip is a chip set on the motherboard for initializing and detecting various hardware during the startup process of the computing device. The BIOS chip includes a flash memory area.
[0059] The out-of-band management module is located in the out-of-band controller, and the operating system is located in the processor.
[0060] The out-of-band management module may be a management unit of a non-business module. For example, the out-of-band management module may remotely maintain and manage a computing device through a dedicated data channel. The out-of-band management module is completely independent of the operating system of the computing device and may communicate with the BIOS and the operating system through the out-of-band management interface of the computing device.
[0061] Exemplarily, the out-of-band management module may include a management unit for computing device operation status, a management system in a management chip, a computing device motherboard management unit (baseboard management controller, BMC), a system management module (system management mode, SMM), etc. It should be noted that the embodiments of the present application do not limit the specific form of the out-of-band management module, and the above is only an exemplary description.
[0062] OS is a computer program that manages and controls the hardware and software resources of a computing device. Any other software must be supported by the operating system to run. After the computing device is powered on, the BIOS first starts a series of operations such as self-test and initialization, and then guides the OS to start, so that the user can use the computing device normally.
[0063] BIOS is a set of programs that are fixed on the BIOS chip in the computing device. The main function of BIOS is to provide the lowest-level and most direct hardware settings and controls for the computing device.
[0064] Memory, also called internal storage or main memory, is installed in memory slots on the motherboard of a computing device.
[0065] It should be noted that the embodiments of the present application do not limit the device form factor of the computing device, and the above description is merely an exemplary description.
[0066] It should be noted that the system architecture and application scenarios described in the embodiments of the present application are intended to more clearly illustrate the technical solutions of the embodiments of the present application, and do not constitute a limitation on the technical solutions provided in the embodiments of the present application. Ordinary technicians in this field can know that with the evolution of the system architecture and the emergence of new application scenarios, the technical solutions provided in the embodiments of the present application are also applicable to similar technical problems.
[0067] For ease of understanding, the resource scheduling method provided in the embodiment of the present application is exemplarily introduced below in combination with the above system architecture and accompanying drawings.
[0068] like Figure 4 As shown, an embodiment of the present application provides a resource scheduling method, which can be executed by a computing device where a resource scheduling platform is located. The method includes:
[0069] S401, Obtain a resource application request.
[0070] Among them, the resource application request includes the number of processors.
[0071] It can be understood that the resource application request is sent by the resource requester to the resource scheduling platform. The number of processors represents the number of processors that the resource requester needs to use. For example, the resource application request can specifically be: the number of GPUs to be used is 2.
[0072] S402, Determine a set of processors according to the resource application request.
[0073] The set of processors refers to: all sets formed by the available processors in the processor resource pool according to the number of processors (in the resource application request). For example, as shown in Figure 1 If the available processors in the processor resource pool are processors 1-4 (the number is 4), and the number of processors in the resource application request is 2 at this time, then the sets of processors are respectively: processors 1 and 2, processors 1 and 3, processors 1 and 4, processors 2 and 3, processors 2 and 4, processors 3 and 4. It can be understood that the available processors are the idle processors in the processing resource pool. In this way, according to the required number of processors, all sets among the available processors are determined from the processor resource pool, which is convenient for efficiently and accurately determining the target set of processors later and better meeting the needs of the resource requester.
[0074] S403, Traverse the set of processors, and determine the scheduling weight of each set of processors according to the first communication performance index and the second communication performance index corresponding to each set of processors.
[0075] Among them, the first communication performance index represents the communication performance between different processors within the set of processors, and the second communication performance index represents the communication performance between each processor within the set of processors and each processor outside the set of processors. That is to say, the first communication performance represents the communication performance situation within each set of processors, and the second communication performance represents the communication performance situation between each set of processors and external processors.
[0076] The scheduling weight is used to represent the probability that the set of processors is scheduled. Exemplarily, the scheduling weight is positively correlated with the first communication performance index, and the scheduling weight is negatively correlated with the second communication performance index.
[0077] It can be understood that since the scheduling weight is positively correlated with the first communication performance index and negatively correlated with the second communication performance index, this indicates that the larger the first communication performance index and the smaller the second communication performance index, the larger the scheduling weight. On the one hand, the larger the first communication performance index and the smaller the second communication performance index indicate that the communication performance within the processor set is better, while the communication performance with external processors is worse. On the other hand, since the scheduling weight represents the probability that the processor set is scheduled, the processor set with better internal communication performance and worse communication performance with external processors has a greater probability of being scheduled. In this way, it is possible to avoid selecting the processor set with prominent second communication performance index as the target in this resource scheduling, so as to reserve available processor resources with better communication performance for subsequent resource scheduling. While achieving the optimal performance of this resource scheduling (i.e., the communication performance within the processor set), it also avoids having an adverse impact on subsequent processor resource scheduling.
[0078] In some embodiments, the communication performance includes bandwidth. In this way, the bandwidth can intuitively quantify the data transmission rate between processors, thus accurately reflecting the communication performance between processors.
[0079] In one implementation, the first communication performance index is positively correlated with the sum of the bandwidths of the first type of physical links.
[0080] Among them, the first type of physical links are the physical links between different processors within each processor set.
[0081] In the embodiments of the present application, since the first communication performance index is positively correlated with the sum of the bandwidths of the first type of physical links, the larger the sum of the bandwidths of the first type of physical links, the larger the first communication performance index. Thus, according to the size of the first communication performance index, the communication performance situation within each processor set can be reflected, which is conducive to efficiently screening out the processor sets with larger first communication performance index (i.e., larger sum of the bandwidths of the first type of physical links).
[0082] Combined Figure 1 As shown, an exemplary description of the first type of physical links is given for different processor sets (hereinafter referred to as sets): Set 1 (processors 2 and 6), the corresponding first type of physical link is Link 1; Set 2 (processors 6 and 7), the corresponding first type of physical link is Link 2; Set 3 (processors 3 and 7), the corresponding first type of physical link is Link 3. Set 6 (processors 2, 6 and 7), the corresponding first type of physical links are Links 3, 4 and 11 respectively. Further, the sum of the bandwidths of the first type of physical links of Set 6 is the sum of the bandwidths of Links 3, 4 and 11.
[0083] Optionally, the first communication performance metric can be the sum of the bandwidths of the first type of physical links. In this way, there is no need to convert the sum of the bandwidths of the first type of physical links into other values, making the measurement of the communication performance within the processor set more intuitive and convenient.
[0084] Optionally, the first communication performance metric can be the sum of the weights corresponding to the bandwidths of the first type of physical links. In this way, by setting weights corresponding to the bandwidths of the first type of physical links and taking the sum of the weights corresponding to the bandwidths of the first type of physical links as the first communication performance metric, it is possible to perform differential evaluation and quantification of the communication performance within the processor set according to actual needs, which is more flexible.
[0085] Exemplarily, the bandwidth of the first type of physical links is positively correlated with the weight corresponding to the bandwidth of the first type of physical links. That is to say, for each first type of physical link, a weight is set according to its bandwidth, and the larger the bandwidth of the first type of physical link, the larger the set weight. In this way, the first communication performance metric determined according to the sum of the weights corresponding to the bandwidths of the first type of physical links can accurately reflect the communication performance situation within the processor set.
[0086] In one implementation, the second communication performance metric is positively correlated with the sum of the bandwidths of the second type of physical links.
[0087] Among them, the second type of physical links are the physical links between each processor within each processor set and each idle processor outside the processor set. It can be understood that the idle processor can be a processor available in the processor resource pool.
[0088] In the embodiments of the present application, since the second communication performance metric is positively correlated with the sum of the bandwidths of the second type of physical links, the smaller the sum of the bandwidths of the second type of physical links, the smaller the second communication performance metric. Thus, according to the magnitude of the second communication performance metric, the communication performance situation between each processor set and external processors can be reflected, which is beneficial to efficiently screening out the processor sets with a smaller second communication performance metric (i.e., a smaller sum of the bandwidths of the second type of physical links).
[0089] Exemplarily, such as Figure 1As shown, taking Set 1 (Processors 2 and 6) as an example, the second type of physical link and the sum of the bandwidths of the second type of physical link are described. The second type of physical link corresponding to Set 1 includes: physical links between Processor 2 and each idle processor outside Set 1, such as Links 1, 2, 4, 5, etc.; and physical links between Processor 6 and each idle processor outside Set 1, such as Links 6, 7, 8, 9, 10, 11, 12, etc. Further, calculate: D1 = D12 + D16, where D1 is the sum of the bandwidths of the second type of physical link corresponding to Set 1, D12 is the sum of the bandwidths of the physical links between Processor 2 and each idle processor outside Set 1, and D16 is the sum of the bandwidths of the physical links between Processor 6 and each idle processor outside Set 1.
[0090] For example, if the bandwidths of the physical links between Processor 2 and 10 idle processors outside Set 1 (Processor 1, Processors 3 - 5, and Processors 7 - 12) are all 3.0, then the sum of the bandwidths of the physical links between Processor 2 and each idle processor outside Set 1 is 30.0. For the 10 idle processors outside Set 1 and Processor 6, the bandwidths of 9 physical links are 3.0 and the bandwidth of 1 physical link is 10.0, then the sum of the bandwidths of the physical links between Processor 6 and each idle processor outside Set 1 is 37.0. Based on this, the sum of the bandwidths of the second type of physical link corresponding to Set 1 is 67.0.
[0091] It should be noted that the specific values such as the bandwidth shown in the embodiments of the present application are only for exemplary reference.
[0092] Optionally, the second communication performance metric can be the sum of the bandwidths of the second type of physical link. In this way, there is no need to convert the sum of the bandwidths of the second type of physical link into other values, making the measurement of the communication performance between the processor set and the external processor more intuitive and simple.
[0093] For example, if the sum of the bandwidths of the second type of physical link corresponding to Set 1 is 67.0, then the second communication performance metric corresponding to Set 1 is also 67.0.
[0094] Optionally, the second communication performance metric can be the sum of the weights corresponding to the bandwidths of the second type of physical link. In this way, by setting the corresponding weights based on the bandwidths of the second type of physical link and taking the sum of the weights corresponding to the bandwidths of the second type of physical link as the second communication performance metric, it is possible to differentially evaluate and quantify the communication performance between each processor set and the external processor according to actual needs, which is more flexible.
[0095] Exemplarily, the bandwidth of the second type of physical link is positively correlated with the weight value corresponding to the bandwidth of the second type of physical link. That is, for each second type of physical link, a weight value is set according to its bandwidth, and the smaller the bandwidth of the second type of physical link, the smaller the set weight value. In this way, the second communication performance index determined according to the sum of the weight values corresponding to the bandwidth of the second type of physical link can accurately reflect the communication performance between the processor set and the external processor.
[0096] For example, taking Set 1 (Processors 2 and 6) as an example, if the bandwidth of the physical link between Processor 2 and 10 idle processors (Processors 1, 3 - 5, and 7 - 12) outside Set 1 is 3.0 for all, then the weight values corresponding to the bandwidth of the physical link between Processor 2 and the 10 idle processors outside Set 1 can all be set to 30. In this way, the sum of the weight values corresponding to the bandwidth of the physical link between Processor 2 and each idle processor outside Set 1 is 300. Between Processor 6 and the 10 idle processors outside Set 1, the bandwidth of 9 physical links is 3.0 and the bandwidth of 1 physical link is 10.0. Then, the weight values corresponding to the bandwidth of 9 physical links can be set to 30, and the weight value corresponding to the bandwidth of 1 physical link can be set to 100. In this way, the sum of the weight values corresponding to the bandwidth of the physical link between Processor 6 and each idle processor outside Set 1 is 370. Based on this, the sum of the weight values corresponding to the bandwidth of the second type of physical link corresponding to Set 1 is 670, that is, the second communication performance index corresponding to Set 1 is 670.
[0097] In some other embodiments, the communication performance includes latency. In this way, the latency can accurately reflect the time cost consumed for data exchange between processors, thereby accurately reflecting the communication performance between processors.
[0098] In one implementation manner, the first communication performance index is negatively correlated with the sum of the latencies of the first type of physical link.
[0099] In the embodiments of the present application, since the first communication performance index is negatively correlated with the sum of the latencies of the first type of physical link, the smaller the sum of the latencies of the first type of physical link, the larger the first communication performance index. Thus, according to the magnitude of the first communication performance index, the communication performance within each processor set can be fed back, which is beneficial for efficiently screening out the processor sets with larger first communication performance indices (i.e., smaller sums of the latencies of the first type of physical link).
[0100] Exemplarily, as Figure 1 shown, taking Set 6 (Processors 2, 6, and 7) as an example, the sum of the latencies of the first type of physical link is described: The first type of physical links corresponding to Set 6 are Link 3, Link 4, and Link 11 respectively. Further, the sum of the latencies of the first type of physical link is the sum of the latencies of Link 3, Link 4, and Link 11.
[0101] Optionally, the first communication performance metric can be the sum of the reciprocals of the delays of the first type of physical links. In this way, by specifically setting the first communication performance metric as the sum of the reciprocals of the delays of the first type of physical links, it more clearly reflects that the first communication performance metric is negatively correlated with the sum of the delays of the first type of physical links, making the measurement of the communication performance within the processor set more intuitive and simple.
[0102] For example, taking set 6 (processors 2, 6, and 7) as an example, illustrate the sum of the reciprocals of the delays of the first type of physical links: The first type of physical links corresponding to set 6 are links 3, 4, and 11 respectively. Further, calculate the reciprocals of the delays of links 3, 4, and 11 respectively, and based on this, sum up the reciprocals of the three, so as to obtain the first communication performance metric.
[0103] Optionally, the first communication performance metric can be the sum of the weights corresponding to the delays of the first type of physical links. In this way, by setting the corresponding weights based on the delays of the first type of physical links and taking the sum of the weights corresponding to the delays of the first type of physical links as the first communication performance metric, it is possible to conduct differential evaluation and quantification of the communication performance within the processor set according to actual needs, which is more flexible.
[0104] Exemplarily, the delay of the first type of physical link is negatively correlated with the weight corresponding to the delay of the first type of physical link. That is to say, for each first type of physical link, a weight is set according to its delay, and the greater the delay of the first type of physical link, the smaller the set weight. In this way, the first communication performance metric determined according to the sum of the weights corresponding to the delays of the first type of physical links can accurately reflect the communication performance situation within the processor set.
[0105] For example, taking set 6 (processors 2, 6, and 7) as an example, illustrate the sum of the weights corresponding to the delays of the first type of physical links: The first type of physical links corresponding to set 6 are links 3, 4, and 11 respectively. Further, determine the weights corresponding to links 3, 4, and 11 respectively according to the delays of links 3, 4, and 11, and based on this, sum up the weights of the three, so as to obtain the first communication performance metric.
[0106] In one implementation, the second communication performance metric is negatively correlated with the sum of the delays of the second type of physical links.
[0107] In the embodiments of the present application, since the second communication performance index is negatively correlated with the sum of the delays of the second type of physical links, the larger the sum of the delays of the second type of physical links, the smaller the second communication performance index. Therefore, according to the size of the second communication performance index, the communication performance of each processor set with external processors can be reflected, which is conducive to efficiently screening out the processor sets with a smaller second communication performance index (i.e., a larger sum of the delays of the second type of physical links).
[0108] Exemplarily, as Figure 1 shown, taking set 6 (processors 2, 6, and 7) as an example, the sum of the delays of the second type of physical links is described. The sum of the delays of the second type of physical links corresponding to set 6: D6 = D62 + D66 + D67, where D62 is the sum of the delays of the physical links between processor 2 and each idle processor other than set 6, D66 is the sum of the delays of the physical links between processor 6 and each idle processor other than set 6, and D67 is the sum of the delays of the physical links between processor 7 and each idle processor other than set 6.
[0109] Optionally, the second communication performance index can be the sum of the reciprocals of the delays of the second type of physical links. In this way, the measurement of the communication performance of the processor set with external processors is made more intuitive and simple.
[0110] Exemplarily, taking set 1 (processors 2 and 6) as an example, the sum of the reciprocals of the delays of the second type of physical links is described. If the delays of the physical links between processor 2 and 10 idle processors (processors 1, 3 - 5, and 7 - 12) outside set 1 are all 10.0, then the sum of the reciprocals of the delays of the physical links between processor 2 and each idle processor other than set 1 is 1. Among the 10 idle processors outside set 1 for processor 6, the delays of 9 physical links are 13.0 and the delay of 1 physical link is 10.0. Then, the sum of the reciprocals of the delays of the physical links between processor 6 and each idle processor other than set 1 is 103 / 130. Based on this, the sum of the reciprocals of the delays of the second type of physical links corresponding to set 1 is 233 / 130.
[0111] It should be noted that the specific values of the delays shown in the embodiments of the present application are only for exemplary reference.
[0112] Optionally, the second communication performance index can be the sum of the weights corresponding to the delays of the second type of physical links. In this way, by setting corresponding weights based on the delays of the second type of physical links and taking the sum of the weights corresponding to the delays of the second type of physical links as the second communication performance index, the communication performance of the processor set with external processors can be evaluated and quantified differently according to actual needs, which is more flexible.
[0113] Exemplarily, the latency of the second type of physical link is negatively correlated with the weight value corresponding to the latency of the second type of physical link. That is to say, for each second type of physical link, a weight value is set according to its latency, and the greater the latency of the second type of physical link, the smaller the set weight value. In this way, the second communication performance index determined according to the sum of the weight values corresponding to the latency of the second type of physical link can accurately reflect the communication performance of the processor set and the external processor.
[0114] For example, taking Set 1 (Processors 2 and 6) as an example, the sum of the weight values corresponding to the latency of the second type of physical link is described. If the latency of the physical link between Processor 2 and 10 idle processors outside Set 1 (Processors 1, 3 - 5, and 7 - 12) is 10.0 for all, then the weight value corresponding to the latency of the physical link between Processor 2 and each idle processor outside Set 1 can be set to 1. For the 10 idle processors outside Set 1 and Processor 6, among which the latency of 9 physical links is 13.0 and the latency of 1 physical link is 10.0, then the weight values corresponding to the latency of the 9 physical links can all be 10 / 13, and the weight value corresponding to the latency of 1 physical link can be set to 1. Based on this, the sum of the weight values corresponding to the latency of the second type of physical link corresponding to Set 1 is 233 / 13.
[0115] S404. Determine the target processor set according to the scheduling weight of each processor set.
[0116] Among them, the target processor set is selected from the processor sets in which the first communication performance index is greater than or equal to the preset index threshold among all processor sets.
[0117] Exemplarily, the preset index threshold can be set to 0.85X, 0.9X, or 0.95X. Among them, X is the maximum value among the first communication performance indexes corresponding to all processor sets.
[0118] Specifically, the preset index threshold can be set to 0.9X. For example, if the first communication performance index corresponding to Set 1 (Processors 2 and 6) is 10, and the first communication performance index corresponding to Set 2 (Processors 6 and 7) is 10.1, and the first communication performance index corresponding to Set 2 (10.1) is the maximum value among the first communication performance indexes of all processor sets. In this way, X is 10.1. Based on this, 0.9X is 9.09, then all processor sets within [9.09, 10.1] are the processor sets in which the first communication performance index is greater than or equal to the preset index threshold among all processor sets, such as Set 1 and Set 2. Since the preset index threshold has a small gap with the maximum value of the first communication performance index and belongs to a reasonable performance fluctuation range, such a setting is more reasonable and flexible.
[0119] In some embodiments, determining a target processor set according to the scheduling weight value of each processor set includes: determining, from the processor sets in which the first communication performance metric is greater than or equal to a preset metric threshold among all processor sets, the processor set with the largest scheduling weight value. The processor set with the largest scheduling weight value is determined as the target processor set.
[0120] In the embodiments of the present application, determining, from the processor sets in which the first communication performance metric is greater than or equal to a preset metric threshold among all processor sets, the processor set with the largest scheduling weight value means that this processor set is a processor set with a relatively large first communication performance metric and a relatively small second communication performance metric, that is, a processor set with better internal communication performance and poorer communication performance with external processors. In this way, by determining the processor set with the largest scheduling weight value as the target processor set, it is avoided to select a processor set with better communication performance with external processors as the target, which not only achieves the optimal performance of this resource scheduling, but also avoids adverse effects on subsequent processor resource scheduling. For example, a processor set with the optimal communication performance is reserved for the next scheduling.
[0121] In one implementation, the scheduling weight value is the difference between the first communication performance metric and the second communication performance metric.
[0122] In the embodiments of the present application, by calculating the difference between the first communication performance metric and the second communication performance metric, the scheduling weight value is obtained to accurately obtain a processor set with a relatively large first communication performance metric and a relatively small second communication performance metric, so as to more accurately select the target processor set.
[0123] In some embodiments, the method further includes: if there are multiple processor sets with the largest scheduling weight value, determining the processor set that meets the preset conditions as the target processor set.
[0124] In the embodiments of the present application, when there are multiple processor sets with the largest scheduling weight value, accurate screening can be performed according to the preset conditions, further improving the rationality of resource scheduling.
[0125] Exemplarily, the preset conditions include that the sum of the resource utilization rates of the processors in the processor set is the highest. In this way, in the case of multiple processor sets with the largest scheduling weight value, the processor set with the optimal overall resource utilization efficiency can be selected, further improving the performance of resource scheduling.
[0126] Exemplarily, the preset conditions include that the priority of the processor set is the highest. In this way, by the preset priority of each processor set, the processor set with the highest priority is selected as the target processor set, and the target processor set can be efficiently selected according to the preset priority rules, which better meets the actual needs of users.
[0127] In some other embodiments, the method further includes: if there are multiple sets of processors with the largest scheduling weight values, randomly select one set of processors and determine it as the target set of processors. In this way, the target set of processors can be obtained more efficiently.
[0128] In some other embodiments, the method further includes: if there is only one set of processors among all sets of processors whose first communication performance metric is greater than or equal to a preset metric threshold, use the set of processors whose first communication performance metric is greater than or equal to the preset metric threshold as the target set of processors.
[0129] In the embodiments of the present application, if there is only one set of processors among all sets of processors whose first communication performance metric is greater than or equal to the preset metric threshold, it indicates that there is only one set of processors with a relatively large first communication performance metric. In this case, it is not necessary to determine the second communication performance metric, and it is even less necessary to determine the scheduling weight value. By directly using this set of processors as the target set of processors, while ensuring that the communication performance of the set of processors scheduled this time is optimal, it also improves the efficiency of scheduling the target set of processors, effectively avoiding resource waste caused by unnecessary calculations and screenings.
[0130] In some embodiments, the method further includes: obtaining the communication performance between different processors in the processor resource pool. According to the communication performance between different processors in the processor resource pool, determine the first communication performance metric and the second communication performance metric corresponding to each set of processors.
[0131] Among them, the processor resource pool includes processors inside and outside the set of processors.
[0132] In the embodiments of the present application, obtaining the first communication performance metric and the second communication performance metric based on the communication performance between different processors in the processor resource pool, rather than a fixed value set according to the type of physical link, can better conform to the actual situation of the physical environment where the processors are located. Even in the case of a large number of processors and link types, the target set of processors with the best actual communication performance can be determined.
[0133] In one implementation manner, obtaining the communication performance between different processors in the processor resource pool includes: obtaining the processor topology. Test the communication performance of each physical link in the processor topology to obtain the communication performance between different processors in the processor resource pool.
[0134] Among them, the processor topology is used to represent the physical links between different processors in the processor resource pool.
[0135] In the embodiments of the present application, based on the processor topology, by testing the communication performance of each physical link in the processor topology, the actual communication performance between different processors in the processor resource pool can be obtained in real time and accurately, which can better conform to the actual situation of the physical environment where the processor is located. Based on the test, the first communication performance index and the second communication performance index are determined, further improving the accuracy and rationality of resource scheduling and better meeting the resource scheduling requirements.
[0136] Combined with Figure 5 As shown, the embodiments of the present application provide a schematic flowchart of another resource scheduling method. As Figure 5 shown, the computing device where the resource scheduling platform is located executes this method, and this method includes the following steps:
[0137] S501, obtain a resource application request.
[0138] S502, determine a processor set according to the resource application request.
[0139] S503, traverse the processor set and determine the first communication performance index corresponding to each processor set.
[0140] S504, determine the processor sets in all processor sets whose first communication performance index is greater than or equal to a preset index threshold, and use them as candidate processor sets.
[0141] S505, determine whether the number of candidate processor sets is greater than 1. If so, go to S506; if not, go to S509.
[0142] S506, determine the second communication performance index corresponding to each candidate processor set.
[0143] S507, determine the scheduling weight value of each candidate processor set according to the first communication performance index and the second communication performance index of each candidate processor set.
[0144] S508, determine the candidate processor set with the largest scheduling weight value as the target processor set.
[0145] S509, determine the candidate processor set as the target processor set.
[0146] In the embodiments of the present application, first, all available processor sets in the processor resource pool are determined, and then the candidate processor sets with relatively large first communication performance metrics (i.e., better internal communication performance) are focused on. Based on this, in the case where there are multiple candidate processor sets, not only is the communication performance between different processors within the processor set fully considered, but also the communication performance between each processor within the processor set and each processor outside the processor set is fully considered. In this way, the target processor set with a relatively high scheduling probability is the candidate processor set with a relatively large first communication performance metric and a relatively small second communication performance metric. Further, it is avoided that the processor set with a prominent second communication performance metric is selected as the target in this resource scheduling, so as to reserve available processor resources with better communication performance for subsequent resource scheduling. This not only achieves the optimal performance of this resource scheduling, but also avoids adverse effects on subsequent processor resource scheduling. For example, a processor set with the optimal communication performance is reserved for the next scheduling.
[0147] In some embodiments, the processor includes a GPU. In this way, for the application scenario of GPU scheduling, it is possible to avoid selecting the GPU set with a prominent second communication performance metric as the target in this resource scheduling, so as to reserve available GPU resources with better communication performance for subsequent resource scheduling. This not only achieves the optimal communication performance between GPUs in this resource scheduling, but also reserves a GPU set with the optimal communication performance for the next scheduling, improving the performance of subsequent GPU resource scheduling.
[0148] As Figure 6 shown, the embodiments of the present application provide a resource scheduling device 200, and the resource scheduling device 200 includes: an acquisition module 21, a first determination module 22, a second determination module 23, and a third determination module 24. The acquisition module 21 acquires a resource application request, and the resource application request includes the number of processors. The first determination module 22 is configured to determine a processor set according to the resource application request. The second determination module 23 is configured to traverse the processor sets and determine the scheduling weight of each processor set according to the first communication performance metric and the second communication performance metric corresponding to each processor set; the first communication performance metric represents the communication performance between different processors within the processor set, and the second communication performance metric represents the communication performance between each processor within the processor set and each processor outside the processor set. The third determination module 24 is configured to determine a target processor set according to the scheduling weight of each processor set. The target processor set is selected from the processor sets in which the first communication performance metric is greater than or equal to a preset metric threshold among all processor sets.
[0149] As Figure 7As shown, an embodiment of the present application provides another computing device 500. The computing device 500 includes a processor 510 (specifically, a central processing unit) and a memory 520 for storing processor-executable instructions. When the processor 510 is configured to execute instructions, the computing device 500 implements the resource scheduling method as described above.
[0150] Figure 7 The shown computing device 500 is merely an example and should not impose any limitations on the functions and usage scope of the embodiments of the present application.
[0151] The computing device 500 is presented in the form of a general-purpose computing device. The components of the computing device 500 may include, but are not limited to: one or more processors 510, a memory 520, a communication bus 540 connecting different system components (including the memory 520 and the processor 510), and a communication interface 530.
[0152] The communication bus 540 represents one or more of several types of bus structures, including a memory bus or a memory controller, a peripheral bus, a graphics acceleration port, a processor, or a local bus using any bus structure in a variety of bus structures. For example, these architectures include, but are not limited to, Industry Standard Architecture (hereinafter referred to as: ISA) bus, Micro Channel Architecture (hereinafter referred to as: MAC) bus, Enhanced ISA bus, Video Electronics Standards Association (hereinafter referred to as: VESA) local bus, and Peripheral Component Interconnection (hereinafter referred to as: PCI) bus.
[0153] The computing device 500 typically includes a variety of computer system-readable media. These media can be any available media accessible by the computing device, including volatile and non-volatile media, removable and non-removable media.
[0154] The memory 520 may include computer system-readable media in the form of volatile memory, such as random access memory (hereinafter referred to as: RAM) and / or cache memory. The computing device may further include other removable / non-removable, volatile / non-volatile computer system storage media. Although Figure 7Not shown in the figure, a disk drive for reading and writing to a removable non-volatile disk (such as a "floppy disk") and an optical disk drive for reading and writing to a removable non-volatile optical disk (such as Compact Disc Read Only Memory (hereinafter referred to as: CD-ROM), Digital Video Disc Read Only Memory (hereinafter referred to as: DVD-ROM) or other optical media) may be provided. In these cases, each drive may be connected to the communication bus 540 through one or more data medium interfaces. The memory 520 may include at least one program product having a set (such as at least one) of program modules configured to perform the functions of the embodiments of the present application.
[0155] A program / utility with a set (at least one) of program modules may be stored in the memory 520. Such program modules include - but are not limited to - an operating system, one or more application programs, other program modules, and program data. Each or some combination of these examples may include the implementation of a network environment. The program modules generally perform the functions and / or methods in the embodiments described in the present application.
[0156] The computing device 500 may also communicate with one or more external devices (such as a keyboard, a pointing device, a display, etc.), and may also communicate with one or more devices that enable a user to interact with the computing device, and / or communicate with any device that enables the computing device to communicate with one or more other computing devices (such as a network card, a modem, etc.). Such communication may be carried out through the communication interface 530. And, the computing device 500 may also communicate with one or more networks (such as a Local Area Network (hereinafter referred to as: LAN), a Wide Area Network (hereinafter referred to as: WAN) and / or a public network, such as the Internet) through a network adapter ( Figure 7 not shown in the figure), and the above-mentioned network adapter may communicate with other modules of the computing device through the communication bus 540. It should be understood that although Figure 7 not shown in the figure, other hardware and / or software modules may be used in combination with the computing device 500, including but not limited to: microcode, device drivers, redundant processing units, external disk drive arrays, Redundant Arrays of Independent Drives (hereinafter referred to as: RAID) systems, tape drives, and data backup storage systems, etc.
[0157] The processor 510 executes various functional applications and data processing by running programs stored in the memory 520, such as implementing the resource scheduling method provided in the embodiments of the present application.
[0158] It can be understood that the interface connection relationships between the modules illustrated in the embodiments of the present application are only illustrative descriptions and do not constitute a structural limitation on the computing device 500. In other embodiments of the present application, the computing device 500 may also adopt different interface connection methods in the above embodiments, or a combination of multiple interface connection methods.
[0159] It can be understood that in order to implement the above functions, the above computing device and the like include corresponding hardware structures and / or software modules for executing each function. Those skilled in the art should easily realize that, in combination with the exemplary units and algorithm steps described in the embodiments disclosed herein, the embodiments of the present application can be implemented in the form of hardware or a combination of hardware and computer software. Whether a certain function is executed in the form of hardware or computer software driving the hardware depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the embodiments of the present application.
[0160] The embodiments of the present application can perform functional module division on the above computing device and the like according to the above method examples. For example, each functional module can be divided corresponding to each function, or two or more functions can be integrated into one processing module. The above integrated modules can be implemented in the form of hardware or in the form of software functional modules. It should be noted that the division of modules in the embodiments of the present application is illustrative, merely a logical functional division, and there may be other division methods in actual implementation.
[0161] The embodiments of the present application further provide a storage medium, on which computer program instructions are stored. When the computer program instructions are executed by a computing device, the computing device is enabled to implement the resource scheduling method as described above.
[0162] The embodiments of the present application further provide a program product, which includes a computer program. When at least one processor executes the computer program, at least one processor is enabled to execute the resource scheduling method provided in the embodiments of the present application.
[0163] The computing device, storage medium or computer program product provided in the embodiments of the present application are all used to execute the corresponding methods provided above. Therefore, the beneficial effects that can be achieved can refer to the beneficial effects in the corresponding methods provided above, and will not be elaborated here.
[0164] Through the description of the above embodiments, those skilled in the art can clearly understand that for the convenience and brevity of description, only the division of the above functional modules is used as an example. In actual applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device is divided into different functional modules to complete all or part of the functions described above. For the specific working processes of the system, device, and unit described above, reference can be made to the corresponding processes in the foregoing method embodiments, which will not be elaborated here.
[0165] In each embodiment of this application, each functional unit can be integrated into a processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above integrated unit can be implemented in the form of hardware or in the form of a software functional unit.
[0166] If the above integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application embodiment, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) or a processor to execute all or part of the steps of the method described in each embodiment of this application. The foregoing storage medium includes: various media such as flash memory, mobile hard disk, read-only memory, random access memory, magnetic disk, or optical disc that can store program codes.
[0167] The above is only the specific implementation manner of this application, but the protection scope of this application is not limited thereto. Any changes or substitutions within the technical scope disclosed in this application should be covered by the protection scope of this application. Therefore, the protection scope of this application should be subject to the protection scope of the claimed rights.
Claims
1. A resource scheduling method, characterized in that, The method includes: Obtaining a resource application request, where the resource application request includes the number of processors; Determining a processor set according to the resource application request; Traversing the processor set, and determining a scheduling weight value for each processor set according to a first communication performance metric and a second communication performance metric corresponding to each processor set; the first communication performance metric represents the communication performance between different processors within the processor set, and the second communication performance metric represents the communication performance between each processor within the processor set and each processor outside the processor set; Determining a target processor set according to the scheduling weight value of each processor set, where the target processor set is selected from the processor sets in which the first communication performance metric is greater than or equal to a preset metric threshold among all processor sets.
2. The method according to claim 1, characterized in that, Before the step of traversing the processor set and determining a scheduling weight value for each processor set according to a first communication performance metric and a second communication performance metric corresponding to each processor set, it further includes: Obtaining the communication performance between different processors in a processor resource pool; the processor resource pool includes processors within and outside the processor set; Determining the first communication performance metric and the second communication performance metric corresponding to each processor set according to the communication performance between different processors in the processor resource pool.
3. The method according to claim 2, characterized in that, The obtaining of the communication performance between different processors in the processor resource pool includes: Obtaining a processor topology, where the processor topology is used to represent the physical links between different processors in the processor resource pool; Testing the communication performance of each physical link in the processor topology to obtain the communication performance between different processors in the processor resource pool.
4. The method according to claim 3, wherein The communication performance includes bandwidth; the first communication performance metric is positively correlated with the sum of the bandwidths of the first type of physical links, where the first type of physical links are the physical links between different processors within each processor set; And / or The communication performance includes latency; the first communication performance metric is negatively correlated with the sum of the latencies of the first type of physical links.
5. The method according to claim 3, wherein The communication performance includes bandwidth; the second communication performance metric is positively correlated with the sum of the bandwidths of the second type of physical links, where the second type of physical links are the physical links between each processor within each processor set and each idle processor outside the processor set; And / or The communication performance includes latency; the second communication performance metric is negatively correlated with the sum of the latencies of the second type of physical links.
6. The method according to any one of claims 1 to 5, characterized in that, The determining of the target processor set according to the scheduling weight value of each processor set includes: Determining the processor set with the largest scheduling weight value from the processor sets in which the first communication performance metric is greater than or equal to a preset metric threshold among all processor sets; Determining the processor set with the largest scheduling weight value as the target processor set.
7. The method according to claim 6, wherein The scheduling weight is the difference between the first communication performance metric and the second communication performance metric.
8. The method according to any one of claims 1-7, characterized in that, The processor includes a graphics processing unit (GPU).
9. A resource scheduling device, characterized in that, The device includes: an acquisition module, configured to acquire a resource application request, where the resource application request includes the number of processors; a first determination module, configured to determine a processor set according to the resource application request; a second determination module, configured to traverse the processor set, and determine the scheduling weight of each processor set according to the first communication performance metric and the second communication performance metric corresponding to each processor set; the first communication performance metric represents the communication performance between different processors within the processor set, and the second communication performance metric represents the communication performance between each processor within the processor set and each processor outside the processor set; a third determination module, configured to determine a target processor set according to the scheduling weight of each processor set, where the target processor set is selected from the processor sets in which the first communication performance metric is greater than or equal to a preset metric threshold among all processor sets.
10. A computing device, characterized in that, including: a processor, a memory for storing processor-executable instructions; When the processor is configured to execute the instructions, the computing device implements the method according to any one of claims 1-8.