Computing power resource allocation method and device, nonvolatile storage medium and electronic equipment
By optimizing computing power resource allocation through a quadratic random selection algorithm and Gini coefficient analysis, the problem of high resource scheduling complexity in existing technologies is solved, and efficient and balanced resource allocation is achieved.
Patent Information
- Application Number
- CN202511217646.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-28
- Publication Date
- 2025-09-30
- Estimated Expiration
- 2045-08-28
AI Technical Summary
The computing resource scheduling methods in the prior art are too complex, resulting in resource allocation delays and low allocation efficiency, especially when there are a large number of computing devices.
A quadratic random selection algorithm is used to randomly select two computing devices and compare their load information. The device with a smaller load is selected as the target computing device. The resource allocation strategy is optimized by combining the Gini coefficient and cluster analysis to reduce computational complexity and improve allocation efficiency.
It reduces the complexity of the computing resource allocation process, improves allocation efficiency, ensures resource balance and stability, and adapts to high-concurrency request scenarios.
Smart Images

Figure CN120723477A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of cloud computing, and more specifically, to a computing resource allocation method, device, non-volatile storage medium, and electronic device. Background Art
[0002] In the related art, scheduling computing devices typically involves scoring each available device and selecting the best one to handle the computing task. This approach is complex when there are a large number of computing devices, leading to delayed and inefficient allocation of computing resources.
[0003] To address the above-mentioned problems, no effective solutions have been proposed so far. Summary of the Invention
[0004] The embodiments of the present application provide a computing power resource allocation method, device, non-volatile storage medium and electronic device to at least solve the technical problems of computing power resource allocation delay and low allocation efficiency caused by the high complexity of computing power resource scheduling methods in related technologies.
[0005] According to one aspect of an embodiment of the present application, a computing power resource allocation method is provided, including: determining a request type of a resource request; when the request type is a load balancing type, randomly selecting a first alternative computing device from a computing power resource pool, and after selecting the first alternative computing device, randomly selecting a second alternative computing device from the computing power resource pool; comparing load information of the first alternative computing device and the second alternative computing device, and determining that the computing device with a smaller load between the first alternative computing device and the second alternative computing device is the target computing device corresponding to the resource request.
[0006] Optionally, randomly selecting a first alternative computing device from the computing power resource pool, and after selecting the first alternative computing device, randomly selecting a second alternative computing device from the computing power resource pool includes: determining the selection probability of each computing device based on the load task situation of each computing device in the computing power resource pool; randomly selecting the first alternative computing device from the computing power resource pool based on the selection probability, and after selecting the first alternative computing device, randomly selecting the second alternative computing device from the computing power resource pool based on the selection probability.
[0007] Optionally, determining the probability of each computing device being selected based on the load task conditions of each computing device in the computing power resource pool includes: determining the number of load tasks of each computing device within a first preset time period; determining the probability of the computing device being selected based on the number of load tasks of the computing device within the first preset time period, wherein the probability of being selected is negatively correlated with the number of load tasks.
[0008] Optionally, the computing power resource allocation method also includes: determining the Gini coefficient corresponding to the computing power resource pool within a second preset time period, wherein the Gini coefficient is used to reflect the degree of balance of the load distribution of each computing device in the computing power resource pool, and the higher the Gini coefficient, the more unbalanced the load distribution; when the Gini coefficient is greater than a preset coefficient threshold, determining the target computing device in the computing power resource pool, wherein the target computing device is a computing device whose number of load tasks within the second preset time period is lower than a preset threshold; and adjusting the selection method of the first alternative computing device and the second alternative computing device based on the target computing device.
[0009] Optionally, adjusting the selection method of the first candidate computing device and the second candidate computing device according to the target computing device includes: increasing the probability that the target computing device is selected as the first candidate computing device or the second candidate computing device.
[0010] Optionally, adjusting the selection method of the first alternative computing device and the second alternative computing device based on the target computing device includes: clustering the target computing device; determining target hardware characteristics of the target computing device based on the clustering result, wherein the hardware characteristics are hardware characteristics shared by the target computing devices in the same cluster; and adjusting the selection method of the first alternative computing device and the second alternative computing device based on the target hardware characteristics.
[0011] Optionally, the load information includes at least one of the following: power consumption of the computing device, core usage of the computing device, number of tasks carried by the computing device, and video memory usage of the computing device.
[0012] According to another aspect of an embodiment of the present application, a computing power resource allocation device is also provided, including: a first processing module for determining the request type of a resource request; a second processing module for randomly selecting a first alternative computing device from a computing power resource pool when the request type is a load balancing type, and after selecting the first alternative computing device, randomly selecting a second alternative computing device from the computing power resource pool; a third processing module for comparing the load information of the first alternative computing device and the second alternative computing device, and determining that the computing device with a smaller load between the first alternative computing device and the second alternative computing device is the target computing device corresponding to the resource request.
[0013] According to another aspect of an embodiment of the present application, a non-volatile storage medium is provided, in which a program is stored. When the program is running, the device where the non-volatile storage medium is located is controlled to execute the computing power resource allocation method.
[0014] According to another aspect of an embodiment of the present application, an electronic device is further provided, including: a memory and a processor, the processor being configured to run a program stored in the memory, wherein the computing power resource allocation method is executed when the program is run.
[0015] According to another aspect of an embodiment of the present application, a computer program product is also provided, including a computer program, which implements a computing power resource allocation method when executed by a processor.
[0016] In an embodiment of the present application, a method is adopted to determine the request type of a resource request; when the request type is a load balancing type, a first alternative computing device is randomly selected from the computing power resource pool, and after the first alternative computing device is selected, a second alternative computing device is randomly selected from the computing power resource pool; the load information of the first alternative computing device and the second alternative computing device is compared, and the computing device with the smaller load in the first alternative computing device and the second alternative computing device is determined to be the target computing device corresponding to the resource request. By randomly selecting alternative computing devices twice and determining the target alternative computing device according to the load information, the purpose of reducing the complexity of the computing power resource allocation process is achieved, thereby achieving the technical effect of reducing allocation delay and improving allocation efficiency, and further solving the technical problem of computing power resource allocation delay and low allocation efficiency caused by the high complexity of the computing power resource scheduling method in the related art. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation on the present application. In the drawings:
[0018] Figure 1 This is a schematic diagram of the structure of a computer terminal (or mobile device) provided according to an embodiment of the present application;
[0019] Figure 2 This is a flowchart of a computing resource allocation method provided in accordance with an embodiment of the present application;
[0020] Figure 3 This is a schematic diagram comparing the calculation results of the Gini coefficient provided in an embodiment of the present application;
[0021] Figure 4 This is a flowchart of a computing resource allocation process provided according to an embodiment of the present application;
[0022] Figure 5 It is a structural diagram of a computing power resource allocation device provided according to an embodiment of the present application. DETAILED DESCRIPTION
[0023] In order to enable those skilled in the art to better understand the present invention, the following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments in the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts should fall within the scope of protection of this application.
[0024] It should be noted that the terms "first", "second", etc. in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequential order. It should be understood that the data used in this way can be interchangeable where appropriate, so that the embodiments of the present application described herein can be implemented in a sequence other than those illustrated or described herein. In addition, the terms "including" and "having" and any of their variations are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device comprising a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.
[0025] In order to better understand the embodiments of the present application, the technical terms involved in the embodiments of the present application are explained as follows:
[0026] Resource preselection: One of the two phases of computing power scheduling, this phase aims to screen all nodes that meet basic scheduling criteria and generate a "viable node list" (which may be empty or contain multiple nodes). This is a hard condition check; nodes that fail any of the conditions are excluded from subsequent sorting. The preselection phase primarily considers whether the node meets the basic resource requirements, network connectivity requirements, hardware constraints, and other factors for running the task.
[0027] Resource Optimization: One of the two phases of computing power scheduling, this phase prioritizes the list of viable nodes generated in the pre-selection phase and selects the most suitable nodes (typically those with the highest scores or lowest loads). This is a soft-condition evaluation, quantifying the "goodness" of nodes through weighted scoring, allowing for flexible policy configuration. The optimization phase focuses on the node's current load, resource utilization efficiency, and historical performance.
[0028] Power of Two Random Choices (P2RC) is an efficient load balancing algorithm. Its core idea is to significantly reduce the maximum load in the system by randomly selecting two candidate nodes and choosing the one with the lighter load. This patented innovation applies this algorithm to resource allocation for tasks that require evenly distributed computing power during the computing power scheduling optimization phase, optimizing the complexity from O(n*m) to O(1) while maintaining balanced resource allocation.
[0029] Spread strategy: This is a computing power optimization strategy that aims to distribute computing tasks as evenly as possible across all computing nodes or cards to avoid hotspots and bottlenecks caused by resource concentration. This strategy is suitable for scenarios that require balanced load distribution and improved system fault tolerance and stability.
[0030] Gini coefficient: A statistical indicator that quantifies the degree of distribution imbalance. It was originally used in economics to measure the degree of income inequality. In this patent, it is used to evaluate the balance of computing resource allocation. The closer the Gini coefficient is to 0, the more balanced the resource allocation.
[0031] Resource scheduling algorithms in related technologies typically divide resource scheduling into two phases: pre-selection and optimization. During the optimization phase, each pre-selected node must be scored. In large-scale cluster environments, the time complexity of this resource calculation for each node is O(n). When considering multiple compute cards within a node (such as GPUs and NPUs), the complexity increases to O(n*m), where n is the number of nodes and m is the number of compute cards per node. This linearly increasing computational complexity significantly degrades scheduler performance in large-scale cluster environments, leading to delayed and inefficient resource allocation.
[0032] Moreover, although the fast scheduling methods in related technologies (such as random selection) can reduce computational complexity, it is difficult to ensure the balance of resource allocation, which can easily lead to overload of some nodes and idleness of other nodes, reducing overall resource utilization and shortening the service life of hardware; while algorithms that pursue absolute balance need to consider the global resource status, have high computational complexity, and are difficult to cope with high-concurrency scheduling requests.
[0033] In addition, the scheduling methods in related technologies often require the use of completely different algorithms when implementing different scheduling strategies (such as average allocation, centralized allocation, etc.), lacking unified architectural support, which increases the difficulty of system maintenance and the complexity of strategy switching.
[0034] In order to solve the above problems, relevant solutions are provided in the embodiments of the present application, which are described in detail below.
[0035] According to an embodiment of the present application, a method embodiment of a computing power resource allocation method is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer executable instructions, and although a logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in an order different from that shown here.
[0036] The method embodiments provided in the embodiments of the present application can be executed in a mobile terminal, a computer terminal or a similar computing device. Figure 1 The following is a hardware block diagram of a computer terminal (or mobile device) for implementing a computing resource allocation method. Figure 1 As shown, the computer terminal 10 (or mobile device 10) may include one or more processors 102 (illustrated as 102a, 102b, ..., 102n in the figure) (the processor 102 may include, but is not limited to, a processing device such as a microprocessor MCU or a programmable logic device FPGA), a memory 104 for storing data, and a transmission device 106 for communication functions. In addition, it may also include: a display, an input / output interface (I / O interface), a universal serial bus (USB) port (which may be included as one of the ports of the BUS bus), a network interface, a power supply, and / or a camera. Those skilled in the art will understand that Figure 1 The structure shown is only for illustration and does not limit the structure of the above electronic device. Figure 1 More or fewer components than shown, or with Figure 1 Different configurations shown.
[0037] It should be noted that the one or more processors 102 and / or other data processing circuits described above may generally be referred to herein as "data processing circuitry." This data processing circuitry may be embodied in whole or in part as software, hardware, firmware, or any other combination thereof. Furthermore, the data processing circuitry may be a single, independent processing module, or may be fully or partially integrated into any of the other components of the computer terminal 10 (or mobile device 10). As discussed in the embodiments of this application, this data processing circuitry serves as a processor control (e.g., selecting a variable resistor terminal path connected to an interface).
[0038] The memory 104 may be used to store software programs and modules of application software, such as the program instructions / data storage device corresponding to the computing power resource allocation method in the embodiment of the present application. The processor 102 executes various functional applications and data processing by running the software programs and modules stored in the memory 104, that is, implementing the above-mentioned computing power resource allocation method. The memory 104 may include a high-speed random access memory and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some examples, the memory 104 may further include a memory remotely located relative to the processor 102, and these remote memories may be connected to the computer terminal 10 via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.
[0039] Transmission device 106 is configured to receive or transmit data via a network. A specific example of the aforementioned network may include a wireless network provided by the communications provider of computer terminal 10. In one embodiment, transmission device 106 includes a network interface controller (NIC), which can be connected to other network devices via a base station to enable communication with the Internet. In another embodiment, transmission device 106 may be a radio frequency (RF) module configured to communicate with the Internet wirelessly.
[0040] The display may be, for example, a touch screen liquid crystal display (LCD) that enables a user to interact with a user interface of the computer terminal 10 (or mobile device 10 ).
[0041] In the above operating environment, the embodiment of the present application provides a method for allocating computing resources, such as Figure 2 As shown, the method includes the following steps:
[0042] Step S202, determining the request type of the resource request;
[0043] In some embodiments of the present application, for executing Figure 2 The computing resource allocation method system shown includes a resource pre-selection engine, a resource optimization engine, a resource status manager, and a balance assessment system. The resource pre-selection engine is responsible for selecting nodes that meet basic requirements based on task requirements and hard constraints, generating a list of feasible nodes. This includes functional modules such as resource requirement matching, network connectivity checking, and hardware architecture compatibility verification.
[0044] The resource optimization engine is responsible for selecting the optimal node based on the pre-selection results and the configured scheduling strategy (such as Spread strategy). Figure 2The computing resource allocation method shown achieves load balancing with O(1) complexity.
[0045] The Resource Status Manager is used to maintain and update the resource usage of each node and compute card in the cluster, providing data support for scheduling decisions. Each node in the cluster contains multiple compute cards, which are a type of computing device.
[0046] The balance assessment system is used to quantitatively evaluate the degree of balance in resource allocation based on mathematical models such as weighted entropy and the Gini coefficient, providing a basis for optimizing scheduling strategies.
[0047] Step S204: If the request type is a load balancing type, randomly select a first candidate computing device from the computing resource pool, and after the first candidate computing device is selected, randomly select a second candidate computing device from the computing resource pool;
[0048] The above load balancing types include the spread type.
[0049] In the technical solution provided in step S204, the steps of randomly selecting a first alternative computing device from the computing power resource pool, and after selecting the first alternative computing device, randomly selecting a second alternative computing device from the computing power resource pool include: determining the selection probability of each computing device based on the load task situation of each computing device in the computing power resource pool; randomly selecting the first alternative computing device from the computing power resource pool based on the selection probability, and after selecting the first alternative computing device, randomly selecting the second alternative computing device from the computing power resource pool based on the selection probability.
[0050] As an optional implementation, determining the probability of each computing device being selected based on the load task conditions of each computing device in the computing power resource pool includes: determining the number of load tasks of each computing device within a first preset time period; determining the probability of the computing device being selected based on the number of load tasks of the computing device within the first preset time period, wherein the probability of selection is negatively correlated with the number of load tasks.
[0051] In some embodiments of the present application, when processing a resource request of the load balancing type, the current load task situation of each computing device in the computing resource pool is first analyzed. The analysis process involves collecting real-time resource usage data of the computing device, including but not limited to information such as the utilization rate of the computing core, memory usage, the number of tasks running, and device power consumption. Afterwards, the load level of each computing device can be calculated based on these data, and further converted into a probability of being selected. The specific conversion logic can be flexibly designed according to actual needs, such as using an inverse proportional function so that a computing device with a lower load has a higher probability of being selected, and vice versa.
[0052] Subsequently, the first candidate computing device is randomly selected from the computing resource pool based on the calculated selection probability, and the resource pool status is updated immediately to avoid repeated selections at the same time. Next, the second candidate computing device is randomly selected again based on the updated resource pool status and selection probability. In this way, the selection probability of a computing device changes dynamically during the two selection processes, reflecting the real-time resource utilization and task allocation, thereby improving the flexibility and responsiveness of resource scheduling.
[0053] Ultimately, by comparing the load information of the first and second candidate computing devices, the less-loaded computing device is selected as the target computing device to execute the current resource request. This approach not only ensures efficient resource utilization but also effectively balances task allocation between computing devices by dynamically adjusting the probability of selection, avoiding excessive resource concentration and further improving the overall stability and efficiency of the computing resource allocation process. This approach demonstrates particularly significant performance advantages when processing large-scale concurrent requests.
[0054] Step S206 , comparing the load information of the first candidate computing device and the second candidate computing device, and determining that the computing device with the smaller load among the first candidate computing device and the second candidate computing device is the target computing device corresponding to the resource request.
[0055] In the technical solution provided in step S206, the load information includes at least one of the following: computing device power consumption, computing device core usage, number of tasks carried by the computing device, and computing device video memory usage. Furthermore, if the computing device is a computing card, the load information may also include the computing card model.
[0056] In some embodiments of the present application, the process of allocating computing resources includes the following steps:
[0057] The first step is to define the function: First, define a function called power_of_two_random_choices, which receives a parameter cards, which is a list of all currently available calculation cards.
[0058] The second step is to check the length of the list: Check whether the length of the cards list is greater than or equal to 2. This is a prerequisite for performing a secondary random selection. If the list length is less than 2, go directly to step 4 to handle special cases.
[0059] Step 3: Randomly select two calculation cards (i.e., calculation devices): When the cards list contains at least two calculation cards, randomly select two different calculation cards from the list, denoted as card_a and card_b. Using random.sample(cards, k=2) ensures that each selection is random and the two cards are different.
[0060] Step 4: Handle special cases: If there is only one calculation card in the cards list or it is empty, directly return this card (if it exists) or return None (if the list is empty), indicating that there is no suitable calculation card for task assignment.
[0061] Step 5: Get the compute card load information: For two randomly selected compute cards, call the get_current_load() method to obtain their current load status. The load information obtained in this step is used for subsequent comparison and decision-making.
[0062] Step 6: Compare the loads and make a decision: Compare the load information of card_a and card_b, load_a and load_b. If load_a is less than load_b, card_a has a lower load, and the function returns card_a. If load_b is less than or equal to load_a, the function returns card_b, which is the card with the lower or equal load.
[0063] Step 7: Return the optimal computing card: Finally, the function returns the information of the computing card with lower load, completing a resource optimization process based on secondary random selection.
[0064] In the above process, by simplifying the decision-making process, the time complexity of large-scale cluster scheduling is optimized from O(n*m) to O(1), and significant advantages are shown in ensuring balanced resource allocation, thereby improving the efficiency of computing power scheduling and the overall performance of the system.
[0065] In some embodiments of the present application, the computing power resource allocation method also includes: determining the Gini coefficient corresponding to the computing power resource pool within a second preset time period, wherein the Gini coefficient is used to reflect the degree of balance of the load distribution of each computing device in the computing power resource pool, and the higher the Gini coefficient, the more unbalanced the load distribution; when the Gini coefficient is greater than the preset coefficient threshold, determining the target computing device in the computing power resource pool, wherein the target computing device is a computing device whose number of load tasks within the second preset time period is lower than the preset threshold; and adjusting the selection method of the first alternative computing device and the second alternative computing device based on the target computing device.
[0066] As an optional implementation, the step of adjusting the selection method of the first candidate computing device and the second candidate computing device according to the target computing device includes: increasing the probability of the target computing device being selected as the first candidate computing device or the second candidate computing device.
[0067] In some embodiments of the present application, the step of adjusting the selection method of the first alternative computing device and the second alternative computing device based on the target computing device includes: clustering the target computing device; determining the target hardware characteristics of the target computing device based on the clustering result, wherein the hardware characteristics are the hardware characteristics shared by the target computing devices in the same cluster; and adjusting the selection method of the first alternative computing device and the second alternative computing device based on the target hardware characteristics.
[0068] When it is determined that the Gini coefficient of the computing power resource pool exceeds the preset threshold, indicating that the load distribution imbalance among the computing devices has reached a level that requires intervention, the system further identifies the computing devices whose number of load tasks is lower than the preset threshold within the second preset time period, and uses them as target computing devices, in order to improve the overall load balance by optimizing the resource allocation strategy.
[0069] In some embodiments of the present application, the target computing devices may be clustered based on their hardware characteristics by executing the aforementioned clustering process. Hardware characteristics may include the model, performance level, power consumption, available memory size, network bandwidth, etc. The clustering process aims to identify and form clusters of computing devices with similar hardware characteristics. This allows for a more detailed understanding of the composition and capability distribution of computing devices in the resource pool.
[0070] Based on the target hardware characteristics determined by the clustering results, the selection method for the first and second candidate computing devices can be adjusted. Specifically, for a cluster of computing devices with the target hardware characteristics, the probability of that computing device being selected as a candidate can be increased. This adjustment strategy not only promotes the full utilization of hardware resources but also effectively reduces the Gini coefficient in the resource pool by directing tasks to computing devices with lower loads and similar hardware, thereby improving the balance of load distribution.
[0071] By dynamically adjusting the selection probability of computing devices based on the Gini coefficient and hardware characteristics, not only can the refined management and scheduling of computing resources be achieved, the flexibility and efficiency of resource scheduling can be improved, but also the stability and reliability of the system can be enhanced, ensuring that resource allocation can be maintained in the optimal state under changing load conditions, avoiding resource waste and the risk of equipment overload, and providing strong support for high concurrency and complex task processing.
[0072] Identifying the common features of devices with a low probability of being selected through clustering also has the following benefits:
[0073] Identifying resource bottlenecks: Cluster analysis can quickly identify which hardware characteristics (such as GPU model, number of CPU cores, memory capacity, etc.) are associated with low utilization. This provides clear direction for the optimization process, focusing on improving or prioritizing devices that utilize these specific hardware characteristics.
[0074] Precise Scheduling: Based on the clustering characteristics of computing devices, targeted resource scheduling strategies can be developed. For example, if a certain type of GPU is found to have low utilization due to compatibility issues, subsequent scheduling can prioritize compatible workloads for this type of device, or the algorithm can be adjusted to appropriately increase the number of scheduled devices while meeting other conditions, thereby improving overall resource utilization.
[0075] Improved balance: By analyzing and leveraging the common characteristics of devices, resource allocation can be more precisely controlled, preventing certain devices from being consistently neglected and reducing imbalances in resource allocation. For example, if a group of devices has a low load due to network latency, the Spread strategy can be used to increase the load on these devices by adjusting network connection optimization or priority settings, thereby balancing the network load across the entire cluster.
[0076] Policy Tuning: Cluster analysis provides data support for policy tuning. Administrators or automated policies can set different scheduling rules for different device groups based on clustering results, thereby better adapting to changing business environments and resource requirements, and enabling dynamic adjustment and optimization of policies.
[0077] Preventive measures: Understanding common device issues helps you take preventive measures in advance, such as performance optimization, software upgrades, or hardware replacement for devices with specific hardware characteristics, thereby fundamentally resolving the problem of low utilization.
[0078] In summary, identifying the common characteristics of devices with a low probability of being selected through clustering not only helps to deeply understand the heterogeneity and device utilization within the resource pool, but also serves as a basis for adjusting resource scheduling strategies to further improve resource utilization and load balancing.
[0079] In some embodiments of the present application, in order to verify the effectiveness of the computing power resource allocation method of the secondary random selection computing device in terms of resource allocation balance, a quantitative comparison is made between the single random allocation method in the related art and the computing power resource allocation method provided in the present application in the following manner:
[0080] 1. Large-scale simulation test: This simulates a scenario where 1 million tasks are assigned to 1,000 nodes, comparing the resource allocation results of the single randomization and quadratic randomization strategies. Note that the specific number of tasks and node data volumes used here are for illustrative purposes only and do not represent a limitation on the solution. Different numbers of tasks and nodes can be used based on actual circumstances.
[0081] 2. Distribution range analysis: By counting the number of tasks assigned to each node, we can analyze the fluctuation range of resource allocation under different strategies.
[0082] 3. Gini coefficient calculation: The Gini coefficient is used to quantify the degree of imbalance in resource allocation. The closer the value is to 0, the more balanced the allocation.
[0083] 4. Long-term operational stability: Evaluate whether the balance of resource allocation will change significantly during the long-term operation of the system.
[0084] For the above comparison test methods, the test results are as follows:
[0085] 1. The distribution of single random scheduling results is wide, with the number of tasks on a node ranging from 900 to 1200, and the maximum difference reaching 300 tasks.
[0086] 2. The distribution range of the method provided in the embodiment of the present application is significantly narrowed, with the number of tasks ranging between 995 and 1002, with the maximum difference being only 7 tasks, which is relatively more balanced.
[0087] 3. The calculated Gini coefficient is approximately 0.0173 for the single random method and approximately 0.0005 for the quadratic random method. The latter improves balance by approximately 97.19%.
[0088] 4. Through lightweight load snapshot comparison (two random selections + one comparison), it can still maintain a response time of seconds (1.60 seconds) even with a task volume of millions, meeting the real-time requirements of high-throughput scenarios.
[0089] In some embodiments of the present application, the final Gini coefficient result is as follows: Figure 3 shown. Figure 3The left side shows the distribution of resource allocation using the single-shot randomization strategy. Under this strategy, each task is randomly assigned to any compute node in the cluster, regardless of load differences between nodes. The horizontal axis of the chart represents "requests / server," or the number of tasks assigned to each server; the vertical axis represents "probability density," reflecting the distribution density of requests across different servers. The data points are concentrated between 975 and 1025, with a peak near 1000. This indicates that under the single-shot randomization strategy, the number of tasks distributed across servers is relatively wide. Although the average distribution is roughly around 1000, there is significant fluctuation, indicating poor resource allocation balance, with a Gini coefficient of 0.0173.
[0090] Figure 3 The right side shows the distribution of resource allocation when the method provided in this application is adopted. This strategy randomly selects two computing nodes for comparison in the optimization stage, and selects the node with lower load to allocate tasks. The chart also uses the horizontal axis to represent "number of requests / server" and the vertical axis to represent "probability density". The data points are concentrated between 999 and 1001, and the peak is also located near 1000, but the distribution range is significantly narrowed. This means that under the quadratic random strategy, the difference in the number of tasks between servers is greatly reduced, and resource allocation is more balanced. The Gini coefficient is reduced to 0.005, which shows that compared with the single random strategy, the quadratic random strategy has significantly improved the balance of resource allocation, showing its superiority in resource allocation.
[0091] pass Figure 3 , we can intuitively see the significant effect of the computing power resource allocation method provided by the embodiment of this application in terms of balanced resource allocation. It not only effectively reduces the difference in the number of tasks between servers, but also maintains a fast allocation speed. It is a powerful tool for improving resource utilization and system stability in large-scale cluster environments. The application of this strategy helps to avoid excessive concentration of resources in a few nodes, thereby preventing the formation of system bottlenecks and ensuring the rational and efficient use of resources in the cluster.
[0092] In some embodiments of the present application, the following simulation test process is also provided:
[0093] The first step is to initialize the environment:
[0094] First, set the parameters of the simulation environment, including the total number of task requests (NUM_REQUESTS), the number of servers or computing nodes (NUM_SERVERS), and the random seed (SEED) to ensure reproducible results.
[0095] The second step is a single random strategy test:
[0096] Outputs a message indicating the start of a single random strategy execution. Use the time module to record the start time of the test. Use the np.random.choice function to randomly select NUM_REQUESTS servers from NUM_SERVERS to simulate the task allocation process. Use the np.bincount function to count the number of times each server is assigned a task, generating a task allocation count. Record the time again, and calculate and output the execution time of the single random strategy.
[0097] The third step is to test the secondary random strategy (that is, the computing power resource allocation method provided in the embodiment of the present application):
[0098] After completing a single random test, output a message announcing the start of the secondary randomization strategy. Also record the start time to prepare for calculating the execution time of the secondary randomization strategy. Use the np.random.choice function to generate an array of size (NUM_REQUESTS, 2) to simulate the process of randomly selecting different servers twice. Initialize an array called counts_twice to count the number of tasks assigned to each server. Loop through each pair of candidate servers, comparing their current number of assigned tasks and assigning the task to the server with fewer tasks. Calculate and output the execution time of the above process.
[0099] Step 4: Analysis and Evaluation:
[0100] After the test is complete, the results of the two strategies are analyzed. This includes comparing the distribution of tasks across servers using histograms, observing the concentration and distribution range of data points, and then evaluating the balance of resource allocation.
[0101] The Gini coefficients of resource allocation under the two strategies were calculated. This is a statistical indicator of unequal distribution. The closer the Gini coefficient is to 0, the more balanced the resource allocation. By comparing the Gini coefficients of the two strategies, we evaluate the effectiveness of the quadratic random selection strategy in improving resource allocation balance and demonstrate its superiority over the single random selection strategy.
[0102] Step 5: Draw conclusions:
[0103] Finally, by observing and analyzing the test results, such as the changes in execution time and Gini coefficient, we conclude that the quadratic random strategy not only performs well in execution efficiency, but also significantly improves the balance of resource allocation, and is an effective method to optimize the scheduling performance of large-scale clusters.
[0104] In some embodiments of the present application, there is also provided a Figure 4 The computing power resource allocation process shown. The components involved in executing this computing power resource allocation process may include:
[0105] Resource preselection engine: This engine is responsible for selecting nodes or computing cards in the cluster that meet basic scheduling requirements based on task requirements and hard constraints, generating a "feasible node list." This stage focuses on matching nodes' resource requirements, network connectivity, and hardware architecture compatibility.
[0106] Resource Optimization Engine: Based on pre-selection results, it selects the optimal resource node according to the configured scheduling strategy. For Spread-type scheduling strategies, the optimization engine implements a quadratic random selection algorithm, which can achieve highly efficient load balancing with O(1) time complexity.
[0107] Resource Status Manager: Continuously monitors and updates the real-time resource usage of each node and compute card in the cluster, providing real-time data support for scheduling decisions. This includes key indicators such as CPU usage, GPU load, memory usage, and network bandwidth.
[0108] Balance Assessment System: This system uses mathematical models, such as weighted entropy and the Gini coefficient, to quantitatively assess the balance of resource allocation. This system provides data for optimizing scheduling strategies, ensuring efficient and balanced resource allocation within the cluster.
[0109] Figure 4 The computing power resource allocation process shown includes the following steps:
[0110] The first step is scheduling request reception: the scheduling system receives computing power scheduling requests from different sources. These requests usually contain the type and quantity of required resources and any specific hardware requirements.
[0111] The second step is resource preselection: The resource preselection engine selects nodes that meet the requirements from all nodes in the cluster based on the conditions in the scheduling request, generating a "feasible node list." This process ensures that only nodes with the hardware and resources required to execute the task are considered.
[0112] Step 3: Strategy determination: The system determines whether the current scheduling request is a Spread strategy requirement. If so, it invokes the quadratic random selection algorithm; if not, it invokes a resource optimization algorithm suitable for other strategies.
[0113] Step 4: Secondary random selection execution: For requests using the Spread strategy, the resource optimization engine randomly selects two compute cards from a pre-selected node list, compares their load information, and selects the one with the lower load for task allocation.
[0114] Step 5: Scheduling result output: After the optimization process is completed, the system will output a node information representing the optimal resource allocation for the actual deployment of subsequent tasks.
[0115] Step 6, resource status update: After each task is assigned, the resource status manager will update the resource usage status of the assigned node to ensure that subsequent scheduling decisions are based on the latest resource information.
[0116] Step 7: Balance Assessment and Feedback: The balance assessment system regularly calculates cluster resource allocation balance indicators, such as the Gini coefficient, to evaluate the effectiveness of the scheduling algorithm. This data can be used to optimize the algorithm and adjust the strategy to ensure long-term balanced resource allocation.
[0117] By adopting a method of determining the request type of a resource request; when the request type is a load balancing type, randomly selecting a first alternative computing device from the computing power resource pool, and after selecting the first alternative computing device, randomly selecting a second alternative computing device from the computing power resource pool; comparing the load information of the first alternative computing device and the second alternative computing device, and determining that the computing device with the smaller load in the first alternative computing device and the second alternative computing device is the target computing device corresponding to the resource request, the purpose of reducing the complexity of the computing power resource allocation process is achieved by randomly selecting alternative computing devices twice and determining the target alternative computing device according to the load information, thereby achieving the technical effect of reducing allocation delay and improving allocation efficiency, and further solving the technical problem of computing power resource allocation delay and low allocation efficiency caused by the high complexity of the computing power resource scheduling method in the related art.
[0118] In addition, compared with related technologies, the method provided in the embodiments of the present application also has the following technical effects: 1) optimizing the computing power scheduling algorithm based on secondary randomness to improve the platform computing power scheduling capability; 2) a balance evaluation system based on a mathematical model; 3) a cross-domain resource coordination mechanism; 4) intelligent processing of boundary conditions and a high concurrency security mechanism.
[0119] The practical effects of the methods provided in the embodiments of this application include: 1) reducing the computational complexity of the spread operator in the computing power scheduling algorithm from linear complexity O(N*M) to constant time complexity O(1), reducing platform scheduling latency and improving overall scheduling performance and throughput. 2) establishing a quantitative evaluation system for resource allocation balance based on the Gini coefficient. Simulation experiments verified that the quadratic random selection method significantly reduced the resource distribution range from the traditional 900-1200 to 995-1002, improving balance by approximately 85%.
[0120] The present invention provides a computing resource allocation device. Figure 5 It is a schematic diagram of the structure of the device. Figure 5It can be seen that the device includes: a first processing module 50, which is used to determine the request type of the resource request; a second processing module 52, which is used to randomly select a first alternative computing device from the computing power resource pool when the request type is a load balancing type, and after selecting the first alternative computing device, randomly select a second alternative computing device from the computing power resource pool; a third processing module 54, which is used to compare the load information of the first alternative computing device and the second alternative computing device, and determine that the computing device with a smaller load in the first alternative computing device and the second alternative computing device is the target computing device corresponding to the resource request.
[0121] In some embodiments of the present application, the first processing module 50 randomly selects a first alternative computing device from the computing power resource pool, and after selecting the first alternative computing device, randomly selects a second alternative computing device from the computing power resource pool. The steps include: determining the selection probability of each computing device based on the load task situation of each computing device in the computing power resource pool; randomly selecting the first alternative computing device from the computing power resource pool based on the selection probability, and after selecting the first alternative computing device, randomly selecting the second alternative computing device from the computing power resource pool based on the selection probability.
[0122] In some embodiments of the present application, the first processing module 50 determines the probability of each computing device being selected based on the load task conditions of each computing device in the computing power resource pool, including the following steps: determining the number of load tasks of each computing device within a first preset time period; determining the probability of the computing device being selected based on the number of load tasks of the computing device within the first preset time period, wherein the probability of selection is negatively correlated with the number of load tasks.
[0123] In some embodiments of the present application, the load information includes at least one of the following: power consumption of the computing device, core usage of the computing device, number of tasks carried by the computing device, and video memory usage of the computing device.
[0124] In some embodiments of the present application, the computing power resource allocation device is also used to: determine the Gini coefficient corresponding to the computing power resource pool within a second preset time period, wherein the Gini coefficient is used to reflect the degree of balance of the load distribution of each computing device in the computing power resource pool, and the higher the Gini coefficient, the more unbalanced the load distribution; when the Gini coefficient is greater than the preset coefficient threshold, determine the target computing device in the computing power resource pool, wherein the target computing device is a computing device whose number of load tasks within the second preset time period is lower than the preset threshold; and adjust the selection method of the first alternative computing device and the second alternative computing device based on the target computing device.
[0125] In some embodiments of the present application, the step of the computing resource allocation device adjusting the selection method of the first alternative computing device and the second alternative computing device based on the target computing device includes: increasing the probability of the target computing device being selected as the first alternative computing device or the second alternative computing device.
[0126] In some embodiments of the present application, the steps of the computing resource allocation device adjusting the selection method of the first alternative computing device and the second alternative computing device based on the target computing device include: clustering the target computing device; determining the target hardware characteristics of the target computing device based on the clustering results, wherein the hardware characteristics are the hardware characteristics common to the target computing devices in the same cluster; and adjusting the selection method of the first alternative computing device and the second alternative computing device based on the target hardware characteristics.
[0127] It should be noted that the various modules in the above-mentioned computing resource allocation device can be program modules (for example, a set of program instructions that implement a certain specific function) or hardware modules. For the latter, it can be expressed in the following forms, but is not limited to this: the expression form of each of the above-mentioned modules is a processor, or the functions of each of the above-mentioned modules are implemented by a processor.
[0128] According to an embodiment of the present application, a non-volatile storage medium is also provided, in which a program is stored, wherein when the program is running, the device where the non-volatile storage medium is located is controlled to execute the following computing power resource allocation method: determine the request type of the resource request; when the request type is a load balancing type, randomly select a first alternative computing device from the computing power resource pool, and after selecting the first alternative computing device, randomly select a second alternative computing device from the computing power resource pool; compare the load information of the first alternative computing device and the second alternative computing device, and determine that the computing device with the smaller load in the first alternative computing device and the second alternative computing device is the target computing device corresponding to the resource request.
[0129] According to an embodiment of the present application, an electronic device is also provided, including a memory and a processor, the processor being used to run a program stored in the memory, wherein the following computing power resource allocation method is executed when the program is running: determining the request type of the resource request; when the request type is a load balancing type, randomly selecting a first alternative computing device from the computing power resource pool, and after selecting the first alternative computing device, randomly selecting a second alternative computing device from the computing power resource pool; comparing the load information of the first alternative computing device and the second alternative computing device, and determining that the computing device with the smaller load among the first alternative computing device and the second alternative computing device is the target computing device corresponding to the resource request.
[0130] According to an embodiment of the present application, a computer program product is also provided, including a computer program, which implements the following computing power resource allocation method when executed by a processor: determining the request type of the resource request; when the request type is a load balancing type, randomly selecting a first alternative computing device from the computing power resource pool, and after selecting the first alternative computing device, randomly selecting a second alternative computing device from the computing power resource pool; comparing the load information of the first alternative computing device and the second alternative computing device, and determining that the computing device with the smaller load in the first alternative computing device and the second alternative computing device is the target computing device corresponding to the resource request.
[0131] In the above embodiments of the present application, the description of each embodiment has its own focus. For parts that are not described in detail in a certain embodiment, please refer to the relevant description of other embodiments.
[0132] In the several embodiments provided in this application, it should be understood that the disclosed technical content can be implemented in other ways. Among them, the device embodiments described above are only exemplary. For example, the division of the units can be a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of units or modules, which can be electrical or other forms.
[0133] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple units. Some or all of the units may be selected according to actual needs to achieve the purpose of the present embodiment.
[0134] In addition, the functional units in the various embodiments of the present application may be integrated into a single processing unit, or each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.
[0135] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, or the part that contributes to the relevant technology, or all or part of the technical solution can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes a number of instructions for enabling a computer device (which can be a personal computer, server or network device, etc.) to execute all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes: U disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), mobile hard disk, magnetic disk or optical disk, etc. Various media that can store program code.
[0136] The above is only a preferred embodiment of the present application. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present application. These improvements and modifications should also be regarded as the scope of protection of the present application.
Claims
1. A computing resource allocation method, characterized in that: include: Determine the request type for the resource request; When the request type is a load balancing type, randomly selecting a first candidate computing device from a computing resource pool, and after selecting the first candidate computing device, randomly selecting a second candidate computing device from the computing resource pool; The load information of the first candidate computing device and the second candidate computing device are compared, and the computing device with the smaller load among the first candidate computing device and the second candidate computing device is determined to be the target computing device corresponding to the resource request.
2. The computing resource allocation method according to claim 1, characterized in that: Randomly selecting a first candidate computing device from a computing power resource pool, and after selecting the first candidate computing device, randomly selecting a second candidate computing device from the computing power resource pool includes: Determining the probability of each computing device being selected based on the load task status of each computing device in the computing power resource pool; A first candidate computing device is randomly selected from the computing power resource pool according to the selection probability, and after the first candidate computing device is selected, a second candidate computing device is randomly selected from the computing power resource pool according to the selection probability.
3. The computing resource allocation method according to claim 2, characterized in that: Determining the selection probability of each computing device according to the load task status of each computing device in the computing power resource pool includes: Determining the number of load tasks of each of the computing devices within a first preset time period; The selection probability of the computing device is determined according to the number of the load tasks of the computing device within the first preset time period, wherein the selection probability is negatively correlated with the number of the load tasks.
4. The computing resource allocation method according to claim 1, characterized in that: The computing power resource allocation method further includes: Determining a Gini coefficient corresponding to the computing power resource pool within a second preset time period, wherein the Gini coefficient is used to reflect the degree of load balance of each computing device in the computing power resource pool, and a higher Gini coefficient indicates a more unbalanced load distribution; When the Gini coefficient is greater than a preset coefficient threshold, determining a target computing device in the computing resource pool, wherein the target computing device is a computing device whose number of load tasks in the second preset time period is lower than a preset threshold; A method for selecting the first candidate computing device and the second candidate computing device is adjusted according to the target computing device.
5. The computing resource allocation method according to claim 4, characterized in that: The method of adjusting and selecting the first candidate computing device and the second candidate computing device according to the target computing device includes: The probability of the target computing device being selected as the first candidate computing device or the second candidate computing device is increased.
6. The computing resource allocation method according to claim 4, characterized in that: The method of adjusting and selecting the first candidate computing device and the second candidate computing device according to the target computing device includes: performing clustering processing on the target computing device; Determining target hardware features of the target computing device according to the clustering result, wherein the hardware features are hardware features shared by the target computing devices in the same cluster; The method for selecting the first candidate computing device and the second candidate computing device is adjusted according to the target hardware characteristics.
7. The computing resource allocation method according to claim 1, characterized in that: The load information includes at least one of the following: power consumption of the computing device, core usage of the computing device, number of tasks carried by the computing device, and video memory usage of the computing device.
8. A computing resource allocation device, characterized in that: include: A first processing module, configured to determine a request type of a resource request; a second processing module, configured to randomly select a first candidate computing device from a computing resource pool when the request type is a load balancing type, and after selecting the first candidate computing device, randomly select a second candidate computing device from the computing resource pool; The third processing module is used to compare the load information of the first candidate computing device and the second candidate computing device, and determine that the computing device with a smaller load between the first candidate computing device and the second candidate computing device is the target computing device corresponding to the resource request.
9. A non-volatile storage medium, characterized in that: The non-volatile storage medium stores a program, wherein when the program is running, the device where the non-volatile storage medium is located is controlled to execute the computing power resource allocation method according to any one of claims 1 to 7.
10. An electronic device, characterized in that: include: A memory and a processor, wherein the processor is used to run a program stored in the memory, wherein the computing power resource allocation method described in any one of claims 1 to 7 is executed when the program is run.
11. A computer program product, characterized in that It comprises a computer program which, when executed by a processor, implements the computing power resource allocation method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Cloud computing resource scheduling method and device, equipment and medium
CN115065685A
Energy intelligent management analysis method and system based on computing power service engine
CN118333432A
Fault diagnosis method and system for power distribution terminal
CN118707257A
Scheduling method and system based on cloud resource computing power distribution
CN119883633A
Computing power resource allocation method and device, storage medium and program product
CN120162150A
Cited By
Distributed cloud server load balancing optimization system and method based on edge computing
CN121334166A