Online resource allocation method and system based on multiplicative weight update

By using an online resource allocation method with multiplicative weight updates, the resource allocation weights are dynamically adjusted, resolving the conflict between resource utilization and QoS in cloud computing platforms. This achieves efficient and fair resource allocation, adapts to unsteady load changes, and improves system stability and resource utilization.

CN120909763AActive Publication Date: 2025-11-07INSTITUTE OF INFORMATION ENGINEERING CHINESE ACADEMY OF SCIENCES

Patent Information

Application Number
CN202510869928.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-26
Publication Date
2025-11-07
Estimated Expiration
2045-06-26

AI Technical Summary

Technical Problem

In existing cloud computing platforms, the irreconcilable contradiction between resource utilization and quality of service (QoS), insufficient dynamic feedback information, poor adaptability to non-steady-state demands, and lack of theoretical performance guarantees make it difficult for resource allocation strategies to balance efficiency and fairness in multi-user scenarios.

Method used

An online resource allocation method based on multiplicative weight updates is adopted. By obtaining binary feedback on whether the queue is empty, the resource allocation weight is dynamically adjusted. Combined with exponential weight growth and truncated simplex constraints, the dynamic scheduling and reasonable allocation of resources among multiple users are realized, ensuring service quality satisfaction and system stability.

Benefits of technology

It improves the overall resource utilization of the system, responds quickly to changes in non-steady load, avoids QoS violations and resource waste, enhances the fairness and stability of scheduling, and approaches the optimal resource utilization efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120909763A_ABST
    Figure CN120909763A_ABST
Patent Text Reader

Abstract

The invention discloses an online resource allocation method and system based on multiplicative weight updating, and belongs to the technical field of computer resource scheduling. In order to solve the problems that the resource utilization rate is low, the service quality is difficult to guarantee and the feedback information is limited in a multi-user sharing scene, a multiplicative weight updating strategy based on a binary queue state is mainly adopted, and the resource allocation proportion is dynamically adjusted in combination with a truncated simplex projection mechanism. According to the method, collaborative optimization of resource utilization efficiency maximization and service quality guarantee can be realized under the condition of only depending on the empty / non-empty state of the queue, and the method has good efficiency, robustness and fairness.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application belongs to the technical field of computer resource scheduling, and particularly relates to an online resource allocation method and system based on multiplicative weight update. BACKGROUND

[0002] In the current cloud computing platform, under the background of multi-tenant sharing of physical resources (such as CPU, memory, storage and network bandwidth), mainstream resource allocation methods focus on balancing efficiency and Quality of Service (QoS) through different strategies. Static allocation strategies meet user needs by predefining fixed quotas, for example, reserving 30%, 30% and 40% of CPU resources for three users respectively to ensure the basic performance guarantee of their task processing; greedy strategies dynamically allocate resources based on real-time queue status, preferentially tilting resources to current non-empty queues to maximize instantaneous utilization, for example, in a Kubernetes cluster, the scheduler allocates idle resources to high-load tasks based on queue activity; and proportional allocation strategies combine preset weights and real-time status to adjust resource allocation, for example, when the QoS weights of users A, B and C are 0.3, 0.3 and 0.4, the system dynamically allocates resources in a ratio of 3:3:4, which adapts to user priorities and accommodates dynamic needs, especially in high-volatility scenarios such as video streaming, through the linkage of weight ratio and queue activity to achieve flexible resource adaptation. These methods are all centered around improving resource utilization efficiency and guaranteeing service quality, and meet the needs of diverse scenarios through different dimensions of design.

[0003] The core problems of existing technologies can be summarized as follows: First, the contradiction between resource utilization and QoS makes it impossible for static strategies and greedy strategies to balance efficiency and fairness. For example, when a certain financial transaction platform uses static allocation, the resource utilization rate is only 65%, but after switching to a greedy strategy, the transaction delay of low-priority users increases significantly, and the QoS violation rate increases significantly. Second, the lack of dynamic feedback information severely limits the adjustment ability of the algorithm. Traditional algorithms rely on queue length or task arrival volume, but actual systems can only provide binary feedback (queue empty / non-empty). In edge computing scenarios, due to sensor precision and communication bandwidth limitations, controllers cannot obtain real-time queue backlog, limiting dynamic adjustment capabilities. Third, the poor adaptability of non-steady-state demand makes it difficult for existing algorithms to respond to sudden traffic or periodic peaks. For example, the instantaneous request volume of an e-commerce platform during a promotion period can reach tens of times that of daily traffic, and traditional proportional allocation strategies cannot quickly respond to such fluctuations, resulting in a sharp drop in resource utilization and causing large-scale queue backlog. Finally, the lack of theoretical performance guarantees makes it difficult for existing methods to quantify the gap between them and the ideal optimal solution. SUMMARY

[0004] The application aims to provide an online resource allocation method and system based on multiplicative weight update, which can realize dynamic scheduling and reasonable allocation of resources among multiple users under the condition of limited feedback of only obtaining whether the queue is empty, improve the overall resource utilization of the system, and have the adaptability to non-steady load changes and the stability of long-term operation while meeting the minimum requirements of the service quality of each user.

[0005] To achieve the above-mentioned purpose, the technical scheme adopted by the application is as follows:

[0006] An online resource allocation method based on multiplicative weight update comprises the following steps:

[0007] 1) At each time step, the task load L t (n) of each user is obtained, and the task queue length Q t (n) of each user is updated according to the resource allocation at the previous time step and the task load;

[0008] 2) Based on the updated task queue length Q t (n), the active user queue set A t and the non-active user queue set B t of the current time step are divided, and the active user refers to the user whose queue length is greater than zero;

[0009] 3) For each active user in the active user queue, the service quality satisfaction is judged according to the current resource allocation H t (n) and the service quality demand η(n), and the user whose service quality is not satisfied is classified into set and the user whose service quality is satisfied is classified into set

[0010] 4) According to the queue classification result of step 3), the gain value g t (n) of each user is generated, which is used to measure the current activity and service quality satisfaction of the user;

[0011] 5) According to the resource allocation vector H t (n) at the previous time step and the current gain value g t (n), the resource allocation weight of each user is exponentially weighted updated to obtain the resource allocation temporary vector

[0012] 6) The resource allocation temporary vector is projected to a truncated simplex feasible region to obtain the formal resource allocation vector H t+1 (n);

[0013] 7) The above process is repeated to perform resource allocation at the next time step.

[0014] Further, step 1) initializes parameters before proceeding, including: setting the quality of service requirement of each user ρ(n), the truncation parameter δ, the learning rate η, the gain factor λ, and the initial resource allocation vector H1, where, O(·) represents the asymptotic upper bound, T represents the total time; H1 is composed of the resource allocation amounts of different users.

[0015] Further, the initial resource allocation vector H1 is randomly sampled from the truncated simplex feasible region , satisfying and where n represents a user, and N represents the number of users.

[0016] Further, the update mode of the task queue length Q t (n) in step 1) is:

[0017] Q t (n) = max{L t (n) + Q t-1 (n) - H t (n), 0}

[0018] where L t (n) represents the task load of user n at time step t, and H t (n) represents the current resource allocation amount for user n.

[0019] Further, the judgment mode of service quality satisfaction in step 3) is: for active user n, if the current resource allocation amount H t (n) is less than , it is determined that the user does not meet its quality of service requirements, where is the total quality of service requirement of all active users.

[0020] Further, the calculation rule of the gain value g t (n) in step 4) includes:

[0021]

[0022] where λ represents the gain factor, and n represents a user.

[0023] Further, the calculation mode of the intermediate resource vector in step 5) is:

[0024]

[0025] where η represents the learning rate, and g t(n) represents the gain value of the current user.

[0026] Further, the resource allocation temporary vector H is projected to a truncated simplex feasible region to obtain the formal resource allocation vector H t+1 (n) is represented as follows:

[0027]

[0028] The constraint condition is n represents a user, N represents the number of users, and δ represents a truncation parameter.

[0029] Further, the truncated simplex feasible region in step 6) is represented as follows:

[0030]

[0031] where x(n) represents the resource share allocated to user n, represents that the resource share allocated to user n is greater than or equal to the minimum resource guarantee threshold N represents the number of users, and δ represents a truncation parameter.

[0032] An online resource allocation system based on multiplicative weight update includes a memory for storing a computer program and a processor for executing the computer program to implement the steps of the above method.

[0033] The beneficial effects obtained by the present application are as follows:

[0034] 1. The present application converts the only available binary feedback (queue empty / non-empty) into an effective optimization signal by dynamically adjusting the resource allocation weight, maximizes the cumulative task completion amount, and strictly meets the quality of service (QoS) constraints of all users.

[0035] 2. The present application prioritizes the scheduling of continuous unsatisfied resource allocation queues through an exponential weight growth mechanism, effectively avoiding QoS violation phenomena caused by insufficient resources.

[0036] 3. The present application ensures that the resource allocation of each user is not lower than the set lower limit by introducing a truncated simplex constraint, maintains scheduling stability in the case of severe system load fluctuations, and improves the robustness of the algorithm.

[0037] 4. The present application realizes efficient dynamic resource scheduling through a multiplicative weight update mechanism, can quickly respond to changes in the state of each queue in the system, prioritizes the resource needs of high-load users, and at the same time takes into account the minimum resource allocation of low-priority users, improving the overall scheduling fairness.

[0038] 5. The application avoids resource waste caused by static quota strategy and resource preemption under greedy strategy by adjusting resource allocation proportion in real time, effectively improves resource utilization in multi-user mixed load scenarios, and makes system running efficiency close to the optimal level.

[0039] 6. The application constrains the weight update result by projection mechanism, ensures that each queue can obtain the minimum resource guarantee even in low load state, prevents "hunger phenomenon" of low priority users, and improves long-term service quality guarantee capability. BRIEF DESCRIPTION OF DRAWINGS

[0040] Figure 1 is a flowchart of an online resource allocation method based on multiplicative weight update.

[0041] Figure 2 is a queue work total amount diagram of experimental results.

[0042] Figure 3 is a queue work total amount diagram of experimental results.

[0043] Figure 4 is a queue length diagram of experimental results. DETAILED DESCRIPTION

[0044] In order to make the technical features and advantages or technical effects of the above technical solutions of the application more obvious and easy to understand, the following embodiments are described in detail.

[0045] The embodiment of the application provides an online resource allocation method based on multiplicative weight update (QoS-MWU), as shown in Figure 1 The method is suitable for a resource allocation system meeting multi-user quality of service (QoS) constraints, and the core idea is to map the user queue state into a gain signal in the form of binary feedback, realize dynamic adjustment of resource allocation through a multiplicative weight update mechanism, maximize the overall task completion amount, and strictly guarantee the QoS requirements of each user in the whole process.

[0046] The method comprises the following steps:

[0047] I. Initialization phase

[0048] Step 1: Parameter and space definition

[0049] Define the user set as [N] = {1, 2,..., N}, and each user n corresponds to a resource queue.

[0050] Set the minimum resource requirement vector of the quality of service (QoS) of each user as:

[0051] ρ = (ρ(1), ρ(2),..., ρ(N)) satisfies

[0052] To avoid a user being assigned extremely low resources for a long time, a truncation parameter is defined as is a parameter converging inversely proportional to the total time step T, indicating that δ gradually decreases as T increases, and its order of magnitude does not exceed the order of magnitude of 1 / T; where O(·) represents the upper bound of the growth speed of the function as a variable approaches infinity, T is the total time step.

[0053] Based on the truncation parameter, a truncated simplex feasible region is constructed and defined as

[0054]

[0055] wherein, represents a truncated simplex set with lower bound constraints, which is an extended set of the original simplex set (i.e., a set of vectors with non-negative components and a sum of 1) by further setting a minimum lower bound for each component; x is an arbitrary variable used to define the set; x(n) represents the resource share allocated to user n; represents that a minimum resource guarantee threshold is set for each user n: cannot be less than The truncated simplex feasible region ensures that the resource allocation of each user is not less than Thus, the resource allocation tends to zero, effectively suppressing the resource "hunger" phenomenon.

[0056] Step 2: Initial resource allocation and parameter setting

[0057] From the above defined truncated simplex feasible region a random initial resource allocation vector is sampled:

[0058] H1 = (H1(1), H1(2),..., H1(N))

[0059] wherein, is the resource allocation amount,

[0060] Set the gain amplification factor and the learning rate η ∈ (0, 1). Initialize the time step t = 1, and set the initial length of all queues to zero, i.e.,

[0061]

[0062] II. Dynamic resource allocation loop (every time step t)

[0063] Step 3: Current resource allocation execution

[0064] Resource allocation vector H according to current time step t t = (H t (1), H t (2),..., H t (N)) allocates resources to each user's resource queue:

[0065] The amount of resources each user n obtains at this time step is H t (n) to process its pending task load.

[0066] Step 4: Update queue length

[0067] Each user queue receives a new task load L t (n) at the current time step, combined with the currently allocated resources H t (n), to calculate the updated queue length:

[0068] Q t (n) = max{L t (n) + Q t-1 (n) - H t (n), 0}

[0069] where L t (n) is the unobservable task load; the above equation indicates that the next step is determined based only on whether the updated queue is empty or not.

[0070] Step 5: Binary queue state observation

[0071] All users are divided into two sets by the non-zero or zero state of the queue length Q t (n):

[0072] Active user queue set (i.e., there is a task backlog):

[0073] A t = {n ∈ [N] | Q t (n) > 0}

[0074] Non-active user queue set (queue is empty):

[0075] B t = [N] \ A t

[0076] Step 6: Quality of service determination and active user queue classification

[0077] Calculate the total QoS demand of the current active user queue:

[0078]

[0079] For each active user queue n∈A t , the following QoS allocation judgment is made:

[0080] If then determine that n does not meet the QoS requirement, and classify it into the set:

[0081]

[0082] Otherwise, classify it into the set of users that meet the QoS requirement:

[0083]

[0084] The above threshold is the minimum reasonable resource allocation ratio for fault tolerance, and δ is a relaxation factor to improve the robustness of the method.

[0085] Step 7: Gain function generation

[0086] According to the queue classification result, calculate the gain value g t (n) for each user n to guide the weight update in the next step:

[0087]

[0088] where λ is the gain factor. The three cases in the above formula are as follows:

[0089] Queues that do not meet the QoS requirement will obtain an amplified gain, increasing their proportion in subsequent allocation;

[0090] Queues that meet the QoS requirement maintain the current allocation ratio;

[0091] Non-active user queues B t do not participate in additional resource competition for the time being.

[0092] Step 8: Multiplicative weight update

[0093] According to the above gain, perform exponential weighted update on the current resource allocation vector to obtain the updated resource allocation temporary vector:

[0094]

[0095] where η∈(0,1) is a hyperparameter representing the learning rate.

[0096] Step 9: KL projection to feasible region

[0097] In order to make the updated resource allocation temporary vector Satisfy the truncated simplex constraints, project them back to the feasible region via KL divergence Get the formal resource allocation vector H for the next time step t+1 That is,

[0098]

[0099] The constraints are as follows:

[0100] The total resource sum is 1:

[0101] The single queue allocation lower limit is:

[0102] That is, ensure that the total amount of resources is 1, and each user obtains no less than Constrain the fairness of the allocation while preserving the optimization direction.

[0103] III. Termination judgment

[0104] Step 10: Iteration control

[0105] If the current time step t < T, T is the preset maximum time step, then update t = t + 1, return to step 3 to continue the next round of resource allocation; otherwise, terminate the algorithm process.

[0106] Experimental test:

[0107] To verify the feasibility and effectiveness of the method of the application, a simulation experiment was conducted to compare the online resource allocation method based on multiplicative weight update (referred to as QoS-MWU strategy) proposed in the application with three baseline algorithms, including a static quality of service strategy, an online proportional strategy, and a greedy strategy.

[0108] The experiment is based on the following three evaluation indicators: resource utilization, service quality guarantee capability, and queue length dynamic change.

[0109] The first indicator is the cumulative completed workload, that is, the total amount of tasks completed by each algorithm on all queues;

[0110] The second indicator is the service quality satisfaction, which is compared by comparing the cumulative completed workload of each algorithm on each queue with the completed amount of the static quality of service strategy;

[0111] The third indicator is the system stability, which evaluates the queue management performance of each algorithm through the dynamic change of the queue length over time.

[0112] The simulation environment adopts the synthetic generated adversarial load sequence to simulate the challenging resource allocation scene. The experimental system includes 3 user queues, and the quality of service demand vector of each queue is ρ=(0.3, 0.3, 0.4). The simulation time step is set to T=10,000.

[0113] Experimental results:

[0114] Figure 2 The cumulative completed work results of each algorithm during the simulation are shown. The results show that the cumulative work of the QoS-MWU strategy is higher than that of the other three baseline strategies in the entire time range. Under the adversarial load, the online proportional strategy performs worse than the static quality of service strategy, which shows that the static quality of service strategy has certain advantages in strictly meeting the QoS constraint, and the online proportional strategy has deviation in task allocation, which affects the overall completed task amount.

[0115] Figure 3 The performance of each algorithm in the quality of service satisfaction is shown, which is embodied by the comparison of the cumulative completed work of each user with the static quality of service strategy. The experiment shows that the QoS-MWU strategy has advantages in meeting the quality of service demand of each user, and its cumulative completed work is stably higher than that of the static quality of service strategy, reflecting strong service guarantee capability. In contrast, the online proportional strategy has over-allocation of resources in some queues, causing insufficient resources in other queues, which affects the overall quality of service balance; the greedy strategy improves the overall task completion amount, but the quality of service demand of some queues cannot be guaranteed, leading to uneven service.

[0116] Figure 4 The change of the length of each user queue with time under each algorithm is shown. The experimental results show that the QoS-MWU strategy can effectively control the fluctuation of the queue length under dynamic load, and the system as a whole maintains a stable running state. Although the greedy strategy can also maintain the stability of the queue to some extent, it has the problem of insufficient resources in some queues due to the lack of targeted control of the quality of service requirement.

[0117] Although the present application has been disclosed as above with examples, it is not intended to limit the present application, and appropriate modifications or equivalent replacements of the technical solutions of the present application made by those skilled in the art shall be covered within the protection scope of the present application, and the protection scope of the present application is defined by the claims.

Claims

1. An online resource allocation method based on multiplicative weight updates, characterized in that, comprising the steps of: 1) At each time step, obtain the task load L t (n) of each user, and update the task queue length Q t (n) of each user according to the resource allocation and the task load at the previous time step. 2) based on the updated task queue length Q t (n), partitioning the active user queue set A of the current time step t and the inactive user queue set B t , active users refer to users whose queue length is greater than zero; 3) for each active user in the active user queue, according to its current resource allocation amount H t (n) make a service quality satisfaction judgment with the service quality requirement p(n), and put the users whose service quality is not satisfied into set has been satisfied into set 4) From the queue classification result of step 3), a gain value g is generated for each user t (n), which is used to measure the current activity level of the user and the quality of service satisfaction level; 5) The resource allocation vector H according to the previous time step t (n) with the current gain value g t (n), the resource allocation weights for each user are exponentially weighted updated, resulting in a temporary vector of resource allocations 6) temporary vector of resource allocations projection onto a truncated simplex feasible region obtaining the formal resource allocation vector H t+1 (n); 7) repeating the above procedure for the next time step of resource allocation.

2. The method of claim 1, wherein, Step 1) Initialize parameters before proceeding, including: setting the quality of service requirement of each user ρ(n), the truncation parameter δ, the learning rate η, the gain factor λ, and the initial resource allocation vector H1, where, O(·) denotes the asymptotic upper bound, T denotes the total time; H1 consists of the amount of resource allocation of different users.

3. The method of claim 2, wherein, The initial resource allocation vector H1 is obtained from the truncated simplex feasible region with random sampling, satisfying and where n denotes a user and N denotes the number of users.

4. The method of claim 1, wherein, The task queue length Q in step 1) t The update mode of (n) is: Q t (n) = max{L t (n) + Q t-1 (n) - H t (n), 0} where L t (n) denotes the task load of user n at time step t, H t (n) denotes the current amount of resource allocated to user n.

5. The method of claim 2, wherein, The judgment of the service quality satisfaction in step 3) is as follows: for the active user n, if the current resource allocation amount H t (n) is less than then it is determined that the user does not satisfy its service quality requirement, wherein is the total demand of the service quality of all the active users.

6. The method of claim 1, wherein, The gain value g in step 4) t The calculation rule of (n) includes: where λ denotes a gain factor and n denotes a user.

7. The method of claim 1, wherein, Resource allocation intermediate resource vector in step 5 The calculation is as follows: wherein η denotes a learning rate, g t (n) denotes a gain value of the current user.

8. The method of claim 1, wherein, Resource allocation tentative vector by Kullback-Leibler divergence in step 6) Projection to a truncated simplex feasible region Obtaining the formal resource allocation vector H t+1 (n), is represented as follows: The constraint is n denotes a user, N denotes the number of users, δ denotes a truncation parameter, and x(n) denotes the resource share allocated to user n.

9. The method of claim 1 or 8, wherein, The truncated simplex feasible region in step 6) is represented as follows: where x(n) represents the resource share allocated to user n, represents that the resource share allocated to user n is greater than or equal to the minimum resource guarantee threshold N represents the number of users, and δ represents a truncation parameter.

10. An online resource allocation system based on multiplicative weight updates, characterized in that, comprising a memory for storing a computer program and a processor for executing the computer program to implement the steps of the method of any of claims 1-9.

Citation Information

Patent Citations

  • Intelligent access control and resource allocation method based on distributed A-C

    CN112887999A

  • Low-carbon plan scheduling method for ship segmented coating

    CN118735146A

  • Deep learning task hybrid deployment method and system

    CN119987974A

  • Method and system for managing load balancing in data processing system

    US20060212873A1

  • Method and system for solving a dynamic programming problem

    US20200349453A1

Cited By

  • Dynamic weight resource allocation method based on FIFO queue

    CN122137802A