Systems and Methods for Optimizing Resource Allocation Using GPUs
Through GPU parallel processing and CPU aggregation, combined with dual multiplication optimization, the resource allocation optimization problem of large-scale e-commerce platforms is solved and efficient resource allocation decisions are achieved.
Patent Information
- Application Number
- CN202180007599.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2020-01-09
- Filing Date
- 2021-01-11
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2041-01-11
AI Technical Summary
The existing technology is difficult to efficiently deal with the resource allocation optimization problem of large-scale e-commerce platforms, especially at the scale of millions of users, the computational complexity of solving backpack problems is too high.
Multiple objective functions are processed in parallel by using graphics processor (GPU), combined with the central processor (CPU) aggregation calculation results, and resource allocation is optimized using dual multipliers, and the objective functions are transformed and decomposed through the Lagrangian method to improve efficiency.
It realizes efficient determination of the optimal resource allocation plan under large-scale users, reduces computing complexity and improves processing speed, and can optimize resource allocation decisions on a large scale.
Smart Images

Figure CN114902273B_ABST
Abstract
Description
Technical Field
[0001] This application generally relates to optimizing resource allocation in a recommendation system. Background Art
[0002] E-commerce platforms often (e.g., daily, hourly, or even near real-time) make resource allocation decisions. These resource allocation problems are typically solved by solving Knapsack Problems (KPs), which can only be handled on a relatively small scale. Optimizing large-scale resource allocation decisions has been an open technical challenge. Summary of the Invention
[0003] Various embodiments of this specification include systems, methods, and non-transitory computer-readable media for optimizing resource allocation.
[0004] According to one aspect, a method for optimizing resource allocation may include: processing multiple first objective functions in parallel to determine multiple allocation plans, where each of the allocation plans corresponds to allocating zero or more of a plurality of resources associated with a platform to a user, and the multiple first objective functions share one or more dual multipliers; determining multiple revenues and costs of the platform in parallel based on the multiple allocation plans; aggregating the calculated revenues and costs using parallel reduction; updating one or more dual multipliers based on the aggregated costs to determine whether an exit condition is met; and, in response to not meeting the exit condition, repeating the processing of the multiple first objective functions based on the updated one or more dual multipliers.
[0005] In some embodiments, the processing the multiple first objective functions in parallel may include processing the multiple first objective functions in parallel on a Graphics Processing Unit (GPU); the determining the multiple revenues and costs of the platform in parallel may include determining the multiple revenues and costs in parallel on the GPU; the aggregating the calculated revenues and costs using parallel reduction may include aggregating the calculated revenues and costs through the GPU and a Central Processing Unit (CPU); and the updating the one or more dual multipliers may include updating the one or more dual multipliers by the CPU.
[0006] In some embodiments, each of the multiple first objective functions includes N×M coefficients, where N is the number of multiple users and M is the number of multiple resources; before processing the multiple first objective functions, the method further includes: storing non-zero coefficients of the N×M coefficients in the memory of the GPU using a value table and an index table so that the GPU accesses the memory with a constant time complexity each time it reads; the value table uses a resource identifier as the main dimension; and the index table uses a user identifier as the main dimension.
[0007] In some embodiments, storing non-zero values of the N×M coefficients in the memory of the GPU using a value table and an index table may include: storing resource identifier values mapped to the non-zero coefficients in the value table; and storing one or more user identifier values mapped to one or more indices in the value table, wherein for each of the user identifier values, the corresponding index may point to one of the non-zero coefficients associated with the user identified by each of the user identifier values.
[0008] In some embodiments, each of the plurality of first objective functions may be subject to K constraints and may include N×M×K coefficients; and before processing the plurality of first objective functions, the method further includes: storing non-zero values of the N×M×K coefficients in row-major format including at least three dimensions in the memory of the GPU so that the GPU can access the memory with a constant time complexity each time it reads.
[0009] In some embodiments, the method may further include: in response to satisfying the exit condition, allocating the plurality of resources according to the plurality of allocation plans, where the number of the plurality of users is N, the number of the plurality of resources is M, and the i th corresponding to the user the allocation plan i th is represented as a vector containing M elements each x ij represents whether resource j th is being allocated to user i th .
[0010] In some embodiments, the exit condition may include whether one or more of the dual multipliers converge.
[0011] In some embodiments, before parallel processing the plurality of first objective functions, the method may further include: based on the Lagrangian method of dual problem transformation, transforming the original objective function for optimizing the resource allocation into a dual objective function; and decomposing the dual objective function into the plurality of first objective functions.
[0012] In some embodiments, the exit condition may include whether the values of the original objective function and the dual objective function converge, and the method may further include: determining the value of the original objective function based on the aggregated benefits and costs; and determining the value of the dual objective function based on the one or more dual multipliers, aggregated benefits and costs.
[0013] According to another aspect, a system for optimizing resource allocation may include a plurality of sensors and a computer system including a first computing device and a second computing device, the computer system including a processor and a non-transitory computer-readable storage medium storing instructions executable by the processor to cause the system to perform the following operations: processing a plurality of first objective functions in parallel to determine a plurality of allocation plans, wherein each of the allocation plans corresponds to allocating zero or more of a plurality of resources associated with a platform to a user, and the plurality of first objective functions share one or more dual multipliers; determining a plurality of revenues and costs of the platform in parallel based on the plurality of allocation plans; aggregating the calculated revenues and costs using parallel reduction; updating the one or more dual multipliers based on the aggregated costs to determine whether an exit condition is satisfied; and, in response to the exit condition not being satisfied, repeating the processing of the plurality of first objective functions based on the updated one or more dual multipliers.
[0014] According to yet another aspect, a non-transitory computer-readable storage medium for optimizing resource allocation may be configured with instructions executable by one or more processors to cause the one or more processors to perform the following operations: processing a plurality of first objective functions in parallel to determine a plurality of allocation plans, wherein each of the allocation plans corresponds to allocating zero or more of a plurality of resources associated with a platform to a user, and the plurality of first objective functions share one or more dual multipliers; determining a plurality of revenues and costs of the platform in parallel based on the plurality of allocation plans; aggregating the calculated revenues and costs using parallel reduction; updating the one or more dual multipliers based on the aggregated costs to determine whether an exit condition is satisfied; and, in response to the exit condition not being satisfied, repeating the processing of the plurality of first objective functions based on the updated one or more dual multipliers.
[0015] The embodiments disclosed in this specification have one or more technical effects. In some embodiments, one or more GPUs are used to perform parallel computing to determine an optimal resource allocation scheme for a platform (e.g., an e-commerce platform) to allocate resources (e.g., monetary or non-monetary resources) to users (e.g., customers, employees, departments). In one embodiment, the GPU can solve multiple objective functions in parallel, with each objective function corresponding to a user. The parallel computing power provided by the GPU can effectively handle the most time-consuming and computationally intensive parts in finding the optimal resource allocation scheme. In some embodiments, parallel reduction (e.g., aggregation) of the solutions of multiple objective functions can be performed by using a subset of one or more GPUs and one or more CPUs (e.g., the Compute Unified Device Architecture (CUDA) framework for parallel reduction). In some embodiments, the above parallel computing and reduction enable the platform to determine the optimal resource allocation scheme on a large scale (e.g., allocate resources to millions of users). In some embodiments, when coefficients need to be read from the memory, the memory layout of the GPU designed for parallel computing of multiple objective functions can avoid binary search and provide memory access with a constant time complexity for each read during the parallel computing process. In some embodiments, the memory layout can also reduce the memory footprint by avoiding storing duplicate values such as user identifiers.
[0016] Considering the following description and the appended claims in reference to the accompanying drawings, these features of the systems, methods, and non-transitory computer-readable media disclosed herein, as well as the functions of the operational methods and related structural elements, and the combination and manufacturing economy of the components, will become more apparent, all of which form a part of this document, wherein, in the various drawings, like reference numerals represent corresponding components. However, it should be clearly understood that the drawings are for illustrative and exemplary purposes only and are not intended to limit the present invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] Figure 1 Shows an applicable example environment of techniques for determining optimal resource allocation according to various embodiments.
[0018] Figure 2 Shows an example device for determining optimal resource allocation according to various embodiments.
[0019] Figure 3 Shows an example workflow for determining optimal resource allocation according to various embodiments.
[0020] Figure 4 Shows an example GPU memory layout for determining optimal resource allocation according to various embodiments.
[0021] Figure 5Illustrates an example computer system for determining optimal resource allocation according to various embodiments.
[0022] Figure 6 Illustrates a method for optimizing resource allocation according to various embodiments.
[0023] Figure 7 Illustrates an example electronic device for determining optimal resource allocation according to various embodiments. Detailed Description
[0024] The embodiments disclosed herein may help an allocation system (e.g., an e-commerce platform) optimize resource allocation decisions by using parallel processing methods. These resources may include financial budgets (e.g., loans, marketing promotions, advertising expenditures, portfolios of assets) to be allocated among users / user groups, non-monetary resources such as internet user traffic (e.g., impressions, clicks, dwell time) to be allocated across different business channels, and so on. The purpose of these resource allocation decisions may be to optimize a common goal, such as maximizing the conversion of expected users in a marketing campaign or maximizing the number of daily active users.
[0025] In some embodiments, optimizing the resource allocation of a platform may involve seeking an optimal resource allocation plan to allocate resources to multiple users. The platform resources to be allocated may be represented as a set of items. Allocating one of the items to a user may generate benefits and costs for the platform. In some embodiments, some items may be allocated to one or more users more than once. In some embodiments, these resource allocation decisions may be subject to a set of constraints (e.g., limited quantity of resources, budget caps, etc.). These constraints can be divided into global constraints and local constraints, where global constraints can specify the maximum quantity of available resources, while local constraints can impose restrictions on individual users or user groups.
[0026] In some embodiments, the above resource allocation optimization can be represented using the following formulas (1)-(4).
[0027]
[0028]
[0029]
[0030]
[0031] The objective function in formula (1) can be represented as the original objective function, which represents if according to x i,j(For example, a resource allocation plan) allocates resources, which is the total revenue to be obtained by the platform. Solving this original objective function may require finding the optimal resource allocation plan x that maximizes the original objective function in Equation (1). i,j . In one embodiment, the allocation plan determines that if item j th is allocated to user i th (or user group i th ), then x i,j = 1; otherwise, x i,j = 0. In Equations (1)-(4), the number of items to be allocated is denoted as M, and the number of users receiving resources is denoted as N. In some embodiments, there may be multiple global constraints that the resource allocation plan must comply with. For example, the platform may impose a budget limit on the total cost of the resources to be allocated (e.g., the total value of the items to be allocated must be less than one thousand dollars), and a limit on the quantity of the resources to be allocated (e.g., the number of items to be allocated must be less than one hundred). Equation (2) generalizes the global constraints by using K as the number of constraints (i.e., ). For each of the K constraints, each item (e.g., resource) may be associated with a corresponding cost b i,j,k shown in Equation (2). For example, if the first constraint (e.g., k = 1) corresponds to a budget limit of one thousand dollars, then b i,j,1 may represent the cost (e.g., dollar amount) of allocating item j th to user i th . As another example, if the second constraint (e.g., k = 2) corresponds to a limit on the total number of items to be allocated, then b i,j,2 may represent the cost (e.g., counted as one) of allocating item j th to user i th .
[0032] In some embodiments, the U constraint in Equation (3) may refer to a local constraint because it is applied to a single user (or a small group of users). For example, the U constraint in Equation (3) may specify an upper limit on the number of items each user is allowed to receive. In some embodiments, this local constraint may be extended to include multiple constraints. In some embodiments, B k and U may be strictly positive values, p i,j may be non-negative, and b i,j,k may be positive or negative.
[0033] In some embodiments, the x i,j ∈ {0, 1} constraint (e.g., x i,j is a binary integer indicating whether user i th will receive or not receive item j th ) may be selectively relaxed to 0 ≤ x i,j≤ 1. To mitigate the impact of estimation noise in the problem coefficients, the solution x can be regularized by introducing an entropy function as a regularization term. i,j . For example, the entropy function can be defined as -x i, j log x i,j . Therefore, Equation (5) can be used to represent the original objective function (e.g., Equation (1)).
[0034]
[0035] α in Equation (5) can be a predetermined regularization coefficient.
[0036] In some embodiments, to optimize the resource allocation represented in the form of Equation (5), all possible solutions x can be tried i,j to find the optimal solution (e.g., the solution that maximizes the objective function in Equation (5)). However, when the number of users is large (e.g., in the millions), this brute-force solution method may become impractical. Therefore, the objective function in Equation (5) needs to be decomposed and solved using parallel processing (e.g., using a GPU).
[0037] In some embodiments, to perform the decomposition of Equation (5), the original objective function in Equation (5) can first be converted into a dual objective function by introducing a set of dual multipliers λ = (λ1,..., λ K ), where each dual multiplier corresponds to one of the K constraints in Equation (2). In some embodiments, the dual objective function can be represented by Equation (6).
[0038]
[0039] In some embodiments, the maximization problem in Equation (6) can then be decomposed into multiple independent subproblems, with each user (or group of users) having an independent subproblem. For example, the subproblem for user i th can be represented by Equation (7). In some embodiments, B in Equation (6) k can be omitted in Equation (7) because it does not depend on x i,j . These independent subproblems can be solved using the parallel processing capabilities of a GPU to improve efficiency. The parallel processing process is described in more detail below.
[0040]
[0041] Figure 1 Illustrates an applicable example environment of a technique for determining optimal resource allocation according to various embodiments. Figure 1The components shown are illustrative. Depending on the implementation, environment 100 may include additional, fewer, or alternative components.
[0042] As shown, Figure 1 environment 100 in may include a computing system 102. In some embodiments, computing system 102 may be associated with platforms such as an e-commerce platform (e.g., an online marketplace) that hosts millions of users or a small business that serves a local community. Computing system 102 may be implemented in one or more networks (e.g., an enterprise network), one or more endpoints, one or more servers, one or more clouds, or any combination thereof. Computing system 102 may include hardware or software that manages access to centralized resources or services in a network. A cloud may include a cluster of servers and other devices distributed across a network. Computing system 102 may include one or more processors (e.g., a digital processor, an analog processor, a digital circuit designed to process information, a central processing unit, a graphics processing unit, a microcontroller, or a microprocessor, an analog circuit designed to process information, a state machine, and / or other mechanisms for electronically processing information) and one or more memories (e.g., a permanent memory, a temporary memory, a non-transitory computer-readable storage medium). The one or more memories may be configured with instructions executable by the one or more processors. The processor may be configured to perform various operations by interpreting machine-readable instructions stored in the memory. Computing system 102 may be installed with appropriate software (e.g., a platform program, etc.) and / or access hardware of other devices (e.g., wired, wireless connections, etc.).
[0043] In some embodiments, a resource allocation query may include a request for computing system 102 to search for an allocation plan that complies with one or more constraints and maximizes a goal. The query may also include information about the resources 103 to be allocated, information about the users 105 to whom the resources 103 are allocated, and one or more coefficients 107. In some embodiments, the coefficients 107 may include multiple projected (or predetermined / known) revenues of the platform, each revenue corresponding to the allocation of one of the resources to a user. For example, to maximize the number of daily active users (e.g., the objective function), the platform may offer bonuses to users, which can be claimed by logging in to the mobile application of the platform and clicking on one or more buttons. Each user who applies for a bonus by performing these actions can be counted as a daily active user. Thus, allocating a bonus to a user may be associated with a projected revenue (e.g., if the bonus can attract a user to perform an action, there is one daily active user). In some embodiments, the coefficients 107 may also include multiple costs of the platform associated with allocating each resource to a user. For example, each resource may be associated with a dollar amount, so allocating such a resource may result in a dollar amount cost to the platform.
[0044] In some embodiments, the computing system 102 may include an optimization modeling module 112, a parallel processing module 114, an aggregation module 116, and a flow control module 118. The computing system 102 may include other modules. In some embodiments, the optimization modeling module 112 may establish an optimization model in response to a resource allocation query, which may include information about the resources 103 to be allocated, coefficients (e.g., expected or predetermined / known benefits or costs, objectives, budgets, constraints), and information about the users 105 communicating with the computing system 102 via computing devices (e.g., computer 105a, smartphone, or tablet 105b). For example, the optimization modeling module 112 may determine the original objective function of the resource allocation query shown in formula (5), convert the original objective function into the dual objective function shown in formula (6) by using the Lagrangian method for dual problem transformation, and decompose the dual objective function into multiple sub-objective functions for multiple users 105.
[0045] In some embodiments, the parallel processing module 114 may include multiple processing units. In some embodiments, the processing unit may refer to the core or thread of one or more GPUs. Each processing unit may independently solve one of the decomposed sub-objective functions for one user, where each solution corresponds to the resource allocation plan of that user. In some embodiments, multiple decomposed sub-objective functions may share one or more dual multipliers (e.g., λ shown in formula (7) can be adjusted) k to find the solution x i,j ). In some embodiments, the aggregation module 116 may include one or more processing units (e.g., the core or thread of a GPU), which collect the solutions of the processing units of the optimization modeling module 112 in parallel. In some embodiments, the aggregation module 116 may share one or more processing units with the parallel processing module 114. In some embodiments, the aggregated solutions may be used to determine whether the original objective function (e.g., formula (5)) and the dual objective function (e.g., formula (6)) converge. In some embodiments, the aggregated solutions may also be used to adjust the dual multipliers λ in the dual objective function (e.g., formula (6)) and the sub-objective function (e.g., formula (7)) k .
[0046] In some embodiments, the flow control module 118 may manage the operations of the parallel processing module 114 and the aggregation module 116. In some embodiments, in response to the values of the objective functions (e.g., the original objective function and the dual objective function) and the dual multiplier λ k not converging, the flow control module 118 may instruct the parallel processing module 114 to utilize the parameters updated by the aggregation module 116 (e.g., the updated dual multiplier λ k)) Repeat its operation (e.g., solve the sub-objective functions shown in formula (7) in parallel). In response to the aggregation module 116 determining that the exit condition is met, the computing system 102 may terminate the process and confirm the optimal solution 109 (e.g., resource allocation plan) for the resource allocation query. In some embodiments, the optimal solution 109 may be determined based on the aggregation of multiple solutions of the sub-objective functions.
[0047] Figure 2 FIG. 200 shows an example device for determining optimal resource allocation according to various embodiments. Figure 2 The device 200 shown in FIG. 2 may be used to implement Figure 1 the computing system 102 shown in FIG. As shown, Figure 2 the device 200 in FIG. 2 may be equipped with one or more GPUs and CPUs to determine optimal resource allocation. The highly parallel structure of the GPU enables it to efficiently solve independent compute-intensive tasks in parallel. Due to the memory layout in each GPU, each computation (e.g., in each thread) that occurs in the GPU may not support complex logic. Therefore, the CPU may be used for serial instruction processing (e.g., more complex computations) and flow control (e.g., checking whether the exit condition is met).
[0048] In some embodiments, the device 200 may receive coefficients 210 from a resource allocation query. The coefficients 210 may include multiple expected (or predetermined / known) benefits and costs of the platform, each of which may be associated with allocating resources to a user. The coefficients 210 may also include one or more constraints, such as global constraints (e.g., total budget), local constraints (e.g., limits on the amount of resources that can be allocated to a single user).
[0049] In some embodiments, the optimization problem defined by the coefficients 210 may be decomposed into multiple sub-optimization problems. For example, the original objective function of the optimization problem corresponding to the overall objective of the resource allocation query (e.g., formula (5)) may be decomposed into multiple sub-objective functions (e.g., formula (7)) for each individual user. Figure 2The flowchart 201 in [ ] shows an example workflow for determining the optimal resource allocation using the device 200. In some embodiments, at step 220, the device 200 can rely on one or more GPUs to solve multiple sub-objective functions in parallel. For example, one GPU can include multiple thread grids, and each thread grid can contain several thread blocks. Each thread block can host multiple threads, which can be used to process or solve one or more sub-optimization problems (e.g., sub-objective functions). In some embodiments, if the sub-objective function is formed as an Integer Programming (IP) problem, these problems can be solved by using an open-source or commercial IP solver (e.g., Cplex, Gurobi, other suitable tools, or any combination thereof). For example, the solution of each sub-objective function can include a vector (or matrix) representing the user allocation plan.
[0050] In some embodiments, after determining the allocation plans for multiple users at step 220, at step 230, one or more GPUs can continue to perform calculations in parallel based on the allocation plans. In some embodiments, one or more GPUs can calculate the per-user benefit for each allocation plan at step 230 based on the sum of the resource benefits allocated to the users. For example, if the sub-objective function is represented as formula (7), then the per-user benefit for user i th can be represented as formula (8), where x i,j is the determined allocation plan, and p i,j is the expected (or predetermined / known) benefit from the coefficient 210.
[0051]
[0052] In some embodiments, one or more GPUs can calculate the value of the entropy function for each allocation plan at step 230 for each allocation plan (e.g., for the corresponding single user). For example, if the sub-objective function is represented as formula (7), then the value of the entropy function can be represented as formula (9).
[0053]
[0054] In some embodiments, one or more GPUs can also calculate the per-user cost for each allocation plan at step 230 based on the sum of the resource costs allocated to the users. For example, if the sub-objective function is represented as formula (7), and assuming there is only one type of constraint (e.g., K in formula (7) is 1), then the per-user cost of the user can be represented as formula (10).
[0055]
[0056] In some embodiments, a resource may be associated with multiple constraints. For each constraint, allocating the resource may incur a cost. For example, providing a bonus to a user on a given date may result in a cost to the daily total bonus budget (e.g., the first constraint), as well as a cost to the total number of bonuses that are allowed to be provided (e.g., the second constraint). Assuming there are k constraints, one or more GPUs may calculate k per-user costs corresponding to the k constraints for each allocation plan in step 230, as shown in formula (11).
[0057]
[0058] In some embodiments, after the parallel computation in step 230 is completed, multiple per-user revenues, per-user entropy function values, and per-user costs may be aggregated in step 240. For example, these values may first undergo parallel reduction on one or more GPUs (e.g., using CUDA), and ultimately generate a total revenue (e.g., the sum of per-user revenues), a total entropy value (e.g., the sum of per-user entropy function values), and a total cost (e.g., the sum of per-user costs) to the CPU. If there are K constraints, there may be K total costs to aggregate. As another example, these values may be directly aggregated by the CPU. The total revenue, the total entropy value, and the K total costs may be respectively mapped to and formula (6) of
[0059] In some embodiments, based on the total revenue, the total entropy value, and the total cost aggregated in step 240, the CPU in device 200 may perform flow control in step 250. The flow control may include determining whether an exit condition is met. In some embodiments, in the case where the dual objective function is represented by formula (6), if the dual multipliers converge, the exit condition may be met. In some embodiments, if the original objective function value (e.g., formula (5)) and the dual objective function value (e.g., formula (6)) have converged (e.g., the gap between them is lower than a predefined threshold), the exit condition may be met.
[0060] In some embodiments, in response to the exit condition not being met, the CPU may update the dual multipliers in step 250 and send the updated dual multipliers to one or more GPUs to repeat the computation in step 220. In some embodiments, the dual descent algorithm may be utilized to update the dual multipliers, as shown in formula (12).
[0061]
[0062] where the hyperparameter η is the learning rate.
[0063] Figure 3illustrates an example workflow for determining optimal resource allocation according to various embodiments. The workflow 300 may be executed by the computing system 102 in Figure 1 or the device 200 in Figure 2 As shown, the workflow 300 may start by receiving an input (e.g., a resource allocation query) and performing initialization at step 310. In some embodiments, the input may include multiple coefficients, which include information about the resources to be allocated (e.g., associated expected / scheduled / known benefits and costs, global constraints, local constraints), information about the users receiving the resources (e.g., quantities), other suitable information, or any combination thereof. In some embodiments, the workflow 300 may be initialized at step 310. The initialization may include initializing one or more dual multipliers, other parameters, or any combination thereof. Based on the input and the initialized parameters, the original optimization problem (e.g., Equation (5)) may be determined. In some embodiments, the original optimization problem may be transformed into a dual optimization problem (e.g., Equation (6)) and then decomposed into multiple sub-optimization problems (e.g., Equation (7)).
[0064] In some embodiments, multiple sub-optimization problems may be solved in parallel at step 320 to determine multiple resource allocation plans. In some embodiments, this parallel processing step 320 may be implemented using various mechanisms, such as the Compute Unified Device Architecture (CUDA) platform / model that allows direct access to the virtual instruction set and parallel computing elements of a graphics processing unit (GPU), multi-threaded programming, a map / reduce framework, other suitable mechanisms, or any combination thereof. In some embodiments, each solution to the sub-optimization problem may correspond to a resource allocation plan for a user. Each of these sub-optimization problems may be solved in parallel using one or more GPUs in order to provide a vector (or matrix) of {x th} with a given dual multiplier λ for user i i,j In some embodiments, one or more GPUs may also compute in parallel for user i th for each of and These values may subsequently be used to compute the value of one or more objective functions.
[0065] In some embodiments, the values computed in parallel may be reduced (e.g., aggregated) by performing a parallel reduction at step 330. For example, the parallel reduction may collect for each and and compute and to obtain the values of the original objective function and the dual objective function (e.g., equations (5) and (6)). In some embodiments, the parallel reduction may occur on one or more GPUs, and the final result of the reduction may be communicated to a central processing unit (CPU).
[0066] In some embodiments, the final result of the parallel reduction can be used to determine whether the workflow 300 should be terminated at step 340. This termination may occur when one or more exit conditions are met. In some embodiments, one or more exit conditions can include whether the dual multipliers have converged, whether the values of the original objective function and the dual objective function have converged, other suitable conditions, or any combination thereof. In response to no exit conditions being met, the dual multipliers can be updated based on the reduction result, and the workflow 300 can jump back to 320 to repeat the process. In response to at least one exit condition being met, the workflow 300 can output the resource allocation plan determined at step 320 as the final resource allocation solution (e.g., at step 350).
[0067] In some embodiments, the workflow 300 can be represented as the following pseudocode.
[0068]
[0069] Figure 4 An example GPU memory layout for determining optimal resource allocation according to various embodiments is shown. In some embodiments, if one or more GPUs are used for parallel searching (e.g., using multiple threads in a GPU) for resource allocation plans for multiple users, the memory layout of the GPU can be critical to the efficiency of memory access and related computations. In some embodiments, the layout of one or more coefficient vectors, matrices, and tensors (which are typically sparse) can be organized into a compact format. For example, the non-zero entries of a two-dimensional N×M coefficients (e.g., N is the number of users for whom resources are to be allocated, and M is the number of resources to be allocated) can be organized in a row-major format with the user identifier as the major dimension and the resource identifier as the minor dimension. Each user identifier value can identify a user, and each resource identifier value can identify a resource. As another example, a three-dimensional tensor (e.g., b in equation (7)) can use the user identifier as the major dimension, the constraint identifier as the minor dimension, and the resource identifier as the third dimension. ijk ) can use the user identifier as the major dimension, the constraint identifier as the minor dimension, and the resource identifier as the third dimension.
[0070] Figure 4The example memory layout in [reference] can be used to store N×M coefficients 410 in the memory of one or more GPUs. The memory can refer to the global memory of each GPU, the shared memory of each GPU thread block, the local memory of each thread, other suitable memories, or any combination thereof. The N×M coefficients can refer to p in formulas (5), (6), and (7). i,j (e.g., the expected (or pre-determined / known) benefits of allocating multiple resources to multiple users). In some embodiments, the memory layout of coefficients 410 can be represented as Table 420, where the major dimension is user_id (e.g., user identifier), the minor dimension is item_id (e.g., resource identifier), and each pair of user_idi and item_idj maps to a non-zero p. i,j . Figure 4 Table 420 in [reference] includes several example values. For example, for user 0 (e.g., f = 0), there are three p i,j with non-zero values: p 0,3 = 0.1, p 0,5 = 0.2, p 0,6 = 0.3; for user 1 (e.g., i = 1), there are at least two p i,j with non-zero values: p 1,2 = 0.5, p 1,3 = 0.2. In some embodiments, the non-zero coefficients of a user can be stored continuously (e.g., the non-zero values of user 0 are stored as the first three columns in Table 420 of [reference]). Since the calculations in the parallel processing stage (e.g., modules 114 in [reference], steps 220 and 230 in [reference], and step 320 in [reference]) (e.g., for each user i in [reference]) only require the non-zero p Figure 4 values of each user i, the memory layout in Table 420 can allow the GPU to access these values with a constant time complexity (e.g., O(1) per read), rather than using a binary search with a quasi-linear time complexity (e.g., O(logN) per read). Figure 1 in [reference], Figure 2 steps 220 and 230 in [reference], and Figure 3 step 320 in [reference]) ), i,j The memory layout in Table 420 can allow the GPU to access these values with a constant time complexity (e.g., O(1) per read), rather than using a binary search with a quasi-linear time complexity (e.g., O(logN) per read).
[0071] In some embodiments, the storage efficiency of the memory layout of Table 420 can be further improved by using two tables: an index table 430 and a value table 440. As shown in [reference], the index table 430 can use user_id as the major dimension. Each user_id value can be mapped to a start_index in the value table 440. The value table can use item_id as the major dimension. Each item_id value can be mapped to a non-zero p Figure 4 in [reference]. i,jvalue. In one embodiment, start_index can point to the index of the first non-zero p value of the user specified by the user_id value. i,j For example, in Figure 4 , user 0 has three non-zero p values (0.1, 0.2, 0.3) stored in the value table 440 at indices 0, 1, and 2 respectively. The user_id value 0 in the index table 430 can be mapped to start_index 0 which points to index 0 in the value table. Similarly, user 1 is mapped to start_index 3 (e.g., index 3) in the value table storing the first non-zero p i,j value (0.5). When the thread is computing for user 0 i,j , it may first check the index table for the first start_index that points to the starting index of the non-zero p value of user 0. Then the thread can sequentially read the p i,j values until the start_index of the next user (e.g., start_index 3 of user 1). This approach can not only allow one or more GPUs to access the memory with a constant time complexity each time they read, but also save memory space by avoiding storing duplicate entries (e.g., the three user_id 0s in table 420 may only need to be stored once in table 430). i,j
[0072] Figure 5 FIG. shows a block diagram of a computer system 500 for optimizing resource allocation according to some embodiments. The components of the computer system 500 described below are illustrative. Depending on the implementation, the computer system 500 may include additional, fewer, or alternative components.
[0073] The computer system 500 can be an example implementation of one or more modules of the computing system 102. The device 200 and the workflow 300 can be implemented by the computer system 500. The computer system 500 can include one or more processors and one or more non-transitory computer-readable storage media (e.g., one or more memories) that are coupled to the one or more processors and configured with instructions executable by the one or more processors to cause the system or device (e.g., the processor) to perform the above-described methods, such as method 300. The computer system 500 can include various units / modules corresponding to the instructions (e.g., software instructions).
[0074] In some embodiments, the computer system 500 may be referred to as a device for optimizing resource allocation. The device may include an acquisition module 510 configured to acquire a plurality of coefficients corresponding to a particular resource allocation optimization query (e.g., raw resource allocation optimization), where the coefficients may include information about the resources to be allocated (e.g., expected / scheduled / known associated benefits and costs, global constraints, local constraints), information about the users receiving the resources (e.g., quantity), other appropriate information, or any combination thereof; a parallel processing module 520 configured to parallel process a plurality of first objective functions to determine a plurality of allocation plans and, based on the plurality of allocation plans, parallel determine a plurality of benefits and costs of the platform; an aggregation module 530 configured to aggregate the calculated benefits and costs using parallel reduction; and a flow control module 540 configured to update one or more dual multipliers based on the aggregated costs to determine whether an exit condition is met and, in response to the exit condition not being met, repeat the processing of the plurality of first objective functions based on the updated one or more dual multipliers. The acquisition module 510 may correspond to the optimization modeling module 112. The parallel processing module 520 may correspond to the parallel processing module 114. The aggregation module 530 may correspond to the aggregation module 116. The flow control module 540 may correspond to the flow control module 118.
[0075] The techniques described herein may be implemented by one or more special-purpose computing devices. The special-purpose computing devices may be a desktop computer system, a server computer system, a portable computer system, a handheld device, a network device, or any other device or combination of devices that include hardwired and / or program logic to implement the techniques. The special-purpose computing devices may be implemented as a personal computer, a laptop computer, a mobile phone, a camera phone, a smartphone, a personal digital assistant, a media player, a navigation device, an email device, a game console, a tablet computer, a wearable device, or a combination thereof. Computing devices are typically controlled and coordinated by operating system software. Conventional operating systems control and schedule computer processes for execution, perform memory management, provide file systems, networking, I / O services, and provide user interface functions such as a graphical user interface (“GUI”). The various systems, devices, storage media, modules, and units described herein may be implemented in a special-purpose computing device or in one or more computing chips of one or more special-purpose computing devices. In some embodiments, the instructions described herein may be implemented in a virtual machine on a special-purpose computing device. When executed, the instructions may cause the special-purpose computing device to perform the various methods described herein. The virtual machine may include software, hardware, or a combination thereof.
[0076] Figure 6 A method for optimizing resource allocation according to various embodiments is shown. The method 600 may be performed by a device, apparatus, or system for optimizing resource allocation. The method 600 may be performed by Figures 1 - 5performed by one or more modules / components of the environment or system shown, such as computing system 102. The operations of method 600 described below are illustrative. Depending on the implementation, method 600 may include additional, fewer, or alternative steps performed in various orders or in parallel.
[0077] Block 610 includes processing multiple first objective functions in parallel to determine multiple allocation plans, where each allocation plan corresponds to allocating zero or more of the multiple resources associated with the platform to a user, and the multiple first objective functions share one or more dual multipliers. In some embodiments, processing the multiple first objective functions in parallel may include processing the multiple first objective functions in parallel on a graphics processing unit (GPU). In some embodiments, each of the multiple first objective functions includes N×M coefficients, where N is the number of multiple users and M is the number of multiple resources; before processing the multiple first objective functions, the method further includes: storing non-zero coefficients of the N×M coefficients in the memory of the GPU using a value table and an index table so that the GPU can access the memory with a constant time complexity each time it reads; the value table uses a resource identifier as the main dimension; and the index table uses a user identifier as the main dimension. In some embodiments, storing non-zero values of the N×M coefficients in the memory of the GPU using a value table and an index table may include: in the value table, storing resource identifier values mapped to non-zero coefficients; in the index table, storing one or more user identifier values mapped to one or more indexes in the value table, where for each user identifier value, the corresponding index may point to one of the non-zero coefficients associated with the user identified by each user identifier value. In some embodiments, each of the multiple first objective functions may be subject to K constraints and may include N×M×K coefficients; before processing the multiple first objective functions, the method may further include: storing non-zero values of the N×M×K coefficients in the memory of the GPU in a row-major format including at least three dimensions so that the GPU can access the memory with a constant time complexity each time it reads.
[0078] Block 620 includes determining multiple revenues and costs of the platform in parallel based on the multiple allocation plans. In some embodiments, determining the multiple revenues and costs of the platform in parallel may include determining the multiple revenues and costs in parallel on a GPU.
[0079] Block 630 includes aggregating the calculated revenues and costs using parallel reduction. In some embodiments, aggregating the calculated revenues and costs using parallel reduction may include aggregating the calculated revenues and costs through a GPU and a central processing unit (CPU).
[0080] Block 640 includes updating one or more dual multipliers based on the aggregated cost to determine whether an exit condition is satisfied. In some embodiments, updating the one or more dual multipliers may include updating, by the CPU, the one or more dual multipliers. In some embodiments, the method may further include: in response to satisfying the exit condition, allocating the plurality of resources according to a plurality of allocation plans, wherein the number of the plurality of users is N, the number of the plurality of resources is M, and the allocation plan i in the plurality of allocation plans is th With user i th Correspondingly, and distribution plani th It can be represented as containing M elements A vector, each x ij Indicates whether the resource is being th Assigned to user i th . In some embodiments, the exit condition may include whether one or more dual multipliers converge. In some embodiments, the exit condition may include whether the value of the original objective function and the value of the dual objective function converge, and the method may further include: determining the value of the original objective function based on the aggregated benefit and cost; and determining the value of the dual objective function based on the one or more dual multipliers, the aggregated benefit and cost.
[0081] Block 650 includes repeatedly processing the plurality of first objective functions based on the updated one or more dual multipliers in response to the exit condition not being satisfied.
[0082] In some embodiments, before processing the multiple first objective functions in parallel, the method may also include: converting the original objective function for optimizing the resource allocation into a dual objective function based on the Lagrangian technique of dual problem transformation; and decomposing the dual objective function into the multiple first objective functions.
[0083] Figure 7 An example electronic device for optimizing resource allocation is shown. The electronic device may be used to implement Figures 1 - 6 The electronic device 700 may include a bus 702 or other communication mechanism for communicating information, and one or more hardware processors 704 coupled to the bus 702. The hardware processor 704 may be, for example, one or more general-purpose microprocessors.
[0084] The electronic device 700 may also include a main memory 706 coupled to the bus 702 for storing information and instructions to be executed by the processor 704, such as random access memory (RAM), cache, and / or other dynamic storage devices. The main memory 706 may also be used to store temporary variables or other intermediate information during the execution of instructions to be executed by the (one or more) processors 704. When these instructions are stored in a storage medium accessible to the processor 704, these instructions cause the electronic device 700 to appear as a special-purpose machine customized to perform the operations specified in the instructions. The main memory 706 may include non-volatile media and / or volatile media. The non-volatile media may include, for example, optical discs or magnetic disks. The volatile media may include dynamic memory. Common forms of media may include, for example, floppy disks, flexible disks, hard disks, solid state drives, magnetic tapes, or any other magnetic data storage media, CD-ROMs, any other optical data storage media, any physical media with hole patterns, RAM, PROM, and EPROM, FLASH-EPROM, NVRAM, any other memory chip or cartridge, and their network versions.
[0085] The electronic device 700 may implement the techniques described herein using custom hardwired logic, one or more ASICs or FPGAs, firmware, and / or program logic, where the firmware and / or program logic in combination with the electronic device may cause the electronic device 700 to be a special-purpose machine or program the electronic device 700 as a special-purpose machine. According to one embodiment, the techniques herein are performed by the electronic device 700 in response to execution of one or more sequences of one or more instructions contained in the main memory 706 by the (one or more) processors 704. These instructions may be read from another storage medium, such as the storage device 707, into the main memory 706. Execution of the instruction sequences contained in the main memory 706 may cause the (one or more) processors 704 to perform the processing steps described herein. For example, the processes / methods disclosed herein may be implemented by computer program instructions stored in the main memory 706. When these instructions are executed by the (one or more) processors 704, they may perform the steps shown in the corresponding figures and described above. In alternative embodiments, hardwired circuitry may be used in place of or in combination with software instructions.
[0086] The electronic device 700 further includes a communication interface 710 coupled to the bus 702. The communication interface 710 may provide two-way data communication coupled to one or more network links that connect to one or more networks. As another example, the communication interface 710 may be a local area network (LAN) card for providing a data communication connection to a compatible LAN (or a WAN component communicating with a WAN). A wireless link may also be implemented.
[0087] The performance of certain operations can be distributed among processors, not only residing within a single machine but deployed across multiple machines. In some example embodiments, the processor or processor implementation engine can be located in a single geographical location (e.g., within a home environment, office environment, or server farm). In other example embodiments, the processor or processor-implemented engine can be distributed across multiple geographical locations.
[0088] Each of the processes, methods, and algorithms described in the foregoing section can be implemented in code modules executed by one or more computer systems or computer processors including computer hardware, and are fully or partially implemented automatically by the code modules. The processes and algorithms can be partially or fully implemented in dedicated circuitry.
[0089] When the functions disclosed herein are implemented in the form of software functional units and sold or used as independent products, they can be stored in a non-volatile computer-readable storage medium executable by a processor. The specific technical solutions (in whole or in part) disclosed herein or aspects contributing to the current technology can be embodied in the form of a software product. The software product can be stored in a storage medium, including some instructions to enable a computing device (which can be a personal computer, server, network device, etc.) to execute all or part of the steps of the methods of the embodiments of this application. The storage medium can include a flash drive, a portable hard disk drive, ROM, RAM, a magnetic disk, an optical disk, other media that can be used to store program code, or any combination thereof.
[0090] Certain embodiments also provide a system that includes a processor and a non-transitory computer-readable storage medium storing instructions executable by the processor to cause the system to perform operations corresponding to the steps in any of the methods of the foregoing embodiments. Certain embodiments also provide a non-transitory computer-readable storage medium configured with instructions executable by one or more processors to cause the one or more processors to perform operations corresponding to the steps in any of the methods of the foregoing embodiments.
[0091] The embodiments disclosed herein can be implemented by a cloud platform, a server, or a server group (collectively referred to as a "service system") that interacts with a client. The client can be a terminal device or a client registered by a user on the platform, where the terminal device can be a mobile terminal, a personal computer (PC), and any device that can install a platform application.
[0092] The various features and processes described above can be used independently of each other or can be combined in various ways. All possible combinations and sub - combinations will fall within the scope of this application. Additionally, in some embodiments, certain method or process blocks may be omitted. The methods and processes described herein are also not limited to any particular sequence, and the blocks or states associated therewith can be executed in other appropriate orders. For example, the described blocks or states can be executed in an order different from that specifically disclosed, or multiple blocks or states can be combined in a single block or state. The example blocks or states can be executed serially, in parallel, or in some other manner. Blocks or states can be added to or removed from the disclosed exemplary embodiments. The exemplary systems and components described herein can be configured differently from those described. For example, elements can be added, removed, or rearranged compared to the disclosed example embodiments.
[0093] The various operations of the exemplary methods described herein can be performed at least in part by an algorithm. The algorithm can be included in program code or instructions stored in a memory (e.g., the non - transitory computer - readable storage medium described above). Such an algorithm can include a machine - learning algorithm. In some embodiments, a machine - learning algorithm may not explicitly program a computer to perform a function but can learn from training data to establish a predictive model for performing the function.
[0094] The various operations of the exemplary methods described herein can be performed at least in part by one or more processors that are (e.g., by software) temporarily configured or permanently configured to perform the relevant operations. Whether temporarily configured or permanently configured, such processors can constitute a processor - implemented engine that runs to perform one or more of the operations or functions described herein.
[0095] Similarly, the methods described herein can be at least in part implemented by a processor, where a particular processor or multiple processors are examples of hardware. For example, at least some of the operations of the method can be performed by one or more processors or processor - implemented engines. Additionally, one or more processors can also be operable to support the performance of relevant operations in a “cloud computing” environment or as “software as a service” (SaaS). For example, at least some operations can be performed by a set of computers (e.g., machines including processors), and these operations can be accessed via a network (e.g., the Internet) and via one or more appropriate interfaces (e.g., application programming interfaces (APIs)).
[0096] The performance of certain operations can be distributed among processors, not only residing within a single machine but deployed across multiple machines. In some example embodiments, the processor or processor implementation engine can be located in a single geographical location (e.g., within a home environment, an office environment, or a server farm). In other example embodiments, the processor or processor-implemented engine can be distributed across multiple geographical locations.
[0097] In this document, multiple instances can implement components, operations, or structures described as a single instance. Although the individual operations of one or more methods are shown and described as separate operations, one or more of these separate operations can be performed simultaneously and these operations are not required to be performed in the order shown. Structures and functions presented as separate components in example configurations can be implemented as a combined structure or component. Similarly, structures and functions presented as a single component can be implemented as separate components. These and other variations, modifications, additions, and improvements fall within the scope of the subject matter of this document.
[0098] Although the overview of the subject matter has been described with reference to specific example embodiments, various modifications and changes can be made to these embodiments without departing from the broader scope of the embodiments of this application. If more than one disclosure or concept is actually disclosed, these embodiments of the subject matter can be referred to herein individually or collectively by the term "invention" merely for convenience and not to actively limit the scope of this application to any single disclosure or concept.
[0099] The embodiments shown herein are described in sufficient detail to enable those skilled in the art to practice the disclosed teachings. Other embodiments can be used and other embodiments can be derived therefrom such that structural and logical substitutions and modifications can be made without departing from the scope of this application. Accordingly, the detailed description should not be construed as limiting and the scope of each embodiment is defined only by the appended claims and the full scope of equivalents to which those claims are entitled.
[0100] Any process descriptions, elements, or boxes in the flowcharts described herein and / or depicted in the figures should be understood as potentially representing modules, segments, or portions of code that include one or more executable instructions for implementing specific logical functions or steps in the process. Alternative implementations are included within the scope of the embodiments described herein, where, as would be understood by those skilled in the art, depending on the functionality involved, elements or functions shown or discussed in the embodiments described herein can be deleted, executed out of order, including substantially simultaneously or in reverse order.
[0101] Unless otherwise expressly stated or the context otherwise indicates, the "or" used in this document is inclusive rather than exclusive. Thus, unless otherwise expressly stated or the context otherwise indicates, in this document, "A, B, or C" means "A, B, A and B, A and C, B and C, or, A and B and C". Additionally, unless otherwise expressly stated or the context otherwise indicates, "and" can be used both conjunctively and disjunctively. Thus, in this document, unless otherwise expressly stated or the context otherwise indicates, "A and B" means "A and B conjunctively, or A disjunctively and B disjunctively". Further, multiple separate instances can be provided for a resource, operation, or structure described herein as a single instance. Additionally, the boundaries between various resources, operations, engines, and data stores are somewhat arbitrary, and specific operations are illustrated in the context of specific illustrative configurations. Other allocations of functionality are envisioned and may fall within the scope of various embodiments of this application. Generally, structures and functions presented as separate resources in an example configuration can be implemented as a combined structure or resource. Similarly, structures and functions presented as a single resource can be implemented as separate resources. These and other variations, modifications, additions, and improvements are intended to fall within the scope of this application as represented by the appended claims. Accordingly, the specification and drawings are to be regarded as illustrative rather than restrictive.
[0102] The terms "comprising" or "including" are used to denote the presence of the subsequently recited feature, but do not preclude the addition of other features. Unless otherwise expressly stated or otherwise understood in the context in which it is used, conditional language such as "can", "could", "might", or "may" generally is intended to convey that certain embodiments include, while other embodiments do not include, certain features, elements, and / or steps. Thus, such conditional language generally does not imply that the features, elements, and / or steps are in any way required for one or more embodiments or that one or more embodiments necessarily include logic for deciding, with or without user input or prompting, whether or not these features, elements, and / or steps are included or are to be performed in any particular embodiment.
Claims
1. A computer-implemented method for optimizing resource allocation, comprising: Processing multiple first objective functions in parallel to determine multiple allocation plans, where each of the allocation plans corresponds to allocating zero or more of multiple resources associated with a platform to users, and the multiple first objective functions share one or more dual multipliers; Based on the multiple allocation plans, determining multiple revenues and costs of the platform in parallel; Using parallel reduction to aggregate the calculated revenues and costs respectively to obtain an aggregated revenue and an aggregated cost; Updating the one or more dual multipliers based on the aggregated cost to determine whether an exit condition is satisfied; And In response to the exit condition not being satisfied, repeating the processing of the multiple first objective functions based on the updated one or more dual multipliers to re-determine multiple allocation plans.
2. The method according to claim 1, wherein: The processing of the multiple first objective functions in parallel includes processing the multiple first objective functions in parallel on a graphics processing unit (GPU); The determining of the multiple revenues and costs of the platform in parallel includes determining the multiple revenues and costs in parallel on the GPU; The aggregating of the calculated revenues and costs using parallel reduction includes aggregating the calculated revenues and costs through the GPU and a central processing unit (CPU); The updating of the one or more dual multipliers includes updating the one or more dual multipliers by the CPU.
3. The method according to claim 2, wherein: Each of the multiple first objective functions includes N×M coefficients, where N is the number of multiple users and M is the number of the multiple resources; Before processing the multiple first objective functions, the method further includes: storing non-zero coefficients among the N×M coefficients in the memory of the GPU using a value table and an index table, so that the GPU accesses the memory with a constant time complexity each time it reads; The value table uses a resource identifier as the main dimension; and The index table uses a user identifier as the main dimension.
4. The method according to claim 3, wherein, Storing non-zero values of the N×M coefficients in the memory of the GPU using a value table and an index table includes: Storing resource identifier values mapped to the non-zero coefficients in the value table; and Storing one or more user identifier values mapped to one or more indexes in the value table in the index table, where for each of the user identifier values, the corresponding index points to one of the non-zero coefficients associated with the user identified by each of the user identifier values.
5. The method according to claim 2, wherein: Each of the multiple first objective functions is subject to K constraints and includes N×M×K coefficients; And Before processing the multiple first objective functions, the method further includes: storing non-zero values of the N×M×K coefficients in the memory of the GPU in a row-major format including at least three dimensions, so that the GPU accesses the memory with a constant time complexity each time it reads.
6. The method according to claim 1, further comprising: In response to satisfying the exit condition, allocate the plurality of resources according to the plurality of allocation plans, where the number of the plurality of users is N, the number of the plurality of resources is M, and the i of the plurality of allocation plans th corresponds to the user The allocation plan i th is represented as a vector containing M elements where each x ij indicates whether resource j th is being allocated to the user i th .
7. The method according to claim 1, wherein the exit condition includes whether the one or more dual multipliers converge.
8. The method according to claim 1, further comprising, before parallelly processing the plurality of first objective functions: Converting an original objective function for optimizing the resource allocation into a dual objective function based on a Lagrangian technique for dual problem transformation; And Decomposing the dual objective function into the plurality of first objective functions.
9. The method according to claim 8, wherein the exit condition includes whether the values of the original objective function and the dual objective function converge, and The method further comprises: Determining the value of the original objective function based on the aggregated revenue and the aggregated cost; And Determining the value of the dual objective function based on the one or more dual multipliers and the aggregated revenue and the aggregated cost.
10. A system for optimizing resource allocation, comprising one or more processors and one or more non-transitory computer-readable memories coupled to the one or more processors and configured with instructions, the instructions being executable by the one or more processors to cause the system to perform operations, the operations including: Parallelly processing a plurality of first objective functions to determine a plurality of allocation plans, wherein each of the allocation plans corresponds to allocating zero or more of the plurality of resources associated with a platform to a user, and the plurality of first objective functions share one or more dual multipliers; Parallelly determining a plurality of revenues and costs of the platform based on the plurality of allocation plans; Using parallel reduction to aggregate the calculated revenues and costs respectively to obtain an aggregated revenue and an aggregated cost; Updating the one or more dual multipliers based on the aggregated cost to determine whether an exit condition is met; And In response to the exit condition not being met, repeating the processing of the plurality of first objective functions based on the updated one or more dual multipliers to re-determine the plurality of allocation plans.
11. The system according to claim 10, wherein: The parallelly processing the plurality of first objective functions includes parallelly processing the plurality of first objective functions on a graphics processing unit (GPU); The parallelly determining the plurality of revenues and costs of the platform includes parallelly determining the plurality of revenues and costs on the GPU; The aggregating the calculated revenues and costs using parallel reduction includes aggregating the calculated revenues and costs by the GPU and a central processing unit (CPU); and The updating the one or more dual multipliers includes updating the one or more dual multipliers by the CPU.
12. The system according to claim 11, wherein: Each of the plurality of first objective functions includes N×M coefficients, N being the number of a plurality of users and M being the number of a plurality of resources; Before processing the plurality of first objective functions, the operations further include: storing non-zero coefficients of the N×M coefficients in a memory of the GPU using a value table and an index table so that the GPU accesses the memory with a constant time complexity each time it reads; The value table uses a resource identifier as a main dimension; and The index table uses the user identifier as the main dimension.
13. The system according to claim 12, wherein, Storing the non-zero values of the N×M coefficients in the memory of the GPU using a value table and an index table includes: Storing in the value table the resource identifier values mapped to the non-zero coefficients; and Storing in the index table one or more user identifier values mapped to one or more indices in the value table, wherein for each of the user identifier values, the corresponding index points to one of the non-zero coefficients associated with the user identified by each of the user identifier values.
14. The system according to claim 11, wherein: Each of the plurality of first objective functions is subject to K constraints and includes N×M×K coefficients; And Before processing the plurality of first objective functions, the operation further includes: storing the non-zero values of the N×M×K coefficients in the memory of the GPU in a row-major format including at least three dimensions so that the GPU can access the memory with a constant time complexity each time it reads.
15. The system according to claim 10, wherein the exit condition includes whether the one or more dual multipliers converge.
16. The system according to claim 10, wherein, Before parallel processing of the plurality of first objective functions, the operation further includes: Converting the original objective function for optimizing the resource allocation into a dual objective function based on the Lagrangian method of dual problem transformation; and Decomposing the dual objective function into the plurality of first objective functions.
17. The system according to claim 16, wherein the exit condition includes whether the values of the original objective function and the dual objective function converge, The operation further includes: Determining the value of the original objective function based on the aggregated revenue and the aggregated cost; And Determining the value of the dual objective function based on the one or more dual multipliers and the aggregated revenue and the aggregated cost.
18. A non-transitory computer-readable storage medium for optimizing resource allocation, configured with instructions executable by one or more processors to cause the one or more processors to perform the following operations: Parallel processing of a plurality of first objective functions to determine a plurality of allocation plans, wherein each of the allocation plans corresponds to allocating zero or more of the plurality of resources associated with a platform to users, and the plurality of first objective functions share one or more dual multipliers; Based on the plurality of allocation plans, determining the plurality of revenues and costs of the platform in parallel; Using parallel reduction to aggregate the calculated revenues and costs respectively to obtain an aggregated revenue and an aggregated cost; Updating the one or more dual multipliers based on the aggregated cost to determine whether an exit condition is met; And In response to not meeting the exit condition, repeating the processing of the plurality of first objective functions based on the updated one or more dual multipliers to re-determine a plurality of allocation plans.
19. The storage medium according to claim 18, wherein: The parallel processing of the plurality of first objective functions includes parallel processing of the plurality of first objective functions on a graphics processing unit (GPU); Said determining a plurality of revenues and costs of the platform in parallel includes determining the plurality of revenues and costs in parallel on the GPU; Said aggregating the computed revenues and costs using parallel reduction includes aggregating the computed revenues and costs by the GPU and a central processing unit (CPU); and Said updating the one or more dual multipliers includes updating the one or more dual multipliers by the CPU.
20. The storage medium according to claim 18, wherein the exit condition includes whether the one or more dual multipliers converge.
Citation Information
Patent Citations
Resource allocation method of unmanned aerial vehicle assisted wireless charging edge computing network
CN108924936A
A method and a device for determining cloud computing test resource allocation
CN109032858A