Online distribution method and device for non-stationary request

By dividing the total time of online requests into multiple time periods, and making resource allocation decisions based on history and online data in each time period, the problem of insufficient robustness of non-stationary request online allocation in the prior art is solved, and good results are achieved under various request distributions.

CN119922239APending Publication Date: 2025-05-02ALIPAY (HANGZHOU) INFORMATION TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510080646.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-17
Publication Date
2025-05-02

AI Technical Summary

Technical Problem

The prior art is poorly robust to the distribution when handling online allocation of non-stationary requests and cannot maintain good results under various possible distributions of the request.

Method used

By dividing the total time of online requests into T periods, each period is divided into exploration stages and utilization stages. During the exploration stage, based on historical data, resource allocation vectors are calculated based on historical data, decisions are made on the first N0 online requests, and samples are collected. In the utilization stage, the resource posterior allocation vector is calculated based on the collected samples, and updated based on the accumulated resources and estimated resources to obtain the target resource allocation vector for decision-making.

Benefits of technology

In a non-stationary environment, by combining historical data and online data information, the prior-based decision bias can be corrected to ensure that there are relatively good results under various possible distributions of requests, and the robustness of the allocation is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119922239A_ABST
    Figure CN119922239A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides an online allocation method and device for a non-stationary request. The method comprises the following steps: acquiring preset parameters of current online allocation; the total duration corresponding to the n online requests is divided into T time periods, and each time period is divided into an exploration stage and a utilization stage; the method comprises the following steps: in an exploration stage, based on a resource prior allocation vector and a preset parameter of a current time period calculated based on historical data, making a decision on previous N0 online requests of the current time period, and collecting N0 online request samples of the stage at the same time; in the utilization stage, based on the collected N0 online request samples, a resource posteriori allocation vector of the current time period is calculated, the resource posteriori allocation vector of the current time period is updated according to accumulated resources of all time periods before the current time period and estimated accumulated resources, and a target resource allocation vector is obtained; and making a decision on the remaining online requests based on the target resource allocation vector and a preset parameter.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] One or more embodiments of the present specification relate to the field of computers, and more particularly, to an online allocation method and apparatus for non-stationary requests. Background Art

[0002] The online allocation problem or online matching problem is a common problem in various businesses today, that is, the problem of how to make decisions with limited resources. For example, in business areas such as search, recommendation, and advertising, online decision-making needs to consider preference indicators such as click-through rate and conversion rate on the one hand, and may encounter limitations in terms of funds, costs, and traffic on the other. How to make real-time decisions on online user requests under the premise of limited resources to maximize the ROI of the overall decision-making system. Decision-making may involve the use of privacy data.

[0003] Stationary requests refer to the assumption that the arrival of user requests is independent, that is, during the entire online allocation process, the number of requests arriving at each time point, the revenue of the requests, and the resource consumption are identically distributed. The situation that does not meet this condition is called non-stationary request.

[0004] In the prior art, online allocation for non-stationary requests often cannot guarantee good results under various possible distributions of requests, that is, the robustness of the allocation is poor. Summary of the invention

[0005] One or more embodiments of the present specification describe an online allocation method and device for non-stationary requests, which can ensure relatively good results under various possible distributions of requests, that is, the allocation has good robustness.

[0006] In a first aspect, an online allocation method for non-stationary requests is provided, the method comprising:

[0007] Get the preset parameters currently allocated online;

[0008] The total duration corresponding to n online requests is divided into T time periods, and each time period is divided into an exploration phase and an exploitation phase;

[0009] In the exploration phase, based on the resource prior allocation vector of the current period calculated from historical data and the preset parameters, a decision is made on the first N0 online requests of the current period to obtain a decision result on whether to accept the request, and N0 online request samples of this phase are collected;

[0010] In the utilization stage, based on the collected N0 online request samples, the resource posterior allocation vector of the current time period is calculated, and at the same time, according to the accumulated resources of each time period before the current time period and the estimated accumulated resources, the resource posterior allocation vector of the current time period is updated to obtain the target resource allocation vector, the estimated accumulated resources are determined based on the resource posterior allocation vectors of each time period before the current time period, and the remaining online requests are decided based on the target resource allocation vector and the preset parameters to obtain a decision result of whether to accept the request.

[0011] In one possible implementation, the preset parameters include at least a resource upper bound vector; each online request has a respective benefit and resource consumption vector; the online allocation target corresponding to the decision result is to maximize the total benefit of each request accepted after the online decision, while making the total resource consumption vector of each accepted request satisfy the constraint of the resource upper bound vector.

[0012] Furthermore, the resource prior allocation vector of the current period calculated based on historical data and the preset parameters make a decision on the first N0 online requests in the current period, including:

[0013] According to the resource upper bound vector, the dual variable corresponding to the current j-th online request, the benefit, the resource consumption vector and the accumulated resource amount at the time of the j-th request, a decision is made on the j-th online request to obtain a decision result on whether to accept the request;

[0014] According to the decision result, the accumulated resource amount at the j+1th request is updated;

[0015] Based on the resource prior allocation vector of the current period calculated based on historical data and the decision result, the dual variable corresponding to the j-th online request is updated to obtain the dual variable corresponding to the j+1-th online request.

[0016] In a possible implementation manner, the calculating the resource posterior allocation vector for the current period based on the collected N0 online request samples includes:

[0017] Determine the online distribution of the collected N0 online request samples;

[0018] Based on the weighted sum of the historical distribution and online distribution of the current period, the updated distribution of the current period is obtained;

[0019] According to the update distribution of the current period, the update distribution of each period before the current period, and the historical distribution of each period after the current period, the expected resource allocation vector of each period is obtained;

[0020] The expected resource allocation vector of the current period is used as the posterior resource allocation vector of the current period.

[0021] In a possible implementation manner, updating the resource posterior allocation vector of the current period according to the accumulated resources of each period before the current period and the estimated accumulated resources includes:

[0022] Determining an output allocation strategy according to a magnitude relationship between the estimated cumulative resource allocation vector and the actual cumulative resource allocation vector;

[0023] The target resource allocation vector is determined according to the resource posterior allocation vector of the current period, the output allocation strategy and the preset fuzzy set radius.

[0024] In a possible implementation manner, updating the resource posterior allocation vector of the current period according to the accumulated resources of each period before the current period and the estimated accumulated resources includes:

[0025] According to the size relationship between the estimated cumulative resource allocation vector and the actual cumulative resource allocation vector, the product of the output allocation strategy and the fuzzy set radius is determined;

[0026] The target resource allocation vector is determined according to the resource posterior allocation vector of the current period and the product.

[0027] Further, the making a decision on the remaining online requests based on the target resource allocation vector and the preset parameters includes:

[0028] According to the resource upper bound vector, the dual variable corresponding to the current j-th online request, the benefit, the resource consumption vector and the accumulated resource amount at the time of the j-th request, a decision is made on the j-th online request to obtain a decision result on whether to accept the request;

[0029] According to the decision result, the accumulated resource amount at the j+1th request is updated;

[0030] Based on the actual resource allocation vector of the current period and the decision result, the dual variable corresponding to the j-th online request is updated to obtain the dual variable corresponding to the j+1-th online request.

[0031] In a second aspect, an online allocation device for non-stationary requests is provided, the device comprising:

[0032] An acquisition unit, used for acquiring preset parameters currently allocated online;

[0033] A partitioning unit is used to divide the total duration corresponding to n online requests into T time periods, and each time period is divided into an exploration phase and an utilization phase;

[0034] The first processing unit is used to make a decision on the first N0 online requests in the current period based on the resource prior allocation vector of the current period calculated by historical data and the preset parameters in the exploration phase, obtain a decision result on whether to accept the request, and collect N0 online request samples in this phase;

[0035] The second processing unit is used to calculate the resource posterior allocation vector of the current time period based on the N0 online request samples collected by the first processing unit during the utilization phase, and update the resource posterior allocation vector of the current time period according to the accumulated resources of each time period before the current time period and the estimated accumulated resources to obtain a target resource allocation vector, wherein the estimated accumulated resources are determined based on the resource posterior allocation vectors of each time period before the current time period, and make decisions on the remaining online requests based on the target resource allocation vector and the preset parameters to obtain a decision result on whether to accept the request.

[0036] According to a third aspect, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed in a computer, the computer is caused to execute the method of the first aspect.

[0037] According to a fourth aspect, a computing device is provided, including a memory and a processor, wherein the memory stores executable code, and when the processor executes the executable code, the method of the first aspect is implemented.

[0038] Through the method and device provided in the embodiments of this specification, first, the preset parameters of the current online allocation are obtained; then the total duration corresponding to the n online requests is divided into T time periods, and each time period is divided into an exploration phase and a utilization phase; then in the exploration phase, based on the resource a priori allocation vector of the current time period calculated from historical data and the preset parameters, a decision is made on the first N0 online requests in the current time period to obtain a decision result on whether to accept the request, and N0 online request samples of this phase are collected at the same time; finally, in the utilization phase, based on the collected N0 online request samples, a posteriori resource allocation vector of the current time period is calculated, and at the same time, according to the accumulated resources of each time period before the current time period and the estimated accumulated resources, the posteriori resource allocation vector of the current time period is updated to obtain a target resource allocation vector, the estimated accumulated resources are determined based on the posteriori resource allocation vectors of each time period before the current time period, and a decision is made on the remaining online requests based on the target resource allocation vector and the preset parameters to obtain a decision result on whether to accept the request. As can be seen from the above, the embodiments of this specification combine the information of historical data and online data to make online allocation decisions, wherein, in a non-stationary environment, historical data can guide future decisions to a certain extent, and online data can correct decision deviations based on a priori, thereby ensuring that relatively good results are achieved under various possible distributions of requests, that is, the allocation has good robustness. BRIEF DESCRIPTION OF THE DRAWINGS

[0039] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings required for use in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other accompanying drawings can be obtained based on these accompanying drawings without paying creative work.

[0040] Figure 1 A schematic diagram of an implementation scenario of an embodiment disclosed in this specification;

[0041] Figure 2 A flow chart of an online allocation method for non-stationary requests according to an embodiment is shown;

[0042] Figure 3 A two-stage processing algorithm framework diagram according to one embodiment is shown;

[0043] Figure 4 A schematic block diagram of an online allocation device for non-stationary requests according to an embodiment is shown. DETAILED DESCRIPTION

[0044] The solution provided in this specification is described below in conjunction with the accompanying drawings.

[0045] Figure 1 A schematic diagram of an implementation scenario of an embodiment disclosed in this specification. The implementation scenario involves online allocation for non-stationary requests, in which distributionally robust optimization is performed to improve the robustness of the allocation. Generally, many real-world decision-making problems that arise in engineering and management have uncertain parameters. This parameter uncertainty may be caused by limited data observability, measurement noise, implementation and prediction errors. Distributionally robust optimization is an optimization method that seeks better solutions in uncertain and complex environments. It is designed to handle optimization problems with random factors or data inaccuracies, ensuring that the solution has better results under various possible distribution conditions.

[0046] Online allocation problems are widely present in industrial applications. In typical Internet applications, such as online advertising scenarios, the platform needs to distribute advertisements to each incoming traffic and collect corresponding revenue from advertisers based on click-through conversions. There are usually restrictions on volume guarantees between the platform and advertisers, which requires the platform to allocate sufficient traffic to advertisers while ensuring the high efficiency of this part of traffic. In addition, in cloud computing scenarios, computing tasks need to be assigned to different machines online for execution. However, due to the limitations of the computing power of the machines, how to allocate corresponding machines to different tasks to maximize the overall computing efficiency is also the optimization focus of the platform.

[0047] The mathematical model of the online allocation problem is as follows: Assume there are m types of resources, Refers to the upper bound vector of resources. n requests arrive online one after another, and the coefficient of the jth request is (r j , a j ),in refers to the revenue brought by the jth request, is the resource consumption vector of the jth request. When request j arrives, the decision maker needs to immediately choose x j =0 (representing that the request is not accepted) or x j =1 (representing acceptance of the request). The decision maker's goal is to maximize the total benefit of online decision making while satisfying the resource upper bound constraint. It can be expressed as the following problem (P):

[0048]

[0049] Among them, a j =[a j1 , …, a jm ] T .

[0050] Typically, most online assignment problems assume that the request (r j , a j ) are all generated independently and identically. However, when the request (r j , a j ) is non-stationary and has unpredictable offsets or fluctuations, how to generate a more robust allocation of requests is a problem that this application is concerned with. Figure 1 This application considers the mathematical model of the online allocation problem in a non-stationary environment: Assume that the distribution of n online requests in the time series direction is non-stationary. According to some time series analysis algorithms, n online requests can be divided into T periods, and the number of requests in period t is n / T. Assume that requests in different periods t follow different real distributions. Through the historical data of the same period, a historical sample set of period t can be obtained The size of this sample set is N0, that is

[0051] It should be noted that there are various ways to divide time periods, and the number of requests contained in different time periods can be the same or different. In the embodiments of this specification, the requests can be considered to have periodicity, for example, the periodicity is reflected as one day or one week. The same time period can be understood as two time periods with the same time point in different cycles.

[0052] A common approach to solving online allocation problems is to use a primal-dual algorithm framework to transform the original problem (P) into n allocation variables x. j The optimization of is transformed into the optimization of the dual variable p in its dual problem (D). Often the dimension of the dual variable, that is, the number of resource types m, is much smaller than the number of allocation variables n. Therefore, after the transformation, the complexity of solving the optimization of the entire problem is greatly reduced. Since the original problem (P) is a linear programming problem, the strong duality holds, that is, the original problem (P) and the dual problem (D) are completely equivalent. It can be expressed as the following problem (D):

[0053]

[0054] Based on the better dual variable p * , for the jth requested real-time data (r j , a j ), if the execution policy x j (p * ):

[0055] like x j (p * )=1;

[0056] like x j (p * )=0;

[0057] Based on the better dual variable p * Strategy x j (p * ) is better, that is, a better solution to the original problem (P) However, due to the online allocation process, requests and data arrive online and cannot be obtained in advance. * Therefore, it is necessary to estimate p efficiently. * The method based on solving the original problem (P) generally calculates the dual problem or the original problem once after accumulating a certain amount of request data to obtain the estimated dual variable For subsequent decision making; the method based on solving the dual problem (D) generally uses a first-order gradient descent algorithm to solve Perform iterative updates.

[0058] The embodiments of this specification adopt a method based on solving the dual problem (D). The usual methods have certain advantages and limitations. In a non-stationary environment, historical data can guide future decisions to a certain extent, and online data can correct decision-making biases based on prior knowledge. The embodiments of this specification combine information from historical data and online data to make online allocation decisions in order to improve the robustness of the allocation.

[0059] Figure 2 A flowchart of an online allocation method for non-stationary requests according to an embodiment is shown. The method can be based on Figure 1 The implementation scenario shown in Figure 1 is as follows. Figure 2 As shown, the online allocation method for non-stationary requests in this embodiment includes the following steps: Step 21, obtaining the preset parameters of the current online allocation; Step 22, dividing the total duration corresponding to n online requests into T time periods, each of which is divided into an exploration phase and a utilization phase; Step 23, in the exploration phase, based on the resource a priori allocation vector of the current time period calculated from historical data and the preset parameters, making a decision on the first N0 online requests of the current time period, obtaining a decision result on whether to accept the request, and collecting N0 online request samples of this phase; Step 24, in the utilization phase, based on the collected N0 online request samples, calculating the resource posterior allocation vector of the current time period, and updating the resource posterior allocation vector of the current time period according to the accumulated resources of each time period before the current time period and the estimated accumulated resources, obtaining a target resource allocation vector, the estimated accumulated resources are determined based on the resource posterior allocation vector of each time period before the current time period, making a decision on the remaining online requests based on the target resource allocation vector and the preset parameters, obtaining a decision result on whether to accept the request. The specific execution method of each of the above steps is described below.

[0060] First, in step 21, the preset parameters of the current online allocation are obtained. It can be understood that the above preset parameters are different from the parameters of the online request, and are the same for each online request.

[0061] In one example, the preset parameters include at least a resource upper bound vector; each online request has a respective benefit and resource consumption vector; the online allocation target corresponding to the decision result is to maximize the total benefit of each request accepted after the online decision, while making the total resource consumption vector of each accepted request satisfy the constraint of the resource upper bound vector.

[0062] This example corresponds to the original problem (P) mentioned above. The resource upper bound vector can be expressed as b, and the coefficient of the jth request is (r j , a j ),in refers to the revenue brought by the jth request, is the resource consumption vector of the jth request.

[0063] In addition, the above preset parameters may also include at least one of the following, for example, a historical distribution sequence of T periods: in, is the historical distribution of time period t; the threshold δ in the robust deviation correction algorithm, the fuzzy set radius ε in the robust deviation correction algorithm ti, the weight α of the historical distribution when updating the distribution, the total number of requests n, the total number of periods T, and the length of the exploration phase in each period N0.

[0064] It should be noted that the total number of requests n is assumed to be known and can be determined by, but not limited to, prediction.

[0065] Then, in step 22, the total duration corresponding to the n online requests is divided into T time periods, and each time period is divided into an exploration phase and an utilization phase. It can be understood that a time period can include multiple online requests, and different time periods can include the same number of online requests or different numbers of online requests.

[0066] In the embodiment of this specification, in the exploration phase, decisions are made for the first several online requests of the time period based on historical data, and in the utilization phase, decisions are made for the remaining online requests of the time period by correcting the prior decision deviation based on online data.

[0067] Figure 3 A two-stage processing algorithm framework diagram according to one embodiment is shown. Figure 3 , in the exploration phase: the prior distribution calculated based on historical data Make decisions on the first N0 online requests in time period t, and collect N0 online request samples in this stage; in the utilization stage: update the distribution of time period t based on the collected online samples, and calculate the posterior distribution of time period t At the same time, according to the cumulative resource satisfaction of the previous period, the robust strategy used in the current period is determined to make decisions on the remaining online requests. Among them, the sum of the posterior allocations corresponding to period 1 to period t-1 can be solved The sum of actual allocations from period 1 to period t-1 is C t , to determine the robust strategy for time period t.

[0068] Then, in step 23, in the exploration phase, based on the resource prior allocation vector of the current period calculated from historical data and the preset parameters, a decision is made on the first N0 online requests of the current period to obtain a decision result on whether to accept the request, and N0 online request samples of this phase are collected. It can be understood that there are multiple types of resources, so the resource prior allocations of multiple types constitute a vector.

[0069] Furthermore, the resource prior allocation vector of the current period calculated based on historical data and the preset parameters make a decision on the first N0 online requests in the current period, including:

[0070] According to the resource upper bound vector, the dual variable corresponding to the current j-th online request, the benefit, the resource consumption vector and the accumulated resource amount at the time of the j-th request, a decision is made on the j-th online request to obtain a decision result on whether to accept the request;

[0071] According to the decision result, the accumulated resource amount at the j+1th request is updated;

[0072] Based on the resource prior allocation vector of the current period calculated based on historical data and the decision result, the dual variable corresponding to the j-th online request is updated to obtain the dual variable corresponding to the j+1-th online request.

[0073] In this example, the dual variable can be updated by stochastic gradient descent.

[0074] In the embodiment of this specification, initialization may be performed first, and then the processing of each time period may be performed in sequence. For example, the index of the time period may be initialized to t=1; the dual variable of the jth request may be p j , initialized to p1 = 0; the cumulative resource amount B at the jth request j , initialized to B1 = [0] m ; The cumulative resource volume before the tth period is C t , initialize C1 = [0] m ; According to the historical distribution sequence The pre-allocation for period 1 can be calculated, or called the prior allocation

[0075] In the exploration phase, the following processing may be included: for the first N0 requests in time period t, first, according to the current dual variable p j , determine whether to accept the current request, if accepted, let x j =1, otherwise let x j =0, and update the cumulative resource amount B j , then based on the prior distribution of the current period t For the dual variable p j Perform stochastic gradient descent.

[0076] The exploration phase can be implemented with the following code:

[0077]

[0078] Among them, t represents the index of the time period, and its value ranges from 1 to T; j represents the index of the request, assuming that the number of requests in each time period is the same.

[0079] Finally, in step 24, in the utilization phase, based on the collected N0 online request samples, the resource posterior allocation vector of the current period is calculated, and at the same time, according to the accumulated resources of each period before the current period and the estimated accumulated resources, the resource posterior allocation vector of the current period is updated to obtain the target resource allocation vector, the estimated accumulated resources are determined based on the resource posterior allocation vectors of each period before the current period, and the remaining online requests are decided based on the target resource allocation vector and the preset parameters to obtain the decision result of whether to accept the request. It can be understood that the above target resource allocation vector is conducive to achieving the above online allocation goal.

[0080] In one example, the calculating of the resource posterior allocation vector for the current period based on the collected N0 online request samples includes:

[0081] Determine the online distribution of the collected N0 online request samples;

[0082] Based on the weighted sum of the historical distribution and online distribution of the current period, the updated distribution of the current period is obtained;

[0083] According to the update distribution of the current period, the update distribution of each period before the current period, and the historical distribution of each period after the current period, the expected resource allocation vector of each period is obtained;

[0084] The expected resource allocation vector of the current period is used as the posterior resource allocation vector of the current period.

[0085] In the embodiment of the present specification, the above-mentioned utilization phase may include an update distribution sub-phase, and the update distribution sub-phase may include the following processing process: let the distribution of N0 online samples collected in the exploration phase be Historical distribution based on time period t The weight α and online distribution The weight 1-α is used to obtain the updated distribution of time period t Then according to the updated distribution of time period 1, ..., t and the historical distribution of time periods t+1,…,T that have not been updated Get the expected resource allocation vector for each period in the tth round of calculation Let the posterior distribution of the current period t be Let the prior distribution of the next period t+1 be And calculate the estimated cumulative resource allocation vector for time period 1, ..., t-1

[0086] In the embodiments of the present specification, a robust correction sub-algorithm can be called to obtain an output allocation strategy, and then the target resource allocation vector can be determined based on the output allocation strategy; or another robust correction sub-algorithm can be called to obtain the product of the output allocation strategy and the fuzzy set radius, and then the target resource allocation vector can be determined based on the product.

[0087] In one example, updating the resource posterior allocation vector of the current period according to the accumulated resources of each period before the current period and the estimated accumulated resources includes:

[0088] Determining an output allocation strategy according to a magnitude relationship between the estimated cumulative resource allocation vector and the actual cumulative resource allocation vector;

[0089] The target resource allocation vector is determined according to the resource posterior allocation vector of the current period, the output allocation strategy and the preset fuzzy set radius.

[0090] In this example, by calling Thus, the output allocation strategy e is obtained ti , and then calculate the target resource allocation vector It is understandable that the fuzzy set radius ε needs to be determined in advance ti .

[0091] Among them, the sub-algorithm Input parameters include

[0092] This sub-algorithm can be implemented by the following code:

[0093]

[0094] The sub-algorithm outputs {e i}.

[0095] This sub-algorithm is mainly used for, for each dimension i = 1, ..., m, if This means that the estimated real allocation exceeds the actual allocation by a large amount, which means that the actual allocation in the past was too small, so more allocation should be made at this time. i =-1; if This means that the actual allocation exceeds the estimated real allocation by a large amount, which means that the actual allocation in the past was large, so less allocation should be made at this time. i =1; in other cases, no adjustment is required, that is, e i =0.

[0096] In another example, updating the resource posterior allocation vector of the current period according to the accumulated resources of each period before the current period and the estimated accumulated resources includes:

[0097] According to the size relationship between the estimated cumulative resource allocation vector and the actual cumulative resource allocation vector, the product of the output allocation strategy and the fuzzy set radius is determined;

[0098] The target resource allocation vector is determined according to the resource posterior allocation vector of the current period and the product.

[0099] In this example, by calling Thus we get the aforementioned product ε ti e ti , and then calculate the target resource allocation vector It is understandable that the fuzzy set radius ε does not need to be determined in advance ti .

[0100] Among them, the sub-algorithm Input parameters include

[0101] This sub-algorithm can be implemented by the following code:

[0102]

[0103] This sub-algorithm outputs {ε i e i}.

[0104] This sub-algorithm is mainly used for, for each dimension i = 1, ..., m, if This means that the estimated real allocation exceeds the actual allocation by a large amount, which means that the actual allocation in the past was too small, so more allocation should be made at this time. i = -1, and directly gives ε i The value selection strategy is: if This means that the actual allocation exceeds the estimated real allocation by a large amount, which means that the actual allocation in the past was large, so less allocation should be made at this time. i = 1, giving ε i The value of In other cases, no adjustment is required, that is, e i =0.

[0105] Further, the making a decision on the remaining online requests based on the target resource allocation vector and the preset parameters includes:

[0106] According to the resource upper bound vector, the dual variable corresponding to the current j-th online request, the benefit, the resource consumption vector and the accumulated resource amount at the time of the j-th request, a decision is made on the j-th online request to obtain a decision result on whether to accept the request;

[0107] According to the decision result, the accumulated resource amount at the j+1th request is updated;

[0108] Based on the actual resource allocation vector of the current period and the decision result, the dual variable corresponding to the j-th online request is updated to obtain the dual variable corresponding to the j+1-th online request.

[0109] In the utilization phase, the following processing may be included: for all remaining requests in time period t, the current dual variable p j , determine whether to accept the current request, if accepted, let x j =1, otherwise let x j =0, and update the cumulative resource amount B j , then based on the robust posterior allocation l in the current period t t , for the dual variable p j Perform stochastic gradient descent.

[0110] The exploitation phase can be achieved through the following code:

[0111]

[0112] Among them, t represents the index of the time period, and its value ranges from 1 to T; j represents the index of the request, assuming that the number of requests in each time period is the same.

[0113] In addition, C can be updated after the exploitation phase. t , denoted by C t+1 =B j+1 .

[0114] Finally, the decision results corresponding to the n requests are obtained, that is, x = (x1, ..., x n ).

[0115] The effects achieved by the embodiments of this specification are analyzed and explained below.

[0116] First, the bias based on prior allocation and the suboptimality based on online allocation are overcome.

[0117] We need to introduce some mathematical concepts. Based on the Wassersetin distance, we can define a nonstationary measure called nonstationary-budget:

[0118]

[0119] in, is the real distribution sequence, is the prior distribution sequence based on historical samples, for and The Wasserstein distance between . Given a prior distribution sequence Assume that the fuzzy set where the true distribution sequence lies is defined by the budget W T Decide:

[0120]

[0121] Known: For the sequence based on the prior distribution The algorithm EMP for distribution is used if the real distribution sequence Falling in the fuzzy set B(W T ), then the upper bound of the worst-regret of the EMP algorithm is W T .

[0122] Assuming a prior distribution The empirical distribution of N0 historical samples, is the empirical distribution of N0 online samples collected in the exploration phase. The method of the embodiment of this specification is to let α∈[0,1] to generate a mixed distribution of historical and online samples. The hyperparameter α represents the degree of trust in history and priors. Next, we need to explain the mixed distribution The operation is conducive to online decision-making. It can be proved that With high probability:

[0123] When α = 1, Heng established;

[0124] When α∈[0,1), given the confidence β∈(0,1), the number of samples N0 and other constants c1, c2, a, when

[0125]

[0126] It is established with high probability of 1-β.

[0127] Then, the distribution sequence can be updated as follows:

[0128] True distribution sequence

[0129] Historical distribution series

[0130] Period 1

[0131] Period 2 ............

[0133] Time period t ............

[0135] Period T-1

[0136] Time period T

[0137] It can be immediately obtained that the following formula is established with high probability:

[0138]

[0139] As the distribution sequence is updated, the non-stationary budget continues to decrease, and the upper bound of the worst-regret is also continuously reduced. Therefore, it can be considered that based on the current The allocation is better than that based on of allocation.

[0140] Second, it overcomes the challenge of uneven resource distribution.

[0141] Because before the utilization phase of each period t, it is necessary to Recalculate the posterior allocation. Note that the allocations for time periods 1, ..., t-1 are still being updated and are closer to the actual allocations, but time periods 1, ..., t-1 have already been allocated. Therefore, the actual cumulative allocations C for time periods 1, ..., t-1 are t and pre-accumulated allocations There may be a certain gap between them, resulting in poor decision-making results. Therefore, consider the actual cumulative allocation amount C according to time period 1, ..., t-1 t and pre-accumulated allocations The gap between the posterior distribution of the current period t Make robust decisions.

[0142] First, the allocation amount b is modeled by distributed robust optimization. According to the form of the dual problem (D):

[0143]

[0144] For a certain period of time, it can be considered that the allocation of each resource is d1, ..., d m are independent random variables. Therefore, let the dual variable p = [p1, ..., p m ] T , the dual problem (D) is equivalent to:

[0145]

[0146] Consider constructing an inner model of distributed robust optimization for resource i, and the center of the fuzzy set is the equally divided resource amount b i / n, the fuzzy set radius is set to ε i :

[0147] dual-worst: in Equivalent to the primal-best strategy:

[0148]

[0149] dual-best: in Equivalent to the primal-worst strategy:

[0150]

[0151] Assume that d i p i -ε i e i p i To unify the solvable forms of the two types of inner problems, where e i ∈{0, 1, -1} represents the strategy used, 0 represents the equal-split strategy, 1 represents primal-worst, and -1 represents primal-best. Then, given ε i , e i , the distributed robust optimization model of the allocation amount b is as follows:

[0152]

[0153] where l = [l1, ..., l m ],l i =d i -ε i e i The fuzzy set radius ε corresponding to resource i i is an adjustable hyperparameter, but what kind of robust strategy should be used for each resource? i Need to be determined based on prior information.

[0154] Back to the specific problem, based on the above analysis framework, the posterior allocation of the i-th resource is the fuzzy set center, and ε i is the fuzzy set radius of resource i, with e i For the robust strategy used by resource i, a solvable form of the distributed robust optimization model can be established for the requests in time period t:

[0155]

[0156] in For (Robust-Dual), the online gradient descent method is used to iteratively solve the dual variables:

[0157] p j+1 =p j +αj (a j x j -1)

[0158] The next step is to determine e i , a natural idea is to use the real cumulative allocation C of time periods 1, ..., t-1 t and pre-accumulated allocations The gap between and determines the robust strategy to be used in the current period:

[0159] if This means that too many resources were allocated in the previous period, and primal-worst should be used at this time;

[0160] if This indicates that too few resources were allocated in the previous period, and primal-best should be used at this time;

[0161] In other cases, there is no need to make up for the previous amount of resources, and an equal distribution strategy is adopted.

[0162] Among them, δ is the tolerance threshold input by the decision maker, C t =[C t1 , ..., C tm ] T .

[0163] The solution provided in the embodiments of this specification has been proven to have good robustness through simulation experiments.

[0164] According to the method provided in the embodiments of the present specification, first, the preset parameters of the current online allocation are obtained; then the total duration corresponding to the n online requests is divided into T time periods, and each time period is divided into an exploration phase and a utilization phase; then in the exploration phase, based on the resource a priori allocation vector of the current time period calculated from historical data and the preset parameters, a decision is made on the first N0 online requests of the current time period to obtain a decision result on whether to accept the request, and at the same time, N0 online request samples of this phase are collected; finally, in the utilization phase, based on the collected N0 online request samples, a resource posterior allocation vector of the current time period is calculated, and at the same time, according to the accumulated resources of each time period before the current time period and the estimated accumulated resources, the resource posterior allocation vector of the current time period is updated to obtain a target resource allocation vector, the estimated accumulated resources are determined based on the resource posterior allocation vectors of each time period before the current time period, and a decision is made on the remaining online requests based on the target resource allocation vector and the preset parameters to obtain a decision result on whether to accept the request. As can be seen from the above, the embodiments of this specification combine the information of historical data and online data to make online allocation decisions, wherein, in a non-stationary environment, historical data can guide future decisions to a certain extent, and online data can correct decision deviations based on a priori, thereby ensuring that relatively good results are achieved under various possible distributions of requests, that is, the allocation has good robustness.

[0165] According to another aspect of the embodiment, there is also provided an online allocation device for non-stationary requests, wherein the device is used to execute the method provided in the embodiment of this specification. Figure 4 FIG. 1 is a schematic block diagram of an online allocation device for non-stationary requests according to an embodiment. Figure 4 As shown, the device 400 includes:

[0166] An acquisition unit 41 is used to acquire preset parameters currently allocated online;

[0167] A division unit 42 is used to divide the total duration corresponding to the n online requests into T time periods, each time period is divided into an exploration phase and an utilization phase;

[0168] The first processing unit 43 is used to make a decision on the first N0 online requests in the current period based on the resource prior allocation vector of the current period calculated by historical data and the preset parameters in the exploration phase, obtain a decision result on whether to accept the request, and collect N0 online request samples in this phase;

[0169] The second processing unit 44 is used to calculate the resource posterior allocation vector of the current time period based on the N0 online request samples collected by the first processing unit 43 during the utilization phase, and update the resource posterior allocation vector of the current time period according to the accumulated resources of each time period before the current time period and the estimated accumulated resources to obtain a target resource allocation vector, wherein the estimated accumulated resources are determined based on the resource posterior allocation vectors of each time period before the current time period, and make decisions on the remaining online requests based on the target resource allocation vector and the preset parameters to obtain a decision result on whether to accept the request.

[0170] Optionally, as an embodiment, the preset parameters include at least a resource upper bound vector; each online request has a respective benefit and resource consumption vector; the online allocation target corresponding to the decision result is to maximize the total benefit of each request accepted after the online decision, while making the total resource consumption vector of each accepted request satisfy the constraint of the resource upper bound vector.

[0171] Furthermore, the first processing unit 43 includes:

[0172] The decision subunit is used to make a decision on the j-th online request based on the resource upper bound vector, the dual variable corresponding to the current j-th online request, the benefit, the resource consumption vector and the accumulated resource amount at the time of the j-th request, and obtain a decision result on whether to accept the request;

[0173] A first updating subunit, configured to update the accumulated resource amount at the j+1th request according to the decision result obtained by the decision subunit;

[0174] The second updating subunit is used to update the dual variable corresponding to the j-th online request based on the resource prior allocation vector of the current time period calculated by historical data and the decision result, so as to obtain the dual variable corresponding to the j+1-th online request.

[0175] Optionally, as an embodiment, the second processing unit 44 includes:

[0176] A distribution determination subunit, used to determine the online distribution of the collected N0 online request samples;

[0177] A distribution updating subunit, used to perform weighted summation based on the historical distribution of the current period and the online distribution obtained by the distribution determining subunit, to obtain an updated distribution of the current period;

[0178] The expected allocation subunit is used to obtain the expected resource allocation vector for each time period according to the updated distribution of the current time period, the updated distribution of each time period before the current time period and the historical distribution of each time period after the current time period obtained by the distribution update subunit;

[0179] The a posteriori allocation subunit is used to use the expected resource allocation vector of the current period obtained by the expected allocation subunit as the a posteriori resource allocation vector of the current period.

[0180] Optionally, as an embodiment, the second processing unit 44 includes:

[0181] A strategy determination subunit, used to determine an output allocation strategy according to a magnitude relationship between the estimated cumulative resource allocation vector and the actual cumulative resource allocation vector;

[0182] The first target allocation subunit is used to determine the target resource allocation vector according to the resource posterior allocation vector of the current period, the output allocation strategy obtained by the strategy determination subunit and the preset fuzzy set radius.

[0183] Optionally, as an embodiment, the second processing unit 44 includes:

[0184] A product determination subunit, used to determine the product of the output allocation strategy and the fuzzy set radius according to the size relationship between the estimated cumulative resource allocation vector and the actual cumulative resource allocation vector;

[0185] The second target allocation subunit is used to determine the target resource allocation vector according to the resource posteriori allocation vector of the current time period and the product obtained by the product determination subunit.

[0186] Further, the second processing unit 44 includes:

[0187] The decision subunit is used to make a decision on the j-th online request based on the resource upper bound vector, the dual variable corresponding to the current j-th online request, the benefit, the resource consumption vector and the accumulated resource amount at the time of the j-th request, and obtain a decision result on whether to accept the request;

[0188] A third updating subunit is used to update the accumulated resource amount at the j+1th request according to the decision result obtained by the decision subunit;

[0189] The fourth updating subunit is used to update the dual variable corresponding to the j-th online request based on the actual resource allocation vector of the current time period and the decision result, so as to obtain the dual variable corresponding to the j+1-th online request.

[0190] Through the device provided by the embodiment of the present specification, first, the acquisition unit 41 acquires the preset parameters of the current online allocation; then the division unit 42 divides the total duration corresponding to the n online requests into T time periods, and each time period is divided into an exploration phase and a utilization phase; then the first processing unit 43 makes a decision on the first N0 online requests of the current time period based on the resource a priori allocation vector of the current time period calculated by historical data and the preset parameters in the exploration phase, and obtains a decision result of whether to accept the request, and collects N0 online request samples of this phase; finally, the second processing unit 44 calculates the resource posterior allocation vector of the current time period based on the collected N0 online request samples in the utilization phase, and updates the resource posterior allocation vector of the current time period according to the accumulated resources of each time period before the current time period and the estimated accumulated resources, and obtains the target resource allocation vector, the estimated accumulated resources are determined based on the resource posterior allocation vectors of each time period before the current time period, and makes a decision on the remaining online requests based on the target resource allocation vector and the preset parameters, and obtains a decision result of whether to accept the request. As can be seen from the above, the embodiments of this specification combine the information of historical data and online data to make online allocation decisions, wherein, in a non-stationary environment, historical data can guide future decisions to a certain extent, and online data can correct decision deviations based on a priori, thereby ensuring that relatively good results are achieved under various possible distributions of requests, that is, the allocation has good robustness.

[0191] According to another embodiment, there is also provided a computer-readable storage medium having a computer program stored thereon, which, when executed in a computer, causes the computer to execute a combination of Figure 2 The method described.

[0192] According to another embodiment of the present invention, a computing device is provided, including a memory and a processor, wherein the memory stores executable code, and when the processor executes the executable code, the Figure 2 The method described.

[0193] Those skilled in the art should be aware that in one or more of the above examples, the functions described in the present invention can be implemented by hardware, software, firmware or any combination thereof. When implemented by software, these functions can be stored in a computer-readable medium or transmitted as one or more instructions or codes on a computer-readable medium.

[0194] The specific implementation methods described above further illustrate the objectives, technical solutions and beneficial effects of the present invention in detail. It should be understood that the above description is only a specific implementation method of the present invention and is not intended to limit the scope of protection of the present invention. Any modifications, equivalent substitutions, improvements, etc. made on the basis of the technical solution of the present invention should be included in the scope of protection of the present invention.

Claims

1. An online allocation method for non-stationary requests, the method comprising: Get the preset parameters currently allocated online; The total duration corresponding to n online requests is divided into T time periods, and each time period is divided into an exploration phase and an exploitation phase; In the exploration phase, based on the resource prior allocation vector of the current period calculated from historical data and the preset parameters, a decision is made on the first N0 online requests of the current period to obtain a decision result on whether to accept the request, and N0 online request samples of this phase are collected; In the utilization stage, based on the collected N0 online request samples, the resource posterior allocation vector of the current time period is calculated, and at the same time, according to the accumulated resources of each time period before the current time period and the estimated accumulated resources, the resource posterior allocation vector of the current time period is updated to obtain the target resource allocation vector, the estimated accumulated resources are determined based on the resource posterior allocation vectors of each time period before the current time period, and the remaining online requests are decided based on the target resource allocation vector and the preset parameters to obtain a decision result of whether to accept the request.

2. The method of claim 1, wherein: The preset parameters include at least a resource upper bound vector; each online request has a respective benefit and resource consumption vector; the online allocation target corresponding to the decision result is to maximize the total benefit of each request accepted after the online decision, while making the total resource consumption vector of each accepted request satisfy the constraint of the resource upper bound vector.

3. The method of claim 2, wherein: The resource prior allocation vector of the current period calculated based on historical data and the preset parameters make a decision on the first N0 online requests in the current period, including: According to the resource upper bound vector, the dual variable corresponding to the current j-th online request, the benefit, the resource consumption vector and the accumulated resource amount at the time of the j-th request, a decision is made on the j-th online request to obtain a decision result on whether to accept the request; According to the decision result, the accumulated resource amount at the j+1th request is updated; Based on the resource prior allocation vector of the current period calculated based on historical data and the decision result, the dual variable corresponding to the j-th online request is updated to obtain the dual variable corresponding to the j+1-th online request.

4. The method of claim 1, wherein: The calculating of the resource posterior allocation vector of the current period based on the collected N0 online request samples includes: Determine the online distribution of the collected N0 online request samples; Based on the weighted sum of the historical distribution and online distribution of the current period, the updated distribution of the current period is obtained; According to the update distribution of the current period, the update distribution of each period before the current period, and the historical distribution of each period after the current period, the expected resource allocation vector of each period is obtained; The expected resource allocation vector of the current period is used as the posterior resource allocation vector of the current period.

5. The method of claim 1, wherein: The updating of the resource posterior allocation vector of the current period according to the accumulated resources of each period before the current period and the estimated accumulated resources includes: Determining an output allocation strategy according to a magnitude relationship between the estimated cumulative resource allocation vector and the actual cumulative resource allocation vector; The target resource allocation vector is determined according to the resource posterior allocation vector of the current period, the output allocation strategy and the preset fuzzy set radius.

6. The method of claim 1, wherein: The updating of the resource posterior allocation vector of the current period according to the accumulated resources of each period before the current period and the estimated accumulated resources includes: According to the size relationship between the estimated cumulative resource allocation vector and the actual cumulative resource allocation vector, the product of the output allocation strategy and the fuzzy set radius is determined; The target resource allocation vector is determined according to the resource posterior allocation vector of the current period and the product.

7. The method of claim 2, wherein: The making a decision on the remaining online requests based on the target resource allocation vector and the preset parameters includes: According to the resource upper bound vector, the dual variable corresponding to the current j-th online request, the benefit, the resource consumption vector and the accumulated resource amount at the time of the j-th request, a decision is made on the j-th online request to obtain a decision result on whether to accept the request; According to the decision result, the accumulated resource amount at the j+1th request is updated; Based on the actual resource allocation vector of the current period and the decision result, the dual variable corresponding to the j-th online request is updated to obtain the dual variable corresponding to the j+1-th online request.

8. An online allocation device for non-stationary requests, the device comprising: An acquisition unit, used for acquiring preset parameters currently allocated online; A partitioning unit is used to divide the total duration corresponding to n online requests into T time periods, and each time period is divided into an exploration phase and an utilization phase; The first processing unit is used to make a decision on the first N0 online requests in the current period based on the resource prior allocation vector of the current period calculated by historical data and the preset parameters in the exploration phase, obtain a decision result on whether to accept the request, and collect N0 online request samples in this phase; The second processing unit is used to calculate the resource posterior allocation vector of the current time period based on the N0 online request samples collected by the first processing unit during the utilization phase, and update the resource posterior allocation vector of the current time period according to the accumulated resources of each time period before the current time period and the estimated accumulated resources to obtain a target resource allocation vector, wherein the estimated accumulated resources are determined based on the resource posterior allocation vectors of each time period before the current time period, and make decisions on the remaining online requests based on the target resource allocation vector and the preset parameters to obtain a decision result on whether to accept the request.

9. A computer-readable storage medium having a computer program stored thereon, which, when executed in a computer, causes the computer to execute the method according to any one of claims 1 to 7.

10. A computing device comprising a memory and a processor, wherein the memory stores executable codes, and when the processor executes the executable codes, the method according to any one of claims 1 to 7 is implemented.