Multi-tiered driver allocation optimization framework for omnichannel commerce

The multi-tiered driver allocation framework leverages machine learning and stochastic modeling to optimize driver allocation across different planning horizons, improving delivery efficiency and profitability by addressing demand fluctuations and disruptions in e-commerce logistics.

WO2025227026A9PCT designated stage Publication Date: 2026-01-02CLEVELAND STATE UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/US2025/026353
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-04-25
Filing Date
2025-04-25
Publication Date
2026-01-02

AI Technical Summary

Technical Problem

E-commerce delivery companies face challenges in driver allocation due to the shift from crowdsourced models to third-party services, requiring advanced systems for long-term, short-term, and real-time planning to optimize driver allocation and meet demand fluctuations effectively.

Method used

A multi-tiered driver allocation optimization framework using machine learning and stochastic modeling, integrating long-term planning with mixture density networks, short-term planning with Light Gradient Boosting Machine (LGBM) and Neural Hierarchical Interpolation for Time Series Forecasting (N-HiTS), and real-time planning with a Markov decision process to dynamically allocate drivers.

Benefits of technology

The framework enhances delivery efficiency and profitability by up to 17.9%, ensuring reduced delivery times, improved operational efficiency, and elevated service standards, while addressing demand fluctuations and disruptions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure IMGF000016_0001
    Figure IMGF000016_0001
  • Figure IMGF000016_0002
    Figure IMGF000016_0002
  • Figure IMGF000016_0003
    Figure IMGF000016_0003
Patent Text Reader

Abstract

A multi-tiered optimization and machine learning framework is provided in which operational decisions as far as driver capacity allocation are taken at three different levels or stages, namely, long-term, short-term, and real-time. While decisions at each stage have individual importance for the operations at different time horizons, this framework utilizes decisions upstream to inform better results downstream. This synergy helps improve allocations when compared to solutions that assign driver capacity for each stage individually.
Need to check novelty before this filing date? Find Prior Art

Description

MULTI-TIERED DRIVER ALLOCATION OPTIMIZATION FRAMEWORK FOR OMNICHANNEL COMMERCE

[0001] This application claims priority to and the benefit of U.S. provisional patent application Serial No. 63 / 638,668, filed on April 25, 2024, the full disclosure of which, including appendices, is incorporated herein by reference in its entirety. TECHNICAL FIELD

[0002] The present exemplary embodiments relate to driver allocation and find particular application in conjunction with omnichannel commerce, and will be described with particular reference thereto. However, it is to be appreciated that the present exemplary embodiments are also amenable to other like applications. BACKGROUND

[0003] As global regulations shift towards favoring third-party delivery services over crowdsourced models, a new demand for advanced driver allocation systems emerges. This shift creates a challenge for e-commerce delivery companies, as they can no longer rely on the highly flexible supply of crowdsourced drivers and must request the number of drivers needed ahead of time. BRIEF DESCRIPTION

[0004] In accordance with one aspect of the present exemplary embodiments, a system comprises at least one processor, and, at least one memory having stored thereon instructions or code that, when executed by the at least one processor, causes the system to implement a multi-tiered driver allocation optimization framework for omnichannel commerce by at least performing long-term planning based on a mixture density network to obtain an estimate of demand distribution and generate distribution parameter covariates, short-term planning using the distribution parameter covariates in deep learning ensemble forecasting to predict demand variation and generate current state parameter allocations, real-time planning based on the current state demand variation to capture impact of currentstate on predictions and a Markov decision process to generate a driver allocation, and, outputting the driver allocation to fulfill deliveries by drivers.

[0005] In accordance with another aspect of the present exemplary embodiments, the long-term planning system is based on the mixture density network connected to a newsvendor model.

[0006] In accordance with another aspect of the present exemplary embodiments, the short-term planning system comprises a combination of Light Gradient Boosting Machine (LGBM) and Neural Hierarchical Interpolation for Time Series Forecasting (N-HiTS) forecasting.

[0007] In accordance with another aspect of the present exemplary embodiments, the current state comprises temporal spillover of driver shortages and backorders.

[0008] In accordance with another aspect of the present exemplary embodiments, the driver allocation system includes hourly responsive reallocations.

[0009] In accordance with one aspect of the present exemplary embodiments, a method for implementing a multi-tiered driver allocation optimization framework for omnichannel commerce comprises at least one processor, and, at least one memory having stored thereon instructions or code that, when executed by the at least one processor, causes the system to implement a multi-tiered driver allocation optimization framework for omnichannel commerce by at least performing long-term planning based on a mixture density network to obtain an estimate of demand distribution and generate distribution parameter covariates, short-term planning using the distribution parameter covariates in deep learning ensemble forecasting to predict demand variation and generate current state parameter allocations, real-time planning based on the current state demand variation to capture impact of current state on predictions and a Markov decision process to generate a driver allocation;, and, outputting the driver allocation to fulfill deliveries by drivers.

[0010] In accordance with another aspect of the present exemplary embodiments, the long-term planning method is based on the mixture density network connected to a newsvendor model.

[0011] In accordance with another aspect of the present exemplary embodiments, the short-term planning method comprises a combination of Light Gradient Boosting Machine(LGBM) and Neural Hierarchical Interpolation for Time Series Forecasting (N-HiTS) forecasting.

[0012] In accordance with another aspect of the present exemplary embodiments, the current state comprises temporal spillover of driver shortages and backorders.

[0013] In accordance with another aspect of the present exemplary embodiments, the driver allocation method includes hourly responsive reallocations.

[0014] In accordance with one aspect of the present exemplary embodiments, a non- transitory computer readable medium comprises instructions stored thereon that, when executed by a processor, cause an apparatus to perform long-term planning based on a mixture density network to obtain an estimate of demand distribution and generate distribution parameter covariates, short-term planning using the distribution parameter covariates in deep learning ensemble forecasting to predict demand variation and generate current state parameter allocations, real-time planning based on the current state demand variation to capture impact of current state on predictions and a Markov decision process to generate a driver allocation, and, outputting the driver allocation to fulfill deliveries by drivers.

[0015] In accordance with another aspect of the present exemplary embodiments, the long-term planning is based on the mixture density network connected to a newsvendor model.

[0016] In accordance with another aspect of the present exemplary embodiments, the short-term planning comprises a combination of Light Gradient Boosting Machine (LGBM) and Neural Hierarchical Interpolation for Time Series Forecasting (N-HiTS) forecasting.

[0017] In accordance with another aspect of the present exemplary embodiments, the current state comprises temporal spillover of driver shortages and backorders.

[0018] In accordance with another aspect of the present exemplary embodiments, the driver allocation includes hourly responsive reallocations. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] FIGURE 1 is an example flow diagram illustrating an example method according to the presently described embodiments;

[0020] FIGURE 2 is an example flow diagram illustrating an example method according to the presently described embodiments;

[0021] FIGURE 3 is an example flow diagram illustrating an example mixture distribution network according to the presently described embodiments;

[0022] FIGURES 4(a)-4(d) are graphs showing results;

[0023] FIGURE 5 is a graph showing results;

[0024] FIGURE 6(a)-6(b) are graphs showing results;

[0025] FIGURES 7(a)-7(b) are graphs showing results;

[0026] FIGURES 8(a)-8(b) are graphs showing results;

[0027] FIGURE 9 is a graph showing results; and,

[0028] FIGURE 10 illustrates an example system according to the presently described embodiments. DETAILED DESCRIPTION

[0029] The presently described embodiments, in at least one form, implement a multi- tiered framework designed to optimize driver allocations for last-mile deliveries. This framework, integrating advanced methodologies such as, for example, machine learning and stochastic modeling, marks a departure from conventional models.

[0030] The presently described embodiments offer a dynamic and responsive approach to driver allocation to allow drivers to fulfill deliveries in an improved, e.g., optimized, manner. This includes long-term strategic planning, mid-term adjustments, and real-time operational decisions. Validated using data of a leading e-commerce company, the presently described embodiments demonstrate substantial improvements in delivery efficiency and profitability, increasing a reward function accounting for revenue and delivery costs, for example, by up to 17.9%. This versatile framework offers e-commerce entities a compelling solution to navigate the complexities of third-party delivery logistics, ensuring reduced delivery times, enhanced operational efficiency, and elevated service standards. It presents a significant step forward in sustainable and reliable e-commerce logistics management.

[0031] According to the presently describe embodiments, in at least one form, in long- term driver allocation planning, predictive analytics are leveraged using extensivehistorical data to discern overarching trends and project the required number of drivers for the next one to two years. This model learns from long-horizon historical data, such as past delivery volumes and weekday variations, enabling the system to accurately anticipate future demand distributions. The insights gained from this analysis foster better strategic decision-making, facilitating more informed budgetary allocations, effective contract negotiations with third-party providers, and efficient resource planning.

[0032] In a short-term driver allocation application, the presently described embodiments effectively utilize historical data up to a few days or weeks prior to the specific hour of driver allocation. This data-driven approach enables refinement and updating of long-term models, providing a more accurate and responsive understanding of imminent demand patterns. By doing so, the presently described embodiments ensure efficient scheduling that aligns closely with real-time demand fluctuations and inventory requirements. This not only facilitates more effective communication with third-party providers but also significantly reduces the likelihood of last-minute operational changes. The result is a streamlined process that enhances the ability to meet demand promptly, ultimately leading to improved operational efficiency and customer satisfaction.

[0033] In a real-time driver allocation feature, the system is specifically designed to adeptly handle last-minute changes, enhancing operational flexibility depending on the contract terms with third-party providers. This includes situations like driver no-shows or unexpected spikes in demand. It does so by meticulously balancing the likelihood of backorders against the potential for lost demand, while also considering the possibility of continuing outlier demand trends. This real-time decision-making capability not only ensures a high level of responsiveness to dynamic market conditions but also contributes significantly to maintaining operational efficiency and customer satisfaction. By minimizing the impacts of unforeseen disruptions, the presently described embodiments aid in reducing operational costs and improving service reliability.

[0034] Advantages of the presently described embodiments include Integrated Multi- Horizon Planning including unique integration of long-term, short-term, and real-time planning. Competing systems often use separate models for different planning horizons, leading to disjointed and less efficient decision-making.

[0035] Advanced Long-Term Forecasting of the presently described embodiments utilizes cutting-edge machine learning (ML) techniques, surpassing traditional methods like ARIMA in performance. Research indicates a significant performance edge of ML- based methods, like ours, over standard time series approaches

[0036] Enhanced Short-Term Forecasting with a Decision-Making Pipeline of the presently described embodiments utilizes advanced ML forecasting methods (outperforming common ML approaches like Gradient Boosting Machines and Neural Networks) but also integrates these forecasts into a robust decision-making tool. This integration is a key differentiator, as many companies implement ML-based forecasting without a subsequent decision-making process.

[0037] Innovative Real-Time Allocation Solutions of the presently described embodiments addresses the often-overlooked real-time allocation challenge, which most companies handle ad-hoc, based on intuition and experience. This approach is data-driven, providing a systematic and reliable method for making last-minute decisions, thereby supporting on-the-ground decision-making with empirical evidence.

[0038] Thus, the presently described embodiments, in at least one form, comprise a multi-tiered optimization and machine learning framework in which operational decisions as far as driver capacity allocation are taken at three different levels or stages, namely, long-term, short-term, and real-time. While decisions at each stage have individual importance for the operations at different time horizons, this framework utilizes decisions upstream to inform better results downstream. This synergy helps improve allocations when compared to solutions that assign driver capacity for each stage individually.

[0039] In at least one form, the framework comprises a series of three models that are run sequentially. With reference to FIGURE 1, a graphical description of this process or method 100 is provided. The first model is a newsvendor model that uses demand distributions derived from mixture density networks to produce a first driver allocation at the long-term planning stage 102. The covariates from the stationary demand distribution are used as inputs in a deep learning ensemble forecasting consisting of GBM and N-HiTs submodels, along with recent demand information to attain demand forecasts at the short- term planning level 104. This gives way to a second driver allocation that modifies the one produced for long-term decision-making. Finally, in real-time or operational planning 106,this short-term allocation is used in conjunction with last-minute information about delayed or cancelled orders to produce a final, real-time driver allocation. This last action 106 occurs in the context of a finite-state Markov decision process.

[0040] With continuing reference to FIGURE 1, it will be appreciated that, in at least one form, the presently described embodiments may be implemented as a system, as a method, and / or as a non-transitory computer readable medium comprising instructions stored thereon that, when executed by a processor, cause an apparatus to perform. As a system, a variety of forms may be implemented but, in at least one form, the system comprises at least one processor and at least one memory having stored thereon instructions or code that, when executed by the at least one processor, cause the system to implement the presently described embodiments.

[0041] In these implementations, it should be appreciated that the presently described embodiments, in at least one form, implement, apply, perform, and / or execute long-term planning based on a mixture density network to obtain an estimate of demand distribution and generate distribution parameter covariates, short-term planning using the distribution parameter covariates in deep learning ensemble forecasting to predict demand variation and generate current state parameter allocations, real-time planning based on the current state demand variation to capture impact of current state on predictions and a Markov decision process to generate a driver allocation, and, outputting the driver allocation to fulfill deliveries by drivers.

[0042] In an example of the present exemplary embodiments, the long-term planning system is based on the mixture density network connected to a newsvendor model.

[0043] In an example of the present exemplary embodiments, the short-term planning system comprises a combination of Light Gradient Boosting Machine (LGBM) and Neural Hierarchical Interpolation for Time Series Forecasting (N-HiTS) forecasting.

[0044] In an example of the present exemplary embodiments, the current state comprises temporal spillover of driver shortages and backorders.

[0045] In an example of the present exemplary embodiments, the driver allocation system includes hourly responsive reallocations.

[0046] It should be appreciated that the improved, e.g., optimized, driver allocation that is generated by the presently described embodiments is output and, for example, usedby entities or organizations that manage drivers and scheduled deliveries. This improved driver allocation is used by the managing organization or entity to generate and / or adapt or change a delivery schedule according to the input data and implemented features of the presently described embodiments. This use can be fully automated or include aspects of human intervention as may be desired or appropriate in the circumstances.

[0047] More particularly, the presently described embodiments, in at least one form, implement a multi-tiered framework for optimizing driver allocation in on-demand delivery services. At each planning horizon, the model uses the available information to produce a policy to determine the optimal number of drivers to be allocated at each hour in the future.

[0048] Mathematically, we encode the information currently available for each planning level into states, which include temporal variables and demand forecasts. As the planning horizons shorten, states may also incorporate previous beliefs and decisions. A policy is defined as a mapping between states and driver allocations. A goal is to identify improved, e.g., the optimal, driver allocation policy for each planning horizon – the policy that maximizes a specific reward function.

[0049] We start with the definition and explanation of this reward function. It is crucial not only for determining the optimal policy within our model but also as a metric to evaluate the effectiveness of our approach and compare it to known current practices.

[0050] The suitability of the different driver allocation policies will be measured by a reward function that weights income and costs. The main costs associated with online to offline (O2O) dispatch are the following (a summary is detailed in Table 1): • Driver salary: drivers draw an hourly salary to compensate them for their work. • Cost of missed orders: there are on average a certain number of orders that a driver can fulfill per hour. Requests to fulfill more than this number will result in orders being missed (canceled), and they will no longer be delivered. When this happens, a cost is associated with a lack of profit and poor customer service and satisfaction. • Cost of delayed orders: there are, on average, a certain number of orders that a driver can deliver in an assignment. Because it is known to deliver groceries within 15 minutes, customers are used to timely deliveries. When an order can be delivered within the hour (and is therefore not missed) but is outside the time window expected by thecustomer, this order’s value is penalized by a delay cost. This cost includes potential complaints, deteriorated brand value, and coupons that can be generated as an apology to the customer. Although these delays also carry a cost in terms of customer satisfaction and level of service, they are deemed smaller than the cost incurred if an order has been delayed beyond the hour or canceled.

[0051] With these costs in mind, we define the hourly reward, r : R2+→ R, as a function of the hourly demand, x, and a driver allocation, a, as follows:

[0052] In this formulation, an order is fulfilled once the driver returns to the store from his or her assignment. It is assumed that a driver can fulfill up to v ≥ 1 deliveries per hour. Consequently, a fleet of a drivers can fulfill va orders per hour, and any demand greater than va is not fulfilled in that hour and is considered missing. For that reason, the hourly revenue is limited by the demand or the drivers’ capacity to deliver within the hour, whichever is less, and the number of missed orders is the excess of demand over that capacity if there is such excess. Furthermore, if a driver needs to deliver more than the average number of orders in an assignment, h, the orders between h and v will be considered delayed but delivered within that hour. A driver performs at least one assignment per hour, and thus 1 ≤ h ≤ v. Therefore, the total number of delayed orders is max{min{x − ha, a (v − h)}, 0}. In this expression, x − ha is the portion of the demand that exceeds the non-delay capacity, and a (v − h) is the maximum possible amount of delayed orders, as any orders after that are considered missed as explained before. The average income per order is p, the hourly wages per driver is ca, and the unit cost of missed and delayed orders are cmand cd, respectively. Due to the severity of missed orders versus delayed orders, we assume that cm ≥ cd. Finally, we assume that the company operates with average positive hourly margins and that the marginal income generated by a driver is higher than the wages drawn (i.e., p ^ v > ca).

[0053] Companies that operate in the delivery sector make many strategic decisions, or long-term decisions, based on the number of drivers needed per period. Notably, a known example company depends on long-term driver allocation decisions to budget for onboarding materials, training resources, and equipment. The steady-state condition of the system drives these decisions. Therefore, we study demand behavior conditioned only on static temporal and spatial-related factors and assume that this distribution remains in a steady state.

[0054] We will proceed with a state-based analysis. We define the sets S, D, and T of stores, days, and hours of the day, respectively. These sets are indexed by the indices s, d, and t, such that a state θ is given by the tuple (s, d, t). As an example, the tuple (0, 1, 23) represents the state that comprises store 0, on Monday (day 1), at 11 pm. Table 2 summarizes the notation used at this stage.

[0055] For each state θ, we assume that the demand Xθ is random, with support [0, ∞), finite mean and variance, probability density function (pdf) fθ, and cumulative distribution function (cdf) Fθ. The realized demand is indexed by x. The objective is to determine a policy π∗that maps states to actions and maximizes the expected reward. In other words,this policy will provide the optimal number of drivers (action), ^^ௌ∗ఏ that maximizes the expected reward in each state, E[rθ(X, a)].

[0056] Consider a decision maker aiming to allocate drivers for a state θ. This decision aims to maximize expected revenue and can be modeled “à la newsvendor”. Starting from (1), for a given demand realization x and an action a, we can write the revenue function in state θ = (s, d, t) as and the selected policy for driver allocation will be Proposition 1. At the strategic, or long term, planning level, the optimal driver allocation in state θ, ^^ௌ∗ఏ , is the unique solution to the equation

[0057] The proof of all propositions in this work can be found in the Appendix. We also provide a sensitivity analysis for each parameter.

[0058] Solving the optimality condition (3) requires knowledge of the probability distribution of the demand. While we can pick standard distributions, demand may not be unimodal, and complex trends may exist at the hourly demand level. For this reason, weopt for a data-driven approach based on MDNs that can characterize distributions more flexibly and accurately.

[0059] A mixture density distribution is as a convex combination of a series of K random variables pdfs. These variables are typically selected as part of the same family, e.g., Laplace or normal distributions. We consider a family of K normal distributions with pdfs gk, and cdfs Gk, each with parameters µk and σk, k = 1, . . . , K. Because the demand distribution depends on the state θ, so will the mixing distributions, i.e., pdf and cdf of the mixture distribution are thus given by

[0060] Unlike traditional neural networks, MDNs do not predict point estimates. Instead, they output the parameters of the mixture distribution that maximize the likelihood of the data occurring. To this end, a common loss function that prevents arithmetic underflow is the negative log-likelihood: . Hence, the network yields the mixture density function most likely to represent the underlying distribution of the demand observed in the dataset. The pdf g obtained from the MDN (see (4)) is defined in the support (−∞, ∞). However, the newsvendor equation (2) assumes that the demand has support in [0, ∞) instead. For this reason, we use truncated mixing distributions with pdfs fk, where

[0061] Finally, we obtain the truncated mixture distribution of the demand under state θ as f(x)=

[0062] The critical tactical, or short term, planning decision is how many drivers to contract to meet demand in the next few weeks. The higher the demand, the more drivers are needed, ceteris paribus. However, demand is uncertain and dependent on a multitude of factors. In the long term, we assumed the demand distribution is steady (although seasonal), consistent with the need for high-level decisions. However, variations in demand can occur over time, which impacts tactical decisions. Forecasting at the tactical planning level is often carried out with a more granular forecast frequency and horizon, such as weekly and daily demand. External information, such as leading indicators, is generally unavailable at this sampling frequency. Therefore, developing accurate statistical models that rely on univariate time series is crucial.

[0063] We propose an ensemble deep-learning model to predict potential deviations from strategic level demand forecasts with a two-week lead time. The model leverages various features, including historical demand with weekly lags. We incorporate cyclical covariates (day of the week and hour) to account for demand fluctuations throughout the week and day. Additionally, static categorical covariates (facility and city) capture demand variations across locations. Operating hours (as a Boolean covariate) further refine predictions by specifying active demand periods. Finally, to leverage the predictions from the previous section, we integrate the Strategic distributional predictions to provide a reference point for the deep-learning forecasts. These additional features are the means and variances of the mixture distributions. For a particular mixture distribution whose pdf , the mean and variance are computed as:

[0064] In these equations, µTand ^^்ଶrepresent the mean and variance of the mixture distribution as a function of the zero-truncated normal mixing distributions with means µ^^and variance ^^^ଶ^ . These can be directly obtained from the un-truncated normal distributions parameters.

[0065] Gradient Boosting Machine (GBM) is an ensemble method that builds a strong learner by combining weak regressors (usually decision trees). GBM minimizes the loss function by iteratively adding weak learners, each fit to the residuals of the previous ones. Gradient-boosting decision-tree methods consistently outperform alternative methods when trained on tabular data, but the model choice is not yet evident for time-series forecasting. We use the LightGBM implementation for demand forecasting due to its efficiency, scalability, and overfitting-prevention features.

[0066] Recurrent Neural Networks (RNNs) are traditionally favored for sequential data due to their effectiveness in processing time series information. However, more recent developments like the Neural Basis Expansion Analysis for Time Series (N-BEATS) and Neural Hierarchical Interpolation for Time Series Forecasting (N-HiTS) models have outperformed RNNs in terms of accuracy and efficiency. We choose the N-HiTS model because it stands out for its accuracy and computational efficiency. It employs multilayer perceptrons (MLPs) to identify nonlinear patterns by analyzing both past and future data points. It then uses a hierarchical structure of blocks that interpret the time series at varying granularity, each targeting different cyclic patterns with specific rate parameters. For our purposes, the N-HiTS model generates hourly forecasts up to two weeks in advance, but we focus only on the two-week prediction, ignoring intermediate forecasts for performance evaluation.

[0067] An ensemble learning method is a strategy used in ML to improve prediction performance by effectively combining multiple learning algorithms. A single model may not accurately approximate the true predictive function, but combining independent models can reduce generalization error by seeking the wisdom of the. In this section, we consider a regression ensemble model that merges the predictions from GBM and N-HiTS. The predictions in the training set are combined using a weighted average, and a linear regression model is fitted to determine the weights that minimize prediction error. We thenuse the weights to obtain the predictions for the test set and compute the performance metric.

[0068] Mapping predictions to actions Table 3 summarizes this model’s sets and parameters. The optimal allocation at the tactical level, ^^்∗ఏ , has a closed form, a given by the next proposition.

[0069] Thefrom this proposition suggests that optimal driver allocation depends on how the hourly salary of drivers, ca, compares to the cost of delaying orders, cdh. In high-salary scenarios (ca > cdh), it is best to allocate just as many drivers as needed to deliver the forecasted demand on average, i.e., ^^^ఏ / v. For low-salary scenarios (ca < cdh), we allocate the drivers needed to deliver the forecasted demand even if they were less efficient and only delivered h orders per hour on average (instead of v), i.e., ^^^ఏ / h. Low salaries, thus, drive amplified allocation, where the company overstaffs to diminish the risk of incurring delayed orders. When salaries and total delay costs are the same, any driver allocation in the range [^^^ఏ / v, ^^^ఏ / h] yields the same reward, (p − ca / h) ^^^ఏ. This closed-form solution is agnostic of the method used for forecasting, as it only use the point-estimatedetermine the optimal allocation. A diagram of the ensemble method is shown in FIGURE 2.

[0070] With reference to FIGURE 2, a flow 200 is illustrated. As shown, the flow 200 relates to short-term planning of the presently described embodiments, including input 210, models 220, an ensemble 230, and an output 240. Input 210 to the models 220 (including a combination of LGBM and N-HiTS forecasting models) includes demand time series, a regressiona demand prediction (2 weeks ahead) and a driver allocation policy).

[0071] We understand that drivers may not be contracted at the same rate ca at this level. We study this situation by including an additional cost ce that must be paid for every driver in excess of the strategic allocation. We show that the optimal allocation will only change in very narrow circumstances.

[0072] While the tactical planning modeldemand behavior two weeks ahead, real-time scenarios may deviate from forecasts due to unforeseen events and operational or real-time planning is we propose a Markov decision process of driver availability and demand mispredictions MDP model aims to find a policy that maximizes given horizon. This policy maps the current state as a modification of the tactical driver

[0073] In this planning level, we redefine the state θ as a tuple containing the current demand and the tactical level demand forecast for the next period. At each period, we also input the tactical driver allocation for the next period as exogenous information. The action space consists of possible modifications to the tactical driver allocation for the next hour. These modifications are bounded by pre-defined minimum and maximum values (pminand pmax) to reflect the limitations of last-minute adjustments. The reward function (1) remainsconsistent with the one used in the strategic and tactical planning levels, accounting for revenue, driver salaries, missed orders, and delayed orders.

[0074] This framing of the MDP, derived from empirical experiments, demonstrated superior optimization performance from other alternatives of state definitions and action set space. Note that with increased computational resources or larger data sets, the state definition could be further refined to capture more granular aspects of the system dynamics. A summary of the notation is in Table 4.

[0075] In an MDP, a value function stores cumulative rewards across the studied period of length N . Let us define this value function ν as where pi is chosen following policy π, and α is a discount factor. The function r gives the reward attained when allocating ^^்^∗+ pidrivers during period i and observing a realizationof the random demand Xi∈ X. In turn, the allocation of drivers is a modification of the known discretized tactical optimal allocation decision, ^^்^∗∈ A. In this modification, pi ∈ P represents the additional drivers to be hired for the following hour, in addition to the ones obtained from the tactical policy. The set of actions (operational driver allocations) is defined as the nonnegative elements of the Minkowski sum A + P and is indexed by yi= max{0, ^^்^∗+ pi } ∈ Y. The model initializes with the first demand realization x0. Then, it finds an optimal policy π∗that maps states θ = (xi, (^^^^ା^) to modifications by maximizing Equation (5).

[0076] In summary, the optimal policy uses the current knowledge of the system, i.e., the current demand level xi and the tactical level forecast for the next period ^^^^ା^. It provides the optimal modification ^^∗^ା^. to the tactical decision on driver allocation ^^்^ା∗^ .

[0077] With respect to policy gradients, traditionally, optimal policies are computed via the dynamic programming formulation. This method assumes that future state distributions depend solely on the current state and the action taken. This formulationrequires the transition kernel, q(Xi+1, ^^^^ାଶ|xi, ^^^^ା^, yi), that describes the probabilities oftransitioning between states after taking actions. are notknown in this problem because many possible and there is limited empirical information for the ones observed.

[0078] learning agent explores a policy gradientrule. The advantage of the REINFORCE rule is that it does not require modeling transitions. Instead, the transitions are sampled from the environment. The policy is represented by a neural network πϕ(Yi+1|xi, ^^^^ା^)) that maps a state, θ = (xi, ^^^^ା^) ∈ X × X, to a decision pi+1∈ P. From there, the allocation is directly obtained as the sum yi+1 = ^^்^ା∗^ .+ pi+1. In this case, ϕ represents the weights and biases of the network. The goal is to update this mapping to reduce variance maximize the sum of discounted

[0079] Let ri:= r(xi, yi) be the reward attained when observing the demand xi∈ X and making allocation yi∈ Y. Starting from a random policy πϕ, the algorithm computes theobjective function described in Equation (5). Then, it samples a sequence of events from the empirical distribution and applies actions following the policy πϕ. With these selections, the algorithm forms the trajectory τ = ((x0, ^^^^), y0, r0, (x1, ^^^ଶ), y1, r1, ...), and the transitions ((xi, ^^^^ା^), yi, αiri), i = 0, 1, . . ., are stored. Each sampled sequence is called an episode. After each episode, policy gradient ascent is performed on the objective function to improve the policy: where γ is the learning rate and J(ϕ) the total reward obtained during the episode. Note that J(ϕ) is used as the negative loss function for the network. The gradient is estimated from the episodes with the REINFORCE rule where αiri is the discounted return, and b is a baseline to reduce variance and speed up learning. In our implementation, the baseline b is the mean of the discounted returns across the episode. This update increases the probability of actions that lead to higher returns. The neural network weights are updated using backpropagation, with the discounted returns used as sample weights. Over multiple episodes of gathering experience and updating the policy, the parameters ϕ are optimized to maximize the expected cumulative rewards. The model thus learns a policy that maps states to actions while accounting for the potential ripple effects of those actions.

[0080] The results of the implementation of the presently described embodiments are attained with the multi-tiered framework. The data used for validation and comparison of results is based on a data set from a data source that is dedicated to intermediate contracting “on demand” services by electronic means, to connect customers, local businesses, and delivery drivers. Their main activity is developing and managing a technology platform where local stores in various regions can showcase their products or services. Customers can access these offerings through a mobile or web app. Additionally, the company helps arrange immediate or scheduled deliveries for these products. Its driver allocation challenges are representative of companies in the grocery or other sectors where logistics require ascertaining the availability of resources for delivery operations.

[0081] The data used from the data source is based on, at least in part, a vertical business, which offers delivery of its own products. Separated from the core intermediation activity, it provides easy online access to groceries and more products, and its service is known for its speed and convenience. Orders areusers can track them in real-time. With this business vertical, it offers customers on- demand grocery service. It delivers grocery orders from its micro-fulfillment centers (supermarkets completely dedicated to the online-to-offline (O2O) business covering a predetermined service region) through third-party contracted driver fleets. These contracts pre-determine the number of drivers it believed to be needed for every hour.

[0082] An operational problem can be a two-step process. First, thetwo weeks ahead using a gradient boosting machine (GBM). In the second step, other attributes of the couriers are predicted, such as their average delivery time. Finally, all predictions are combined to generate the ideal fleet size for each time interval. This “forecast-then-allocate” two-step approach requires the businessrouting and scheduling models that jointly optimize fleet size, fleet mix, and order assignments, which are applied to same-day deliveries and longer fulfillment horizons, may not be applicable. While this method provides a basis for tactical-level planning, our investigation suggests that this benchmark practice does not capture the complex, multi-tiered planning needed for optimal operations. Reliance on a single predictive model leaves gaps in its strategic and last-minute planning capabilities. In these areas, nuanced strategies could significantly improve efficiency and service quality.

[0083] Motivated by these improvement opportunities, we aim to develop policies that address driver allocation based on short-term demand and look at different planning stages with different available information. We use different historical data at each stage to obtain policies that map the available information to the suggested number of hourly allocated drivers to maximize a function representing the reward obtained given the driver allocation and the demand realized.

[0084] The validation of the models described herein uses data obtained from the noted source in a market between 2021 and 2022. The dataset provided by the companycomprises 27 micro-fulfillment centers (henceforth referred to as “stores”) in 11 cities. Except for two cities with eight stores each, all others contain one store from where all deliveries depart. It contains hourly-time stamped information on the demand and the delivery drivers present for each facility at each hour. Other information provided includes the working hours of each facility, which vary across weekdays. Finally, although additional information is available, including holiday markers and atmospheric conditions, these were not used for predictions, as they did not show improved performance in empirical experiments.

[0085] Although the data were recorded over two years, not all facilities have two full years of records, as some stores were open after the start date of the time series. Orders were placed more frequently in the early evening and late at night. The demand between 1 am and 8 am is considerably lower, and most stores are closed during these hours. Orders were placed roughly constantly between the late morning and late afternoon. Regarding weekly seasonality, orders were placed more often during the weekends.

[0086] The noted data source currently uses a GBM model to predict demand two weeks ahead and make driver allocation decisions based on their forecasts. This planning horizon is considered a short-time horizon under our framework, and thus, the most direct comparison will occur at the tactical level. A fundamental question is whether including strategic-level planning information in the tactical-level model improves demand forecasting, and therefore, driver allocation. If so, we can conclude that a multi-tiered framework is an improvement over a single-horizon planning model. Additionally, we verify if the tactical level demand forecasts can be fine-tuned to respond to real-time demand variations at the operational level. In the following, we discuss how we implemented the different layers of our framework in this case study.

[0087] For evaluating our models and comparing them against the company’s current benchmarks, we need to specify the parameters that will be used in the reward function. At business vertical, drivers aim for a 15-minute delivery, so we assume they can make two deliveries an hour on average, including return (i.e., v = 2). If a driver needs to deliver two orders in an hour, the second one will be considered late (i.e., h = 1) and incur a delay cost;if a driver needs more than two orders, then any orders over two will be missed. In addition, we set the different costs as follows:ca = $15 / (driver^hour), cd = $5 / order, cm = $20 / order, and p = $100 / order.

[0088] We note that these figures do not reflect the real operations but are representative of the true figures. They can be easily adapted for different locations and service level requirements.

[0089] To attain strategic planning-level results, we separate our dataset according to all possible states. With S = 27 stores, D = 7 days per week, and T = 24 hours per day, there are 3,134 states among the more than 220,000 records in the dataset (if we consider only the hours in which the stores are open). Each record is characterized by the tuple θ and the number of orders received.

[0090] The fundamental aspect for deriving the optimal driver allocation in the long- term, as given by Proposition 1 is determining the demand distribution with an MDN. In our case, this network con- sists of three inputs that correspond to the elements of the tuple θ = (s, d, t), two hidden layers with 16 nodes each, and an output layer consisting of 30 outputs (K = 10). The hidden layers used rectified linear units (ReLus) as activation functions, whereas the three elements of each mixing distribution (πk, µk, σk) were trained with softmax, exponential linear units (ELUs), and identity functions, respectively. The implementation took place in Python using Keras and a special implementation of the MDN layer. A representation of the architecture used in our MDN is shown in FIGURE 3.

[0091] The MDN is trained with an 80 / 20 split for a maximum of 1000 epochs using the Adam optimizer. The neural network yields a pdf of the steady-state demand for each state. For each such state, and considering the choice of cost parameters given, we solve the condition given by (3) to compute the optimal number of drivers needed in each case. FIGURE 4 shows the optimal policies attained for two different types of stores: one that is open 24 / 7 (FIGURE 4(a)) and another that opens during regular business hours from Monday to Thursday and 24 hours on Friday, Saturday, and Sunday (FIGURE 4(b)). As expected, both allocate more drivers in those hours with higher demand (typically during the evenings). The allocation at a given hour varies depending on the day of the week, but the number of drivers required in the evenings remains larger as the week advances. These trends persist across all stores, as shown in FIGURE 4(c) and FIGURE 4(d).

[0092] Although backcast periods typically cover the seasonalities of the time series, our dataset only spans two years, which is insufficient to capture annual (includingholidays) and monthly seasonalities. Therefore, we look at the demand from exactly two and three weeks prior for the two-week ahead predictions. We split the dataset containing all the facilities and cities into an 80% train and 20% test dataset and obtain forecast predictions on the test dataset for each facility time series.

[0093] We tune the LightGBM parameters to limit the maximum depth of a tree to 6tree to 5. We also verify empirically that using a learning rate of 0.01 and 5000 iterations leads to good predictions when L1 and L2 regularization is applied. We test including the parameters of the predicted demand distributions from the strategic planning level. The performance is compared to the GBM results withplanning covariates to verify whether the integration across planning level helps improve the model. FIGURE 5 depicts the importance (in proportions of tree splits) of each feature considered for the demand predictions in the GBM model. The mean of the mixture distribution is a key driver of the GBM prediction, with an importance higher than the most recent demand. This contribution highlights the need for a multi-tiered approach, in which results in one stage build upon the results of previous stages.

[0094] For the N-HiTS model, we set the number of stacks as three, blocks as one, and MLPs asAn increase in these parameters significantly increases computational load with minimal accuracy gains. We use the ReLU activation function, the most commonly used activation function for deep neural networks. We add L1 regularization to the training loss function (i.e., Mean Absolute Error Loss) and a random dropout rate of 0.2 to prevent the model from over-fitting. We trained the model with 50 epochs and a stopping rule that stops the iteration with a loss function tolerance limit of 0.01 over five epochs. Our model training stopped at 11 epochs.

[0095] Consistent with the current policy of the data source, we use the daily- aggregated mean absolute percentage error (MAPE) metric. This aggregated metric is more robust than an hourly MAPE because it avoids the division by zero that would occur when no demand happened in an hour. Additionally, this aggregation removes the influence that predictions from the times the stores are closed may have on the performance metric. We work with MAPE on a store-by-store basis, so it is necessary to average this metric over all days in thewhere xswt (^^^swt) represents the actual (forecasted) orders recorded in-store s, day w, and hour t. In Table 5, we compare the values of MAPEs obtained using different methods, with and without incorporating results from the previous tier. The following observations are made based on the obtained results:   Without Planning Belief  With Planning Belief City Na¨ıve  Auto‐ARIMA  Holt‐Winters  N‐HiTS  GBM N‐HiTS  GBM  EnsembleA  43.3%  95.8%  41.9%  35.1% 31.9%  37.4%  31.9%B  34.1%  83.1%  34.6% 28.9% 26.2%  27.6%  25.8% C  54.8%  97.2%  49.4%  40.5%  45.5% 39.9%  41.9%  39.9%D  43.8%  80.2%  42.1%  35.2%  39.5% 34.2%  35.7%  34.3%E  36.1%  81.0%  43.3%  31.9%  31.2% 27.7%  30.0%  27.4%F  39.3%  97.2%  35.4%  32.7%  32.4% 29.7%  31.8%  29.0% G  34.0%  90.0%  34.4%  25.8%  28.7% 27.1%  27.7%  26.1%H  39.4%  97.6%  40.8%  32.6%  34.5% 29.9%  32.4%  30.0%I  50.0%  97.1%  39.9%  35.2%  41.2% 36.3%  38.4%  36.1%J  27.6%  101.0%  30.4%  24.3%  23.3% 21.3%  23.1%  20.4% K  40.4%  93.0%  44.9%  32.9%  34.6% 31.3%  32.2%  31.0%Average  40.3%  92.1%  39.7%  31.9%  34.1% 30.5%  32.6%  30.2%†: benchmark model.  Table 5 Models Forecast Overall MAPEahead demand prediction) in the test set – lower is better.  1. The N-HiTS model exhibits better performance than the GBM model in both implementations (with and without the strategic planning beliefs). 2. The strategic planning belief enhances the forecast accuracy overall. 3. The ensemble model demonstrates further improvements over the individual models. Compared with the GBM benchmark model, our proposed N-HiTS / GBM ensemble model with strategic planning belief improves forecast accuracy across all cities in the data.

[0096] This implementation of the policy gradient algorithm aims to find the optimal change in driver allocation from the tactical level decision for the next hour. We implemented the policy gradient algorithm with the following considerations:We limit these changes to situations where demand is lower than 50. This restriction helps us focus on more frequently occurring events and avoids issues with infrequent outliers affecting how our model learns. We also limit the allowable changes by the bounds pmin= −2 and pmax= 2, because we assume that the decision to add or remove drivers is restricted at this stage. These bounds were chosen empirically based on the data and computational power available, as discussed in EC.3. Each state is formed by the type (xt−1, ^^^t), containing the demand of the previous period, as well as the predicted demand from the tactical level forecasting model.

[0097] We train our model using the REINFORCE algorithm with a learning rate of γ = 0.01. For computational efficiency and to focus on the immediate impact of decisions, we limit each training episode to 12 periods. A discount factor of α = 0.95 is used to further emphasize the importance of short-term rewards. Our training process starts with a random state (xt−1, ^^^t) and chooses an action pt based on a random policy. This means the policy suggests that the drivers allocated for the next period should be ^^்௧∗+ pi, where௧∗is the tactical level driver allocation.

[0098] We then sample a reward from the dataset matching the state-action pair and sample a new state from the dataset, also based on the state-action pair. This process continues for 12 periods, during which we store the observed states, actions, and rewards. Finally, we use the collected data to update the neural network representing the decision policy, which consists of two fully connected layers with 20 nodes each. We repeat this process for 60,000 episodes to train the model.

[0099] The gains may be obtained using our framework in their driver allocation policies. We compare the results attained by each of the three tiers and the benchmark given by the data source currentwe look at how the driver allocation policies at all levels of ourto those generated by the company’s historical driver allocation. In addition, we calculate the rewards that stem from those allocation options and contrast them.

[0100] FIGURE 6(a) shows the evolution of total driver allocations across all stores. The is for allThe figure suggests that the strategic policy recommends a similar number of drivers overall than the benchmark. However, it allocates them differently (less at the beginning, more at the end). This is confirmed in Table 6. On the other hand, FIGURE 6(b) shows that the rewards obtained with the strategic policy are only slightly worse than those obtained with the company’s current policy, despite current policy using much more recent information than we assumed was available at the strategic level. Tier  Drivers  % Change  Rewards  % Change  (Benchmark)  0.881  ‐  $112.73  ‐ Strategic Planning  0.877  ‐0.34  $110.00  ‐2.42 Tactical Planning  0.850  ‐3.48  $126.49  11.81 Operational Planning  1.021  15.96  $135.70  20.37 Table 6 Overall driver-hours and rewards over approximately 1.5 years (expressed in millions).

[0101] The tactical policy allocates more drivers than the strategic policy in the first months of the analysis when the number of stores is constant. It differs as well from the allocations produced by current policy during this time, and, notably, its rewards outperform the benchmark’s virtually every week, with considerable gains in the first few months. Finally, the operational policy further increases the allocation produced by the tactical policy, especially in the second half of the time span. FIGURE 6(b) shows how the adaptability of the operational model to last-minute changes in day-to-day operations results in consistently increased rewards compared to the tactical policy. We note that the tactical policy focused on saving cost by being more conservative on the number of drivers allocated, while the operational policy leveraged the more recent information to allocate more drivers where revenue could be increased.

[0102] Table 6 presents a quantitative overview and comparison of these results, including the overall number of driver-hours allocated and the total reward accrued over the study duration of approximately 1.5 years. As mentioned above, the strategic performance does not surpass existing policy due to the assumption of limited information at this level. However, despite this limitation, the benchmark outperforms the strategic planning level in terms of rewards only by 2.4%, with roughly the same overall number of drivers. However, it is at the tactical planning level, where the information knowledge assumption aligns with current practices, that we observe a substantial uptick in the rewardsattained with a decrease in drivers (an 11.8% increase in rewards with a 3.5% decrease in driver-hours). This improvement can be attributed to reduced driver idleness due to improved demand forecasts. In turn, the operational planning level model yields even more significant gains (+20.4%) by employing additional drivers in critical periods (+16%).

[0103] Mathematical Proofs of Propositions as provided as follows:

[0104] Proof of Proposition 1. From the reward function given in (1), we can derive its expected value:where (7) is greater than 0 because cm≥ cdand because we assume that the company operates with positive average margins (v ^ p − ca> 0). It follows then that the optimal policy satisfies the first-order condition, is unique, and maximizes the expected reward.

[0105] Proof of Proposition 2. Let us particularlize Equation (2) for x = ^^^θ: where we took into account that v ≥ h when solving for the max and mins from Equation (2). The resulting function is linear piecewise. The slope of the first piece is positive, as pv ≥ ca and cm ≥ cd. The slope of the third piece is negative. The maximum of rθ(^^) will be attained either on the vertex joining the first and the second piece (at x = ^^^θ / v) or on the vertex joining the second and the third piece (at x = ^^^θ / h), depending on the slope of the second piece. This slope is positive if ca < cdh, in which case ^^∗ఏ = ^^^θ / h and negative if ca > cdh, in which case ^^∗ఏ = ^^^θ / v. The second piece plateaus if ca= cdh, in this case, any allocation between these two vertices is optimal. The rewards derived from these optimal allocations follow from direct substitution of ^^ = ^^∗ఏ in rθ(^^). EC.1 Sensitivity Analysis at the Strategic Level

[0106] Next, we develop and comment on some theoretical results that pertain to the sensitivity of the optimal driver allocation in the strategic planning model. In Propositions EC.1 and EC.2, we analyze the changes in optimal allocations when varying the parameters of the newsvendor model and show the results of numerical experiments carried with the data. Proposition EC.1. In the strategic planning model, the optimal allocation of drivers in a state θ, ^^ௌ∗ఏ ,

[0107] Proposition EC.1 provides analytical support to the changes we would expect from the optimal allocation of drivers when changing one parameter of the problem, ceteris paribus. Its proof can be found in EC.4.1. We would expect to increase the number of drivers if missing an order becomes too expensive; likewise, we would expect to reduce hiring if the drivers’ salaries increase or they become more efficient (either because they can deliver more orders in an hour or they can fulfill more orders without delay within the hour). Now, increasing the cost of delaying an order within the hour cd induces an increase in the optimal allocation of drivers only if the ratio between the probability that the demand exceeds the hourly capacity and the probability that the demand exceeds the hourly timely capacity is smaller than the percentage of orders that can be delivered without delay in an hour, h / v. This ratio, bounded above by 1 is a priori unpredictable when h ≠ v, as it depends on the demand distribution. When h = v the allocation is trivially unaffected by changes in the delay cost, as such cost is never triggered and only the cost of missing orders, cm, is relevant in that case. For other cases, the next proposition sheds more light on the behavior of aθwhen cdincreases.

[0108] Proposition EC.2. Assume that h ≠ v. If the distribution of the demand Xθ has the increasing hazard rate property, then there is an optimal driver allocation ^^^∗ఏ above (below) which any increase in the order delay cost induces and increase (decrease) in the optimal allocation.

[0109] Please see EC.4.2 for the proof. The condition above is not necessary, but sufficient, to guarantee that the threshold ^^^∗ఏ is unique. When that is the case, this threshold serves as an equilibrium point for allocations (see FIGURE 7(a)). Deviations from that equilibrium point induce more extreme driver allocations when the order delay cost increases. The increasing hazard rate property guarantees that the ratio is always decreasing and then crosses h / v just once. When this property does not exist, this ratio may increase and decrease and produce multiple thresholdvalues. FIGURE 7(b) shows the case with three threshold points. The first and third act as unstable equilibria, whereas the second acts as an attractor (stable equilibrium point).

[0110] On a more practical level, decision-makers will likely use the results in an aggregated format. FIGURE 8(a) and FIGURE 8(b) showcase the average number of driver-hours over all stores that should be allocated daily to maximize the expected reward in the steady state over different values of missing costs and driver costs. We emphasize that the figures depict the aggregate solutions given by the newsvendor problem, which are yielded with hourly granularity. FIGURE 8(a) plots the optimal solutions after varying the costs ca, and FIGURE 8(b) plots the optimal solutions after varying the costs cm. The number of optimal drivers drops significantly when cagoes from $1 to $5. FIGURE 8(b) suggests that the optimal solution is not as sensitive to changes in cm as in ca. However, this may not always be the case as discussed above. Similar trends when varying h and v are discussed in this same proposition. EC.2. Adding an Excess Cost for Additional Drivers at the Tactical Level

[0111] In this appendix, we incorporate an additional cost in the tactical driver allocation decision and examine the sensitivity of the optimal allocation decision to this cost. Specifically, we introduce the excess driver cost parameter, ce, that increases the total allocation cost when the tactical decision leads to excess drivers beyond the strategic optimal allocation. This cost accounts for any short-term increases in planned budgets for driver onboarding materials, training resources, and equipment. We assume that such cost does not offset the marginal profit that would be obtained for an extra driver (i.e., we assume that pv − ca− ce≥ 0). We can redefine the reward function as: where ^^ௌ∗ఏ denotes the strategic optimal allocation decision. The optimal allocation at the tactical level, ^^்∗ఏ , is characterized by the proposition below.

[0112] Proposition EC.3. Consider the modified reward function (EC.1). At the tactical level, the optimal allocation of drivers is:

[0113] The proof for proposition EC.3 can be found in EC.4.3. The interpretation of the assignments above can be made in view of the assignments made at the strategic level and the marginal costs generated by hiring extra drivers. For example, when the strategic assignment is high (^^ௌ∗ఏ > ^^^θ / h), we could call this low-efficiency allocations), the inclusion of an excess driver cost does not change the assignment at the tactical level with respect to that made when this cost is not taken into account (see Proposition 2). In mid- efficiency allocations at the strategic level (^^^θ / v ≤ ^^ௌ∗ఏ ≤ ^^^θ / h), the tactical allocation matches the strategic allocation when the cost of delaying h orders is between the cost of hiring a contracted driver, ca, and the cost of hiring a driver at a later stage ca+ ceOtherwise, the allocation is again the same with respect to that made when this ce is not considered. Finally, in high-efficiency allocations at the strategic level (^^ௌ∗ఏ < ^^^θ / v), the change in the value of the optimal allocation from ^^^θ / v to ^^^θ / h takes place at the new hiring cost, ca + ce, instead of at the old hiring cost, ca, as in Proposition 2.

[0114] Corollary EC.1. When comparing the reward models with and without the excess driver cost parameter, the optimal allocation at the tactical level only changes when we make high and mid- efficiency allocations at the strategic level and the cost of delaying orders cdh is within the change in the hiring costs (i.e., in the range [ca, ca + ce]).

[0115] The addition of ce, thus, will not even produce changes in ^^்ఏ∗if the allocation at the strategic level ^^ௌ∗ఏ is already high. No matter the extra cost of hiring a new driver(whether it is very expensive or free), the company does not find it worthy to change the original allocation at the tactical level. In other cases, these changes, if they appear, will occur in a reduced range of situations (cdh ∈ [ca, ca+ ce]) and will always be bounded below (above) by ^^^θ / v (^^^θ / h). Therefore, their magnitude is limited by how h and v compare. EC.3. Operational Level Action Set Limitations

[0116] Reinforcement learning models are known for their large computational and data imposed on the allowableavailability and computational resources when determining parameters for their MDP framework.

[0117] In the paper, we allow that the tactical policy may be changed by at most 2 drivers. This value was chosen after testing different boundaries so as to determine the amount of action possibilities the data would support. To further illustrate, we vary the boundaries pmaxto 2, 5, 7, and 10 drivers (and pminits negative counterparts). We run the models for 40,000 episodes of length 12 hours. The rewards are displayed in FIGURE 9.

[0118] The results show that allowing for a wider range of actions did not lead to an increase in rewards as expected. This can be explained by two factors. First, the larger the action set, the longer the required training time to explore the entire range of possibilities. It suggests that 40,000 episodes may have been insufficient for the model to explore. for the pairs the andsets.EC. 4 Proofs of Propositions EC.4.1. Proof of Proposition EC.1

[0121] Because v ≥ h, both ratios are at most 1. The ratio on the left-hand side, however, varies in size depending on the distribution of the demand and therefore a general assessment on whether the optimal allocation will change with the cost of delayed orders cannot be made beforehand.EC.4.2. Proof of Proposition EC.2.EC.4.3. Proof of Proposition EC.3

[0122] With reference now to FIGURE 10, the above-described method and other methods according to the presently described embodiments, as well as suitable architecture such as system components useful to implement a suitable system in connection with embodiments described herein, can be implemented on a computer using well-known computer processors, memory units, storage devices, computer software, and other components. A high-level block diagram of such a computer is illustrated in FIGURE 10. Computer 300 contains at least one processor 350, which controls the overall operation of the computer 300 by executing computer program instructions which define such operation. The computer program instructions may be stored in at least one storage device or memory 380 (e.g., a magnetic disk or any other suitable non-transitory computer readable medium or memory device) and loaded into another memory 380 (e.g., a magneticdisk or any other suitable non-transitory computer readable medium or memory device), or another segment of memory 370, when execution of the computer program instructions is desired. Thus, the methods described herein (such as the method of FIGURE 1) may be defined by the computer program instructions stored in the memory 380 and controlled by the processor 350 executing the computer program instructions. The computer 300 may include unicating with other d ce that enables user in O devices (e.g., keyboa the computer. Such in et of computer program. interface also includes a display for displaying images and maps or other useful tools to the user.

[0123] According to various embodiments, FIGURE 10 is a high-level representation of possible components of a computer for illustrative purposes and the computer may contain other components. Also, the computer 300 is illustrated as a single device or system. However, the computer 300 may be implemented as more than one device or system and, in some forms, may be a distributed system with components or functions suitably distributed in, for example, a network or in various locations.

[0124] The presently described embodiments, in at least one form, implement, a comprehensive framework for the optimal driver allocation in the omnichannel grocery market, integrating strategic (or long-term), tactical (or short-term), and operational (or real-time) planning levels. This effort is motivated by the growing challenges posed by omnichannel practices in the grocery industry, including the need to adapt to more demanding customers and an increasing volume of orders. The result is an environment that requires accurate demand forecasts and dynamic driver allocations.

[0125] The application of this framework to a business comprises, for example, a combination of three distinct planning horizons that enables the company to pinpoint improvement areas, strategically allocate resources with ample lead time, and ultimately optimize its financial resources to achieve maximum efficiency. Likewise, this framework considers the unique characteristics of omnichannel operations, such as the expectation of quick delivery, their potentially multi-modal demand patterns, or last-minute challenges.This is done with an intuitive reward function that business stakeholders can easily understand and helps the company illustrate potential gains across diverse scenarios, accommodating varying cost parameters. While the concept of planning horizons may vary across different companies, this approach can handle any planning horizon as long as there are sufficient data to support the analysis. The results confirm that the gradual incorporation of information about past demand and upper-level driver allocations in a multi-tiered context produces a consistent pattern of improvement over single-horizon alternatives and over to the company’s current policies. Beyond enhancing the existing tactical level decision making, the presently described embodiments provide a comprehensive framework benefiting both business and technical operations. In strategic planning, our strategic forecasts and policy offered valuable quantitative insights for discussions on earnings, return on investment, and budget allocation. Moreover, the tactical level demand predictions emerged as an invaluable tool, facilitating the rapid assimilation and utilization of recent historical information. Finally, the operational level model helped us maximize the use of real-time data for immediate operational benefits.

[0126] The above description illustrates the efficacy of the framework in optimizing driver allocation for the omnichannel grocery market. Our approach, grounded in both theoretical rigor and real- world applicability, demonstrates a scalable solution adaptable to various omnichannel challenges. Ultimately, this approach is poised to significantly improve operational efficiency, driver allocation processes, and customer satisfaction, marking a notable contribution to the field.

[0127] It will be appreciated that variants of the above-disclosed and other features and functions, or alternatives thereof, may be combined into many other different systems or applications. Various presently unforeseen or unanticipated alternatives, modifications, variations or improvements therein may be subsequently made by those skilled in the art which are also intended to be encompassed by the following claims.where we used thata, a} fg (x)dx = f (x - ha)fg (x)dx + f a(v - h) f (x)dx. By means of Leibniz’s Rule, differentiating with respect to a yields E'[r0(X, tz)] = vFg(va)(p + cm- cd) + cdhFg(ha) - ca, whence we obtain the first-order optimality conditionThe equation above has only one solution, which is a maximizer of the expected revenue, as E[re(X, u)]is a concave function in a0.The derivatives of the expected reward at the limits of the support [0,oo) are, respectively,E'[re(X, 0)] = v(p + cm) - cd(v - h) ~ ca= (vp - ca) + (cmv - cd(v - h.)) > 0.G) lim a- cowhere (7) is greater than 0 because cm> Cd and because we assume that the company operates with positive average margins (v • p - ca> 0). It follows then that the optimal policy satisfies the first-order condition, is unique, and maximizes the expected reward.

[0105] Proof of Proposition 2. Let us particularlize Equation (2) for x = xewhere we took into account that v>h when solving for the max and mins from Equation (2). The resulting function is linear piecewise. The slope of the first piece is positive, as pv > caand cm> Cd. The slope of the third piece is negative. The maximum of ro(a) will be attained either on the vertex joining the first and the second piece (at x = xe / v) or on the vertex joining the second and the third piece (at x = xelh), depending on the slope of the second piece. This slope is positive if ca< Cdh, in which case a g = xelh and negative if ca> Cdh, in which case a g = x / v. The second piece plateaus if ca= Cdh, in this case, anyallocation between these two vertices is optimal. The rewards derived from these optimal allocations follow from direct substitution of a = aBin r&(a).EC.l Sensitivity Analysis at the Strategic Level

[0106] Next, we develop and comment on some theoretical results that pertain to the sensitivity of the optimal driver allocation in the strategic planning model. In Propositions EC.l and EC.2, we analyze the changes in optimal allocations when varying the parameters of the newsvendor model and show the results of numerical experiments carried with the data.Proposition EC.l. In the strategic planning model, the optimal allocation of drivers in a state 6, asf ,• strictly increases with the cost of missed orders, cm;• strictly decreases with the salary cost of drivers, ca, the average number of orders fulfilled per hour, v, and the number of orders that can be fulfilled on time within the hour, h;• strictly increases (decreases) with the cost of delayed orders, Cd, only if Fevaf) / Fehae* ) < h / v (> h / v) and remains constant when Fe(ya^ / FB(ha ') = h / v.

[0107] Proposition EC.l provides analytical support to the changes we would expect from the optimal allocation of drivers when changing one parameter of the problem, ceteris paribus. Its proof can be found in EC.4.1. We would expect to increase the number of drivers if missing an order becomes too expensive; likewise, we would expect to reduce hiring if the drivers’ salaries increase or they become more efficient (either because they can deliver more orders in an hour or they can fulfill more orders without delay within the hour). Now, increasing the cost of delaying an order withinthe hour Cd induces an increase in the optimal allocation of drivers only if the ratio between the probability that the demand exceeds the hourly capacity and the probability that the demand exceeds the hourly timely capacity is smaller than the percentage of orders that can be delivered without delay in an hour, h / v. This ratio, bounded above by 1 is a priori unpredictable when h v, as it depends on the demand distribution. When h = v the allocation is trivially unaffected by changes in the delay cost, as such cost is never triggered and only the cost of missing orders, cm, is relevant in that case. For other cases, the next proposition sheds more light on the behavior of ae when Cd increases.

[0108] Proposition EC.2. Assume that hv. If the distribution of the demand Xe has the increasing hazard rate property, then there is an optimal driver allocation ae* above (below) which any increase in the order delay cost induces and increase (decrease) in the optimal allocation.

[0109] Please see EC.4.2 for the proof. The condition above is not necessary, but sufficient, to guarantee that the threshold aeis unique. When that is the case, this threshold serves as an equilibrium point for allocations (see FIGURE 7(a)). Deviations from that equilibrium point induce more extreme driver allocations when the order delay cost increases. The increasing hazard rate property guarantees that the ratio F0(vnJ) / Fe( / iciJ) is always decreasing and then crosses h / v just once. When this property does not exist, this ratio may increase and decrease and produce multiple threshold values. FIGURE 7(b) shows the case with three threshold points. The first and third act as unstable equilibria, whereas the second acts as an attractor (stable equilibrium point).

[0110] On a more practical level, decision-makers will likely use the results in an aggregated format. FIGURE 8(a) and FIGURE 8(b) showcase the average number of driver-hours over all stores that should be allocated dailyto maximize the expected reward in the steady state over different values of missing costs and driver costs. We emphasize that the figures depict the aggregate solutions given by the newsvendor problem, which are yielded with hourly granularity. FIGURE 8(a) plots the optimal solutions after varying the costs ca, and FIGURE 8(b) plots the optimal solutions after varying the costs cm. The number of optimal drivers drops significantly when cagoes from $1 to $5. FIGURE 8(b) suggests that the optimal solution is not as sensitive to changes in cmas in ca. However, this may not always be the case as discussed above. Similar trends when varying h and v are discussed in this same proposition.EC.2. Adding an Excess Cost for Additional Drivers at the Tactical Level

[0111] In this appendix, we incorporate an additional cost in the tactical driver allocation decision and examine the sensitivity of the optimal allocation decision to this cost. Specifically, we introduce the excess driver cost parameter, ce, that increases the total allocation cost when the tactical decision leads to excess drivers beyond the strategic optimal allocation. This cost accounts for any short-term increases in planned budgets for driver onboarding materials, training resources, and equipment. We assume that such cost does not offset the marginal profit that would be obtained for an extra driver (i.e., we assume that pv - ca- ce> 0). We can redefine the reward function as: r0(%0, a) = p min{xe, va] - caa - cmmax{%e- va, 0} - cdmax{min{ xdwhere ase* denotes the strategic optimal allocation decision. The optimal allocation at the tactical level, aTeis characterized by the proposition below.

[0112] Proposition EC.3. Consider the modified reward function (EC.1). At the tactical level, the optimal allocation of drivers is:1. If a g* < xe / v:In addition, when cdh = ca+ ce, all allocations in the segment [xg / v, Xg / h\ are optimal.In addition, when cdh - ca(cdh = ca+ ce), all allocations in the segment [XQ / V, (ZQ*] ([a *, xe / h]) are optimal3. If a * > xe / h:In addition, when cdh = ca, all allocations in the segment [XQ / V, xe / h] are optimal.

[0113] The proof for proposition EC.3 can be found in EC.4.3. The interpretation of the assignments above can be made in view of the assignments made at the strategic level and the marginal costs generated by hiring extra drivers. For example, when the strategic assignment is high (a * > xolh we could call this low-efficiency allocations), the inclusion of an excess driver cost does not change the assignment at the tactical level withrespect to that made when this cost is not taken into account (see Proposition 2). In mid-efficiency allocations at the strategic level (xelv < a < xolh the tactical allocation matches the strategic allocation when the cost of delaying h orders is between the cost of hiring a contracted driver, ca, and the cost of hiring a driver at a later stage ca+ ceOtherwise, the allocation is again the same with respect to that made when this ceis not considered. Finally, in high- efficiency allocations at the strategic level (ase* < xolv the change in the value of the optimal allocation from xe / v to !h takes place at the new hiring cost, ca+ ce, instead of at the old hiring cost, ca, as in Proposition 2.

[0114] Corollary EC.l. When comparing the reward models with and without the excess driver cost parameter, the optimal allocation at the tactical level only changes when we make high and mid- efficiency allocations at the strategic level and the cost of delaying orders Cdh is within the change in the hiring costs (i.e., in the range [ca, ca+ ce]).

[0115] The addition of ce, thus, will not even produce changes in aeif the allocation at the strategic level ase* is already high. No matter the extra cost of hiring a new driver (whether it is very expensive or free), the company does not find it worthy to change the original allocation at the tactical level. In other cases, these changes, if they appear, will occur in a reduced range of situations (cdh E [ca, ca+ Ce]) and will always be bounded below (above) by xe / v xolh Therefore, their magnitude is limited by how h and v compare.EC.3. Operational Level Action Set Limitations

[0116] Reinforcement learning models are known for their large computational and data requirements. We implemented the model varying the constraints imposed on the allowable change. It is important for companies toverify their data availability and computational resources when determining parameters for their MDP framework.

[0117] In the paper, we allow that the tactical policy may be changed by at most 2 drivers. This value was chosen after testing different boundaries so as to determine the amount of action possibilities the data would support. To further illustrate, we vary the boundariesto 2, 5, 7, and 10 drivers (and pminits negative counterparts). We run the models for 40,000 episodes of length 12 hours. The rewards are displayed in FIGURE 9 .

[0118] The results show that allowing for a wider range of actions did not lead to an increase in rewards as expected. This can be explained by two factors. First, the larger the action set, the longer the required training time to explore the entire range of possibilities. It suggests that 40,000 episodes may have been insufficient for the model to explore. Second, the model needs to understand the variation of rewards that can be obtained for being in different states and choosing different actions.

[0119] By making the action space larger in the same dataset, we thin out the possibilities of rewards that can be sampled in each episode. This suggests that the dataset used is not large enough for the model to observe accurate rewards and learn for all pairs of states and actions.

[0120] In sum, the decrease in the reward observed as we increase the range of action possibilities is likely due to a combination of limited computational power and insufficient data availability to model more complex system dynamics than the one described in the paper. Consequently, decision-makers must carefully balance their data availability and computation resources to determine how large they should make their state and action sets.EC. 4 Proofs of PropositionsEC.4.1. Proof of Proposition EC.lThe proof follows form the Implicit Function Theorem by definingrefers to a parameter from the list (ca, cd, cm, v, h). The optimal policy ctg* changes according to any of these parameters as:(EC.2)Recalling that cm> cd, it is straightforward to check that,Given the signs of these partial derivatives, the sign of (EC.2) in each case follows. For the last partial derivative, we have that dQ(a.g*, cd) / dcd> 0 whenwhich only holds if

[0121] Because v > h, both ratios are at most 1. The ratio on the left-hand side, however, varies in size depending on the distribution of the demand and therefore a general assessment on whether the optimal allocation will change with the cost of delayed orders cannot be made beforehand.EC.4.2. Proof of Proposition EC.2.Since the left-hand side of Equation (EC.3) equals 1 when a * = 0 and lim Fe( g* Fe( iase*') = 0, this ratio will cross h / v only once if it is ag'^co strictly decreasing. This occurs only if the derivative of the ratio is negative. As the denominator of the derivative is strictly positive, it suffices that r, equivalently, ifthe hazard rate of the distribution of Xg.Since h / v < 1 by definition, this condition holds automatically when A(vci0*) / A( a0*) > 1, or equivalently, ( a@* > (ha^), which is guaranteed if Z(-) is increasing.EC.4.3. Proof of Proposition EC.3Consider the reward function (EC.l). Let us make a case-by-case analysis based on the value of the optimal strategic allocation. When CLQS< X / V the reward is the following piece- wise linear function:The first and second pieces have positive slopes since pv > ca+ ceand cm> cd. The fourth piece has a negative slop. The third piece has a positive slope if cdh > ca+ ceif cdh < ca+ ce, its slope is negative. In the former case, the maximum is attained at the end of the third interval, where a = x0 / h in the latter it is attached at the end of the second interval, where a = x0lv. If cdh = ca+ ce, the third piece plateaus and the optimal assignment can be found in the entire interval [xs / v, x0 / h] .A similar analysis when x0 / v <> x0 / v completes the proof.

[0122] With reference now to FIGURE 10, the above-described method and other methods according to the presently described embodiments, as well as suitable architecture such as system components useful to implement a suitable system in connection with embodiments described herein, can be implemented on a computer using well-known computer processors, memory units, storage devices, computer software, and other components. A high-level block diagram of such a computer is illustrated in FIGURE 10. Computer 300 contains at least one processor 350, which controls the overall operation of the computer 300 by executing computer program instructions which define such operation. The computer program instructions may be stored in at least one storage device or memory 380 (e.g., a magnetic disk or any other suitable non-transitory computer readable medium or memory device) and loaded intoanother memory 380 (e.g., a magnetic disk or any other suitable non-transitory computer readable medium or memory device), or another segment of memory 370, when execution of the computer program instructions is desired. Thus, the methods described herein (such as the method of FIGURE 1) may be defined by the computer program instructions stored in the memory 380 and controlled by the processor 350 executing the computer program instructions. The computer 300 may include one or more input elements 310 and output elements 320 for communicating with other devices via a network. The computer 300 also includes a user interface that enables user interaction with the computer 300. The user interface may include I / O devices (e.g., keyboard, mouse, speakers, buttons, etc.) to allow the user to interact with the computer. Such input / output devices or elements may be used in conjunction with a set of computer programs in accordance with embodiments described herein. The user interface also includes a display for displaying images and maps or other useful tools to the user.

[0123] According to various embodiments, FIGURE 10 is a high-level representation of possible components of a computer for illustrative purposes and the computer may contain other components. Also, the computer 300 is illustrated as a single device or system. However, the computer 300 may be implemented as more than one device or system and, in some forms, may be a distributed system with components or functions suitably distributed in, for example, a network or in various locations.

[0124] The presently described embodiments, in at least one form, implement, a comprehensive framework for the optimal driver allocation in the omnichannel grocery market, integrating strategic (or long-term), tactical (or short-term), and operational (or real-time) planning levels. This effort is motivated by the growing challenges posed by omnichannel practices in thegrocery industry, including the need to adapt to more demanding customers and an increasing volume of orders. The result is an environment that requires accurate demand forecasts and dynamic driver allocations.

[0125] The application of this framework to a business comprises, for example, a combination of three distinct planning horizons that enables the company to pinpoint improvement areas, strategically allocate resources with ample lead time, and ultimately optimize its financial resources to achieve maximum efficiency. Likewise, this framework considers the unique characteristics of omnichannel operations, such as the expectation of quick delivery, their potentially multi-modal demand patterns, or last-minute challenges. This is done with an intuitive reward function that business stakeholders can easily understand and helps the company illustrate potential gains across diverse scenarios, accommodating varying cost parameters. While the concept of planning horizons may vary across different companies, this approach can handle any planning horizon as long as there are sufficient data to support the analysis. The results confirm that the gradual incorporation of information about past demand and upper-level driver allocations in a multi-tiered context produces a consistent pattern of improvement over singlehorizon alternatives and over to the company’s current policies. Beyond enhancing the existing tactical level decision making, the presently described embodiments provide a comprehensive framework benefiting both business and technical operations. In strategic planning, our strategic forecasts and policy offered valuable quantitative insights for discussions on earnings, return on investment, and budget allocation. Moreover, the tactical level demand predictions emerged as an invaluable tool, facilitating the rapid assimilation and utilization of recent historical information. Finally, theoperational level model helped us maximize the use of real-time data for immediate operational benefits.

[0126] The above description illustrates the efficacy of the framework in optimizing driver allocation for the omnichannel grocery market. Our approach, grounded in both theoretical rigor and real- world applicability, demonstrates a scalable solution adaptable to various omnichannel challenges. Ultimately, this approach is poised to significantly improve operational efficiency, driver allocation processes, and customer satisfaction, marking a notable contribution to the field.

[0127] It will be appreciated that variants of the above-disclosed and other features and functions, or alternatives thereof, may be combined into many other different systems or applications. Various presently unforeseen or unanticipated alternatives, modifications, variations or improvements therein may be subsequently made by those skilled in the art which are also intended to be encompassed by the following claims.

Claims

CLAIMS:

1. A system comprising: at least one processor; and, at least one memory having stored thereon instructions or code that, when executed by the at least one processor, causes the system to implement a multi-tiered driver allocation optimization framework for omnichannel commerce by at least performing: long-term planning based on a mixture density network to obtain an estimate of demand distribution and generate distribution parameter covariates; short-term planning using the distribution parameter covariates in deep learning ensemble forecasting to predict demand variation and generate current state parameter allocations; real-time planning based on the current state demand variation to capture impact of current state on predictions and a Markov decision process to generate a driver allocation; and, outputting the driver allocation to fulfill deliveries by drivers.

2. The system as set forth in claim 1, wherein the long-term planning is based on the mixture density network connected to a newsvendor model.

3. The system as set forth in claim 1, wherein the short-term planning comprises a combination of Light Gradient Boosting Machine (LGBM) and Neural Hierarchical Interpolation for Time Series Forecasting (N-HiTS) forecasting.

4. The system as set forth in claim 1, wherein the current state comprises temporal spillover of driver shortages and backorders.

5. The system as set forth in claim 1, wherein the driver allocation includes hourly responsive reallocations.

376. A method for implementing a multi-tiered driver allocation optimization framework for omnichannel commerce the method comprising: long-term planning based on a mixture density network to obtain an estimate of demand distribution and generate distribution parameter covariates; short-term planning using the distribution parameter covariates in deep learning ensemble forecasting to predict demand variation and generate current state parameter allocations; real-time planning based on the current state demand variation to capture impact of current state on predictions and a Markov decision process to generate a driver allocation; and, outputting the driver allocation to fulfill deliveries by drivers.

7. The method as set forth in claim 6, wherein the long-term planning is based on the mixture density network connected to a newsvendor model.

8. The method as set forth in claim 6, wherein the short-term planning comprises a combination of Light Gradient Boosting Machine (LGBM) and Neural Hierarchical Interpolation for Time Series Forecasting (N-HiTS) forecasting.

9. The method as set forth in claim 6, wherein the current state comprises temporal spillover of driver shortages and backorders.

10. The method as set forth in claim 6, wherein the driver allocation includes hourly responsive reallocations.

11. A non-transitory computer readable medium comprising instructions stored thereon that, when executed by a processor, cause an apparatus to perform: long-term planning based on a mixture density network to obtain an estimate of demand distribution and generate distribution parameter covariates;38short-term planning using the distribution parameter covariates in deep learning ensemble forecasting to predict demand variation and generate current state parameter allocations; real-time planning based on the current state demand variation to capture impact of current state on predictions and a Markov decision process to generate a driver allocation; and, outputting the driver allocation to fulfill deliveries by drivers.

12. The non-transitory computer readable medium as set forth in claim 11, wherein the long-term planning is based on the mixture density network connected to a newsvendor model.

13. The non-transitory computer readable medium as set forth in claim 11, wherein the short-term planning comprises a combination of Light Gradient Boosting Machine (LGBM) and Neural Hierarchical Interpolation for Time Series Forecasting (N- HiTS) forecasting.

14. The non-transitory computer readable medium as set forth in claim 11, wherein the current state comprises temporal spillover of driver shortages and backorders.

15. The non-transitory computer readable medium as set forth in claim 11, wherein the driver allocation includes hourly responsive reallocations.39