Subway mobile warehouse supply strategy optimization method and device based on reinforcement learning

By setting up mobile warehouses in subway stations and building reinforcement learning models, optimizing subway logistics and transportation strategies, the problems of idle subway capacity and fixed courier paths are solved, and efficient logistics distribution and resource utilization are achieved.

CN120494329APending Publication Date: 2025-08-15SOUTHWEST JIAOTONG UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510478842.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-16
Publication Date
2025-08-15

AI Technical Summary

Technical Problem

The existing subway logistics and transportation fails to make full use of the mobility and networking advantages of subways, and fails to adjust the delivery path of couriers according to real-time order needs, resulting in low transportation efficiency, high cost and mutual interference between couriers.

Method used

Based on the method of reinforcement learning, parameter information is obtained through the subway mobile warehouse, reinforcement learning model for subway city distribution centers is constructed, supply strategies are optimized, and order allocation and delivery paths of couriers are dynamically adjusted. Combined with sequential scheduling and Markov decision-making, it reduces user waiting time and labor waste for couriers.

Benefits of technology

It improves the utilization rate of subway capacity, reduces warehousing costs, optimizes the order allocation and distribution paths of couriers, improves the delivery efficiency and order completion rate, and reduces user waiting time.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120494329A_ABST
    Figure CN120494329A_ABST
Patent Text Reader

Abstract

The invention provides a subway mobile warehouse supply strategy optimization method and device based on reinforcement learning, and relates to the technical field of logistics distribution, and the method comprises the steps: obtaining parameter information based on a subway mobile warehouse, and the parameter information comprises road network information, order demand information and courier information; performing model construction through the parameter information, and obtaining a subway city distribution single-center reinforcement learning model based on sequential scheduling and Markov decision; and subway mobile warehouse supply strategy optimization is carried out through the subway city distribution single-center reinforcement learning model, an optimal supply strategy is obtained, and tasks and distribution orders are distributed according to the optimal supply strategy. According to the invention, the problems that the delivery path of the courier is fixed and the supply strategy cannot be adjusted according to the real-time order demand due to the fact that the mobility and networking advantages of the subway are not fully utilized and the cooperation problem between the courier is not considered in the existing subway logistics transportation attempt are solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of logistics and distribution technology, and in particular to a method and device for optimizing a supply strategy for a subway mobile warehouse based on reinforcement learning. Background Art

[0002] With the rapid development of the e-commerce economy, the demand for urban logistics and distribution has increased dramatically. The traditional logistics and distribution model mainly relies on ground transportation, which has problems such as low transportation efficiency, high costs, and heavy storage pressure.

[0003] In recent years, subways, as a vital component of urban rail transit, have seen significant idle capacity at night. Integrating this idle capacity with urban logistics and distribution to develop a supply strategy for mobile subway warehouses has become a pressing issue. Furthermore, supply strategies are essentially the dynamic allocation and scheduling of logistics resources. While there have been attempts to optimize logistics and distribution using reinforcement learning, there is no systematic solution for integrating subway nighttime capacity with urban logistics. Furthermore, existing attempts at subway logistics transportation have failed to fully leverage the subway's mobility and networking advantages, nor have they considered collaboration between couriers. Consequently, delivery routes remain fixed and cannot be adjusted to meet real-time order demands. Summary of the Invention

[0004] The purpose of this invention is to provide a method and device for optimizing the supply strategy of a subway mobile warehouse based on reinforcement learning to improve the above-mentioned problems. To achieve the above-mentioned purpose, the technical solutions adopted by the present invention are as follows:

[0005] In a first aspect, the present application provides a method for optimizing a subway mobile warehouse supply strategy based on reinforcement learning, comprising:

[0006] Acquiring parameter information based on the subway mobile warehouse, the parameter information including road network information, order demand information and courier information;

[0007] The model is constructed through parameter information, and a single-center reinforcement learning model for subway city distribution is obtained based on sequential scheduling and Markov decision making.

[0008] The subway mobile warehouse supply strategy is optimized through the subway city distribution center reinforcement learning model to obtain the optimal supply strategy, and tasks and delivery orders are allocated according to the optimal supply strategy.

[0009] In a second aspect, the present application also provides a subway mobile warehouse supply strategy optimization device based on reinforcement learning, comprising:

[0010] An acquisition module is used to acquire parameter information based on the subway mobile warehouse, wherein the parameter information includes road network information, order demand information and courier information;

[0011] The construction module is used to construct the model through parameter information and obtain the subway city distribution single center reinforcement learning model based on sequential scheduling and Markov decision making;

[0012] The optimization module is used to optimize the supply strategy of the subway mobile warehouse through the subway city distribution center reinforcement learning model, obtain the optimal supply strategy, and allocate tasks and delivery orders according to the optimal supply strategy.

[0013] The beneficial effects of this invention are as follows: By installing mobile subway warehouses at subway stations and utilizing idle subway capacity at night to transport goods, this invention solves the problems of idle subway capacity and high storage costs. Furthermore, by combining sequential scheduling and Markov decision making to construct a reinforcement learning model for the subway city distribution center, this model can dynamically optimize supply strategies, adjust couriers' order allocation and delivery routes based on real-time order demand, reduce user waiting time, improve delivery efficiency, and avoid mutual interference and labor waste between couriers, significantly improving order completion rates. Furthermore, through sequential scheduling configuration, the decision-making speed of the subway city distribution center reinforcement learning model is improved.

[0014] Other features and advantages of the present invention will be described in the following description, and in part will become apparent from the description, or understood by practicing the embodiments of the present invention. The purposes and other advantages of the present invention can be realized and obtained by the structures particularly pointed out in the written description, claims, and drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0015] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for use in the embodiments. It should be understood that the following drawings only illustrate certain embodiments of the present invention and therefore should not be regarded as limiting the scope. For ordinary technicians in this field, other relevant drawings can be obtained based on these drawings without paying any creative work.

[0016] Figure 1 This is a flow chart of a method for optimizing a supply strategy for a subway mobile warehouse based on reinforcement learning according to an embodiment of the present invention;

[0017] Figure 2 Schematic diagram of the flow of the sequential scheduling configuration method in an embodiment of the present invention;

[0018] Figure 3 is a schematic diagram of grid labels of a grid area in an embodiment of the present invention;

[0019] Figure 4 This is a schematic diagram of the locations of all couriers in an embodiment of the present invention;

[0020] Figure 5Schematic diagram of the amount of pending orders in a grid in an embodiment of the present invention;

[0021] Figure 6 Schematic diagram of the positions of other couriers compared to courier c1 in an embodiment of the present invention;

[0022] Figure 7 Schematic diagram of the positions of all couriers after courier c1 performs the optimal action in an embodiment of the present invention;

[0023] Figure 8 This is a schematic diagram of the number of pending orders in the grid after courier c1 performs the optimal action in an embodiment of the present invention;

[0024] Figure 9 Schematic diagram of the positions of other couriers compared to courier c2 in an embodiment of the present invention. DETAILED DESCRIPTION

[0025] In order to make the purpose, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. The components of the embodiments of the present invention generally described and shown in the drawings herein can be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of the present invention provided in the drawings is not intended to limit the scope of the claimed invention, but merely represents selected embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present invention.

[0026] It should be noted that similar reference numerals and letters represent similar items in the following drawings. Therefore, once an item is defined in one drawing, it does not need to be further defined or explained in subsequent drawings. At the same time, in the description of the present invention, the terms "first", "second", etc. are only used to distinguish the description and should not be understood as indicating or implying relative importance.

[0027] Example 1:

[0028] This embodiment provides a method for optimizing the supply strategy of a subway mobile warehouse based on reinforcement learning.

[0029] See also Figure 1 , the figure shows that the method includes step S100, step S200 and step S300.

[0030] Step S100: Acquiring parameter information based on the subway mobile warehouse, the parameter information including road network information, order demand information and courier information;

[0031] The step S100 includes:

[0032] Step S101: Setting up a subway mobile warehouse at a subway station. The subway mobile warehouse is used to store goods transported by the subway at night. The goods are allocated based on the order forecast demand of the previous day.

[0033] In this embodiment, the subway mobile warehouse serves as the initial distribution center for couriers, and the subway cars serve as temporary storage points for goods at night. Goods are transported to various subway mobile warehouses via the night subway, and couriers depart from the subway mobile warehouse when they start work every day.

[0034] At the same time, the goods in the subway mobile warehouse are mainly allocated based on the predicted demand for orders from the previous day. When the subway is idle at night, the goods are transported in batches from the suburban warehouse to the subway mobile warehouse for storage, covering most of the orders for the next day. In the case that real-time orders cannot be met, the mobility of the subway network can be utilized to allocate goods from the subway mobile warehouse at other stations during non-peak hours during the day, or obtain them from the suburban warehouse by road transportation.

[0035] Step S102: Divide the urban area into a plurality of grid areas centered on the subway mobile depot, each of the grid areas being composed of a plurality of grids of the same size;

[0036] Step S103: Obtain the road network information and courier information of each grid area, and obtain the order demand of each grid in each grid area to obtain order demand information, wherein the order demand information includes the order generation location and order generation time.

[0037] In this embodiment, because e-commerce order demand is dynamic and users can place orders at any time, this places high demands on the service efficiency of the e-commerce delivery system. Furthermore, e-commerce delivery systems often need to pursue long-term optimization. When a courier departs from a transfer station to perform a delivery, each step must consider not only short-term demand fulfillment but also long-term delivery service efficiency.

[0038] Therefore, by zoning different subway stations, the city area is divided into independent grid areas based on each subway station, and each grid area is focused on separately. In addition to reducing the complexity of the problem, the urban area division also has practical operational benefits, allowing for better utilization of the advantages of subway mobile warehouses while reducing the construction costs of transfer stations.

[0039] In this step, the road network information includes the distance between each node in the corresponding grid area. This information allows us to calculate the distance between any two orders during the delivery process. The order location in the order demand information actually indicates the specific grid area where the order was generated and the location of the specific node within that grid area.

[0040] Step S200: constructing a model using parameter information, and obtaining a single-center reinforcement learning model for subway city distribution based on sequential scheduling and Markov decision making;

[0041] Step S300: Optimize the supply strategy of the subway mobile warehouse through the subway city distribution center reinforcement learning model to obtain the optimal supply strategy, and allocate tasks and delivery orders according to the optimal supply strategy.

[0042] The step S300 includes:

[0043] Step S301: Divide the working time into multiple time periods;

[0044] Step S302: Initialize the global state space, which includes the current states of all couriers;

[0045] In this embodiment, the courier's current status includes the number of pending orders, the current grid positions of other couriers, the current grid position of the courier, and the current time period, where the number of pending orders represents the number of orders that have been accepted but not delivered.

[0046] Step S303: Selecting actions for the courier based on the global state space to obtain the optimal supply strategy for the current time period;

[0047] In this example, each courier has nine possible preset actions in each time period. Specifically, the courier can choose to move to eight adjacent grids or remain in their current location. Therefore, this step also sets a global action space, which includes these nine preset actions.

[0048] Since there are hundreds or even thousands of couriers delivering services at the same time within the operating range of a subway entrance, considering the individual actions of all couriers at the same time will result in 9 n There are different situations, where n represents the number of couriers. Based on the above situation, using reinforcement learning algorithm to directly select actions will have a very low selection efficiency.

[0049] Therefore, this embodiment proposes a sequential scheduling configuration method, that is, in each time period, tasks are assigned to each courier in sequence until all couriers have their own tasks and then start to act at the same time at the beginning of each time period, and then work in their grid area until the end of this time period. At this time, the above process is repeated to continue to determine the tasks for the next time period, where the tasks actually include the courier's next actions and the orders to be delivered. Figure 2 As shown, it is a schematic diagram of selecting actions for three couriers in the same grid area in sequence, where t represents the time period, c1, c2 and c3 represent courier 1, courier 2 and courier 3 respectively.

[0050] The courier's current status indicates that even within the same time period, the locations of other couriers may differ from the courier's current location due to the different couriers. Furthermore, since a courier's current status is influenced by the actions of other couriers who took action before them, the number of pending orders faced by each courier will also vary.

[0051] The step S303 includes:

[0052] Step A100: After sorting all couriers according to the preset sorting rules, couriers are selected in sequence;

[0053] In this embodiment, there are many couriers delivering in the same time period. When sorting, they can be sorted randomly, sorted according to the order of the couriers' work IDs, and sorted according to the couriers' delivery efficiency, etc. The specific settings can be made according to the actual situation of the grid area.

[0054] Step A200: For the selected courier, calculate the expected function value of each preset action and the number of orders processed based on the current state;

[0055] The step A200 includes:

[0056] Step A201: Obtain the number of pending orders for each grid based on the current state, and calculate the number of orders processed for each preset action using the service time, available time, and the number of pending orders;

[0057] In this embodiment, the calculation formula for the order processing quantity is:

[0058]

[0059] In the formula, ds represents the number of orders processed, min(·) represents the minimum value, Indicates rounding down, t k represents the available time, t frepresents the service time, dc represents the number of pending orders, and available time represents the time the courier can actually process orders after deducting the travel time.

[0060] Step A202: Calculating the instant reward value of each preset action based on the number of orders processed for each preset action;

[0061] The step A202 includes:

[0062] Step B100: Obtain the delivery speed of the selected courier;

[0063] In this embodiment, the courier's speed determines the distance he can move in a single time period, which directly affects whether he can reach the target area and specific location of the order in a timely manner.

[0064] Step B200: Allocate orders to each preset action according to the order processing quantity, and obtain the order generation position of each order under each preset action;

[0065] Step B300: Calculate the actual time taken for each order based on the delivery speed and the order generation location, and calculate the order time for each order using the actual time taken and the service time;

[0066] In this embodiment, the delivery distance of the order is calculated based on the location where the order is generated and the location of the courier, and the service time represents the time required for the courier to process a single order when it arrives at the location where the order is generated.

[0067] In this step, the calculation formula for order time is:

[0068]

[0069] Where tx represents the order time, d x represents the delivery distance, vr represents the delivery speed, t ∈ Indicates service time.

[0070] Step B400: Calculate the instant reward for each order based on the actual time, the maximum demand response time, and the remaining time after the move;

[0071] In this embodiment, the maximum demand response time represents the timeout threshold of the order. If the order is not completed within the maximum demand response time, a negative reward will be triggered, which will reduce the immediate reward value and prompt the order that is about to time out to be processed first when calculating the expected function value.

[0072] The actual time spent in the order time and the service time jointly affect the acquisition of instant rewards. If the maximum demand response time is not exceeded and the remaining time after moving within the time period is not enough to complete the service, that is, the remaining time after moving is less than the service time, the order cannot be processed.

[0073] In this embodiment, the formula for instant reward is:

[0074]

[0075] Where R represents the immediate reward, and the specified time is the maximum demand response time.

[0076] In this step, the specific value of the instant reward can be changed according to the actual situation. Here, the instant reward is set by simple value.

[0077] Step B500: Sum up the instant rewards of all orders in each action to obtain the instant reward value of each preset action.

[0078] In this embodiment, when assigning orders to each preset action, multiple order combinations are obtained based on the order processing quantity, and then the instant reward value of each order combination under each preset action is calculated, the largest instant reward value is selected as the instant reward value of the corresponding preset action, and the corresponding order combination is used as the order that needs to be delivered under the preset action.

[0079] Step A203: obtaining the number of couriers in the current action time period according to the courier information, and selecting an expectation function based on the number of couriers, wherein the expectation function is a Q-value iterative function of a Sarsa algorithm or a Q-learning algorithm;

[0080] In this embodiment, since the model-independent reinforcement learning algorithm does not rely on the reward model and the prior transfer of the state, the model-independent reinforcement learning algorithm is applied to the subway city distribution single center reinforcement learning model.

[0081] When the number of couriers is greater than the preset threshold, the Q-value iterative function of the Sarsa algorithm is selected as the expected function. When the number of couriers is not greater than the preset threshold, the Q-value iterative function of the Q-learning algorithm is selected as the expected function.

[0082] In this step, the formula of the expectation function is:

[0083]

[0084] In the formula, Q represents the expected function, n represents the number of couriers, and n s represents the preset threshold, α represents the learning rate, which is used to control the speed of the expected function iteration, γ represents the discount factor, Q(S t ,a t ) indicates that in S t Next take a t The expected function value when Q(S t+1 ,a t+1 ) indicates that in S t+1 Next take at+1 The expected function value when r t Indicates that in S t Next take a t The immediate reward value obtained after S t and S t+1 Represents the current state of time period t and time period t+1 respectively, a t and a t+1 Respectively represent the actions taken in time period t and time period t+1, max a (·) means taking the maximum value after changing action a.

[0085] Step A204: Calculate the expected function value of each preset action using the selected expected function and the immediate reward value.

[0086] Step A300: Selecting a preset action with the largest expected function value and the corresponding order processing quantity and assigning it to the courier to obtain the optimal action for the courier;

[0087] Step A400: After the courier performs the optimal action, the current status of all couriers is updated, and the optimal action of the next selected courier is assigned until the assignment is completed, thereby obtaining the optimal supply strategy for the current action time period.

[0088] Step S304: At the beginning of the current time period, all couriers start delivering orders simultaneously according to the optimal supply strategy;

[0089] In this embodiment, at the beginning of each time period, all couriers obtain all the goods that need to be delivered within the time period from the subway mobile warehouse at one time according to their own optimal actions.

[0090] Step S305: Update the global state space according to the end time of the current time period, and calculate the optimal supply strategy for the next time period.

[0091] In this example, within each grid area, a sequentially scheduled subway city distribution center reinforcement learning model assigns actions to each courier within each time period t. Therefore, taking a grid area consisting of nine grids and three couriers as an example, the process for obtaining the optimal supply strategy is as follows: Among them, S wt Indicates that courier c in time period t w The current state of a wt Indicates that in time period t, courier c w The optimal action is w = 1, 2, ..., n, where w is the index number of the courier and n is the number of couriers.

[0092] like Figure 3As shown, the grid g i The grid number of the grid area, where i = 1, 2, ..., m, and m represents the number of grids. Assume that there are 3 couriers in Figure 4 In the grid area shown, they are located in grid g3, grid g4 and grid g9 respectively. Figure 5 As shown, it describes the order demand that has been received but not yet processed by the courier, where the number in each grid represents the number of orders to be processed in each grid.

[0093] For courier c1 located in grid g3, the other two couriers are located in grids g4 and g9 respectively. Figure 6 The numbers in the grid represent the number of couriers in that grid. Therefore, for courier c1, the grid g3 he is currently in, the grids other couriers are currently in, the number of pending orders in each grid, and the current time period t constitute his current state S in time period t. 1t In the current state S 1t Based on this, assign the optimal action a to courier c1 1t , the optimal action a 1t Make courier c1 go to grid g2, such as Figure 7 As shown, the couriers are located in grids g2, g4, and g9 respectively. Assume that courier c1 completes the optimal action a 1t After that, and the order delivery within time period t is completed, the current status of courier c2 is considered next.

[0094] Through the optimal action a 1t Get the number of orders processed by courier C1 in grid G2. For example, if the number of orders processed by courier C1 in grid G2 is 3, update the number of pending orders in each grid to obtain Figure 8 The number of pending orders faced by courier c2 is shown in . At the same time, due to the action of courier c1, the positions of other couriers faced by courier c2 also change accordingly, such as Figure 9 Therefore, we can conclude that the current state S of courier c2 in time period t is 2t This process is repeated continuously until all couriers have taken their own actions in time period t. This means that the optimal supply strategy for time period t is obtained, and all couriers deliver orders together according to the optimal supply strategy at the beginning of time period t until entering the next time period t+1.

[0095] Example 2:

[0096] In this example, based on JD Logistics' historical order data in Chengdu, we used a reinforcement learning model from the Metro City Distribution Center to simulate order delivery and optimize supply strategies for the grid area of Chengdu's Chunxi Road subway station. To better align with real-life courier work schedules, we chose the daily working hours of 8:00 AM to 1:00 PM as the delivery time.

[0097] At the same time, this embodiment also introduces a random algorithm and a greedy algorithm for comparison with the subway city distribution center reinforcement learning model using the sequential scheduling configuration method of the present invention. The random algorithm assumes that the courier will randomly select any one of the 9 preset actions, and the greedy algorithm assumes that each courier always chooses the grid with the largest number of uncompleted orders among the 9 preset actions.

[0098] In order to better adapt to the actual delivery process, the courier's delivery speed is set not to exceed 15km / h, and the maximum response time for an order demand shall not be greater than 1 hour. According to the workload in the grid area of Chengdu Chunxi Road Metro Station, a certain number of couriers are allocated to the grid area. In fact, the number of couriers is also determined according to actual conditions, such as the courier's work pressure and topography. However, since the goal of this embodiment is to determine the performance of the subway city distribution center reinforcement learning model under the premise of a given number of couriers, experiments are conducted with different numbers of couriers to test the effectiveness of the subway city distribution center reinforcement learning model. Table 1 shows the proposed parameters of this embodiment.

[0099] Table 1 Proposed parameters

[0100] The extent of the grid zone 1.2km×1.2km Grid division 12×12 Order quantity 5319 Delivery speed of the courier ≤15km / h Maximum demand response time 1h Service time for each order 2min

[0101] Based on the proposed parameters, this example combines the proposed method with the real-life e-commerce courier order delivery problem to conduct simulation experiments. The performance of the proposed method is verified by comparing the experimental results with other classic reinforcement learning algorithms. First, the following assumptions are made: The overlap between couriers and grid boundaries and external factors such as traffic lights are not considered, and only cargo orders are considered as order demands within the grid boundaries.

[0102] Based on these assumptions, we simulated the operation of the method of the present invention using a system simulation. Using road network information, the system simulation quickly determined the distance between two locations. Furthermore, the simulation assumed that the distance between two orders within the same grid follows a normal distribution. Secondly, based on JD Logistics' historical order data, we assumed that the number of orders within each grid and time period follows a normal distribution. Based on these assumptions about distance and order quantity, we conducted the following system simulation.

[0103] Each courier begins their workday at the mobile subway warehouse at Chunxi Road subway station. During each time period t throughout the delivery process, the system first generates order requirements for each grid cell in that time period. Orders that exceed the maximum response time in the previous time period are then eliminated, updating the order quantity. The optimal action for each courier in that time period t is then assigned based on the reinforcement learning model of the subway city distribution center.

[0104] In this embodiment, the order completion rate is used as the standard for evaluating the completion of the delivery task, that is, the proportion of completed orders in the total number of orders. Therefore, the formula for the order completion rate is:

[0105]

[0106] Where PCR represents the order completion rate.

[0107] At the same time, the discount factor γ in the reinforcement learning model of the subway city distribution center, γ∈[0,1], is used to find the optimal discount factor. The impact of the discount factor on the PCR in the reinforcement learning model of the subway city distribution center at Chunxi Road Subway Station is studied with 20, 30, 40, 50, and 60 couriers. The PCR under different discount factors with different numbers of couriers is calculated, and a curve graph of the PCR with different discount factors under different numbers of couriers is obtained. It is found that when the discount factor is around 0.7, the PCR is the largest under different numbers of couriers. Therefore, 0.7 is selected as the discount factor for this embodiment.

[0108] When the expectation function is the Q-value iterative function of the Sarsa algorithm, the greediness value also needs to be considered. In order to ensure the performance of the algorithm, the range of the greediness value is set while keeping other parameters unchanged. Referring to classic research in the field of reinforcement learning, the greediness value of the Sarsa algorithm is often divided into three intervals: (0, 0.1], (0.1, 0.2] and (0.2, 1]. Through designing experiments, the Sarsa algorithm with three different intervals of greediness is applied to the subway city distribution center reinforcement learning model with 30 couriers. 100 iterative trainings are performed, and the iterative convergence speed and PCR are used as evaluation criteria at the same time. The final results are shown in Table 2.

[0109] Table 2 Evaluation results of different greedy values

[0110] Greediness Convergence time / s PCR (0,0.1] 11.5 0.54 (0.1,0.2] 15.3 0.72 (0.2,1] —— 0.65

[0111] As can be seen from Table 2, when the greed value selected by the subway city distribution center reinforcement learning model is in the range of (0, 0.1], the subway city distribution center reinforcement learning model will choose the future state with the largest expected reward value, resulting in slow convergence. Specifically, the algorithm takes a long time to iterate 100 times. At the same time, because it only focuses on the interval with the largest reward value, it is easy to ignore the order demand in other grids, which is specifically manifested as a PCR of only 0.54.

[0112] When the greediness value is in the range of (0.2, 1], the strategy is too conservative. The simulated couriers often stay in one grid for too long for fear of making mistakes, which makes it difficult to converge. At the same time, due to the long delay, the PCR is low. Therefore, based on the experimental data of the above control variables, when the greediness value is in the range of (0.1, 0.2], it can both aggressively select the next strategy and conservatively estimate the optimal result. Therefore, the greediness value selected in this embodiment is 0.1.

[0113] In this embodiment, different numbers of couriers are set for the subway city distribution center reinforcement learning model of Chunxi Road subway station based on the selected discount factor and greedy value, and the random algorithm and the greedy algorithm are selected for comparison, as shown in Table 3.

[0114] Table 3 Comparison of PCR results of strategy optimization using different algorithms

[0115]

[0116]

[0117] As shown in Table 3, the Q-learning and Sarsa algorithms are used to determine the expected function in a reinforcement learning model for a subway city delivery center. Of the two heuristic algorithms, the randomized algorithm performed the worst. This is because it doesn't consider how to achieve the optimal goal, but simply observes the results of the decision-making process under completely random action selection. To some extent, the randomized algorithm can represent the results of a delivery system without any guidance, completely chaotic, and free to act by couriers.

[0118] It is worth noting that when the number of couriers is small, the greedy algorithm and the reinforcement learning model of the subway city distribution center of the present invention perform equally well, with almost no difference. This is because when the number of couriers is very small, there is very little interference between couriers, and the completed orders and the location of other couriers do not have a significant impact on the couriers. At the same time, because the number of orders within the grid area of Chengdu Chunxi Road subway station is too large for 10 couriers, the long-term optimal goal in this case has no strong practical significance. This means that when the number of orders is too large for the couriers, the couriers only need to consider completing the current goal, so there is no labor waste caused by the greedy algorithm.

[0119] However, as the number of couriers working simultaneously within the grid area increases, the greedy algorithm's performance deteriorates, with the PCR dropping sharply compared to the Metro City Distribution Center reinforcement learning model. This is because the Metro City Distribution Center reinforcement learning model accounts for collaboration between couriers. As the number of couriers increases, the greedy algorithm causes couriers to blindly head to the area with the most orders, disregarding the choices of their peers. This can lead to labor waste. The Metro City Distribution Center reinforcement learning model, on the other hand, continuously draws information from the environment and self-learns, allowing it to consider the optimal progress of the entire delivery process, thus avoiding the problem of couriers clumping together in the greedy algorithm.

[0120] In the reinforcement learning model for the subway city distribution center, the Q-learning and Sarsa algorithms can be used to determine the expected function. Table 3 shows that both the Q-learning and Sarsa algorithms achieve significantly higher PCRs than the random and greedy algorithms. Through the information interaction between the agent and the environment in reinforcement learning, as well as the model's self-learning process, the optimal supply strategy is obtained, leading to maximum profit. However, differences in sampling and decision-making logic between the Q-learning and Sarsa algorithms lead to different performance in different environments, as shown in the following:

[0121] When the number of couriers is limited, the PCR of a metro city delivery center reinforcement learning model using the Sarsa algorithm is often similar to or even worse than that of a metro city delivery center reinforcement learning model using the Q-learning algorithm. This is because the Q-learning algorithm does not consider greedy values and therefore makes bold decisions. When the number of couriers is limited, such bold decisions can often solve the vast majority of order problems, thereby improving the overall PCR. However, the Sarsa algorithm's sampling and decision-making process are overly conservative compared to the Q-learning algorithm, and sometimes some profitable decisions are not executed due to risk considerations.

[0122] As the number of couriers increases, the PCR of the reinforcement learning model for metro-city delivery centers using the Sarsa algorithm gradually improves, ultimately surpassing the one using the Q-learning algorithm. This is because with an increase in the number of couriers, the perceived waste of courier labor is no longer as acute as it was when the number of couriers was insufficient. At the same time, it is more likely to cause a decrease in PCR due to aggressive decisions. However, the Sarsa algorithm, due to its greediness considerations, is more conservative in its decisions, thus avoiding many undesirable aggressive decisions. These advantages are reflected in the fact that as the number of couriers increases, the PCR of the reinforcement learning model for metro-city delivery centers using the Sarsa algorithm surpasses that of the reinforcement learning model using the Q-learning algorithm. Therefore, in practical applications, it is possible to select an expectation function based on the number of couriers to obtain a corresponding reinforcement learning model for metro-city delivery centers.

[0123] To sum up, the present invention combines the subway mobile warehouse with the urban logistics distribution system, sets up subway mobile warehouses according to subway stations, and transports goods to various stations at night, so that the nighttime subway capacity can be fully utilized, resource utilization can be improved, and the return on investment of the subway can be increased. It solves the problem of idle subway capacity at night, reduces the demand for storage space, and reduces storage costs.

[0124] At the same time, a reinforcement learning model for the subway city distribution center is constructed through the sequential scheduling configuration method, which dynamically adjusts the supply strategy and optimizes the courier's order allocation and delivery route decision-making, reducing delivery time while improving delivery efficiency. It can also make real-time adjustments based on real-time order demand, reducing user waiting time.

[0125] Therefore, the method of the present invention can provide clear and efficient action allocation when a large number of couriers are faced with massive orders at the same time. It is suitable for e-commerce logistics and distribution environments in large cities with a relatively complex population structure and high requirements for order accuracy. In later practical applications, the corresponding expectation function can be adaptively selected according to the actual number of couriers.

[0126] Example 3:

[0127] This embodiment provides a device for optimizing a supply strategy for a subway mobile warehouse based on reinforcement learning, the device comprising:

[0128] An acquisition module is used to acquire parameter information based on the subway mobile warehouse, wherein the parameter information includes road network information, order demand information and courier information;

[0129] The construction module is used to construct the model through parameter information and obtain the subway city distribution single center reinforcement learning model based on sequential scheduling and Markov decision making;

[0130] The optimization module is used to optimize the supply strategy of the subway mobile warehouse through the subway city distribution center reinforcement learning model, obtain the optimal supply strategy, and allocate tasks and delivery orders according to the optimal supply strategy.

[0131] The acquisition module includes:

[0132] A setting unit is used to set up a subway mobile warehouse at the subway station, wherein the subway mobile warehouse is used to store goods transported by the subway at night, and the goods are allocated according to the order demand forecast of the previous day;

[0133] A first division unit is configured to divide the urban area into a plurality of grid areas centered on the subway mobile depot, each of the grid areas being composed of a plurality of grids of the same size;

[0134] The acquisition unit is used to acquire the road network information and courier information of each grid area, and acquire the order demand of each grid in each grid area to obtain order demand information, wherein the order demand information includes the order generation location and order generation time.

[0135] The optimization module includes:

[0136] A second division unit is used to divide the working time into multiple time periods;

[0137] An initialization unit, configured to initialize a global state space, wherein the global state space includes the current states of all couriers;

[0138] The selection unit is used to select actions for couriers based on the global state space to obtain the optimal supply strategy for the current time period;

[0139] The delivery unit is used to enable all couriers to start delivering orders simultaneously according to the optimal supply strategy at the beginning of the current time period;

[0140] The updating unit is used to update the global state space according to the end time of the current time period and calculate the optimal supply strategy for the next time period.

[0141] The selection unit includes:

[0142] A sorting unit, used to sort all couriers according to a preset sorting rule and select couriers in sequence;

[0143] A calculation unit, configured to calculate the expected function value of each preset action and the number of orders processed for the selected courier based on the current state;

[0144] An allocation unit is used to select a preset action with the largest expected function value and the corresponding order processing quantity and assign it to the courier to obtain the optimal action of the courier;

[0145] The execution unit is used to update the current status of all couriers after the courier performs the optimal action, and assign the optimal action of the next selected courier until the assignment is completed, thereby obtaining the optimal supply strategy for the current action time period.

[0146] It should be noted that, regarding the apparatus in the above embodiment, the specific manner in which each module performs operations has been described in detail in the embodiment of the method, and will not be elaborated on here.

[0147] The foregoing description is merely a preferred embodiment of the present invention and is not intended to limit the present invention. Those skilled in the art will readily appreciate that various modifications and variations of the present invention are possible. Any modifications, equivalent substitutions, or improvements made within the spirit and principles of the present invention are intended to be within the scope of protection of the present invention.

[0148] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any modifications or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in the present invention should be included within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be based on the scope of protection of the claims.

Claims

1. A subway mobile warehouse supply strategy optimization method based on reinforcement learning, characterized in that: include: Acquiring parameter information based on the subway mobile warehouse, the parameter information including road network information, order demand information and courier information; The model is constructed through parameter information, and a single-center reinforcement learning model for subway city distribution is obtained based on sequential scheduling and Markov decision making. The subway mobile warehouse supply strategy is optimized through the subway city distribution center reinforcement learning model to obtain the optimal supply strategy, and tasks and delivery orders are allocated according to the optimal supply strategy.

2. The subway mobile warehouse supply strategy optimization method based on reinforcement learning according to claim 1 is characterized in that ,Based on the subway mobile warehouse, parameter information is obtained, including: A subway mobile warehouse is set up at the subway station. The subway mobile warehouse is used to store goods transported by the subway at night. The goods are allocated according to the order demand forecast of the previous day; Dividing the urban area into a plurality of grid areas centered on the subway mobile depot, each of the grid areas being composed of a plurality of grids of the same size; The road network information and courier information of each grid area are obtained, and the order demand of each grid in each grid area is obtained to obtain order demand information, wherein the order demand information includes the order generation location and the order generation time.

3. The subway mobile warehouse supply strategy optimization method based on reinforcement learning according to claim 1 is characterized in that: The subway mobile warehouse supply strategy is optimized through the subway city distribution center reinforcement learning model to obtain the optimal supply strategy. Tasks and delivery orders are then assigned based on the optimal supply strategy, including: Divide working time into time blocks; Initialize a global state space, which includes the current states of all couriers; Select actions for couriers based on the global state space to obtain the optimal supply strategy for the current time period; At the beginning of the current time period, all couriers start delivering orders simultaneously according to the optimal supply strategy; Update the global state space according to the end time of the current time period and calculate the optimal supply strategy for the next time period.

4. The subway mobile warehouse supply strategy optimization method based on reinforcement learning according to claim 3 is characterized in that: Based on the global state space, we select actions for the couriers and obtain the optimal supply strategy for the current time period, including: After sorting all couriers according to the preset sorting rules, couriers are selected in turn; For the selected courier, calculate the expected function value of each preset action and the number of orders processed based on the current state; Select the preset action with the largest expected function value and the corresponding order processing quantity and assign it to the courier to obtain the optimal action for the courier; After the courier performs the optimal action, the current status of all couriers is updated, and the optimal action of the next selected courier is assigned until the assignment is completed, and the optimal supply strategy for the current action time period is obtained.

5. The subway mobile warehouse supply strategy optimization method based on reinforcement learning according to claim 4 is characterized in that: Calculate the expected function value and order processing quantity of each preset action based on the current state, including: Get the pending order quantity for each grid based on the current status, and calculate the order processing quantity for each preset action using service time, available time, and pending order quantity; Calculate the instant reward value for each preset action based on the number of orders processed for each preset action; Obtaining the number of couriers within the current action time period according to the courier information, and selecting an expectation function based on the number of couriers, wherein the expectation function is a Q-value iterative function of a Sarsa algorithm or a Q-learning algorithm; The expected function value of each preset action is calculated by the selected expected function and the immediate reward value.

6. The subway mobile warehouse supply strategy optimization method based on reinforcement learning according to claim 5 is characterized in that: The instant reward value for each preset action is calculated based on the number of orders processed for each preset action, including: Get the delivery speed of the selected courier; Allocate orders to each preset action according to the order processing quantity, and obtain the order generation position of each order under each preset action; Calculate the actual time taken for each order based on the delivery speed and the location where the order was generated, and calculate the order time for each order using the actual time taken and the service time; Calculate the instant reward for each order based on the actual time, the maximum demand response time and the remaining time after the move; The instant rewards of all orders in each action are summed up to obtain the instant reward value of each preset action.

7. A subway mobile warehouse supply strategy optimization device based on reinforcement learning, characterized in that: include: An acquisition module is used to acquire parameter information based on the subway mobile warehouse, wherein the parameter information includes road network information, order demand information and courier information; The construction module is used to construct the model through parameter information and obtain the subway city distribution single center reinforcement learning model based on sequential scheduling and Markov decision making; The optimization module is used to optimize the supply strategy of the subway mobile warehouse through the subway city distribution center reinforcement learning model, obtain the optimal supply strategy, and allocate tasks and delivery orders according to the optimal supply strategy.

8. The subway mobile warehouse supply strategy optimization device based on reinforcement learning according to claim 7 is characterized in that: The acquisition module includes: A setting unit is used to set up a subway mobile warehouse at the subway station, wherein the subway mobile warehouse is used to store goods transported by the subway at night, and the goods are allocated according to the order demand forecast of the previous day; A first division unit is configured to divide the urban area into a plurality of grid areas centered on the subway mobile depot, each of the grid areas being composed of a plurality of grids of the same size; The acquisition unit is used to acquire the road network information and courier information of each grid area, and acquire the order demand of each grid in each grid area to obtain order demand information, wherein the order demand information includes the order generation location and order generation time.

9. The subway mobile warehouse supply strategy optimization device based on reinforcement learning according to claim 7 is characterized in that: The optimization module includes: A second division unit is used to divide the working time into multiple time periods; An initialization unit, configured to initialize a global state space, wherein the global state space includes the current states of all couriers; The selection unit is used to select actions for couriers based on the global state space to obtain the optimal supply strategy for the current time period; The delivery unit is used to enable all couriers to start delivering orders simultaneously according to the optimal supply strategy at the beginning of the current time period; The updating unit is used to update the global state space according to the end time of the current time period and calculate the optimal supply strategy for the next time period.

10. The subway mobile warehouse supply strategy optimization device based on reinforcement learning according to claim 9 is characterized in that: The selection unit includes: A sorting unit, used to sort all couriers according to a preset sorting rule and select couriers in sequence; A calculation unit, configured to calculate the expected function value of each preset action and the number of orders processed for the selected courier based on the current state; An allocation unit is used to select a preset action with the largest expected function value and the corresponding order processing quantity and assign it to the courier to obtain the optimal action of the courier; The execution unit is used to update the current status of all couriers after the courier performs the optimal action, and assign the optimal action of the next selected courier until the assignment is completed, thereby obtaining the optimal supply strategy for the current action time period.