Delivery planning device, delivery planning method, and program

By using a neural network with an actor-critic method and a masking algorithm to address time frame and time cost constraints, the solution efficiently solves the vehicle routing problem, reducing calculation time and improving route optimization for practical-scale problems.

JP7683716B2Active Publication Date: 2025-05-27NIPPON TELEGRAPH & TELEPHONE CORP
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
JP2023550859
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2021-09-29
Publication Date
2025-05-27
Estimated Expiration
2041-09-29

AI Technical Summary

Technical Problem

Existing vehicle routing problem (VRP) solutions, particularly for practical-scale problems with 100 or more customers, face challenges in obtaining optimal or approximate solutions in a reasonable time due to their NP-hard nature. Additionally, conventional operations research (OR)-based methods require different handcrafted search models and initial conditions for various VRP variations, making them difficult to generalize and apply in real business scenarios.

Method used

The proposed solution employs a neural network that performs reinforcement learning using the actor-critic method to solve the delivery planning problem. This approach includes a masking algorithm that considers time frame and time cost constraints, allowing for efficient calculation of delivery plans that minimize costs and adhere to specified time windows and service durations.

Benefits of technology

The solution significantly reduces the calculation time for vehicle allocation plans, enabling optimization of routes considering time constraints and service hours. It can handle medium-scale datasets efficiently, reducing operational costs and improving service quality, while being adaptable to various VRP variations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007683716000009
    Figure 0007683716000009
  • Figure 0007683716000010
    Figure 0007683716000010
  • Figure 0007683716000011
    Figure 0007683716000011
Patent Text Reader

Abstract

A delivery planning device comprising an algorithm calculation unit that, using a neural network for performing reinforcement learning based on an actor-critic scheme, solves a delivery planning problem to determine a path for providing service to a plurality of customers using a vehicle departing from a service center, the algorithm calculation unit solving the delivery planning problem while employing, as constraints, a time frame that indicates the range of time in which the customers should be reached and a time cost that indicates the length of time required for providing service to the customers.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a technique for solving the distribution planning problem.

Background Art

[0002] The distribution planning problem (VRP: vehicle routing problem) is an optimization problem that considers which service vehicle should visit which customer in what order (to minimize costs) when delivering goods from a goods collection point (service center) to each customer using service vehicles. Note that the "distribution planning problem" may also be referred to as the "vehicle scheduling problem".

[0003] In actual applications, there are many practical business scenarios such as just-in-time delivery in e-commerce, cold chain delivery, and store replenishment, where distribution and service costs can be optimized through the solution of the VRP.

[0004] Therefore, various variations of the VRP have been proposed according to different practical requirements. As a variation of the VRP, for example, there is the VRP with time windows (VRPTW). In VRPTW, a time window for delivering goods to customers is set. Another VRP is the multi-depot distribution planning problem (MDVRP). In MDVRP, there are multiple depots (service centers), and vehicles can depart from or end their trips there.

[0005] Since the VRP and its variations have been proven to be NP-hard problems, various operations research (OR)-based methods that return approximate solutions have been studied for many years.

[0006] Usually, in OR-based algorithms, a search model is defined manually, and the solution of the VRP is obtained by sacrificing the quality of the solution in order to improve efficiency. However, the conventional OR-based methods have two drawbacks.

[0007] As a first drawback, in the case of a practical-scale VRP problem (having 100 or more customers), when using an OR-based algorithm, it takes several days or years for the calculation to obtain an optimal solution or an approximate solution.

[0008] As a second drawback, different variations of VRP require different handcrafted search models and initial search conditions, and thus require different OR algorithms. For example, an inappropriate initial solution may lead to long processing times and local optimal solutions. In such a regard, it is difficult to generalize and use OR-based algorithms in real business scenarios.

[0009] Non-Patent Document 1 discloses a solution for VRP based on the actor-critic method of reinforcement learning, which solves the drawbacks of OR-based algorithms. That is, with a neural network model, particularly when the number of customer nodes is large, the complexity and representational ability can be significantly improved with high precision.

[0010] Furthermore, although it takes time in the learning phase, with a neural network, an approximate solution can be found instantaneously in the inference phase, and the execution efficiency in practical business applications can be significantly improved.

[0011] Also, since a data-driven neural network does not need to define a mathematical model for search, it can be applied to various VRP variations just by supplying new data and adjusting the reward function or other basic engineering tasks, which is also very convenient for practical research and business development.

Prior Art Documents

Non-Patent Documents

[0012]

Non-Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0013] In real-world applications, there are many practical business scenarios where distribution and service costs, such as just-in-time delivery for e-commerce, cold chain delivery, and store replenishment, can be optimized through VRP solutions.

[0014] For example, a communication carrier receives a large number of requests from customers every day and sends technicians from the service center to the customer's home to assist in repairing network failures. The repair time varies depending on the type of failure, and the difference is often very large. From the perspective of the service center, considering the repair time slots specified by customers, planning a reasonable and efficient repair order and route to minimize the number of repair staff and working hours is considered one of the most necessary means to reduce costs and improve service quality.

[0015] The present invention has been made in view of the above points, and an object of the present invention is to provide a technique for realizing a delivery plan under time frame constraints and time cost constraints by solving a delivery plan problem that takes into account time frame constraints and time cost constraints.

Means for Solving the Problems

[0016] According to the disclosed technology, an algorithm calculation unit is provided that solves a delivery plan problem of determining a route for providing services to a plurality of customers by a vehicle departing from a service center, using a neural network that performs reinforcement learning by an actor-critic method. The algorithm calculation unit solves the delivery planning problem with, as constraints, a time frame indicating the range of time when a customer should arrive and a time cost indicating the length of time required for service provision to the customer. A delivery planning device, The algorithm calculation unit masks customers who do not satisfy the time frame constraint with respect to the probability distribution of customers obtained using a decoder in the neural network A delivery planning device is provided.

Advantages of the Invention

[0017] According to the disclosed technology, by solving the delivery planning problem considering the constraints of the time frame and the time cost, a technology for realizing a delivery plan under the constraints of the time frame and the time cost is provided.

Brief Description of the Drawings

[0018]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Modes for Carrying Out the Invention

[0019] Hereinafter, embodiments of the present invention (the present embodiments) will be described with reference to the drawings. The embodiments described below are merely examples, and the embodiments to which the present invention is applied are not limited to the following embodiments.

[0020] (Outline of the Embodiment) First, the outline of the present embodiment will be described. In the present embodiment, a new VRP called VRPTWTC, which is a very practical problem formulation in a business scenario, is introduced.

[0021] In this embodiment, in problem formulation, in addition to the existing constraints in VRP such as demand and load in the optimization process, two new constraints (time window and time cost) are introduced. Note that in this embodiment, "load" is assumed to be "cargo", "loaded goods", etc. mounted on the service vehicle, and "load" may be replaced with "cargo", "loaded goods", etc.

[0022] In this embodiment, in order to solve VRPTWTC, a data-driven, end-to-end policy-based reinforcement learning framework is used. The policy-based reinforcement learning framework includes two neural networks, an actor network and a critic network. The actor network generates the route of VRPTWTC, and the critic network estimates and evaluates the value function.

[0023] Also, in this embodiment, a new masking algorithm combined with the actor network is used. By the masking algorithm, the problem can be solved under the constraints of the time window constraint and the time cost constraint formulated in this embodiment together with the constraints in the conventional VRP.

[0024] Also, in this embodiment, by using the API of the map application based on the implementation map, the route can be calculated under the actual road connection conditions, and the adoptability in the actual industry can be increased.

[0025] (Example of device configuration) FIG. 1 shows a configuration diagram of the delivery planning device 100 in this embodiment. As shown in FIG. 1, the delivery planning device 100 includes a user information collection unit 110, a service vehicle information collection unit 120, an algorithm calculation unit 130, a map API unit 140, and a vehicle allocation unit 150.

[0026] The delivery planning device 100 may be implemented by one device (computer) or by a plurality of devices. For example, the algorithm calculation unit 130 may be implemented by a certain computer, and the other functional units may be implemented by another computer. The general operation of the delivery planning device 100 is as follows.

[0027] The user information collection unit 110 acquires the feature amounts of each user (customer). The feature amounts of each user include, for example, the designated time window of each user, the time cost of the service, etc.

[0028] The service vehicle information collection unit 120 collects the feature amounts of each service vehicle. The feature amounts of each service vehicle include, for example, the departure position of each service vehicle, etc.

[0029] The algorithm calculation unit 130 outputs a delivery plan by solving the VRP problem based on the information of each user (customer) and each service vehicle. Details of the algorithm calculation unit 130 will be described later.

[0030] The map API unit 140 performs a route search based on the information of the delivery plan output from the algorithm calculation unit 130, and draws, for example, the route of the delivery plan of each service vehicle on the map. The vehicle dispatching unit 150 distributes the service route information to each service vehicle (or the terminal of the service center) via the network based on the output result of the map API unit 140. Note that the vehicle dispatching unit 150 may be referred to as an "output unit".

[0031] The map API unit 140 may perform a route search, etc. by accessing, for example, an external map server. Also, the map API unit 140 itself may store a map database and perform a route search using the map database.

[0032] As an example, assume that a delivery plan of "0→2→3→0" is obtained by the algorithm calculation unit 130. Here, 0 indicates the service center, and 2 and 3 indicate the customer numbers respectively. In this case, the map API unit 140 draws the actual road route of "service center→customer 2→customer 3→service center" on the map, and the vehicle allocation unit 150 outputs the map information with the route drawn on it.

[0033] (Configuration example of the algorithm calculation unit 120) Fig. 2 shows a configuration example of the algorithm calculation unit 130. The algorithm calculation unit 130 is a neural network model that performs actor-critic method-based reinforcement learning. This model may also be referred to as a VRPTWTC model.

[0034] As shown in Fig. 2, this model includes two neural networks: an actor network 131 and a critic network 132.

[0035] The actor network 131 has a Dense embedding layer (one layer), an LSTM cell, an Attention layer, a Softmax calculation unit (Softmax), and a masking unit (Masking). These components form an encoder-decoder configuration and a pointer network. The critic network 132 has a Dense embedding layer (three layers).

[0036] The Dense embedding layer, LSTM cell, Attention layer in the actor network 131, and the Dense embedding layer in the critic network 132 have learnable parameters in the neural network.

[0037] In the actor network 131, the feature obtained by the LSTM cell from the hidden state, which is the output from the Dense embedding layer corresponding to the encoder, is input to the Attention layer, and the context is obtained from the output from the Dense embedding layer and the output from the Attention layer. The value calculated by the Softmax from the context is output through Masking and used for reward calculation. In the critic network 132, based on the feature obtained by the Dense embedding layer from the input data and the reward, a loss (Loss function) is obtained, and learning is performed to minimize the loss.

[0038] Note that the arrow lines between the input hidden state, context, Attention layer, LSTM, and Softmax indicate an attention-based pointer network. The loss function is calculated by the reward function and the critic network.

[0039] The algorithm calculation unit 130 is configured to use the neural network shown in FIG. 2 to learn a large amount of simulation learning data and be able to perform tests (delivery plan generation) on both actual data and simulation data.

[0040] Specifically, by using a reinforcement learning model based on actor-critic and a Masking algorithm, it is possible to efficiently calculate and output a delivery plan under the constraints that the service vehicle always arrives at the specified time of the customer (within the specified time frame) and each service vehicle always works within 8 hours a day.

[0041] Hereinafter, the processing content of the algorithm calculation unit 130 will be described in more detail.

[0042] (Overview of the processing of the algorithm calculation unit 130) First, the overview of the VRP problem solved by the algorithm calculation unit 130 will be described. This problem has the following three elements. Note that in this specification, "customer" may also be referred to as "user".

[0043] (1) Provide services to all customers, and the time of service provision (the time when the service vehicle arrives) must be within the time frame (time window) specified by each customer.

[0044] (2) Each customer has a time cost for the service that varies according to the service. This "time cost" is the time required for service provision at the customer's residence. Service at the customer's residence is, for example, the repair of communication facilities.

[0045] (3) When providing services to multiple customers, the service vehicle cannot exceed the total service time limit.

[0046] The above problem is called the "Vehicle Routing Problem with Time Windows and Time Costs" (VRPTWTC).

[0047] In this embodiment, a neural network corresponding to the algorithm calculation unit 130 solves the VRPTWTC in an end-to-end, data-driven manner based on actor-critic based deep reinforcement learning, and outputs a solution (delivery plan) to the above problem.

[0048] The features of the algorithm calculation unit 130 in this embodiment are as follows.

[0049] First, different from the conventional method, there is no need to define elements of a handmade model such as an objective function or initial search conditions, and the solution of VRPTWTC in a medium-scale dataset (up to 100 customers) can be optimized in a very short processing time (less than 10 seconds). This can not only reduce the operation cost of actual business applications, but also make it easier to deploy this method in the actual industry.

[0050] Second, the time frame specified by the customer and the time cost of the service are strictly considered in the optimization process. Violations of the time frame or the total labor time limit are not allowed, which also helps to improve the service quality and protect the rights of the staff.

[0051] Finally, different from other conventional VRP solutions, in this embodiment, the effectiveness of the algorithm is evaluated using the actual map application programming interface (API). For example, it is evaluated whether the service vehicle arrives within the specified time frame. This improves the applicability of the proposed method in the actual industry.

[0052] (Details of the processing of the algorithm calculation unit 130) Hereinafter, the processing content of the algorithm calculation unit 130 will be described in detail.

[0053] <A: Problem setting> Referring to FIG. 3, the problem setting in this embodiment will be described. The set χ = {x 1 , x 2 , ··· x N} of customers is located within a certain range on the map, and each x n in the set is a customer who needs a service. There is also a service center for loading the load for service provision. It is assumed that the positions of the customers and the service center are known. Also, the travel time of the service vehicle between the customer and the service center and between any customers may be known (for example, calculated from a predetermined speed and distance), or may be calculated in consideration of the actual road conditions (traffic jams, etc.) from the map API.

[0054] First, the set of service vehicles is arranged at the service center. Each service vehicle can leave the service center and provide services to the set χ of customers. Each customer is served only once by one of the service vehicles. After the service vehicle visits all the planned customers, it returns to the service center.

[0055] Since each customer in χ has four characteristics, each customer x n is represented as a vector, x n =[x n f1 ,x n f2 ,x n f3 ,x n f4 .x n f1 is the address of the nth customer. x n f2 is the demand of the nth customer, which is the same as the demand characteristic of the classical VRP problem. x n f3 is the time window specified by the nth customer, which means that the customer needs to be visited by the service vehicle within that time window. x n f4 is the time cost of serving the nth customer, which indicates how much time is required to serve the nth customer. For the sake of simplicity in modeling, the service center is regarded as the 0th customer in the problem formulation.

[0056] Here, in this problem, violations of the time window (inability to serve the customer within the time window) and violations of the time cost (the service vehicle working more than 8 hours a day) are not allowed.

[0057] For each service vehicle, define the characteristic of a fixed initial load that indicates the maximum load capacity of the service vehicle providing the service. Specifically, before the service vehicle leaves the service center and provides service to the customer, initialize the load with a value of 1 (adjustable according to the task).

[0058] Also, set the maximum service time for each service vehicle to 8 hours. This means that each service vehicle has a maximum service time of 8 hours. That is, the longest time for the service vehicle to leave the service center and provide service is not allowed to exceed 8 hours (this can be adjusted according to the actual business requirements).

[0059] As conditions for the service vehicle to provide services, the following two conditions (1) and (2) are defined. The service vehicle must return to the service center in the following cases (1) or (2).

[0060] Condition (1) When the load of the service vehicle is close to 0 and the capacity to provide services to the remaining customers ( remaining load) is insufficient Condition (2) When the service time of the service vehicle is close to the maximum of 8 hours Under the above customer information and constraints for optimization, find a solution ζ for VRPTWTC. The solution ζ is a sequence of customers in χ that can be interpreted as the service route or the order of services. For example, when a sequence of ζ = {0, 3, 2, 0, 4, 1, 0} is obtained as a solution, this sequence corresponds to two routes. One is the route that proceeds along 0 → 3 → 2 → 0, and the other is the route that proceeds along 0 → 4 → 1 → 0, which implicitly indicates that two service vehicles are used. Also, this can be interpreted as a case where a certain service vehicle returns to the service center once.

[0061] <B: Pointer Network in Actor Network 131> The solution ζ of VRPTWTC is a Markov decision process (MDP) of a sequence, which is a process of selecting the next action in the sequence (that is, which customer node to target for service next).

[0062] In this embodiment, a PointerNet (Pointer Network) is used for the formulation of the MDP process. Note that the PointerNet itself is an existing technology. First, an encoder with a Dense layer embeds the features of all input customers and depots (service centers) to extract hidden states. Subsequently, the decoder restores the actions of the MDP by using LSTM (Long Short-Term Memory) cells connected one by one and passes them to the Attention layer. Each LSTM cell (action) outputs a pointer representing the probability that the input customer node receives the service.

[0063] The key difference between the technology disclosed in Non-Patent Document 1 and the technology according to this embodiment is that, in this embodiment, a new masking algorithm is designed and incorporated into the actor network 131 to obtain a solution under the constraints of time frames, time costs, and total time limits.

[0064] The Dense embedding layer (encoder) of the actor network and the PointerNet will be described more specifically.

[0065] As described above, each x in χ = {x 1 , x 2 , ···, x N} represents a customer (customer node), and each x n is embedded as a dense representation x n by the encoder as shown in Equation (1). n-dense

[0066]

Equation

[0067] The decoder includes a sequence of LSTM cells. In the decoder, the sequence of LSTM cells is used to model the actions in the MDP. At each step m ∈ (1, 2, …, M) of the decoder part, the hidden state in the LSTM cell with weight θ LSTM is represented by d m . M is the total number of decoder steps.

[0068] In this embodiment, similar to PointerNet, the service order is modeled by calculating the pointer D m . That is, at each step m of the decoder part, to determine which member of χ = {x 1 , x 2 , ···, x N} is pointed to, the Softmax result is calculated.

[0069] Here, p(D m |D 1 , D 2 ··· D m-1 , χ; θ) is modeled by the following equations (2) and (3) using the LSTM cell with parameter θ Pointer .

[0070]

Equation

[0071]

Equation

[0072] The final output of the actor network 131 is the service route ζ, which corresponds to the output of the sequence of all m LSTM cells. Here, multiple LSTMs can be interpreted as an MDP. p(D m │D 1 ,D 2 ···D m-1 ,χ;θ) is abbreviated as p(D m ).

[0073] <C: Masking> As described above, in this embodiment, a new masking algorithm is proposed and combined with the actor network 131 to optimize VRPTWTC. In the masking algorithm, there are three sub-maskings: load-demand masking, time-window masking, and time-cost masking.

[0074] Load-demand masking is used to solve the conventional VRP constraints. Time-window masking and time-cost masking are used to optimize the new constraints formulated in VRPTWTC.

[0075] The masking algorithm is combined with the actor network 131 to output the probability of an action in reinforcement learning. First, each of these three sub-maskings will be explained, and then the method of combining it in the actor network will be explained. Note that both (2) and (3) may be implemented, or either one may be implemented.

[0076] (1) Load-demand sub-masking: Both the service capacity of the service vehicle and the demand of the customers are finite. Since they are limited, when there is no remaining load on the service vehicle, the service vehicle must return to the service center for replenishment.

[0077] Here, load-demand sub-masking is used to model this process. At each decoder step m ∈ (1, 2... M), the remaining demand δ at each customer ∈ (1, 2... N) n,mand the remaining vehicle load Δ m are tracked simultaneously. When m = 1, these are δ n,m = δ n , Δ m = 1 and are initialized, and then updated as follows. Note that π m is the index of the customer selected as the service target at decoder step m.

[0078]

Number

[0079]

Number

[0080] Equation (5) shows that when m + 1, when the service vehicle returns to the service center, the load of the vehicle becomes 1 (the value to be replenished), and otherwise, the load of the vehicle becomes the value obtained by subtracting the demand of the customer to be served from the load at m (when the customer is served by the vehicle). Note that in the formulation of this problem, since the service center is the 0th customer, π m = 0 indicates that the service vehicle has returned to the service center.

[0081] (2) Time window submasking: In the problem setting of this embodiment, since the service vehicle must arrive at each customer at the specified time (within the specified time window), in each step of the decoder, time window submasking is added to set the probability of customers who are unlikely to arrive at the specified time to 0. Thus, setting the probability of a customer to 0 may be called masking or filtering.

[0082] As described above, Equation (3) shows that the pointer (Softmax) normalizes the vector u m to the output probability distribution p(D m ) for all input customers χ. Here, p(D m ) is an n-dimensional vector, representing the probability distribution over all of χ at step m of the decoder.

[0083] At each step m of the decoder, the set of customers that need to be served is denoted by χ´ ∈ χ. The reason for using such a set is that some customers have been served before step m or the service vehicle does not have a sufficient load.

[0084] For the set of customers χ´ with the number of customers N´, the submasking τ n´,m of the time frame is calculated by repeating the following process for each customer n´ in χ´:

[0085]

Equation

[0086]

Number

[0087] Figure 4 shows the masking processing algorithm (Algorithm 1). This is the processing executed by the Masking (masking section) in Figure 2. In the first line, for each customer n ∈ (1, 2... N), the demand x n f2 of the customer, the vehicle capacity Δ 0 , the time frame x n f3 , the time cost x n f4 are input, and t total is initialized to 0.

[0088] The second line means that in each decoder step m = 1, 2.... M, steps 3 to 13 are repeated. In the third line, in step m, if the remaining demand δ n,m = 0 for all customers n ∈ (1, 2... N), the loop processing ends.

[0089] In the fourth line, for each customer n ∈ (1, 2... N), if δ n,m > 0 and δ n,m < Δ m , then msk n,m = 1, otherwise msk n,m = 0. msk n,m = 1 indicates that it is serviceable, and msk n,m = 0 indicates that it has been serviced or the vehicle capacity is insufficient, indicating that the customer is not targeted for service (setting the probability to 0).

[0090] In the 5th line, sort the N members of the vector p(D m ) in descending order, and let p with the sorted index i(1, 2…N) be p sort (D m ).

[0091] In the 6th to 7th lines, for each i-th member p sort (D m ) in p sort,i (D m ), filter (mask) the customers based on Equation (6) (time frame sub-masking).

[0092] In the 8th line, set Softmax(p sort,i (D m )) as the probability of the new action pointer. In the 9th line, perform the check of time cost masking according to Equation (7).

[0093] In the 10th line, update the residual demand δ n,m according to Equation (4). In the 11th line, update the residual load according to Equation (5). In the 12th line, update m to m + 1. In the 13th line, if n is not 0, then t total = t total + t move + x n f4 is set. This means that the total operating time from the service center to the completion of service for a certain customer, plus the travel time from this customer to the next customer and the time cost at the next customer, is set as the total operating time at the next customer. End the process in the 14th line.

[0094] As shown in Figure 4, three sub-maskings are introduced in the masking algorithm shown in Algorithm 1. After data input and initialization, at each step m of the LSTM-based decoder, first, for each demand of the customers, if all demands are 0, that is, if all customers have received service, the decoder loop ends.

[0095] Otherwise, mask all customers with non-zero demand values as 1. Note that the demand value needs to be smaller than the dynamic load of the vehicle.

[0096] Next, sort the members of the vector p(D m ) which is the probability of the pointers generated by the action network 131 in descending order, and set it as p sort (D m ). Then, using Equation (6), considering the time frame and the total time cost of the current service route, filter out customers that cannot be served and set it as p sort,i (D m ), and normalize p sort,i (D m ) using Softmax.

[0097] Furthermore, using Equation (7), check whether the total time cost t total exceeds 8 hours. If it does, return the service vehicle to the service center (customer 0). Finally, update the dynamic demand δ n,m , the dynamic load Δ m , and the total time cost t total , and proceed to the next decoder step m + 1.

[0098] <D: Actor-Critic> In this embodiment, to simultaneously learn both the policy and the value function, deep reinforcement learning based on actor-critic is used. Note that deep reinforcement learning based on actor-critic itself is an existing technology.

[0099] Regarding the actor network 131, as described in A, it has learnable weights θ actor ={θ embedded ,θ LSTM ,θ Pointer}.

[0100] In this embodiment, the pointer parameters θ Pointer ={ν, W 1 , W 2} in the actor network 131 and the LSTM parameters θLSTM We use LSTM to parameterize the stochastic policy π. The stochastic policy π generates a probability distribution over the next action (which customer to visit) at any given decoder step.

[0101] On the other hand, the critic network 132 with learnable parameters θ critic estimates the gradient for any problem instance from a given state in reinforcement learning.

[0102] The critic network 132 consists of three Dense layers, takes static and dynamic states as inputs, and predicts the reward. In this embodiment, by using the output probability of the actor network 131 as a weight and calculating the weighted sum of the embedded inputs (outputs from the Dense layers), a single value is output. This can be interpreted as the output of the value function predicted by the critic network 131.

[0103] Figure 5 shows the actor-critic algorithm (Algorithm 2).

[0104] In the first line, initialize the actor network (Embedding2Seq with PN) with random weights θ actor ={θ embedded , θ LSTM , θ Pointer}, and initialize the critic network with random weights θ critic . Lines 2 and 17 mean repeating lines 3 - 16 for each epoch.

[0105] In the third line, reset the parameter gradients dθ actor and dθ critic to 0 respectively. In the fourth line, sample B instances according to the actor network with the current θ actor . Lines 5 and 14 mean repeating lines 6 - 13 for each sample in B.

[0106] In the 6th row, based on the current θ embedded , perform the processing of the embedding layer to obtain x n-dense (batch). The 7th and 12th rows mean repeating the 8th to 11th rows at each decoder step m ∈ (1, 2, …, M). The 8th row means repeating the 9th to 11th rows as long as the end condition is satisfied.

[0107] In the 9th row, based on the distribution p(D m ), calculate D based on a stochastic decoder m . D m indicates the customer to be served (the destination) at the m-th step.

[0108] In the 10th row, observe the new state sequence D1, …, D m-1 , D m . In the 11th row, update m to m + 1.

[0109] In the 13th row, calculate the reward R. In the 15th row, calculate the policy gradient ∇θ actor by Equation (8) and update θ actor . In the 16th row, calculate the gradient ∇θ critic and update θ critic .

[0110] Actor-Critic Algorithm 2 in the present embodiment shown in FIG. 5 shows the training process. After this training process, it may be possible to perform testing (output of the actual delivery plan), or it may be possible to perform testing while continuing the training.

[0111] As already explained, use two neural networks (actor network and critic network) with weight vectors θ actor and θ critic . θ actor includes θ embedded , θ LSTM , θ Pointer .

[0112] The current weights θ of the actor network actor In each iteration of learning with, B samples are obtained and a sequence that can be realized based on the current policy is generated using Monte Carlo simulation. This means that at each step of the decoder, based on the distribution p(D m ) which is the output of the actor network, the pointer D m is calculated probabilistically.

[0113] When sampling is complete, the reward and the policy gradient are calculated and the actor network is updated in line 15. In this step, V(D m ; θ critic ) is the value function approximated from the critic network.

[0114] Also, in line 16, the critic network is updated in a direction to reduce the difference between the observed reward and the expected reward. Finally, in an end-to-end manner with the same learning rate, using the gradients dθ actor and the gradient dθ critic θ actor and θ critic are updated. The policy gradient and reward are described below.

[0115] (1) Policy gradient: In line 15 of Algorithm 2, the policy gradient of the actor network is approximated by Monte Carlo sampling as follows:

[0116] [Equation] Here, R is the reward of the path instance, which is the reward for the sequence of D m indicating the service path. V(χ; θ critic ) is the value function that predicts the reward for all raw inputs. "R - V(χ; θ critic)」 is used as an advantage function that replaces the cumulative reward of the conventional VRP method based on reinforcement learning. In the actor-critic, the technique of using the advantage function itself is an existing technology.

[0117] 2) Reward: In this embodiment, a reward function based on the length of the tour (total route) is used in the same way as the existing technology. A penalty term that adds a penalty value when a time frame is violated may be included. Note that using the length of the tour is an example, and a reward function other than the length may be used.

[0118] (Hardware configuration example) The delivery planning device 100 can be realized, for example, by causing a computer to execute a program. This computer may be a physical computer or a virtual machine on the cloud.

[0119] That is, the delivery planning device 100 can be realized by using hardware resources such as a CPU and a memory built in the computer to execute a program corresponding to the processing performed by the delivery planning device. The above program can be recorded on a computer-readable recording medium (such as a portable memory), stored, distributed, or provided through a network such as the Internet or e-mail.

[0120] FIG. 6 is a diagram showing a hardware configuration example of the above computer. The computer in FIG. 6 includes a drive device 1000, an auxiliary storage device 1002, a memory device 1003, a CPU 1004, an interface device 1005, a display device 1006, an input device 1007, an output device 1008, etc., which are mutually connected by a bus BS.

[0121] A program for realizing the processing on the computer is provided, for example, by a recording medium 1001 such as a CD-ROM or a memory card. When the recording medium 1001 storing the program is set in the drive device 1000, the program is installed from the recording medium 1001 via the drive device 1000 into the auxiliary storage device 1002. However, the installation of the program does not necessarily have to be performed from the recording medium 1001, and it may be downloaded from another computer via a network. The auxiliary storage device 1002 stores the installed program and also stores necessary files, data, etc.

[0122] When an instruction to start the program is given, the memory device 1003 reads out and stores the program from the auxiliary storage device 1002. The CPU 1004 realizes the functions related to the light touch maintenance device 100 according to the program stored in the memory device 1003. The interface device 1005 is used as an interface for connecting to a network or the like. The display device 1006 displays a GUI (Graphical User Interface) or the like according to the program. The input device 1007 is composed of a keyboard, a mouse, buttons, or a touch panel, etc., and is used to input various operation instructions. The output device 1008 outputs the calculation result.

[0123] (Effects of the Embodiment) As described above, the technology according to the present embodiment has the following effects shown in (1), (2), and (3).

[0124] (1) Compared with creating a vehicle allocation plan for service vehicles manually as in the prior art, the calculation time of the vehicle allocation plan can be significantly reduced. That is, for the NP-hard VRP problem, the larger the number of customers, the more enormous the calculation amount, so it is difficult to calculate manually. Even when there are 50 to 100 customers that cannot be handled by the conventional OR-based method, the technology according to the present embodiment enables calculation within 1 second.

[0125] (2)In the VRP problem, it is possible to optimize the route considering the restrictions of arriving at the customer's specified time and the working hours of each service vehicle within 8 hours a day.

[0126] (3)By utilizing the map API, it is possible to calculate the actual travel route and travel time, and also output the image of the route. Therefore, more accurate experiments and an easy-to-understand vehicle allocation plan can be output.

[0127] (Summary of the Embodiment) This specification discloses at least a delivery planning apparatus, a delivery planning method, and a program according to the following respective items. (Item 1) An algorithm calculation unit that solves a delivery planning problem of determining a route for providing services to a plurality of customers by a vehicle departing from a service center, using a neural network that performs reinforcement learning by the actor-critic method, The algorithm calculation unit solves the delivery planning problem with a time frame indicating the range of time when the customer should be arrived at and a time cost indicating the length of time required for service provision to the customer as constraints. A delivery planning apparatus. (Item 2) The algorithm calculation unit performs masking on the probability distribution of customers obtained using the decoder in the neural network for customers who do not satisfy the time frame constraint. The delivery planning apparatus according to Item 1. (Item 3) When a value based on the total operating time of the vehicle exceeds a threshold value, the algorithm calculation unit performs masking on the probability distribution of customers obtained using the decoder in the neural network so that the vehicle returns to the service center. The delivery planning apparatus according to Item 1 or Item 2. (Item 4) The algorithm calculation unit sets, as the total operation time for the next customer, the value obtained by adding the travel time from the current customer to the next customer and the time cost for the next customer to the total operation time from the service center until service completion for a certain customer. The delivery planning device according to claim 3. (Claim 5) A map API unit that draws, on a map, the route of visits to each customer, which is the delivery plan calculated by the algorithm calculation unit. The delivery planning device according to any one of claims 1 to 4, further comprising the map API unit. (Claim 6) A delivery planning method executed by a delivery planning device, comprising an algorithm calculation step of solving a delivery planning problem of determining a route for providing services to a plurality of customers by a vehicle departing from a service center, using a neural network of reinforcement learning by the actor-critic method. In the algorithm calculation step, the delivery planning problem is solved with a time frame indicating the range of time when arriving at a customer and a time cost indicating the length of time required for service provision at the customer as constraints. Delivery planning method. (Claim 7) A program for causing a computer to function as each unit in the delivery planning device according to any one of claims 1 to 5.

[0128] As described above, the present embodiment has been explained, but the present invention is not limited to such a specific embodiment, and various modifications and changes are possible within the scope of the gist of the present invention described in the claims.

Explanation of Signs

[0129] 100 Delivery planning device 110 User information collection unit 120 Service vehicle information collection unit 130 Algorithm calculation unit 140 Map API unit 150 Vehicle allocation unit 1000 Drive device 1001 Recording medium 1002 Auxiliary storage device 1003 Memory device 1004 CPU 1005 Interface device 1006 Display device 1007 Input device 1008 Output device

Claims

1. An algorithm calculation unit that solves a delivery planning problem of determining a route for providing services to a plurality of customers by a vehicle departing from a service center, using a neural network that performs reinforcement learning by an actor-critic method, wherein the algorithm calculation unit is a delivery planning device that solves the delivery planning problem, with a time frame indicating a range of time when the customer should arrive and a time cost indicating a length of time required for service provision to the customer as constraints, wherein the algorithm calculation unit masks customers who do not satisfy the time frame constraint with respect to the probability distribution of customers obtained using a decoder in the neural network, A delivery planning device.

2. An algorithm calculation unit that solves a delivery planning problem of determining a route for providing services to a plurality of customers by a vehicle departing from a service center, using a neural network that performs reinforcement learning by an actor-critic method, wherein the algorithm calculation unit is a delivery planning device that solves the delivery planning problem, with a time frame indicating a range of time when the customer should arrive and a time cost indicating a length of time required for service provision to the customer as constraints, wherein the algorithm calculation unit masks the probability distribution of customers obtained using a decoder in the neural network so that the vehicle returns to the service center when a value based on the total operating time of the vehicle exceeds a threshold value, A delivery planning device.

3. The algorithm calculation unit sets, as the total operating time at the next customer, the value obtained by adding the moving time from the current customer to the next customer and the time cost at the next customer to the total operating time from the service center to the completion of service at a certain customer. The delivery planning device according to claim 2.

4. A map API unit that draws, on a map, the route of visits to each customer, which is the delivery plan calculated by the algorithm calculation unit The delivery planning device according to any one of claims 1 to 3, further comprising.

5. A delivery planning method executed by a delivery planning device, comprising an algorithm calculation step of solving a delivery planning problem of determining a route for providing services to a plurality of customers by a vehicle departing from a service center, using a neural network of reinforcement learning by an actor-critic method, In the algorithm calculation step, a delivery planning method for solving the delivery planning problem, with a time frame indicating the range of time when a customer should arrive and a time cost indicating the length of time required for service provision to the customer as constraints, In the algorithm calculation step, masking is performed on the probability distribution of customers obtained using the decoder in the neural network for customers who do not satisfy the constraint of the time frame. Delivery planning method.

6. A delivery planning method executed by a delivery planning device, Comprising an algorithm calculation step for solving a delivery planning problem of determining a route for providing services to a plurality of customers by a vehicle departing from a service center using a neural network of reinforcement learning by the actor-critic method, In the algorithm calculation step, a delivery planning method for solving the delivery planning problem, with a time frame indicating the range of time when a customer should arrive and a time cost indicating the length of time required for service provision to the customer as constraints, In the algorithm calculation step, when a value based on the total operating time of the vehicle exceeds a threshold value, masking is performed on the probability distribution of customers obtained using the decoder in the neural network so that the vehicle returns to the service center. Delivery planning method.

7. A program for causing a computer to function as each part in the delivery planning device according to any one of Claims 1 to 4.

Citation Information

Patent Citations

  • Method / System for planning car allocation

    JP1995234997A

  • Device and method for deciding optimum delivery route and delivery vehicle and medium recording program for deciding optimum delivery route and delivery vehicle

    JP1998134300A

  • Delivery management device and delivery management method and delivery management system

    JP2018147108A

  • Route planning method and route planning device

    JP2019114258A

  • Information processing device and information processing program

    JP2020030663A