Delivery planning device, delivery planning method, and program

A neural network-based algorithm with multiple actor networks addresses the inefficiencies of traditional VRP methods by considering vehicle differences, providing real-time optimal route planning for large-scale delivery scenarios.

JP7700962B2Active Publication Date: 2025-07-01NIPPON TELEGRAPH & TELEPHONE CORP
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
JP2024516004
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2022-04-21
Publication Date
2025-07-01
Estimated Expiration
2042-04-21

AI Technical Summary

Technical Problem

Existing optimization methods for vehicle routing problems (VRP) struggle with large-scale instances, taking days or years to calculate optimal solutions and require different search models for variations, failing to consider differences between multiple delivery vehicles.

Method used

A neural network-based algorithm using an actor-critic method with multiple actor networks for each delivery vehicle, considering their unique states and environments to determine optimal routes through reinforcement learning.

Benefits of technology

Enables real-time generation of superior solutions for large-scale VRP problems by considering vehicle-specific characteristics, improving efficiency and reducing calculation times.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007700962000001
    Figure 0007700962000001
  • Figure 0007700962000002
    Figure 0007700962000002
  • Figure 0007700962000003
    Figure 0007700962000003
Patent Text Reader

Abstract

A delivery planning device according to the present invention comprises an algorithm calculation unit that uses a neural network that learns by actor-critic reinforcement learning to solve a vehicle routing problem that determines routes for service provision by a plurality of mobile bodies through a plurality of nodes. The algorithm calculation unit comprises a plurality of actor networks that correspond to the plurality of mobile bodies. Each actor network determines routes on the basis of the state of a particular mobile body and the states of the plurality of nodes.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a technique for solving the distribution planning problem.

Background Art

[0002] The distribution planning problem (VRP: vehicle routing problem) is an optimization problem that considers which delivery vehicle should visit which customer in which order (to minimize costs) when delivering goods to each customer using delivery vehicles from a goods collection point (service center). Note that the "distribution planning problem" may also be referred to as the "vehicle scheduling problem".

[0003] In actual applications, there are many practical business scenarios where distribution and service costs can be optimized through the solution of VRP, such as route generation for drones in e-commerce, just-in-time delivery in e-commerce, cold chain delivery, and store replenishment.

[0004] Therefore, various variations of VRP have been proposed according to different practical requirements. As a variation of VRP, for example, there is the VRP with time windows (VRPTW). In VRPTW, a time window for delivering goods to customers is set. Another VRP is the multi-depot distribution planning problem (MDVRP). In MDVRP, there are multiple depots (service centers) from which vehicles can depart or where the travel can end.

[0005] Since VRP and its variations have been proven to be NP-hard problems, various operations research (OR)-based methods that return approximate solutions have been studied for many years.

[0006] Usually, in OR-based algorithms, a search model is defined manually, and the solution of VRP is obtained by sacrificing the quality of the solution in order to improve efficiency. However, the conventional OR-based methods have two drawbacks.

[0007] As a first drawback, in the case of a practical-scale VRP problem (having 100 or more customers), when using an OR-based algorithm, it takes several days or years for the calculation to obtain an optimal solution or an approximate solution.

[0008] As a second drawback, different variations of VRP require different handmade search models and initial search conditions, and thus require different OR algorithms. For example, an inappropriate initial solution may result in a long processing time and a local optimum. In such a case, it is difficult to generalize and use an OR-based algorithm in a real business scenario.

[0009] Non-Patent Document 1 discloses a solution to the VRP based on the actor-critic method of reinforcement learning, which solves the drawbacks of the OR-based algorithm. That is, the neural network model can significantly improve the complexity and representation ability with high precision, especially when the number of customer nodes is large.

Prior Art Documents

Non-Patent Documents

[0010]

Non-Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0011] However, in both OR and reinforcement learning, there is a problem that differences (variations) between multiple delivery vehicles (which may be human) cannot be considered. That is, when exploring the routes of the VRP problem for multiple delivery vehicles, differences between delivery vehicles such as the departure location, loading capacity, and whether service can be provided to each customer (node) cannot be considered in either OR or reinforcement learning.

[0012] The present invention has been made in view of the above points, and an object thereof is to provide a technology that enables solving a delivery planning problem while considering differences between multiple delivery vehicles.

Means for Solving the Problem

[0013] According to the disclosed technology, there is provided an algorithm calculation unit that solves a delivery planning problem of determining a route for providing service to a plurality of nodes by a plurality of moving bodies, using a neural network that performs reinforcement learning by an actor-critic method. The algorithm calculation unit includes a plurality of actor networks corresponding to the plurality of moving bodies, and each actor network determines the route based on the state of a certain moving body and the states of the plurality of nodes. A delivery planning apparatus is provided.

Effect of the Invention

[0014] According to the disclosed technology, a technology that enables solving a delivery planning problem while considering differences between multiple delivery vehicles is provided.

Brief Description of the Drawings

[0015]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Embodiments for Carrying Out the Invention

[0016] Hereinafter, embodiments (these embodiments) of the present invention will be described with reference to the drawings. The embodiments described below are merely examples, and the embodiments to which the present invention is applied are not limited to the following embodiments.

[0017] In these embodiments, in the VRP problem, it is assumed that a delivery vehicle moves to carry goods to nodes (customers), but this is an example. The entity carrying the goods may be a person, or may be something other than a delivery vehicle and a person. The entity carrying the goods may be generically referred to as a "mobile object". Also, in these embodiments, the delivery of goods is taken as an example of "service", but "service" is not limited to the delivery of goods.

[0018] (Overview of the Embodiment) First, the overview of these embodiments will be described. In these embodiments, in order to solve the VRP problem, a data-driven, end-to-end policy-based reinforcement learning framework is used. The policy-based reinforcement learning framework includes a neural network composed of an actor network and a critic network. The actor network generates a delivery route, and the critic network estimates and evaluates a value function.

[0019] In particular, these embodiments are characterized in that, in the algorithm calculation unit described later, a plurality of actor networks corresponding to a plurality of agents are used. Each agent corresponds to each delivery vehicle. Each agent is formulated as a Markov decision process (MDP) for each time step, and actions are taken alternately while considering the states of other agents and the distribution of nodes.

[0020] The above configuration is combined with, for example, PointerNet disclosed in Non-Patent Document 1 and an actor-critic neural network model to form an overall algorithm.

[0021] As described above, by representing each agent with its respective actor network, it is possible to consider the differences in the feature amounts of each agent.

[0022] Also, as described above, by configuring each agent to act alternately (in order) at each time step, agents can act in order, and changes in the state of the environment (nodes) and updates of the agent states can be considered simultaneously. Therefore, it is possible to train a neural network so that better actions can be taken.

[0023] In addition, due to the configuration of the algorithm calculation unit according to the present embodiment, a cooperative multi-agent can be learned end-to-end as a whole, and a result close to the optimum can be obtained. In large-scale instances, the model according to the present embodiment can generate a solution superior to other commonly used heuristic models and can generate a solution in real time.

[0024] (Example of device configuration) FIG. 1 shows a configuration diagram of a delivery plan device 100 according to the present embodiment. As shown in FIG. 1, the delivery plan device 100 includes a node information collection unit 110, a delivery vehicle information collection unit 120, an algorithm calculation unit 130, a map API unit 140, and a vehicle allocation unit 150.

[0025] The delivery plan device 100 may be implemented by one device (computer) or by a plurality of devices. For example, the algorithm calculation unit 130 may be implemented by a certain computer, and the other functional units may be implemented by another computer. The outline of the operation of the delivery plan device 100 is as follows.

[0026] The node information collection unit 110 acquires feature quantities of each node (customer). The feature quantities of each node include, for example, the position of the node, demand, etc.

[0027] The delivery vehicle information collection unit 120 collects feature quantities of each delivery vehicle. The feature quantities of each delivery vehicle include, for example, the departure position of each delivery vehicle, the loading capacity, whether it can provide services to each node, etc. The feature quantities of each delivery vehicle collected by the delivery vehicle information collection unit 120 reflect the differences between the delivery vehicles.

[0028] The algorithm calculation unit 130 outputs a delivery plan by solving the VRP problem based on the information of each node (customer) and each delivery vehicle. Details of the algorithm calculation unit 130 will be described later.

[0029] The map API unit 140 performs a route search based on the information of the delivery plan output from the algorithm calculation unit 130, and draws, for example, the route of the delivery plan of each delivery vehicle on a map. The vehicle allocation unit 150 distributes the service route information to each delivery vehicle (or the terminal of the service center) via the network based on the output result of the map API unit 140. Note that the vehicle allocation unit 150 may be referred to as an "output unit".

[0030] The map API unit 140 may perform a route search, etc. by accessing, for example, an external map server. Alternatively, the map API unit 140 itself may store a map database and perform a route search using the map database.

[0031] As an example, assume that as a delivery plan, for a certain delivery vehicle j, a delivery plan of "0→2→3→0" is obtained by the algorithm calculation unit 130. Here, 0 indicates the node of the service center, and 2 and 3 indicate the numbers of the nodes corresponding to the customers respectively. In this case, the map API unit 140 draws the actual road route of "service center→customer 2→customer 3→service center" on the map, and the vehicle allocation unit 150 outputs the map information with the route drawn.

[0032] (Configuration example of the algorithm calculation unit 120) Fig. 2 shows a configuration example of the algorithm calculation unit 130. The algorithm calculation unit 130 is a model of a neural network that performs actor-critic method-based reinforcement learning.

[0033] As shown in Fig. 2, this model mainly includes two neural networks: an actor network 131 and a critic network 132.

[0034] The actor network 131 in this embodiment has actor networks for a plurality of agents. That is, it has a plurality of actor networks corresponding to a plurality of agents. Each agent corresponds to each delivery vehicle. The example in Fig. 2 shows an example of having actor networks for three agents. The actor network corresponding to agent 1 includes an LSTM cell 1 and an Attention layer 1, the actor network corresponding to agent 2 includes an LSTM cell 2 and an Attention layer 2, and the actor network corresponding to agent 3 includes an LSTM cell 3 and an Attention layer 3.

[0035] Note that there is a one-layer Dense embedding layer (corresponding to an encoder) in the actor network 131 that embeds the input data (Input data) into an embedded representation and inputs it to each LSTM cell. However, in Fig. 2, the Dense embedding layer is not shown for simplicity of notation.

[0036] In addition to each LSTM cell and each Attention layer, the actor network 131 has a Softmax calculation unit (Softmax) and a masking unit (Masking). These constitute an encoder-decoder configuration and a pointer network. The critic network 132 has a Dense embedding layer (three layers).

[0037] The actor network 131 has learnable parameters for each agent. The critic network 132 (Dense embedding layer) also has learnable parameters.

[0038] In the actor network 131, for each time step, the state of the environment and the state of the delivery vehicle are input to each LSTM cell. In each agent, the feature amount obtained by the LSTM cell is input to the Attention layer. The output from each Attention layer is input to the Softmax calculation unit (Softmax), and the value calculated by the Softmax is output through Masking and used for reward calculation. In the critic network 132, based on the feature amount obtained by the Dense embedding layer from the input data and the reward, a loss (Loss function) is obtained, and learning is performed to reduce the loss.

[0039] The algorithm calculation unit 130 is configured to learn a large amount of simulation learning data using the neural network shown in FIG. 2 and be able to perform testing (delivery plan generation) on both actual data and simulation data. In the algorithm calculation unit 130, the learning method itself in the actor-critic method of reinforcement learning is the same as, for example, the learning method disclosed in Non-Patent Document 1.

[0040] Note that the actor network 131 in FIG. 2 constitutes a pointer network (PtrNet). The pointer network has an encoder and a decoder. The encoder reads the input sequence (node distribution), and the decoder determines which input to select using the Attention mechanism.

[0041] Also, in the VRP, since the order of the input data has no meaning, in the configuration shown in FIG. 2, the RNN encoder is omitted and embedding input is used.

[0042] Also, the learning method is the same as the method disclosed in Non-Patent Document 1, and the policy gradient method is used. In the policy gradient method, each actor network predicts the probability distribution of the next action by PtrNet, and the critic network estimates the reward of the problem instance.

[0043] Hereinafter, the processing content of the algorithm calculation unit 130 will be described in more detail.

[0044] (Details of the processing of the algorithm calculation unit 130)

[0045] <Problem setting> Referring to FIG. 3, the problem setting in the present embodiment will be described. A set of nodes (which may be called customers) is located in a certain range on the map. Each black circle in FIG. 3 indicates a node that requires a service (in this embodiment, the delivery of luggage). There is also a service center (a luggage collection point) where luggage (which may be called a load) for service provision is loaded. It is assumed that the positions of the nodes and the service center are known. Also, the travel time of the delivery vehicle between the node and the service center and between any nodes may be known (for example, calculated from a predetermined speed and distance), or may be calculated in consideration of the actual road conditions (such as traffic jams) from the map API.

[0046] First, a set of delivery vehicles is arranged at the service center. In the optimization problem of the present embodiment, when delivering luggage from the service center to each node using a delivery vehicle, it is considered which delivery vehicle visits which node in which order to be optimal (the cost is the lowest). The above cost is, for example, the travel distance of the delivery vehicle. Note that the starting point may be different for each delivery vehicle. Also, for each delivery vehicle, the nodes that can be visited (or the nodes that cannot be visited) may be determined.

[0047] Each node receives service only once by one of the delivery vehicles. After the delivery vehicle visits all the planned nodes, it returns to the service center.

[0048] In FIG. 3, an image is shown where three delivery trucks (corresponding to three agents) exist and each performs delivery on a different delivery route.

[0049] <Regarding input data> In FIG. 2, the input data to the LSTM shown as "state of the environment" and "state of the delivery truck" will be described. Note that the "state of the environment" corresponds to the "state of the node".

[0050] If the state of node i is denoted as x i and described as such, x i is expressed as follows. t indicates each time in the time step.

[0051] x i :{x t i ={s i ,d t i ), t = 0,....,T} The meaning of each symbol is as follows.

[0052] x i : The state of node i s i : The two - dimensional coordinates (address) of node i d t i : The demand at node i at step t (it is assumed that the demand varies for each node) The above - mentioned demand is the same as the demand characteristics in the classical VRP problem.

[0053] If the state of delivery truck j is denoted as y j and described as such, y j is expressed as follows.

[0054] y j :{y t j ={p t j ,l t j ), t = 0,....,T} The meaning of each symbol is as follows.

[0055] y j : State of delivery vehicle j p j : Two-dimensional coordinates of delivery vehicle j (current position updated at each step) t t j : Vehicle load at step t (a value that can vary depending on the problem setting and the capabilities of the delivery person) Regarding the vehicle load, for each delivery vehicle, the characteristics of a fixed initial load indicating the maximum load capacity of the delivery vehicle may be defined. Specifically, as an example, before the delivery vehicle leaves the service center and provides goods to the nodes, the load is initialized with a value of 1 (adjustable according to the task).

[0056] Also, in the calculation, when the load of the delivery vehicle is close to 0 and the capacity (remaining load) to provide services to the remaining nodes is insufficient, conditions such as returning to the service center may be set.

[0057] Under the above input data and conditions, find a solution ζ for the VRP problem. The solution ζ is a sequence (sequence) of nodes that can be interpreted as the service route or the order of services. For example, if a sequence of ζ = {0, 3, 2, 0, 4, 1, 0} is obtained as a solution, this sequence corresponds to two routes. One is the route that proceeds along 0 → 3 → 2 → 0, and the other is the route that proceeds along 0 → 4 → 1 → 0, which can be interpreted as the case where the corresponding delivery vehicle returns to the service center once.

[0058] <LSTM, Attention in Actor Network 131> The solution ζ of the VRP problem is a Markov decision process (MDP) of a sequence (sequence), which is a process of selecting the next action (that is, which node to deliver to next) within the sequence.

[0059] In each agent, the actions of the MDP are restored from the input data by using LSTM (Long Short - term Memory) cells and passed to the Attention layer. The Attention layer and Softmax output a pointer representing the probability that each input node will receive a delivery. The actor network 131 may determine the next action (the next node to visit), for example, as the node with the highest probability among all nodes, or may determine the next action (the next node to visit) by other methods based on probability.

[0060] In the actor network 131, as a result of the above actions, the input data is updated at each time step. The input data at time step t is as follows.

[0061] I node states: x t :{x t i ={s t i ,d t i ), i = 1,....,I} J delivery vehicle states: y t :{y t j ={p t j ,l t j ), j = 1,....,J} At time step t, the I node states are input to each LSTM cell. In the example of Figure 2, the I node states are input to each of LSTM cells 1 - 3. Regarding the delivery vehicle states, for delivery vehicle j, y t j is input to the LSTM cell j of actor network j corresponding to delivery vehicle j. In the example of Figure 2, y t 1 is input to LSTM cell 1, y t 2 is input to LSTM cell 2, and y t 3 is input to LSTM cell 3.

[0062] Based on the following formula for each agent by the Attention layer j and the softmax layer, a t j is calculated.

[0063] u t j = v a j tanh(w a j [x t ; y t j ; h t ) a t j = softmax(u t j ) Here, h t represents the hidden state of the LSTM cell at time step t, and ";" represents the concatenation process. a t j represents the degree of relevance and represents the conditional probability of being selected next. v a j and w a j for each agent are learnable parameters.

[0064] Note that Softmax normalizes the vector u (of length I) t j to the output distribution (probability distribution) for all nodes. That is, a t j = softmax(u t j ) outputs the probability (probability of being selected as the service target) of each node at time step t.

[0065] <Masking>

[0066] Regarding the masking unit (Masking in Fig. 2) of the present embodiment, the masking technique disclosed in Non-Patent Document 1 can be used. For example, when the remaining load capacity of the delivery vehicle is 0, the masking unit masks all nodes (that is, does not deliver anywhere). Further, the masking unit masks, for example, nodes having a demand greater than the current load capacity (load) of the delivery vehicle (that is, does not deliver to the said nodes). Note that the load capacity of the delivery vehicle decreases by the amount of the load each time the load is delivered to a node.

[0067] <Actor-Critic> In the present embodiment, deep reinforcement learning based on Actor-Critic is used to simultaneously learn both the policy and the value function. Note that deep reinforcement learning based on Actor-Critic itself is an existing technique.

[0068] Regarding the actor network 131, the actor network corresponding to each agent j has weights θj. The policy for each agent to select an action is represented by the actor network. The policy π θj generates a probability distribution for the next action (which node to visit) at any time step.

[0069] The policy π for the delivery vehicle (agent) j θj is as follows.

[0070] π θj (a|x t ,y t )=a t j The critic network 132 predicts how good the selected route is by using a three-layer DenseLayer (weights are φ) from the two-dimensional coordinates (address) of each node. An example of the value function for the prediction is as follows.

[0071] V φ(s) = Relu(W3 · Relu(W2 · Relu(W1 · s + b1) + b2) + b3) W1, W2, and W3 are weights respectively. b1, b2, and b3 are predetermined constants. s is the two-dimensional coordinate of each node.

[0072] The gradient of the policy of the actor network 131 (agent j) is calculated based on the following formula. J(π θj ) is the evaluation function of the policy. ∇ θj J(π θj ) is the derivative with respect to θj.

[0073] ∇ θj J(π θj ) = (1 / BT) Σ b=1 B Σ t=0 T ∇ θj (R - V φ (s)) · logπ θj (a b |x t b , y t b ; θ j ) Note that B is the batch size. R is the reward for the route traveled by the deliverer j. The reward is determined such that, for example, the shorter the route length, the higher the reward.

[0074] The gradient of the value of the critic network 132 is calculated based on the following formula. J(π θj ) is the evaluation function of the value. ∇ φ J(V φ ) is the derivative with respect to φ.

[0075] ∇ φ J(V φ ) = (1 / B) Σ b=1 B (R - V φ (s)) 2 Figure 4 shows the actor-critic algorithm (the processing procedure of the algorithm calculation unit). Note that since the weight update process related to reinforcement learning itself is the same as the existing technology described in Non-Patent Document 1 etc., only an outline of the process related to reinforcement learning is shown.

[0076] In the first line, the actor network (each agent) is initialized with random weights θj, and the critic network is initialized with random weights φ. Note that the character immediately following "random weight" in the first line of Figure 4 is described as φ in the specification text. The second and 18th lines mean that the third to 17th lines are repeated for each epoch.

[0077] In the third line, B instances are sampled according to the actor network with the current θj. The fourth and 17th lines mean that for each sample in B, the fifth to 16th lines are repeated.

[0078] In the fifth line, the process of the embedding layer of the actor network is performed to obtain the initial input data (x i , y j ). The sixth and 11th lines mean that for each decoder step t∈(1,2,….T), the seventh to 10th lines are repeated.

[0079] The seventh and 10th lines mean that for each agent j∈(1,2,….J), the eighth to ninth lines are repeated.

[0080] In the eighth line, an action (behavior, that is, the visited node) is selected based on the policy π θj . In the ninth line, based on the result of the action, the current state (x t , y t ) is updated to a new state (x t+1 , y t+1 ). Specifically, for example, for delivery vehicle j, it becomes the position of the selected node as the delivery destination, the load capacity is reduced by the amount of the goods, and the demand at the corresponding node is reduced by the amount of the delivered goods.

[0081] Lines 12 and 15 mean repeating Lines 13 - 14 for each agent j ∈ (1, 2, …, J). In Lines 13 and 14, the policy gradient ∇ θj is calculated to update the weights θj of the actor network corresponding to agent j. In Line 16, the gradient ∇ φ is calculated and the weights φ of the critic network are updated.

[0082] The actor - critic algorithm in the present embodiment shown in FIG. 4 shows the training process. After this training process, it may be possible to perform testing (actual delivery plan output), or it may be possible to perform testing while continuing learning.

[0083] In the algorithm of FIG. 4, as already described, an actor network with weights θj for each agent j and a critic network with weights φ are used.

[0084] In each iteration of learning with the current weights θj of the actor network (agent j), Monte Carlo simulation is used to generate a feasible sequence based on the current policy. This means that at each step of the decoder, a pointer is probabilistically calculated based on the distribution a t j which is the output of the actor network. When sampling is complete, the reward and the policy gradient are calculated, and θj is updated for each j in Lines 12 - 15. Also, in Line 16, the critic network is updated in a direction to reduce the difference between the observed reward and the expected reward.

[0085] (Hardware Configuration Example) The delivery planning device 100 can be realized, for example, by causing a computer to execute a program. This computer may be a physical computer or a virtual machine on the cloud.

[0086] That is, the delivery planning device 100 can be realized by executing a program corresponding to the processing performed by the delivery planning device by using hardware resources such as a CPU and a memory built into a computer. The above program can be recorded on a computer-readable recording medium (such as a portable memory), saved, or distributed. It is also possible to provide the above program through a network such as the Internet or e-mail.

[0087] FIG. 5 is a diagram showing an example of the hardware configuration of the computer. The computer in FIG. 5 includes a drive device 1000, an auxiliary storage device 1002, a memory device 1003, a CPU 1004, an interface device 1005, a display device 1006, an input device 1007, an output device 1008, etc., which are mutually connected by a bus BS.

[0088] A program for realizing the processing on the computer is provided by a recording medium 1001 such as a CD-ROM or a memory card, for example. When the recording medium 1001 storing the program is set in the drive device 1000, the program is installed from the recording medium 1001 to the auxiliary storage device 1002 via the drive device 1000. However, the installation of the program does not necessarily have to be performed from the recording medium 1001, and it may be downloaded from another computer via a network. The auxiliary storage device 1002 stores the installed program and also stores necessary files, data, etc.

[0089] When an instruction to start the program is given, the memory device 1003 reads out and stores the program from the auxiliary storage device 1002. The CPU 1004 realizes the functions related to the delivery planning device 100 according to the program stored in the memory device 1003. Specifically, for example, it executes calculations for updating the states in the procedure shown in FIG. 4, storing the updated data in the memory device 1003, reading out the updated data from the memory device 1003 for the next time step, weight update calculations, etc.

[0090] The interface device 1005 is used as an interface for connecting to a network or the like. The display device 1006 displays a GUI (Graphical User Interface) or the like by a program. The input device 1007 is composed of a keyboard, a mouse, buttons, a touch panel, or the like, and is used to input various operation instructions. The output device 1008 outputs the calculation result.

[0091] (Effect of the Embodiment) As described above, according to the technology of the present embodiment, it is possible to solve the delivery planning problem in consideration of the differences between a plurality of delivery vehicles.

[0092] (Supplementary Note) This specification discloses at least the following delivery planning device, delivery planning method, and program. (Supplementary Note Item 1) An algorithm calculation unit for solving a delivery planning problem for determining a route for providing services to a plurality of nodes by a plurality of moving bodies using a neural network that performs reinforcement learning by an actor-critic method, The algorithm calculation unit includes a plurality of actor networks corresponding to the plurality of moving bodies, and each actor network determines the route based on the state of a certain moving body and the states of the plurality of nodes Delivery planning device. (Supplementary Note Item 2) The state of each moving body includes at least the position and the loading amount, and the state of each node includes at least the position and the demand The delivery planning device according to Supplementary Note Item 1. (Supplementary Note Item 3) The algorithm calculation unit, The process of determining an action for each moving body and updating the state is repeatedly executed for each time step The delivery planning device according to Supplementary Note Item 1. (Supplementary Note Item 4) A delivery planning method executed by a delivery planning device, An algorithm calculation step for solving a delivery planning problem of determining a route for providing services to a plurality of nodes by a plurality of mobile bodies using a neural network that performs reinforcement learning by an actor-critic method, In the algorithm calculation step, each actor network in the plurality of actor networks corresponding to the plurality of mobile bodies determines the route based on the state of a certain mobile body and the states of the plurality of nodes. Delivery planning method. (Additional item 5) A program for causing a computer to function as each part in the delivery planning apparatus according to any one of claims 1 to 3.

[0093] As described above, the present embodiment has been described, but the present invention is not limited to such a specific embodiment, and various modifications and changes are possible within the scope of the gist of the present invention described in the claims.

Explanation of reference numerals

[0094] 100 Delivery planning apparatus 110 User information collection unit 120 Delivery vehicle information collection unit 130 Algorithm calculation unit 140 Map API unit 150 Vehicle allocation unit 1000 Drive device 1001 Recording medium 1002 Auxiliary storage device 1003 Memory device 1004 CPU 1005 Interface device 1006 Display device 1007 Input device 1008 Output device

Claims

1. An algorithm calculation unit that solves a delivery planning problem of determining a route for providing services to a plurality of nodes by a plurality of mobile bodies using a neural network that performs reinforcement learning by an actor-critic method, wherein the algorithm calculation unit includes a plurality of actor networks corresponding to the plurality of mobile bodies, and each actor network determines the route based on the state of a certain mobile body and the states of the plurality of nodes Delivery planning device.

2. The state of each mobile body includes at least a position and a loading capacity, and the state of each node includes at least a position and a demand The delivery planning device according to claim 1.

3. The algorithm calculation unit repeatedly executes a process of determining an action for each mobile body and updating the state at each time step The delivery planning device according to claim 1.

4. A delivery planning method executed by a delivery planning device, including an algorithm calculation step of solving a delivery planning problem of determining a route for providing services to a plurality of nodes by a plurality of mobile bodies using a neural network that performs reinforcement learning by an actor-critic method, wherein in the algorithm calculation step, each actor network in the plurality of actor networks corresponding to the plurality of mobile bodies determines the route based on the state of a certain mobile body and the states of the plurality of nodes Delivery planning method.

5. A program for causing a computer to function as each part in the delivery planning device according to any one of claims 1 to 3.

Citation Information

Patent Citations

  • Multi-AGV motion planning method, device and system

    CN112015174A

  • A method and neural network trained by reinforcement learning to determine a constraint optimal route using a masking function

    EP3916652A1

  • Information processing device and information processing program

    JP2020030663A