Delivery planning device, delivery planning method, and program

A reinforcement learning-based algorithm addresses the challenge of planning efficient delivery routes under time constraints in VRP by using a policy-based model with an actor network and a critique network, and a masking algorithm, resulting in improved solution quality and reduced computational time.

WO2025094277A1PCT designated stage expired Publication Date: 2025-05-08NT T INC
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
PCT/JP2023/039293
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2023-10-31
Publication Date
2025-05-08

AI Technical Summary

Technical Problem

Existing vehicle routing problem (VRP) solutions struggle to efficiently plan deliveries under time constraints, leading to suboptimal solutions and increased computational time, especially for large-scale problems.

Method used

A reinforcement learning-based algorithm that uses a policy-based model with an actor network and a critique network, combined with a masking algorithm, to determine optimal delivery routes while adhering to time windows and time costs, utilizing a reward function that penalizes late arrivals.

Benefits of technology

The solution enables efficient delivery planning under time constraints, significantly reducing computational time and improving solution quality, while ensuring vehicles arrive within specified time frames and operate within designated working hours.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure JP2023039293_08052025_PF_FP_ABST
    Figure JP2023039293_08052025_PF_FP_ABST
Patent Text Reader

Abstract

A delivery planning device according to the present invention comprises an algorithm calculation unit that uses a reinforcement learning model to determine a route for performing service provision to a plurality of nodes by a moving body which departs from a certain location. The algorithm calculation unit is configured to use, as a reward function in the reinforcement learning model, a reward function that generates a penalty if the moving body arrives at a node at a time which exceeds a predetermined time frame.
Need to check novelty before this filing date? Find Prior Art

Description

Delivery planning device, delivery planning method, and program

[0001] The present invention relates to a technique for solving a vehicle delivery planning problem.

[0002] The vehicle routing problem (VRP) is an optimization problem that considers which service vehicles should visit which customers in what order (to minimize cost) when delivering packages from a package collection point (service center) to each customer (each node) using service vehicles. Note that the "vehicle routing problem" may also be called the "vehicle allocation problem." Another problem similar to the VRP is the Traveling Salesman Problem (TSP).

[0003] In real-world applications, there are many practical business scenarios where distribution and service costs can be optimized through VRP solutions, such as drone route generation, just-in-time delivery in e-commerce, cold chain delivery, and store replenishment.

[0004] Therefore, various variations of VRP have been proposed to meet different practical requirements. For example, a variation of VRP is a time-windowed VRP (VRPTW), in which a time window is set for delivery of goods to customers. Another VRP is a multi-depot delivery planning problem (MDVRP), in which there are multiple depots (service centers) from which vehicles can depart and end their trips.

[0005] Since VRP and its variations have been proven to be NP-hard problems, various operations research (OR)-based methods that return approximate solutions have been studied for many years (see, for example, Non-Patent Document 1). However, for practical-scale TRP and VRP problems (having 100 or more nodes), OR-based algorithms require several days of computation to obtain an optimal or approximate solution.

[0006] Non-Patent Document 2 discloses a solution to VRP based on actor-critic reinforcement learning, which overcomes the drawbacks of OR-based algorithms.

[0007] Solving the Traveling Salesman Problem, https: / / blog.routific.com / blog / travelling-salesman-problem#:~:text=The%20Traveling%20Salesman%20Problem%20(TSP,and%20optionally%20an%20ending%20point, Internet, retrieved October 19, 2023. Nazari, Mohammadreza, Afshin Oroojlooy, Lawrence V. Snyder, and Martin Takac, "Reinforcement Learning for Solving the Vehicle Routing Problem", NIPS, 2018.

[0008] However, prior art VRP solving methods based on reinforcement learning were unable to take into account time constraints, such as the requirement that a vehicle must arrive at a node within a specified time frame.

[0009] The present invention has been made in consideration of the above points, and aims to provide a technology for realizing a delivery plan under a time constraint by solving a delivery plan problem taking into account the time constraint.

[0010] According to the disclosed technology, there is provided a delivery planning device that includes an algorithm calculation unit that uses a reinforcement learning model to determine a route for a mobile object departing from a certain location to provide service to a plurality of nodes, and the algorithm calculation unit uses, as a reward function in the reinforcement learning model, a reward function that imposes a penalty if the mobile object arrives at a node at a time that exceeds a predetermined time frame.

[0011] According to the disclosed technology, a technology for realizing a delivery plan under a time constraint is provided by solving a delivery plan problem taking into account the time constraint.

[0012] FIG. 1 is a diagram illustrating the configuration of an apparatus according to a first embodiment. FIG. 2 is a diagram illustrating the configuration of an algorithm calculation unit 130. FIG. 3 is a diagram illustrating problem setting. FIG. 4 is a diagram illustrating algorithm 1 according to the first embodiment. FIG. 5 is a diagram illustrating algorithm 2 according to the first embodiment. FIG. 6 is a diagram illustrating the configuration of an apparatus according to a second embodiment. FIG. 7 is a diagram illustrating the configuration of an algorithm calculation unit 230. FIG. 8 is a diagram illustrating an overview of processing according to the second embodiment. FIG. 9 is a diagram illustrating problem setting. FIG. 10 is a diagram for explaining graph convolution. FIG. 11 is a diagram illustrating processing of an attention mechanism. FIG. 12 is a diagram illustrating an algorithm according to the second embodiment. FIG. 13 is a diagram illustrating an example of the hardware configuration of an apparatus.

[0013] Hereinafter, an embodiment of the present invention (the present embodiment) will be described with reference to the drawings. The embodiment described below is merely an example, and the embodiment to which the present invention is applied is not limited to the following embodiment. Below, a second embodiment of the first embodiment will be described.

[0014] In the first and second embodiments, the delivery of packages is performed by a vehicle, but this is merely an example. Anything that moves (hereinafter referred to as a "mobile body") may perform the delivery. For example, the mobile body may be a person, a drone, a ship, a bicycle, a motorcycle, an airplane, a spaceship, etc.

[0015] Furthermore, the delivery service is not limited to the delivery of parcels, etc. For example, the vehicle may be an EV (electric vehicle), and the service destination node may be a base station. In this case, the EV travels to each base station and charges each base station. In this case, the power required by the base station corresponds to the demand, and the power held by the EV corresponds to the load (parcel).

[0016] (Outline of First Embodiment) First, an outline of the first embodiment will be described. In this embodiment, a new VRP called VRPTWTC, which is a formulation of a very practical problem in a business scenario, is introduced.

[0017] In the first embodiment, in problem formulation, two new constraints (time window and time cost) are introduced in addition to existing constraints in VRP such as demand and load in the optimization process. Note that in the first embodiment (and the second embodiment), "load" is assumed to be "baggage," "cargo," etc., loaded on a service vehicle, and "load" may also be rephrased as "baggage," "cargo," etc.

[0018] In the first embodiment, a data-driven, end-to-end policy-based reinforcement learning model is used to solve VRPTWTC. The policy-based reinforcement learning model includes two neural networks: an actor network and a critic network. The actor network generates the VRPTWTC path, and the critic network estimates and evaluates the value function.

[0019] Furthermore, the first embodiment uses a novel masking algorithm combined with an actor network. The masking algorithm makes it possible to solve a problem under the time window constraint and time cost constraint formulated in the first embodiment, in addition to the constraints in conventional VRP. Furthermore, the first embodiment (and the second embodiment) uses a reward function in the reinforcement learning model that imposes a penalty if a vehicle arrives at a node outside a predetermined time window.

[0020] Furthermore, in the first embodiment (and the second embodiment), by using an API of a map application based on an actual map, routes can be calculated under actual road connection conditions, thereby increasing the possibility of adoption in actual industries.

[0021] (First embodiment: device configuration example) Fig. 1 shows a configuration diagram of a delivery planning device 100 in the first embodiment. As shown in Fig. 1, the delivery planning device 100 has a user information collection unit 110, a service vehicle information collection unit 120, an algorithm calculation unit 130, a map API unit 140, and a vehicle dispatch unit 150.

[0022] The delivery planning device 100 may be implemented in one device (computer) or in multiple devices. For example, the algorithm calculation unit 130 may be implemented in one computer, and the other functional units may be implemented in another computer. An overview of the operation of the delivery planning device 100 is as follows.

[0023] The user information collection unit 110 acquires information (which may be called features) about each user (which may be called a customer, a node, or the like). The information about each user includes, for example, a location, an initial demand, a specified time window, a time cost of a service, and the like.

[0024] The service vehicle information collection unit 120 collects information about each service vehicle, including the initial load capacity (power capacity if it is electricity), the departure position, etc.

[0025] The algorithm calculation unit 130 outputs a delivery plan by solving a VRP problem based on information on each node and each service vehicle. Details of the algorithm calculation unit 130 will be described later.

[0026] The map API unit 140 performs a route search based on the delivery plan information output from the algorithm calculation unit 130, and, for example, draws the delivery plan route of each service vehicle on a map. The dispatch unit 150 distributes the service route information to each service vehicle (or a terminal at the service center) via a network based on the output result of the map API unit 140. The dispatch unit 150 may also be called an "output unit."

[0027] The map API unit 140 may perform route searches, for example, by accessing an external map server. Alternatively, the map API unit 140 may itself store a map database and perform route searches using the map database.

[0028] As an example, assume that the algorithm calculation unit 130 has obtained a delivery plan of "0 → 2 → 3 → 0." Here, 0 indicates the service center, and 2 and 3 indicate the node numbers, respectively. In this case, the map API unit 140 draws the actual road route of "service center → node 2 → node 3 → service center" on the map, and the dispatch unit 150 outputs map information with the route drawn.

[0029] (Configuration Example of Algorithm Calculation Unit 120) Fig. 2 shows a configuration example of the algorithm calculation unit 130. The algorithm calculation unit 130 includes a neural network model that performs actor-critic reinforcement learning. This model may be called a VRPTWTC model.

[0030] As shown in FIG. 2, this model includes two neural networks: an actor network 131 and a critic network 132.

[0031] The actor network 131 has a dense embedding layer (1 layer), an LSTM cell, an attention layer, a Softmax calculation unit (Softmax), and a masking unit (Masking). These constitute an encoder-decoder configuration and a pointer network. The critic network 132 has a dense embedding layer (3 layers).

[0032] The dense embedding layer, LSTM cell, attention layer in the actor network 131, and the dense embedding layer in the critic network 132 have learnable parameters in the neural network.

[0033] In the actor network 131, features obtained by an LSTM cell from a hidden state, which is the output from a dense embedding layer corresponding to an encoder, are input to an attention layer, a context is obtained from the output from the dense embedding layer and the output from the attention layer, and a value calculated from the context by Softmax is output through masking and used for reward calculation. In the critic network 132, a loss function is obtained based on the features obtained from the input data by the dense embedding layer and the reward, and learning is performed to reduce the loss.

[0034] The arrows between the input hidden state, context, attention layer, LSTM, and softmax indicate an attention-based pointer network. The loss function is calculated by the reward function and the critic network.

[0035] The algorithm calculation unit 130 uses the neural network shown in FIG. 2 to learn a large amount of simulation learning data, and is configured to be able to test (create delivery plans) using both real data and simulation data.

[0036] Specifically, by using an actor-critic based reinforcement learning model and a masking algorithm, it is possible to efficiently calculate and output a delivery plan under the constraints that service vehicles must arrive at the node at the specified time (within the specified time frame) and that each service vehicle must work within eight hours per day.

[0037] It should be noted that by using a reward function that takes into account penalties, as will be described later, it is possible to create a delivery plan that ensures that the service vehicle arrives at the node as close to the designated time (within the designated time frame) as possible.

[0038] The processing contents of the algorithm calculation unit 130 will be explained in more detail below.

[0039] (Outline of Processing by Algorithm Calculation Unit 130) First, in the first embodiment, an outline of the VRP problem solved by the algorithm calculation unit 130 will be described. This problem has the following three elements.

[0040] (1) All nodes must be served, and the time at which the service is provided (the time at which the service vehicle arrives) must be within a time frame (time window) specified by each node.

[0041] (2) Each node has a time cost of service that varies depending on the service. This "time cost" is the time required to provide the service at the node. The service at the node may be, for example, transporting luggage, repairing communication equipment, or charging a base station.

[0042] (3) When serving multiple nodes, a service vehicle cannot exceed the total service time limit.

[0043] The above problem is called the "Vehicle Routing Problem with Time Windows and Time Costs" (VRPTWTC).

[0044] In the first embodiment, a neural network corresponding to the algorithm calculation unit 130 solves the VRPTWTC in an end-to-end, data-driven manner based on actor-critic based deep reinforcement learning, thereby outputting a solution (delivery plan) to the above problem.

[0045] In the first embodiment, the time slots and service time costs specified by the nodes are strictly considered during the optimization process. Violations of the time slots and total working hours are not permitted, which helps improve service quality and protect staff rights. Furthermore, by using a reward function (described later), it is possible to apply time constraints taking importance into account.

[0046] (Details of Processing by Algorithm Calculation Unit 130) Hereinafter, the processing content of the algorithm calculation unit 130 in the first embodiment will be described in detail.

[0047] <A: Problem Setting> The problem setting in the first embodiment will be described with reference to Fig. 3. A set of nodes (specifically, customers) χ = {x 1 , x 2 ,...x N} is located in a certain range on the map, and each x in the set n is a node requiring a service. There is also a service center that loads roads for providing the service. The locations of the nodes and the service centers are assumed to be known. Furthermore, the travel time of the service vehicle between the node and the service center, or between any two nodes, may be known (for example, calculated from a predetermined speed and distance), or may be calculated from a map API taking into account actual road conditions (traffic congestion, etc.).

[0048] First, a set of service vehicles is deployed at a service center. Each service vehicle can leave the service center to service a set of nodes, χ, for example, simultaneously (or sequentially). Each node is serviced exactly once by any service vehicle. After visiting all planned nodes, the service vehicle returns to the service center.

[0049] Each node in χ has four features, so each node x n Let x be a vector. n = [x n f1 , x n f2 , x n f3 , x n f4 ]. n f1 is the address of the nth node. n f2 is the demand of the nth node, which is the same as the demand characteristic of the classical VRP problem. n f3 is the time window specified by the nth node, during which the node must be visited by the service vehicle. n f4is the time cost of the service of the nth node, which indicates how long it takes to service the nth node. For ease of modeling, the service center is set as the 0th node in the problem formulation.

[0050] In this problem, we assume that time window violations (failure to service a node within the time window) and time cost violations (service vehicles working more than 8 hours a day) are unacceptable.

[0051] For each service vehicle, we define a fixed initial load characteristic that indicates the maximum load capacity of the service vehicle providing the service. Specifically, before the service vehicle leaves the service center and provides service to a node, we initialize the load to a value of 1 (which can be adjusted depending on the task).

[0052] In addition, the maximum service time for each service vehicle is set to 8 hours, which means that each service vehicle has a maximum service time of 8 hours, that is, the longest time for a service vehicle to leave the service center to provide service does not exceed 8 hours (this can be adjusted according to actual business demands).

[0053] The following two conditions (1) and (2) are set out as conditions for a service vehicle to provide service: The service vehicle must return to the service center in the following cases (1) or (2).

[0054] Condition (1): The load of the service vehicle is close to zero, and the capacity (remaining load) to service the remaining nodes is insufficient. Condition (2): The service time of the service vehicle is close to the maximum of 8 hours. Given the above node information and optimization constraints, find a solution ζ for VRPTWTC. The solution ζ is a sequence of nodes in χ, which can be interpreted as a service route or service order. For example, if the solution sequence ζ = {0, 3, 2, 0, 4, 1, 0} is obtained, this sequence corresponds to two routes. One route follows the order 0 → 3 → 2 → 0, and the other route follows the order 0 → 4 → 1 → 0, which implies the use of two service vehicles. This can also be interpreted as a case where a service vehicle returns to the service center.

[0055] B: Pointer Network in Actor Network 131 The solution ζ to VRPTWTC is a Markov Decision Process (MDP) of sequences, which is the process of choosing the next action in the sequence (i.e., which node to service next).

[0056] In the first embodiment (and the second embodiment), a pointer network (PointerNet) is used to formulate the MDP process. Note that the pointer network (PointerNet) itself is an existing technology. First, an encoder with a dense layer embeds features of all input nodes and depots (service centers) to extract hidden states. Next, a decoder reconstructs the MDP actions by using Long Short-term Memory (LSTM) cells connected one by one and passes them to the Attention layer. Each LSTM cell (action) outputs a pointer representing the probability that the input node will receive service.

[0057] We will now explain in more detail the dense embedding layer (encoder) of the actor network and the pointer network.

[0058] As mentioned above, χ = {x 1 , x 2 , ..., x N Each x inn represents a node, and each x n is converted by the encoder into a dense representation x n-dense Embed as.

[0059] where θ embedded ={ω embed , b embed} is a learnable parameter represented as a dense layer in the embedding layer of the first embodiment.

[0060] The decoder contains a sequence of LSTM cells. In the decoder, the sequence of LSTM cells is used to model actions in the MDP. At each step m∈(1, 2, ..., M) in the decoder, a weight θ LSTM The hidden state in the LSTM cell with m where M is the total number of decoder steps.

[0061] In the first embodiment, similar to PointerNet, the pointer D m That is, at each step m of the decoder part, χ = {x 1 , x 2 , ..., x N} is pointed to, the Softmax result is calculated.

[0062] Here, p(D m |D 1 , D 2 ...D m-1 , χ;θ) as a pointer at each step of the decoder, and the parameter θ Pointer The model is made by the following equations (2) and (3) using an LSTM cell having:

[0063]

[0064] where Softmax is the vector u (of length N) mis normalized to the output distribution (probability distribution) for all inputs χ. In other words, the probability of each node (probability of being selected as a service target) at the m-th step is output by equation (3). θ Pointer = {v, W 1 , W 2} is a learnable parameter of the pointer.

[0065] The final output of the actor network 131 is a service path ζ, which corresponds to the output of a sequence of all m LSTM cells. Here, multiple LSTMs can be interpreted as an MDP. p(D m │D 1 , D 2 ...D m-1 , χ;θ) to p(D m ) is abbreviated as

[0066] <C: Masking> As mentioned above, in the first embodiment, a new masking algorithm is proposed and combined with the actor network 131 to optimize VRPTWTC. In this masking algorithm, there are three sub-maskings: load-demand masking, time window masking, and time cost masking.

[0067] Load-demand masking is used to solve the traditional VRP constraints. Time window masking and time cost masking are used to optimize the new constraints formulated in VRPTWTC.

[0068] The masking algorithm is combined with the actor network 131 to output probabilities of actions in reinforcement learning. We first describe each of these three sub-masking algorithms, and then how to combine them in an actor network.

[0069] Note that (2) and (3) may be implemented together, or either one of them may be implemented. Also, when using a reward function that generates a time frame penalty (described later), masking of the time frame in (2) may not be performed.

[0070] (1) Load-Demand Sub-Masking: Since both the service capacity of the service vehicle and the demand of the node are finite and limited, when the remaining load on the service vehicle runs out, the service vehicle must return to the service center for replenishment.

[0071] Here we use load-demand submasking to model this process: At each decoder step m ∈ (1, 2...M), the remaining demand δ at each node ∈ (1, 2...N) n,m and the remaining vehicle load Δ m and m = 1, and these are δ n,m = δ n , Δ m = 1 and then updated as follows: m is the index of the node selected for service at decoder step m.

[0072]

[0073] Equation (4) shows that if the nth node is selected at decoder step m, then at the next decoder step m+1, the demand of node n will be the greater of 0 (serviced) or demand minus load (if there are insufficient service vehicles to provide the entire service), and the demands of other nodes other than n will remain unchanged.

[0074] Equation (5) shows that at m+1, if the service vehicle returns to the service center, the vehicle load is 1 (the value to be replenished); otherwise, the vehicle load is the load at m minus the demand of the node being serviced (if the node is serviced by the vehicle). Note that in this problem formulation, since the service center is the 0th node, π m = 0 indicates that the service vehicle has returned to the service center.

[0075] (2) Time-Window Sub-Masking: In the problem setting in the first embodiment, the service vehicle must arrive at each node at a specified time (within a specified time window). Therefore, at each step of the decoder, a time-window sub-masking is added to set the probability of a node that is unlikely to arrive at the specified time to 0. In this way, setting the probability of a node to 0 may be called masking or filtering.

[0076] As mentioned above, equation (3) indicates that the pointer (Softmax) is a vector u m is the output probability distribution p(D m ) where p(D m ) is an n-dimensional vector, giving the probability distribution over χ at decoder step m.

[0077] At each step m of the decoder, we denote by χ′∈χ the set of nodes that need to be serviced. The reason for using such a set is that some nodes may be serviced before step m or the service vehicle may not have enough load.

[0078] For a node set χ′ with a node count of N′, the following process is repeated for each node n′ of χ′ to obtain a sub-masking τ n´,m Calculate:

[0079] Equation (6) is t total +t move But time frame x n f3 If it is not within the range of τ n´,m = 0. In equation (6), t total is the total time to the last node served on the current path, and t move is the travel time from the node that was previously served to n'. Equation (6) expresses the total time cost t total and the time t spent moving from the previous node to the current node n′ move This means that if the value obtained by adding .times. ...

[0080] (3) Time Cost Submasking: Time cost submasking is a method to reduce the total time cost t total This is used to force the service vehicle to return to the service center if the service time exceeds 8 hours, and is expressed by the following equation (7):

[0081] Equation (7) expresses the total time cost t total If the time exceeds 8 hours, p(D m ) is set to 1, and p(D m ) is set to 0. Here, n=0 means that the node is a service center. As mentioned above, the service center is the 0th node in the formulation of this problem. total may be called the total operating time.

[0082] Below, we will explain the masking processing algorithm (Algorithm 1) and the reinforcement learning algorithm (Algorithm 2). In these algorithms, the number of vehicles may be one or more. In the case of more than one vehicle, the vehicles may proceed in parallel (synchronized), or one vehicle may proceed in one decoder step (one time step).

[0083] The masking algorithm (Algorithm 1) is shown in Figure 4. This is the process executed by the Masking (Masking Unit) in Figure 2. In the first line, for each node n ∈ (1, 2...N), the demand x n f2 , vehicle capacity Δ 0 , time frame x n f3 , time cost x n f4 Enter t total is initialized to 0.

[0084] The second line means that steps 3 to 13 are repeated at each decoder step m = 1, 2.... M. The third line means that at step m, the remaining demand δ n,m If it is 0, the loop processing ends.

[0085] In the fourth line, for each node n∈(1, 2...N), n,m >0 and δ n,m <Δ m If so, msk n,m = 1, otherwise msk n,m = 0. msk n,m =1 indicates that the service is available, msk n,m = 0 indicates that the node has already been serviced or that the vehicle capacity is insufficient, and indicates that the node is not to be serviced (the probability is set to 0).

[0086] In the fifth line, the vector p(D m ) in descending order, and p with index i (1, 2...N) after sorting sort (D m ) to

[0087] In lines 6 and 7, p sort (D m ) for each i-th member p sort,i (D m ), nodes are filtered (masked) based on equation (6) (time window sub-masking).

[0088] In line 8, Softmax(p sort,i (D m )) as the probability of the new action pointer. In line 9, we check for temporal cost masking according to equation (7).

[0089] In the 10th line, the remaining demand δ according to equation (4) n,m In line 11, the remaining load is updated by equation (5). In line 12, m is updated to m+1. In line 13, if n is not 0, t total = t total +t move +x n f4 This means that the total operating time at a node is the sum of the total operating time from the service center to the completion of service at that node, the travel time from that node to the next node, and the time cost at that node. The process ends at line 14.

[0090] As shown in Fig. 4, three sub-maskings are introduced in the masking algorithm shown in Algorithm 1. After data input and initialization, at each step m of the LSTM-based decoder, first, for each demand of a node, if all demands are 0, i.e., all nodes have been served, the decoder loop ends.

[0091] If not, mask all nodes with non-zero demand values ​​with 1. Note that the demand value must be less than the vehicle's dynamic load.

[0092] Next, the vector p(D m ) in descending order, and sort (D m ) Then, using equation (6), we filter out the nodes that cannot be served by considering the time window and the total time cost of the current service path, and obtain p sort,i (D m ) and use Softmax to sort,i (D m ) is normalized.

[0093] Furthermore, using equation (7), the total time cost t total Check whether the dynamic demand δ exceeds 8 hours. If so, return the service vehicle to the service center (0th node). Finally, n,m , dynamic load Δ m , and the total time cost t total and proceed to the next decoder step m+1.

[0094] <D: Actor-Critic> In the first embodiment (and the second embodiment), actor-critic based deep reinforcement learning is used to simultaneously learn both the policy (measure) and the value function. Note that actor-critic based deep reinforcement learning itself is an existing technology.

[0095] As for the actor network 131, as explained in A, the learnable weights θ actor = {θ embedded , θ LSTM , θPointer}.

[0096] In the first embodiment, the pointer parameter θ in the actor network 131 Pointer = {ν, W 1 , W 2} and LSTM parameter θ LSTM We use π to parametrize a stochastic policy π, which generates a probability distribution over the next action (which node to visit) at any given decoder step.

[0097] On the other hand, the learnable parameter θ critic The critic network 132 with ∇ ...

[0098] The critic network 132 consists of three dense layers, and takes static and dynamic states as inputs to predict rewards (which may also be called earnings). In the first embodiment, for example, the output probabilities of the actor network 131 are used as weights to calculate the weighted sum of the embedded inputs (outputs from the dense layers), thereby outputting a single value. This can be interpreted as the output of the value function predicted by the critic network 131.

[0099] FIG. 5 shows the actor-critic algorithm (Algorithm 2) executed by the algorithm learning unit 130.

[0100] In the first line, we set the actor network (Embedding2Seq with PN) with random weights θ actor = {θ embedded , θ LSTM , θ Pointer} and initialize the critic network with random weights θ critic Lines 2 and 17 mean that lines 3 to 16 are repeated in each epoch.

[0101] In the third line, the gradient of the parameter dθ actor and dθ critic are reset to 0. From the fourth line onwards, the current θ actorLines 5 and 14 mean that lines 6 to 13 are repeated for each sample in B.

[0102] In the sixth line, the current θ embedded Based on this, the embedding layer is processed to obtain x n-dense (batch). Lines 7 and 12 mean that lines 8 to 11 are repeated for each decoder step m∈(1, 2, .... M). Line 8 means that lines 9 to 11 are repeated until a termination condition (e.g., the demands of all nodes are satisfied) is met.

[0103] In line 9, the distribution p(D m ) based on which the decoder m Calculate the following. m indicates the node to be serviced (visited) in the m-th step.

[0104] In row 10, the new state columns D1, . . . , D m-1 , D m In line 11, m is updated to m+1.

[0105] In line 13, the reward R is calculated. In line 15, the policy gradient ∇θ is calculated using equation (8) below. actor (dθ actor ) and calculate θ by Adam. actor In line 16, the gradient ∇θ critic (dθ critic ) is calculated, and θ critic Update.

[0106] The actor-critic algorithm 2 in this embodiment shown in Figure 5 shows a training process. After this training process, a test (actual delivery plan output) may be performed, or a test may be performed while the training is progressing.

[0107] As already explained, the weight vector θ actor and θ criticWe use two neural networks (actor network and critic network) with θ actor is θ embedded , θ LSTM , θ Pointer Includes:

[0108] The current weight θ of the actor network actor At each training iteration, we take B samples and use Monte Carlo simulation to generate feasible sequences based on the current policy. This means that at each decoder step, we generate a distribution p(D m ) based on the pointer D m This means that the following is calculated (selected) probabilistically.

[0109] Once sampling is complete, the reward and policy gradient are calculated and the actor network is updated in line 15.

[0110] Also, in line 16, the critic network is updated in a direction that reduces the difference between the observed reward and the expected reward. Finally, the gradient dθ is calculated using the same learning speed in the end-to-end method. actor and gradient dθ critic Using θ actor and θ critic Below, we explain the policy gradient and reward.

[0111] (1) Policy gradient: In line 15 of Algorithm 2, the policy gradient of the actor network is approximated by Monte Carlo sampling as follows:

[0112] Here, R is the reward (reward function) for the route instance, and D is the route that the vehicle travels. m is the reward for the sequence of V(χ;θ critic ) is a value function that predicts the reward for all raw inputs. critic ") is used as an advantage function to replace the cumulative reward in the VRP method based on conventional reinforcement learning. The method of using an advantage function in actor-critic is itself an existing technology.

[0113] 2) Reward Function: In the first embodiment (and the second embodiment), a reward function is used that imposes a penalty if the time when a vehicle arrives at a node exceeds a time window defined for that node. The time window has a start time and an end time. Here, the end time of the time window is called the leaving time. The leaving time of node i is l i That is, the reward function is the vehicle's leaving time (l i ) ) ), a penalty is incurred. Specifically, the reward function is as shown in the following equation (9).

[0114] p(c, l) is a penalty function, as shown in the following equation (10).

[0115] In formula (10), c i represents the arrival time at node i. In other words, p(c, l) is the sum, for nodes 1 to N, of the time by which the arrival time exceeds the time window.

[0116] The first term on the right side of equation (9) is "Σ N i=1 c i " represents the travel cost when a vehicle travels to nodes 1 to N. The travel cost may be, for example, the total travel time or the total travel distance when a vehicle travels to nodes 1 to N.

[0117] The same applies when there are multiple vehicles. For example, if there are two vehicles A and B, and vehicle A travels between nodes 1 and 5, and vehicle B travels between nodes 6 and N, then "Σ N i=1 c i " is the sum of the travel cost of vehicle A and the travel cost of vehicle B.

[0118] For each i, the leaving time (l i) is obtained as initial input information. For each i, the arrival time at node i is obtained as the state when the vehicle arrives at node i. The algorithm calculation unit 130 can calculate R using this information.

[0119] In normal reinforcement learning, learning is performed to maximize the reward, but in the first embodiment (and the second embodiment), the opposite is true; learning is performed to minimize (make as small as possible) the reward R (the sum of the movement cost and the penalty).

[0120] In equation (9), ω is a weight (importance), and setting the weight larger will place more importance on the penalty for time violation. In other words, the importance of the penalty can be adjusted by ω.

[0121] (Outline of Second Embodiment) Next, a description will be given of a second embodiment. The second embodiment is basically the same as the first embodiment in that a delivery plan is implemented using a reinforcement learning model, but differs from the first embodiment in that the second embodiment includes a feature for realizing a delivery plan that takes into account the status of edges between nodes.

[0122] That is, in the conventional reinforcement learning-based method disclosed in Non-Patent Document 2, edges between nodes are approximated by a straight-line distance, and therefore situations such as impassable areas or traffic congestion cannot be taken into consideration. Furthermore, it cannot handle cases where edges change (such as sudden traffic congestion). To improve this point, the second embodiment takes into consideration the status of edges between nodes. Furthermore, the second embodiment also uses the reward function described in the first embodiment.

[0123] In the second embodiment, a graph convolutional network (GCN) is used to obtain edge-related feature quantities. However, edge-related feature quantities may be obtained using a method other than the GCN.

[0124] (Second embodiment: device configuration example) Fig. 6 shows a configuration diagram of a delivery planning device 200 in the second embodiment. As shown in Fig. 6, the delivery planning device 200 has a node information collection unit 210, a service vehicle information collection unit 220, an algorithm calculation unit 230, a map API unit 240, a vehicle dispatch unit 250, and an edge calculation unit 260. Note that the map API unit 240 is used for both the output from the node information collection unit 210 and the output from the algorithm calculation unit 230, and is therefore shown in two places for convenience.

[0125] The delivery planning device 200 may be implemented by one device (computer) or by multiple devices. For example, the algorithm calculation unit 230 may be implemented by one computer, and the other functional units may be implemented by another computer. An overview of the operation of the delivery planning device 200 is as follows.

[0126] The node information collector 210 acquires information (feature amounts) about each node, such as its location, initial demand, and designated time frame.

[0127] The edge calculation unit 260 uses the map API unit 240 to calculate the actual travel time (or travel distance) between nodes and outputs the calculation result in the form of a distance matrix. Note that travel distance and travel time may be collectively referred to as travel cost. Note that an element in the distance matrix may be travel time.

[0128] The service vehicle information collection unit 220 collects information about each service vehicle. The information about each service vehicle includes the initial load capacity (power capacity if it is electricity), the departure position, etc. The load (amount of luggage) of the service vehicle may be initialized to 1 at the time of departure of each service vehicle.

[0129] The algorithm calculation unit 230 outputs a delivery plan by solving a VRP problem based on information on each node and each service vehicle. That is, the algorithm calculation unit 230 converts the Distance Matrix calculated by the edge calculation unit 260 into features and generates a delivery plan using a Pointer Network. Also, as in the first embodiment, the algorithm calculation unit 230 uses an ActorCritic reinforcement learning model to perform learning so that an optimal delivery plan can be output. Details of the algorithm calculation unit 230 will be described later.

[0130] The map API unit 240 can acquire map information as well as actual road conditions (e.g., road congestion status, travel time when congested, and road closure information due to accidents or construction) in real time. The map API unit 240 also performs a route search for the output (delivery plan) from the algorithm calculation unit 230, and draws, for example, the routes of the delivery plan for each service vehicle on a map. The vehicle dispatch unit 250 distributes service route information to each service vehicle (or a terminal at the service center) via a network based on the output result of the map API unit 240. The vehicle dispatch unit 250 may also be referred to as an "output unit."

[0131] The map API unit 240 may perform route searches, for example, by accessing an external map server. Alternatively, the map API unit 240 may itself store a map database and perform route searches using the map database.

[0132] As an example, assume that the algorithm calculation unit 130 has obtained a delivery plan of "0 → 2 → 3 → 0." Here, 0 indicates the service center, and 2 and 3 indicate the node numbers, respectively. In this case, the map API unit 240 draws the actual road route of "service center → node 2 → node 3 → service center" on the map, and the dispatch unit 250 outputs map information with the route drawn.

[0133] (Configuration Example of Algorithm Calculation Unit 230) Fig. 7 shows a configuration example of the algorithm calculation unit 230. The algorithm calculation unit 230 is a neural network model that performs actor-critic reinforcement learning.

[0134] As shown in FIG. 7, this model includes two neural networks: an actor network 231 and a critic network 232.

[0135] The actor network 231 also includes an encoder 233 for embedding (feature quantification) input data, and the encoder 233 includes a graph convolutional layer 233. The graph convolutional layer calculates the features of nodes and edges.

[0136] 7 includes an encoder 233 and a decoder 234. The decoder 234 corresponds to a pointer network. The encoder 233 converts the input sequence (node ​​information and edge information) into features and inputs them to the decoder 234 (the decoder uses the features). The decoder 234 uses an attention mechanism to determine which input to select as the next delivery destination.

[0137] 7 includes an embedding layer that embeds input data into an h-dimensional vector, and a graph convolutional layer. As described above, the graph convolutional layer calculates the feature quantities of nodes and edges.

[0138] The graph convolutional layer may be called a GCN (Graph Convolutional Network). A GCN is a network that convolves a graph structure, receives data having a graph structure as input, and outputs feature quantities of the data.

[0139] Furthermore, since the order of input data is irrelevant in VRP, the RNN encoder is omitted from the configuration shown in FIG.

[0140] The learning method is the same as in the first embodiment, using a policy gradient method. In this policy gradient method, an actor network 231 predicts the probability distribution of the next action using a pointer network, and a critic network 232 estimates the reward for the problem instance.

[0141] The critic network 232 has a dense embedding layer. In the critic network 232, a loss function is obtained based on the features obtained from the input data by the dense embedding layer and the reward, and learning is performed to reduce the loss.

[0142] An image of the overall processing in the algorithm calculation unit 230 is shown in Fig. 8. As shown in Fig. 8, the decoder (pointer network) determines the vehicle's behavior at each step (the next node to visit, etc.), and the state is updated based on that behavior (action), and a reward is calculated, and the model is trained based on the reward.

[0143] (Details of Processing by Algorithm Calculation Unit 230) The following describes in detail the processing content of the algorithm calculation unit 230. In addition, in relation to the input to the algorithm calculation unit 230, the operations of the edge calculation unit 260 and the map API unit 240 will also be described.

[0144] <Problem Setting> The problem setting in the second embodiment will be described with reference to Fig. 9. As shown in Fig. 9, a set of nodes is located within a certain range on a map. Each black circle in Fig. 9 indicates a node requiring a service. There is also a service center (luggage collection point) where luggage (load) for providing the service is loaded.

[0145] First, a set of delivery vehicles is deployed at a service center. In the optimization problem of the second embodiment, when delivering packages from the service center to each node using delivery vehicles, the optimal order (the order that results in the lowest cost) for which delivery vehicles should visit which nodes is considered. The above cost is, for example, the travel distance (mileage) or travel time (travel time) of the delivery vehicles. Note that the departure point may be different for each delivery vehicle. Also, nodes that can be visited (or nodes that cannot be visited) may be determined for each delivery vehicle.

[0146] Each node is serviced only once by one of the delivery vehicles. After visiting all planned nodes, the delivery vehicle returns to the service center. Figure 9 shows an example in which there are three delivery vehicles, each making deliveries along a different route. Figure 9 can also be thought of as an example in which one delivery vehicle returns to the service center twice to replenish packages while making deliveries.

[0147] As a condition for the service vehicle to provide service, for example, a condition (constraint) is used such that the service vehicle must return to the service center when "the load on the service vehicle is close to 0 and the capacity (remaining load) to provide service to the remaining nodes is insufficient." The constraint may be the same as the constraint in the first embodiment.

[0148] The algorithm calculation unit 230 finds a solution ζ for the VRP. The solution ζ is a sequence in the set of nodes χ, which can be interpreted as a service route or service order. For example, if the solution sequence ζ = {0, 3, 2, 0, 4, 1, 0} is obtained, this sequence corresponds to two routes. One route follows the order 0 → 3 → 2 → 0, and the other route follows the order 0 → 4 → 1 → 0, which implies the use of two service vehicles. This can also be interpreted as a service vehicle returning to the service center.

[0149] <Input Data> Next, the input data to the algorithm calculation unit 230 will be described. The state of node i (node ​​information) is expressed as x iWhen written as follows, x i is expressed as follows: where t indicates each time in the time step.

[0150] x i :{x t i = (s i , d t i ), t=0,...,T} The meaning of each symbol is as follows.

[0151] x i : State s of node i i : Two-dimensional coordinates (address) of node i t i : Demand at node i at step t (assuming that demand differs for each node) The above demand is the same as the demand characteristics in the classical VRP problem. Also, assume that there are N nodes. The time frame explained in the first embodiment is included as input information on the state of node i.

[0152] In addition, the algorithm calculation unit 230 receives edge information e between two nodes. ij The edge information e for node i is input. ij More specifically, it is expressed as follows:

[0153] e ij : {e t ij , t=0,...., T, j=0,...., n} where e ij is the travel time of the service vehicle from node i to node j. Note that using travel time is an example. ij The travel distance of the service vehicle from node i to node j may be used as .

[0154] e ij Regarding the calculation method of e, for example, the map API unit 240 acquires the actual road conditions between each node or traffic information photographed from the air, and the edge calculation unit 260 calculates each e using information such as the actual road conditions. ij Calculate the following: ij and e ji In addition, since the actual road conditions change depending on the time t, et ij may change depending on the time t. ij wo e ij It may also be expressed as:

[0155] Actor Network 231 The solution ζ to the VRP is a Markov Decision Process (MDP) for sequences, which is the process of choosing the next action in the sequence (i.e., which node to service next).

[0156] In the second embodiment, as in the first embodiment, a pointer network (PointerNet) is used to formulate the MDP process.

[0157] First, the encoder 233 performs feature quantification on the node information and edge information. Then, the decoder 234 restores the behavior of the MDP by using RNN cells connected one by one. The RNN cells are supplied with node position information (s i ) is input. The encoder 233 and the pointer network will be described in more detail below.

[0158] First, the embedding layer in the encoder 233 determines the two-dimensional coordinates s of each node as the feature of each node i. i ∈[0, 1] 2 is embedded into h-dimensional features as follows:

[0159] α i = A 1 s i +b 1 Also, the embedding layer uses the edge value e ij is embedded into h-dimensional features as follows:

[0160] β ij = A 2 e ij +b 2 Here, A 1 ∈R h×2 , A 2 ∈R h×1 , is.

[0161] Next, the graph embedding layer in the encoder 233 calculates the node feature quantity in layer l (l is a lowercase letter L) using the following formula: - x i l and edge features - e ij l In the text of this specification, for convenience of description, the symbol at the beginning of a letter is written at the upper left of the letter. - "x" is an example.

[0162] Depending on the number of layers, the following formula is repeatedly applied to obtain the node features. - x i l and edge features - e ij l The number of layers can be a predetermined value.

[0163]

[0164] where W∈R h×h , - x i l=0 = α i , - e ij l=0 = β ij , and η ij l is a dense attention map, and contains information on anisotropy (direction dependence) on the graph. Figure 10 shows an image of calculating the features of layer l+1 from the features of layer l. That is, the features of node i in layer l+1 are obtained from "the features of node i in layer l, the features of the adjacent nodes adjacent to node i in layer l, and the features of the edge between node i and the adjacent node in layer l." Also, the features of edge ij in layer l+1 are obtained from "the features of node i in layer l, the features of node j in layer l, and the features of edge ij in layer l."

[0165] The decoder 234 contains a sequence of RNN cells (e.g., LSTM cells). In the decoder 134, the sequence of RNN cells is used to model actions in the MDP. At each step t ∈ (0, 1, ..., T) of the decoder, the hidden state in the RNN cell is expressed as h t where T is the total number of decoder steps.

[0166] In the second embodiment, the attention mechanism t+1 The service order is modeled by calculating . That is, it determines which node is pointed to at each step t of the decoder 234. Note that the operation of the attention mechanism here is the same as that of the attention mechanism in the first embodiment, except for the difference in the features used.

[0167] Specifically, in step t, t is calculated by softmax as follows: t indicates how relevant each input data (each node) is (whether it is suitable as a delivery destination) at step t.

[0168] a t = a t ( - e t i , h t )=softmax(u t ) u t i =v a T tanh (W a [ - e t i ;h t ]) - e t i is a feature amount after embedding by GCN, and is a feature amount of an adjacent edge of node i. - e t i is mentioned above. - e ij l is the value at step t. tis the hidden state of the RNN cell at step t, as previously mentioned.

[0169] Next, in the decoder 134, the context vector c is calculated by the following formula: t Calculate.

[0170] Next, in the decoder 234, c t and - e t i Combine - u t i The value is normalized and the next node (y t+1 ) conditional probability P(y t+1 |Y t , E t ) is calculated. Note that normalization is equivalent to finding the probability distribution for all input nodes.

[0171] P(y t+1 |Y t , E t )=softmax( - u t i ) - u t i =v c T tanh (W c [ - e t i ;c t ]) In the above formula, Y t is the sequence of selected nodes up to t, and E t is the edge feature of the input data after GCN at t ( - e) v a , v c , W a , W c are all parameters that can be learned. The image of the above calculation process is shown in Figure 11.

[0172] Furthermore, in the actor network 231 of the algorithm calculation unit 230, for example, of the three masking methods described in the first embodiment, (1) may be applied, or in addition to (1), either or both of (2) and (3) may be applied. For example, if the remaining load of a delivery truck is 0, all nodes are masked (i.e., no delivery is made to any node). Also, for example, a node with a demand greater than the current load of the delivery truck is masked (i.e., no delivery is made to that node). Note that the load of a delivery truck is reduced by the amount of cargo each time a package is delivered to a node.

[0173] <Actor-Critic> In the second embodiment, as in the first embodiment, actor-critic based deep reinforcement learning is used to simultaneously learn both a policy (measure) and a value function.

[0174] The learnable parameter (weight) in the actor network 231 is represented as θ. More specifically, θ={θ GCN , θ RNN , θ PrtNet}.

[0175] In the second embodiment, the actor network parameters θ are used to parameterize a stochastic policy π, which generates a probability distribution over the next action (which node to visit) at any given decoder step t.

[0176] On the other hand, the learnable parameter θ BL The gradient for any problem instance is estimated from a given state in reinforcement learning by a critic network 132 having

[0177] The critic network 132, for example, consists of three dense layers and predicts rewards (profits). In the second embodiment, as in the first embodiment, the output probabilities of the actor network 231 are used as weights to calculate the weighted sum of the embedded inputs (outputs from the dense layers), thereby outputting a single value. This can be interpreted as the output of the value function predicted by the critic network 132.

[0178] The reinforcement learning algorithm will be described below. In this algorithm, the number of vehicles may be one or more. In the case of multiple vehicles, the vehicles may be moved in parallel (synchronized), or one vehicle may be moved in one decoder step (one time step).

[0179] FIG. 12 shows an example of the actor-critic process (algorithm) executed by the algorithm calculation unit 230.

[0180] In the first line, we initialize the actor network with random weights θ and the critic network with random weights θ BL Lines 2 and 18 mean that lines 3 to 17 are repeated at each epoch.

[0181] In the third line, the gradients of the parameters dθ and dθ BL In the fourth line and onward, B instances are sampled according to the actor network with the current θ. In the fifth line, each sample X 0 This means repeating lines 6 to 13 for "X 0 " indicates the initial information of each node (location, demand, time frame, etc.) and the edge information of each edge in the sample.

[0182] In the sixth line, the current weight (θ GCN ) and perform graph embedding etc. to obtain the features of nodes and edges. In line 7, the step counter t is initialized to 0.

[0183] Lines 8 and 12 mean that lines 9 to 10 are repeated until a termination condition (for example, the demands of all nodes are satisfied) is met.

[0184] In line 9, the probability distribution P(y t+1 |Y t , X t ) based on y t+1 Select y t+1 indicates a node to be serviced (visited) in the t+1th step.

[0185] In line 10, the new state X t+1 More specifically, we observe the sequence y of nodes visited by the vehicle. 1 , ..., y t-1 , y t , y t+1 Yt+1 and state X t+1 State X t+1 contains information necessary for calculating the reward function (such as the arrival time of each node), which will be described later.

[0186] In line 11, t is updated with t+1.

[0187] Next, in line 13, the reward R is calculated based on the state, as in the first embodiment. Details of the reward will be described later. Next, in line 14, the policy gradient of the actor network is calculated using the following equation (14), and dθ is updated.

[0188] In the 15th line, the policy gradient of the critic network is calculated by the following equation (15), and dθ BL Update.

[0189] In line 16, θ is updated by Adam using dθ. In line 17, dθ BL θ by Adam using BL Update.

[0190] where R denotes the serving route y t is the reward for the sequence of V(X 0 ;θ BL ) is the sum of all raw inputs (X0 ) is a value function that predicts the reward for 0 ;θ BL ) is used as an advantage function to replace the cumulative reward in the conventional reinforcement learning-based VRP method.

[0191] The process shown in FIG. 12 can be summarized as follows.

[0192] First, as a sampling, Monte Carlo simulation is performed using the current weights θ of the actor network to generate B possible sequences based on the current policy.

[0193] Once sampling is complete, the reward and policy gradient are calculated and the actor network is updated in line 16.

[0194] Also, in line 17, the critic network is updated in a direction that reduces the difference between the observed reward and the expected reward. Finally, the gradient dθ and the gradient dθ are calculated at the same learning speed in the end-to-end method. BL Using θ and θ BL Update.

[0195] Next, the reward function will be explained. As in the first embodiment, the second embodiment also uses a reward function that imposes a penalty if the time when a vehicle arrives at a node exceeds the time frame set for that node. The time frame has a start time and an end time. Here, the end time of the time frame is called the leaving time. The leaving time of node i is l i That is, the reward function is the vehicle's leaving time (l i ) ) ), a penalty is incurred. Specifically, the reward function is as shown in the following equation (16).

[0196] p(c, l) is a penalty function, as shown in the following equation (17).

[0197] In formula (17), c irepresents the arrival time at node i. In other words, p(c, l) is the sum, for nodes 1 to N, of the time by which the arrival time exceeds the time window.

[0198] The first term on the right side of equation (16) is "Σ N i=1 c i " represents the travel cost when a vehicle travels to nodes 1 to N. The travel cost may be, for example, the total travel time or the total travel distance when a vehicle travels to nodes 1 to N.

[0199] The same applies when there are multiple vehicles. For example, if there are two vehicles A and B, and vehicle A travels between nodes 1 and 5, and vehicle B travels between nodes 6 and N, then "Σ N i=1 c i " is the sum of the travel cost of vehicle A and the travel cost of vehicle B.

[0200] For each i, the leaving time (l i ) is obtained as initial input information. For each i, the arrival time at node i is obtained as the state when the vehicle arrives at node i. The algorithm calculation unit 230 can calculate R using this information.

[0201] In normal reinforcement learning, learning is performed to maximize the reward, but in the second embodiment (and the first embodiment), the opposite is true; learning is performed to minimize (make as small as possible) the reward R (the sum of the movement cost and the penalty).

[0202] In equation (16), ω is a weight (importance), and setting the weight to a large value will place more importance on the penalty for a time violation. In other words, the importance of the penalty can be adjusted by ω.

[0203] (Hardware Configuration Example) The devices (delivery planning devices 100 and 200) described in the first and second embodiments can both be realized, for example, by causing a computer to execute a program. This computer may be a physical computer or a virtual machine on the cloud.

[0204] That is, the device can be realized by executing a program corresponding to the processing performed by the device using hardware resources such as a CPU and memory built into a computer. The program can be recorded on a computer-readable recording medium (such as a portable memory) and stored or distributed. The program can also be provided via a network such as the Internet or email.

[0205] Fig. 13 is a diagram showing an example of the hardware configuration of the computer. The computer in Fig. 13 includes a drive device 1000, an auxiliary storage device 1002, a memory device 1003, a CPU 1004, an interface device 1005, a display device 1006, an input device 1007, an output device 1008, and the like, all of which are interconnected by a bus BS.

[0206] The program that realizes the processing on the computer is provided by a recording medium 1001, such as a CD-ROM or a memory card. When the recording medium 1001 storing the program is set in the drive device 1000, the program is installed from the recording medium 1001 to the auxiliary storage device 1002 via the drive device 1000. However, the program does not necessarily have to be installed from the recording medium 1001, but may be downloaded from another computer via a network. The auxiliary storage device 1002 stores the installed program as well as necessary files, data, etc.

[0207] When an instruction to start a program is received, the memory device 1003 reads and stores the program from the auxiliary storage device 1002. The CPU 1004 realizes functions related to the device in accordance with the program stored in the memory device 1003. Specifically, for example, in the execution of the algorithms described in the first and second embodiments, the CPU 1004 performs calculations for updating the state, stores the updated data in the memory device 1003, reads the updated data for the next step from the memory device 1003, calculates rewards, and so on.

[0208] The interface device 1005 is used as an interface for connecting to a network, etc. The display device 1006 displays a GUI (Graphical User Interface) or the like according to a program. The input device 1007 is composed of a keyboard, mouse, buttons, a touch panel, etc., and is used to input various operation instructions. The output device 1008 outputs the results of calculations.

[0209] (Effects of the embodiments) As described above, the technology according to the first and second embodiments makes it possible to solve a delivery plan problem taking time constraints into consideration, thereby making it possible to realize a delivery plan under time constraints.

[0210] Furthermore, by using the parameter ω, it is possible to adjust the trade-off relationship (which is more important) between time constraint violation and travel efficiency, and it is possible to simultaneously consider time constraint violation and travel efficiency in a balanced manner according to the actual usage situation.

[0211] Furthermore, the technologies according to the first and second embodiments combine PointerNet, actor-critic neural network models, GCN, etc. with a mechanism that takes into account violations of time constraints, making it possible to instantly find approximate solutions while taking into account actual traffic conditions during the inference phase and re-planning, thereby significantly improving the execution efficiency in practical business applications.

[0212] (Supplementary Notes) This specification discloses at least the matters described in Supplementary Notes 1 to 3 below.

[0213] <Supplementary Note 1> (Supplementary Item 1) A delivery planning device including: a memory; and at least one processor connected to the memory, wherein the processor determines a route for a mobile object departing from a certain location to provide service to a plurality of nodes using a reinforcement learning model, and the processor uses, as a reward function in the reinforcement learning model, a reward function that imposes a penalty if the mobile object arrives at a node at a time beyond a predetermined time frame. (Supplementary Item 2) The delivery planning device according to Supplementary Item 1, wherein the reward function has a travel cost of the mobile object and a penalty function representing the penalty. (Supplementary Item 3) The delivery planning device according to Supplementary Item 2, wherein a weight is added to the penalty function so that the importance of the penalty can be adjusted. (Supplementary Item 4) A delivery planning method executed by a delivery planning device, comprising: an algorithm calculation step of determining a route for a mobile object departing from a certain location to provide service to a plurality of nodes using a reinforcement learning model, wherein in the algorithm calculation step, a reward function for the reinforcement learning model is used that imposes a penalty if the mobile object arrives at a node at a time beyond a predetermined time frame. (Supplementary Item 5) A non-transitory storage medium that stores a program for causing a computer to function as an algorithm calculation unit in the delivery planning device according to any one of Supplementary Item 1 to 3.

[0214] <Supplementary Note 2> (Supplementary Item 1) A delivery planning device comprising: an algorithm calculation unit that uses a neural network that performs reinforcement learning using an actor-critic method to solve a delivery planning problem that determines a route for a vehicle departing from a service center to provide service to a plurality of customers, wherein the algorithm calculation unit solves the delivery planning problem with constraints of a time frame indicating a range of times when the vehicle should arrive at the customer and a time cost indicating the length of time it takes to provide the service to the customer. (Supplementary Item 2) The delivery planning device according to Supplementary Item 1, wherein the algorithm calculation unit masks customers that do not satisfy the time frame constraint from a probability distribution of customers obtained using a decoder in the neural network. (Supplementary Item 3) The delivery planning device according to Supplementary Item 1 or Supplementary Item 2, wherein the algorithm calculation unit masks the probability distribution of customers obtained using the decoder in the neural network so that the vehicle returns to the service center when a value based on the total operating time of the vehicle exceeds a threshold. (Supplementary Item 4) The delivery planning device according to Supplementary Item 3, wherein the algorithm calculation unit sets the total operation time for the next customer to be the sum of the total operation time from the service center to the completion of service at that customer, the travel time from that customer to the next customer, and the time cost at that next customer. (Supplementary Item 5) The delivery planning device according to any one of Supplementary Item 1 to Supplementary Item 4, further comprising a map API unit that plots on a map a route to visit each customer, which is the delivery plan calculated by the algorithm calculation unit. (Supplementary Item 6) A delivery planning method executed by a delivery planning device, comprising: an algorithm calculation step of solving a delivery planning problem that determines a route for a vehicle departing from a service center to provide service to a plurality of customers, using a neural network of reinforcement learning based on an actor-critic method, wherein the algorithm calculation step solves the delivery planning problem using constraints of a time frame indicating a range of times when the vehicle should arrive at the customer and a time cost indicating the length of time it takes to provide the service at the customer. (Supplementary Item 7) A program for causing a computer to function as each unit in the delivery planning device according to any one of Supplementary Items 1 to 5.

[0215] <Supplementary Note 3> (Supplementary Item 1) A delivery planning device comprising: an edge calculation unit that calculates, for each pair of nodes in a plurality of nodes, edge information that is a travel time or travel distance between the nodes; and an algorithm calculation unit that uses a neural network that performs reinforcement learning using an actor-critic method to solve a delivery planning problem that determines a route for a mobile object departing from a certain location to provide service to the plurality of nodes, wherein the algorithm calculation unit solves the delivery planning problem using feature quantities related to the edge information obtained by the edge calculation unit. (Supplementary Item 2) The delivery planning device according to Supplementary Item 1, further comprising a map API unit, wherein the edge calculation unit calculates the edge information based on road conditions obtained by the map API unit. (Supplementary Item 3) The delivery planning device according to Supplementary Item 1 or 2, wherein the algorithm calculation unit has a graph convolutional layer that calculates the feature quantities using the edge information. (Supplementary Item 4) The delivery planning device of any one of Supplementary Items 1 to 3, wherein the algorithm calculation unit includes an attention mechanism, and the attention mechanism uses the feature amount to calculate the probability that each node will be selected as a delivery destination. (Supplementary Item 5) A delivery planning method executed by a computer, comprising: an edge calculation step of calculating, for each pair of nodes in a plurality of nodes, edge information that is a travel time or travel distance between the nodes; and an algorithm calculation step of using a neural network that performs reinforcement learning using an actor-critic method to solve a delivery planning problem that determines a route for a mobile object departing from a certain location to provide service to the plurality of nodes, wherein the algorithm calculation step solves the delivery planning problem using the feature amount related to the edge information obtained in the edge calculation step. (Supplementary Item 6) A program for causing a computer to function as each unit in the delivery planning device of any one of Supplementary Items 1 to 4.

[0216] Although the present embodiment has been described above, the present invention is not limited to such a specific embodiment, and various modifications and changes are possible within the scope of the gist of the present invention described in the claims.

[0217] REFERENCE SIGNS LIST 100 Delivery planning device 110 User information collection unit 120 Service vehicle information collection unit 130 Algorithm calculation unit 140 Map API unit 150 Vehicle dispatch unit 200 Delivery planning device 210 Node information collection unit 220 Service vehicle information collection unit 230 Algorithm calculation unit 240 Map API unit 250 Vehicle dispatch unit 1000 Drive device 1001 Recording medium 1002 Auxiliary storage device 1003 Memory device 1004 CPU 1005 Interface device 1006 Display device 1007 Input device 1008 Output device

Claims

1. A delivery planning device comprising an algorithm calculation unit that uses a reinforcement learning model to determine a route for a mobile object departing from a certain location to provide service to a plurality of nodes, wherein the algorithm calculation unit uses, as a reward function in the reinforcement learning model, a reward function that generates a penalty if the mobile object arrives at a node at a time beyond a predetermined time frame.

2. The delivery planning device according to claim 1, wherein the reward function has a movement cost of the moving object and a penalty function representing the penalty.

3. The delivery planning device according to claim 2, wherein a weight is added to the penalty function to adjust the importance of the penalty.

4. A delivery planning method executed by a delivery planning device, comprising an algorithm calculation step of determining a route for a mobile object departing from a certain location to provide service to a plurality of nodes using a reinforcement learning model, wherein in the algorithm calculation step, a reward function for the reinforcement learning model is used that generates a penalty if the mobile object arrives at a node at a time beyond a predetermined time frame.

5. A program for causing a computer to function as an algorithm calculation unit in the delivery planning device according to any one of claims 1 to 3.

Citation Information

Patent Citations

  • Power saving in radio access network

    JP2023066415A

  • Delivery planning device, delivery planning method, and program

    WO2023053287A1